Method, apparatus, device, and storage medium for generating a talking video

By acquiring phoneme characteristics and acoustic characteristics, using facial key points to extract networks and generate adversarial networks, and generating natural and real speaking videos, the problem of low lip accuracy in the prior art is solved, and high matching and coherent lip changes are achieved.

CN114093384BActive Publication Date: 2025-07-18SHANGHAI SENSETIME TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111386695.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-22
Publication Date
2025-07-18
Estimated Expiration
2041-11-22

AI Technical Summary

Technical Problem

In the existing speaking video generation technology, the lip accuracy is low and the changes are stiff, making it difficult to generate natural and real speaking videos.

Method used

By obtaining phoneme characteristics and acoustic characteristics of sound driving data, a network is extracted and an adversarial network is generated using face key points, and a target face image is generated, and speaking videos are synthesized based on these images to ensure that the lip shape and sound are matched highly and coherent.

Benefits of technology

In the generated speaking video, the lip shape and voice of the target object have a high degree of matching, the lip shape changes coherently, and the speaking state of the target object is real and natural.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114093384B_ABST
    Figure CN114093384B_ABST
Patent Text Reader

Abstract

A method, apparatus, device, and storage medium for generating a speaking video are disclosed. The method includes: obtaining a phoneme feature and an acoustic feature of voice driving data, where the voice driving data includes at least one of audio and text; obtaining at least one set of facial key point information of a target object in a first image according to the phoneme feature and the acoustic feature; obtaining at least one target facial image corresponding to the voice driving data according to the at least one set of facial key point information and a second image including the face of the target object, where a set area of the mouth of the target object in the second image is occluded; and obtaining a speaking video of the target object according to the voice driving data and the at least one target facial image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and particularly to a method, apparatus, device, and storage medium for generating a talking video. Background Art

[0002] The talking video generation technology is an important type in voice-driven character image and cross-modal video generation tasks, and is also a key technology in the commercialization of virtual digital humans. Currently, it is usually adopted to determine the corresponding mouth shape image according to the speech frame, so as to obtain a series of mouth shape images corresponding to the output speech to generate a talking video. However, the mouth shape accuracy of the speaker in the video generated by this method is relatively low and the mouth shape change is rigid. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for generating a talking video.

[0004] According to a first aspect of the present disclosure, there is provided a method for generating a talking video, the method including: obtaining a phoneme feature and an acoustic feature of voice driving data, where the voice driving data includes at least one of audio and text; obtaining at least one set of face key point information of a target object in a first image according to the phoneme feature and the acoustic feature; obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and a second image including a face of the target object, where a set area including a mouth of the target object in the second image is occluded; and obtaining a talking video of the target object according to the voice driving data and the at least one target face image.

[0005] Combined with any implementation manner provided by the present disclosure, the obtaining a phoneme feature and an acoustic feature of voice driving data includes: obtaining phonemes included in the audio corresponding to the voice driving data and time stamps corresponding to the respective phonemes, so as to obtain the phoneme feature of the voice driving data; and performing feature extraction on the audio corresponding to the voice driving data to obtain the acoustic feature of the voice driving data.

[0006] Combined with any implementation manner provided by the present disclosure, the obtaining at least one set of face key point information of a target object in a first image according to the phoneme feature and the acoustic feature includes: obtaining a plurality of sub-phoneme features included in the phoneme feature and corresponding sub-acoustic features; and inputting the sub-phoneme features and the corresponding sub-acoustic features into a face key point extraction network to obtain face key point information corresponding to the sub-phoneme features and the sub-acoustic features.

[0007] Combined with any one of the embodiments provided in the present disclosure, the face key point information includes 3D face key point information. Before obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and the second image including the face of the target object, the method further includes: projecting the 3D face key point information onto a 2D plane to obtain 2D face key point information corresponding to the 3D face key point information; and updating the face key point information by using the 2D face key point information.

[0008] Combined with any one of the embodiments provided in the present disclosure, before obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and the second image including the face of the target object, the method further includes: performing filtering processing on multiple sets of face key point information to make the change amount between the face key point information of each image frame and the face key point information of the adjacent frame meet a set condition.

[0009] Combined with any one of the embodiments provided in the present disclosure, obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and the second image including the face of the target object includes: inputting each set of face key point information and the second image into a face completion network to obtain a target face image corresponding to the face key point information, where the face completion network is used to complete a set area occluded in the second image according to the face key point information.

[0010] Combined with any one of the embodiments provided in the present disclosure, obtaining the speaking video of the target object according to the voice driving data and the at least one target face image includes: fusing the at least one target face image with a set background image to obtain a first image sequence; and obtaining the speaking video of the target object according to the first image sequence and the audio corresponding to the voice driving data.

[0011] Combined with any one of the embodiments provided in the present disclosure, the face key point extraction network is trained by using phoneme feature samples and corresponding acoustic feature samples, where the phoneme feature samples and the acoustic feature samples include the labeled face key point information of the target object.

[0012] Combined with any one of the embodiments provided in the present disclosure, the face key point extraction network is trained in the following manner: training an initial face key point extraction network according to the phoneme feature samples and the corresponding acoustic feature samples, and completing the training to obtain the face key point extraction network when the change of the network loss meets a convergence condition, where the network loss includes the difference between the face key point information predicted by the initial neural network and the labeled face key point information.

[0013] In combination with any of the embodiments provided in the present disclosure, the phoneme feature sample and the acoustic feature sample are obtained by performing annotation of the face key point information of an object on the phoneme feature and the acoustic feature of the audio of the object.

[0014] In combination with any of the embodiments provided in the present disclosure, the phoneme feature sample and the acoustic feature sample are obtained by the following method: obtaining the speaking video of the object; obtaining a plurality of face images and a plurality of audio frames corresponding to the face images according to the speaking video; obtaining the phoneme feature and the acoustic feature of at least one audio frame corresponding to the face image; obtaining the face key point information according to the face image, and performing annotation on the phoneme feature and the acoustic feature according to the face key point information to obtain the phoneme feature sample and the acoustic feature sample.

[0015] In combination with any of the embodiments provided in the present disclosure, the face completion network is trained by using a generative adversarial network, and the generative adversarial network includes the face completion network and a first discriminative network. The network loss of the training includes: a first loss, which is used to indicate the difference between the face completion image output by the face completion network and the complete face image, where the complete face image is the face image corresponding to the face key point information; a second loss, which is used to indicate the difference between the classification result output by the first discriminative network for the input image and the annotation information of the input image, where the annotation information indicates that the input image is the face completion image output by the face completion network or a real face image.

[0016] In combination with any of the embodiments provided in the present disclosure, the generative adversarial network further includes a second discriminative network, and the network loss of the training further includes: a third loss, which is used to indicate the difference between the discrimination result of the second discriminative network for the face completion image corresponding to the phoneme feature and the real corresponding result.

[0017] According to a second aspect of the present disclosure, there is provided a speaking video generation device, where the device includes: a first obtaining unit, which is used to obtain the phoneme feature and the acoustic feature of voice driving data, and the voice driving data includes at least one of audio and text; a second obtaining unit, which is used to obtain at least one set of face key point information of a target object in a first image according to the phoneme feature and the acoustic feature; a first obtaining unit, which is used to obtain at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and a second image including the face of the target object, where a set area of the mouth of the target object in the second image is blocked; a second obtaining unit, which is used to obtain the speaking video of the target object according to the voice driving data and the at least one target face image.

[0018] In combination with any of the embodiments provided in the present disclosure, the first acquisition unit is specifically configured to: acquire the phonemes included in the audio corresponding to the voice driving data and the timestamps corresponding to each phoneme, so as to obtain the phoneme features of the voice driving data; perform feature extraction on the audio corresponding to the voice driving data to obtain the acoustic features of the voice driving data.

[0019] In combination with any of the embodiments provided in the present disclosure, the second acquisition unit is specifically configured to: acquire a plurality of sub-phoneme features included in the phoneme features and the sub-acoustic features corresponding to the plurality of sub-phoneme features; input the sub-phoneme features and the corresponding sub-acoustic features into a face key point extraction network to obtain face key point information corresponding to the sub-phoneme features and the sub-acoustic features.

[0020] In combination with any of the embodiments provided in the present disclosure, the face key point information includes 3D face key point information, and the device further includes a projection unit, configured to project the 3D face key point information onto a 2D plane to obtain 2D face key point information corresponding to the 3D face key point information before obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and a second image including the face of the target object; update the face key point information by using the 2D face key point information.

[0021] In combination with any of the embodiments provided in the present disclosure, the device further includes a filtering unit, configured to perform filtering processing on multiple sets of face key point information before obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and a second image including the face of the target object, so that the change amount between the face key point information of each image frame and the face key point information of an adjacent frame meets a set condition.

[0022] In combination with any of the embodiments provided in the present disclosure, the first obtaining unit is specifically configured to: input each set of face key point information and the second image into a face completion network to obtain a target face image corresponding to the face key point information, where the face completion network is used to complete a set area occluded in the second image according to the face key point information.

[0023] In combination with any of the embodiments provided in the present disclosure, the second obtaining unit is specifically configured to: fuse the at least one target face image with a set background image to obtain a first image sequence; obtain a speaking video of the target object according to the first image sequence and the audio corresponding to the voice driving data.

[0024] In combination with any of the embodiments provided in the present disclosure, the face key point extraction network is trained using phoneme feature samples and corresponding acoustic feature samples, wherein the phoneme feature samples and the acoustic feature samples include the labeled face key point information of the target object.

[0025] In combination with any of the embodiments provided in the present disclosure, the face key point extraction network is trained in the following manner: training an initial face key point extraction network according to the phoneme feature samples and the corresponding acoustic feature samples, and completing the training to obtain the face key point extraction network when the change of the network loss meets the convergence condition, wherein the network loss includes the difference between the face key point information predicted by the initial neural network and the labeled face key point information.

[0026] In combination with any of the embodiments provided in the present disclosure, the phoneme feature samples and the acoustic feature samples are obtained by annotating the face key point information of an object on the phoneme features and acoustic features of the audio of the object.

[0027] In combination with any of the embodiments provided in the present disclosure, the phoneme feature samples and the acoustic feature samples are obtained in the following manner: obtaining the speaking video of the object; obtaining multiple face images and corresponding multiple audio frames according to the speaking video; obtaining the phoneme features and acoustic features of at least one audio frame corresponding to the face image; obtaining the face key point information according to the face image, and annotating the phoneme features and the acoustic features according to the face key point information to obtain the phoneme feature samples and the acoustic feature samples.

[0028] In combination with any of the embodiments provided in the present disclosure, the face completion network is trained using a generative adversarial network, and the generative adversarial network includes the face completion network and a first discriminator network. The network loss for the training includes: a first loss, which is used to indicate the difference between the face completion image output by the face completion network and the complete face image, wherein the complete face image is the face image corresponding to the face key point information; a second loss, which is used to indicate the difference between the classification result output by the first discriminator network for the input image and the annotation information of the input image, wherein the annotation information indicates that the input image is the face completion image output by the face completion network or a real face image.

[0029] In combination with any of the embodiments provided in the present disclosure, the generative adversarial network further includes a second discriminator network, and the network loss for the training further includes: a third loss, which is used to indicate the difference between the discrimination result of the second discriminator network for the face completion image corresponding to the phoneme feature and the real corresponding result.

[0030] According to a third aspect of the present disclosure, there is provided an electronic device, including a memory and a processor. The memory is configured to store computer instructions that can run on the processor, and the processor is configured to implement the talking video generation method according to any implementation manner provided by the present disclosure when executing the computer instructions.

[0031] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the talking video generation method according to any implementation manner provided by the present disclosure.

[0032] For the talking video generation method, device, equipment and computer-readable storage medium of one or more embodiments of the present disclosure, at least one set of facial key point information of a target object in a first image is obtained according to the phoneme feature and acoustic feature of voice driving data; and at least one target facial image corresponding to the voice driving data is obtained according to the at least one set of facial key point information and a second image including the face of the target object, wherein a set area of the mouth of the target object in the second image is occluded; finally, a talking video of the target object is obtained according to the voice driving data and the at least one target facial image. In the embodiments of the present disclosure, the target facial image is generated according to the facial key information of the target object corresponding to the voice driving data and the image of the target object with the mouth occluded. In the obtained talking video of the target object, the mouth shape of the target object has a high matching degree with the voice driving data, and the mouth shape change is coherent. The talking state of the target object is real and natural. Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in one or more embodiments of the present specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in one or more embodiments of the present specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0034] Figure 1 is a flowchart of a talking video generation method proposed by at least one embodiment of the present disclosure;

[0035] Figure 2 is a flowchart of a facial key point extraction network training method proposed by at least one embodiment of the present disclosure;

[0036] Figure 3 is a flowchart of a sample acquisition method proposed by at least one embodiment of the present disclosure;

[0037] Figure 4It is a flowchart of another method for generating a talking video proposed by at least one embodiment of the present disclosure;

[0038] Figure 5 is Figure 4 a schematic diagram of the method for generating a talking video shown;

[0039] Figure 6 is a schematic diagram of obtaining face key point information in the method for generating a talking video proposed by at least one embodiment of the present disclosure;

[0040] Figure 7 is a schematic structural diagram of a device for generating a talking video proposed by at least one embodiment of the present disclosure;

[0041] Figure 8 is a schematic structural diagram of an electronic device proposed by at least one embodiment of the present disclosure. Detailed implementation manners

[0042] Here, exemplary embodiments will be described in detail, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0043] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" in this article represents any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0044] At least one embodiment of the present disclosure provides a method for generating a talking video. This method can be executed by an electronic device such as a terminal device or a server. The terminal device can be a fixed terminal or a mobile terminal, such as a mobile phone, a tablet computer, a game console, a desktop computer, an advertising machine, an all-in-one machine, a vehicle-mounted terminal, etc. The server includes a local server or a cloud server, etc. This method can also be implemented by a processor calling computer-readable instructions stored in a memory.

[0045] Figure 1 shows a flowchart of the method for generating a talking video according to at least one embodiment of the present disclosure, as Figure 1 shown, the method includes steps 101 to 104.

[0046] In step 101, obtain the phoneme features and acoustic features of the voice drive data.

[0047] A phoneme is the smallest speech unit that makes up a syllable. Phoneme features may include features representing the start and end times of the pronunciation of each phoneme contained in the audio corresponding to the voice drive data. Taking the audio corresponding to the voice drive data as the speech segment of "Hello" as an example, the phoneme features of the voice drive data may include, for example: n[0,0.2], i 3[0.2,0.4], h[0.5,0.7], ao3[0.7,1.2], where the values within [] indicate the start and end times of the pronunciation of each phoneme, and the unit is, for example, seconds. In the embodiments of the present disclosure, the phoneme features of the voice drive data can be obtained by acquiring the phonemes contained in the audio corresponding to the voice drive data and the timestamps corresponding to each phoneme.

[0048] Acoustic features are mainly used to describe the pronunciation characteristics of audio. The acoustic features include but are not limited to at least one of linear prediction parameters, Mel frequency cepstral coefficients, perceptual linear prediction coefficients, etc. In the embodiments of the present disclosure, the acoustic feature is, for example, the Mel frequency cepstral coefficient. In the embodiments of the present disclosure, the acoustic features of the voice drive data can be obtained by performing feature extraction on the audio corresponding to the voice drive data.

[0049] In the embodiments of the present disclosure, the voice drive data may be at least one of audio and text.

[0050] In the case where the voice drive data only includes audio, the text (character information) corresponding to the audio can be determined by performing speech recognition on the audio, so that the audio and text corresponding to the voice drive data can be obtained;

[0051] In the case where the voice drive data only includes text, the character information corresponding to the text can be converted into audio (speech segment) by performing speech synthesis on the text, so that the audio and text corresponding to the voice drive data can be obtained;

[0052] In the case where the voice drive data includes both audio and text, the audio and text correspond to the same pronunciation. For example, when the text is "Hello", the voice in the voice drive data is the speech segment that emits the sound of "Hello".

[0053] In some embodiments, by performing an alignment operation on the audio and text corresponding to the voice driving data, the phoneme features of the voice driving data can be obtained. The alignment operation refers to aligning each speech segment in the audio with the phonemes in the text corresponding to the pronunciation of the speech segment, that is, determining when in the audio the pronunciation corresponding to the text starts. By performing the alignment operation on the audio and text, on the one hand, the phonemes included in the audio are determined, and at the same time, according to the duration of the pronunciation, the timestamps corresponding to each phoneme can be obtained, so that the phoneme features of the voice driving data can be obtained.

[0054] Still taking "hello" as an example, after performing the alignment operation on the speech and text, it can be determined that the sound of the phoneme "n" is emitted from 0 to 0.2 seconds, the sound of "i 3" is emitted from 0.2 to 0.4 seconds, and so on, so that the phoneme features of the voice driving data can be obtained. Those skilled in the art should understand that the pronunciation of the voice driving data can also be obtained by other means, and the embodiments of the present disclosure do not limit this.

[0055] In step 102, at least one set of facial key point information of the target object in the first image is obtained according to the phoneme features and the acoustic features.

[0056] When a person emits different voices, the mouth shape will change accordingly. Correspondingly, the positions of the facial key points in the mouth area or the set area including the mouth area will change accordingly. It can be seen from this that for the target object, the phoneme features and acoustic features of a voice frame correspond to a set of facial key point information. When the target object emits the pronunciation of a certain phoneme, the corresponding facial key point information of its face can be determined. The facial key point information includes the position information of the key points corresponding to the facial features and the facial contour in the first image. In the present disclosure, the information of each facial key point at the same moment can be referred to as a set of facial key point information.

[0057] Taking the generation of the talking video of the target object in the first image as an example, in this step, at least one set of facial key point information of the target object in the first image is obtained according to the phoneme features and the acoustic features of the voice driving data. When the audio corresponding to the voice driving data includes multiple phonemes, a sequence of facial key point information corresponding to these phonemes and the corresponding acoustic features can be obtained. The facial key point sequence includes multiple sets of facial key point information arranged in chronological order.

[0058] The embodiments of the present disclosure also add acoustic features on the basis of the phoneme features of the voice driving data, so that the obtained facial key point information is more matched with the pronunciation features of the audio corresponding to the voice driving data, and the subsequent generated talking video is more realistic.

[0059] In step 103, at least one target face image corresponding to the voice driving data is obtained according to the at least one set of face key point information and a second image including the face of the target object.

[0060] Wherein, the second image is an image including the face of the target object, and the second image may be the same image as the first image or a different image. For example, if the first image is a face image of target object A smiling, the second image may also be this smiling face image of target object A, or other face images of target object A.

[0061] A set area of the mouth of the target object in the second image is blocked. The set area includes the area where the positions of the face key points change when the target object is speaking. For example, it may be the lower half of the face of the target object, or the face area below the forehead, or the mouth area. The specific blocked area is not limited in the embodiments of the present disclosure.

[0062] In some embodiments, the second image with the set area blocked can be generated by filling noise in the set area. Filling noise in the set area means setting each pixel in the set area with randomly generated pixel values. Those skilled in the art should understand that the set area can also be blocked by other means, and the present disclosure does not limit this.

[0063] According to the at least one set of face key point information obtained in step 102, the blocked part in the second image can be completed, so that the distribution of the face key points in the blocked area of the target object in the second image is consistent with the phoneme features and acoustic features of the voice driving data. In this way, in the at least one target face image generated according to the at least one set of face key point information and the second image, the face key information in the blocked set area matches the voice driving data.

[0064] In step 104, a speaking video of the target object is obtained according to the voice driving data and the at least one target face image.

[0065] In the embodiments of the present disclosure, in the obtained speaking video of the target object, the output voice is the audio corresponding to the voice driving data, and in each face image of the speaking video, the face key point information corresponds to the phoneme features and acoustic features of the output speech. Thus, the mouth shape and speaking expression of the generated target object are consistent with the pronunciation, giving the audience the feeling that the target object is speaking.

[0066] Embodiments of the present disclosure generate a target face image based on the key face information of the target object corresponding to the voice driving data and the image of the target object with the mouth occluded. In the obtained speaking video of the target object, the lip-sync of the target object has a high degree of matching with the voice driving data, and the lip movement is coherent. The speaking state of the target object is real and natural.

[0067] In some embodiments, the at least one target face image can be fused with a set background image to obtain a first video, and based on the audio corresponding to the first video and the voice driving data, the speaking video of the target object can be obtained. In one example, the pixels in the face region of the target face image can be used as foreground pixels and superimposed on the set background image to achieve the fusion of the target face image and the set background image. Those skilled in the art should understand that various methods can be used to fuse the target face image and the set background image, and the present disclosure does not limit this.

[0068] Through the above method, a speaking video of the target object in any background can be generated, enriching the application scenarios of the speaking video generation method.

[0069] In some embodiments, a face key point extraction network can be used to obtain at least one set of face key point information of the target object in the first image corresponding to the phoneme feature and the acoustic feature.

[0070] First, obtain a plurality of sub-phoneme features included in the phoneme feature and the sub-acoustic features corresponding to the plurality of sub-phoneme features.

[0071] In one example, the plurality of sub-phoneme features and sub-acoustic features included in the phoneme feature can be obtained by performing a sliding window on the phoneme feature and the acoustic feature of the voice driving data. For example, the phoneme feature and the acoustic feature can be divided into a plurality of sub-phoneme features and sub-acoustic features according to the length of the time window. Specifically, during the process of performing a sliding window on the phoneme feature and the acoustic feature of the voice driving data, the phoneme feature and the acoustic feature obtained in each time window after each sliding window operation can be used as the sub-phoneme feature and the sub-acoustic feature, and the sub-phoneme feature and the sub-acoustic feature in the same time window correspond to the same speech segment.

[0072] Next, input the sub-phoneme features and the corresponding sub-acoustic features into the trained face key point extraction network to obtain the face key point information corresponding to the sub-phoneme features and the sub-acoustic features. Specifically, multiple sub-phoneme features and the corresponding multiple sub-acoustic features can be input into the face key point extraction network in the form of multiple sub-phoneme feature - sub-acoustic feature pairs in chronological order. The face key point extraction network is used to determine a corresponding set of face key point information according to each sub-phoneme feature - sub-acoustic feature pair. After all the sub-phoneme feature - sub-acoustic feature pairs are input into the face key point extraction network, multiple sets of face key point information corresponding to the voice driving data can be obtained.

[0073] In the embodiment of the present disclosure, by using the trained face key point extraction network to obtain the face key point information corresponding to each sub-phoneme feature - sub-acoustic feature pair, a good match between the pronunciation of the target object and the mouth shape and speaking expression can be achieved.

[0074] In the embodiment of the present disclosure, the face key point generation network can be a 3D face key point generation network, that is, the output face key point information is 3D face key point information, which includes not only the position information of the face key points but also the depth information of the face key points; the face key point generation network can also be a 2D face key point generation network, that is, the output face key point information is 2D face key point information.

[0075] In the case where the face key point information is 3D face key point information, before obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and the second image including the face of the target object, the method further includes: projecting the 3D face key point information onto a 2D plane to obtain 2D face key point information corresponding to the 3D face key point information; using the 2D face key point information to update the face key point information. Then, according to at least one set of 2D face key point information and the second image including the face of the target object, at least one target face image corresponding to the voice driving data is obtained; finally, according to the voice driving data and the at least one target face image, the speaking video of the target object is obtained.

[0076] In some embodiments, filtering processing can be performed on multiple sets of face key point information so that the change amount between the face key point information of each image frame and the face key point information of adjacent frames (including the previous frame and / or the next frame) meets a set condition. The set condition can, for example, include that the change amount between the positions of each face key point and the positions of the corresponding face key points in the adjacent frames is less than a set threshold. By the above method, the jitter frames with a large change amplitude of the face key point information can be filtered out, avoiding the situation that the mouth shape suddenly changes in the generated speaking video.

[0077] In one example, the moving average processing of consecutive frames in the multiple face key-point information can be achieved by performing Gaussian filtering on the multiple sets of face key-point information in a time window. The moving average processing refers to performing weighted averaging on the values of the face key-points of each frame and the values of the face key-points of adjacent frames, and updating the values of the face key-points of this frame using the result of the weighted averaging.

[0078] In some embodiments, at least one target face image corresponding to the voice driving data can be obtained in the following manner: inputting each set of face key-point information and the second image into a face completion network to obtain a target face image corresponding to the face key-point information, where the face completion network is used to complete a set area occluded in the second image according to the face key-point information.

[0079] In the embodiments of the present disclosure, by using the face completion network to complete the set area occluded in the second image according to the face key-point information, the face key-point information of the set area can be made consistent with the input face key-point information, so that the mouth shape and speaking expression of the target object match the uttered voice. Moreover, by using the face completion network to complete the set area occluded in the second image, a target face image with high clarity can be generated.

[0080] In some embodiments, the face key-point extraction network can be trained using phoneme feature samples and acoustic feature samples. This training method can be executed by a server, and the server executing this training method and the device executing the above-mentioned speaking video generation method can be different.

[0081] Figure 2 is a method for training a face key-point extraction network proposed in at least one embodiment of the present disclosure, as Figure 2 shown, this training method includes steps 201 to 202.

[0082] In step 201, phoneme feature samples and corresponding acoustic feature samples are obtained, and the phoneme feature samples and the acoustic feature samples include the labeled face key-point information of the target object. Among them, the phoneme feature samples and the corresponding acoustic feature samples are obtained based on the same speech segment, and the face key-point information labeled in the phoneme feature samples and the corresponding acoustic feature samples is the same.

[0083] In step 202, the initial face key point extraction network is trained according to the phoneme feature samples and the corresponding acoustic feature samples, and the face key point extraction network is obtained after the training is completed when the change of the network loss meets the convergence condition, where the network loss includes the difference between the face key point information predicted by the initial neural network and the labeled face key point information.

[0084] In some embodiments, the phoneme feature samples and the acoustic feature samples are obtained by annotating the phoneme features and acoustic features of the audio of an object with the face key point information of the object. In one example, the phoneme feature samples and the corresponding acoustic feature samples can be obtained by the method shown in Figure 3 the figure.

[0085] In step 301, the talking video of the object is obtained. Wherein, the object can be the target object for which the generated talking video is directed, or an object different from the target object.

[0086] In one example, when it is desired to generate a talking video of a certain target object, the existing talking video of the target object is obtained for obtaining phoneme feature samples and acoustic feature samples.

[0087] In step 302, multiple face images and multiple audio frames corresponding to the face images are obtained according to the talking video.

[0088] By splitting the talking video, the corresponding speech segments of the talking video and multiple face images included in the talking video are obtained. Wherein, multiple audio frames in the speech segments have a corresponding relationship with the multiple face images.

[0089] In step 303, the phoneme features and acoustic features of at least one audio frame corresponding to the face image are obtained.

[0090] According to the corresponding relationship between the multiple face images and the multiple audio frames in the speech segment, the phoneme features and acoustic features of at least one audio frame corresponding to any face image are obtained.

[0091] In step 304, face key point information is obtained according to the face image, and the phoneme features and the acoustic features are annotated according to the face key point information to obtain the phoneme feature samples and the acoustic feature samples.

[0092] In an embodiment of the present disclosure, by generating a phoneme feature sample and an acoustic feature sample based on an existing speaking video of a target object for which a speaking video is to be generated, an association between the phoneme feature and the acoustic feature of the speech of the target object's speech and the facial key point information can be accurately established, which is conducive to better implementing the training of the facial key point generation network.

[0093] In some embodiments, the face completion network can be trained using a generative adversarial network. This training method can be executed by a server, and the server executing this training method and the device executing the above-mentioned speaking video generation method can be different.

[0094] Among them, the generative adversarial network includes the face completion network and a first discriminator network. The face completion network is used to complete an input occluded face image according to the facial key point information to generate a face completion image. The occluded face image is obtained by occluding a set area including the mouth in a complete face image, and the complete face image can be the face image corresponding to the facial key point information. The generated face completion image and the real face image are randomly input into the first discriminator network, and the first discriminator network outputs a discrimination result for the input image, that is, determines whether the input image is a face completion image or a real face image.

[0095] The loss for training the face completion network using the generative adversarial network includes:

[0096] The first loss is used to indicate the difference between the face completion image output by the face completion network and the complete face image, where the complete face image is the face image corresponding to the facial key point information;

[0097] The second loss is used to indicate the difference between the classification result output by the first discriminator network for the input image and the annotation information of the input image, where the annotation information indicates that the input image is the face completion image output by the face completion network or a real face image.

[0098] When the change of the loss in the training satisfies the convergence condition, the training is completed to obtain the face completion network.

[0099] In an embodiment of the present disclosure, training the face completion network using a generative adversarial network can improve the accuracy of the face completion image output by the face completion network, which is conducive to improving the image quality of the generated speaking video of the target object.

[0100] In some embodiments, a second discrimination network for determining whether the face completion image is aligned with the phoneme features may be added to assist in the training of the face completion network. In this training method, the face completion image output by the face completion network is input into the second discrimination network.

[0101] The loss of this training, in addition to the above-mentioned first loss and second loss, further includes a third loss, and the third loss is used to indicate the difference between the discrimination result of the second discrimination network for the correspondence between the face completion image and the phoneme features and the true correspondence result.

[0102] By adding the second discrimination network to train the face completion network, the alignment effect between the phoneme features and the face key points is further improved, which is beneficial to improving the quality of the talking video.

[0103] The following combines Figure 4 the flowchart of the talking video generation method shown in Figure 5 the schematic diagram of the talking video generation method shown in Figure 6 and the schematic diagram of obtaining face key point information shown in

[0104] In step 401, an alignment operation is performed on the audio and text corresponding to the voice driving data to obtain the phoneme features of the voice driving data.

[0105] The audio corresponding to the voice driving data is, for example, Figure 6 as shown, the voice segment of "Hello", and the text corresponding to the voice driving data is the text of "Hello". By performing an alignment operation on the audio and text, the phoneme features of the voice driving data are obtained.

[0106] As Figure 5 shown, in the case where the voice driving data only includes audio, the text corresponding to the audio can be determined by performing speech recognition on the audio; in the case where the voice driving data only includes text, the text can be synthesized into audio by performing speech synthesis on the text to convert the text information into audio.

[0107] In step 402, feature extraction is performed on the audio corresponding to the voice driving data to obtain the Mel cepstrum features of the voice driving data, that is, Mel Frequency Cepstral Coefficients.

[0108] In step 403, by performing a sliding window on the phoneme features and acoustic features of the voice driving data, a plurality of sub-phoneme features and sub-acoustic features included in the phoneme features are obtained. The time window is as Figure 6As shown by the dashed box in [the figure], the arrow indicates the sliding direction of the time window. During the activity, the phoneme features and acoustic features within each obtained time window are sub-phoneme features and sub-acoustic features, and the sub-phoneme features and sub-acoustic features within the same time window correspond to the same speech segment.

[0109] In step 404, the sub-phoneme features and the corresponding sub-acoustic features are input into the trained face key point extraction network to obtain face key point information corresponding to the sub-phoneme features and the sub-acoustic features. As Figure 6 shown, for the sub-phoneme features and sub-acoustic features within each obtained time window, the face key point extraction network outputs the face key point information corresponding to this time window.

[0110] Exemplarily, the face key point extraction network is a 3D face key point extraction network. Correspondingly, the obtained face key point information is 3D face key point information.

[0111] In step 405, 2D face key point information corresponding to the 3D face key point information is obtained.

[0112] In step 406, filtering processing is performed on multiple groups of 2D face key point information so that the change amount between the 2D face key point information of each image frame and the face key point information of the adjacent frame meets the set conditions.

[0113] In step 407, each group of filtered 2D face key point information and the second image are input into the face completion network to obtain a target face image corresponding to the 2D face key point information, where the second image is an occluded face, and the lower half of the second image is filled with noise for occlusion.

[0114] In step 408, multiple frames of target face images (speaking face images) obtained in step 207 are fused with the background image to obtain a first image sequence.

[0115] In step 409, according to the first image sequence and the audio corresponding to the voice driving data, the speaking video of the target object is obtained.

[0116] Figure 7 is a schematic structural diagram of a speaking video generation device proposed by at least one embodiment of the present disclosure; as Figure 7As shown in the figure, the device includes: a first acquisition unit 701, configured to acquire the phoneme features and acoustic features of the voice driving data, where the voice driving data includes at least one of audio and text; a second acquisition unit 702, configured to acquire at least one set of face key point information of the target object in the first image according to the phoneme features and the acoustic features; a first obtaining unit 703, configured to obtain at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and a second image including the face of the target object, where a set area of the mouth of the target object is blocked in the second image; a second obtaining unit 704, configured to obtain a speaking video of the target object according to the voice driving data and the at least one target face image.

[0117] Combined with any one of the embodiments provided in the present disclosure, the first acquisition unit is specifically configured to: acquire the phonemes included in the audio corresponding to the voice driving data and the timestamps corresponding to the respective phonemes, to obtain the phoneme features of the voice driving data; perform feature extraction on the audio corresponding to the voice driving data, to obtain the acoustic features of the voice driving data.

[0118] Combined with any one of the embodiments provided in the present disclosure, the second acquisition unit is specifically configured to: acquire a plurality of sub-phoneme features included in the phoneme features and the corresponding sub-acoustic features; input the sub-phoneme features and the corresponding sub-acoustic features into a face key point extraction network, to obtain face key point information corresponding to the sub-phoneme features and the sub-acoustic features.

[0119] Combined with any one of the embodiments provided in the present disclosure, the face key point information includes 3D face key point information, and the device further includes a projection unit, configured to project the 3D face key point information onto a 2D plane to obtain 2D face key point information corresponding to the 3D face key point information before obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and a second image including the face of the target object; update the face key point information by using the 2D face key point information.

[0120] Combined with any one of the embodiments provided in the present disclosure, the device further includes a filtering unit, configured to perform filtering processing on multiple sets of face key point information before obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and a second image including the face of the target object, so that the change amount between the face key point information of each image frame and the face key point information of an adjacent frame satisfies a set condition.

[0121] In combination with any of the embodiments provided in the present disclosure, the first obtaining unit is specifically configured to: input each group of face key point information and the second image into a face completion network to obtain a target face image corresponding to the face key point information, where the face completion network is used to complete a set area occluded in the second image according to the face key point information.

[0122] In combination with any of the embodiments provided in the present disclosure, the second obtaining unit is specifically configured to: fuse the at least one target face image with a set background image to obtain a first image sequence; obtain a talking video of the target object according to the first image sequence and an audio corresponding to the voice driving data.

[0123] In combination with any of the embodiments provided in the present disclosure, the face key point extraction network is trained by using a phoneme feature sample and a corresponding acoustic feature sample, where the phoneme feature sample and the acoustic feature sample include the labeled face key point information of the target object.

[0124] In combination with any of the embodiments provided in the present disclosure, the face key point extraction network is trained in the following manner: according to the phoneme feature sample and the corresponding acoustic feature sample, train an initial face key point extraction network, and complete the training to obtain the face key point extraction network when the change of the network loss satisfies a convergence condition, where the network loss includes the difference between the face key point information predicted by the initial neural network and the labeled face key point information.

[0125] In combination with any of the embodiments provided in the present disclosure, the phoneme feature sample and the acoustic feature sample are obtained by labeling the face key point information of an object for the phoneme feature and the acoustic feature of the audio of the object.

[0126] In combination with any of the embodiments provided in the present disclosure, the phoneme feature sample and the acoustic feature sample are obtained in the following manner: obtain a talking video of the object; obtain multiple face images according to the talking video, and multiple audio frames corresponding to the face images; obtain the phoneme feature and the acoustic feature of at least one audio frame corresponding to the face image; obtain the face key point information according to the face image, and label the phoneme feature and the acoustic feature according to the face key point information to obtain the phoneme feature sample and the acoustic feature sample.

[0127] In combination with any of the embodiments provided by the present disclosure, the face completion network is trained using a generative adversarial network, and the generative adversarial network includes the face completion network and a first discriminator network. The network loss for the training includes: a first loss, which is used to indicate the difference between the face completion image output by the face completion network and the complete face image, where the complete face image is the face image corresponding to the face key point information; a second loss, which is used to indicate the difference between the classification result output by the first discriminator network for the input image and the annotation information of the input image, where the annotation information indicates that the input image is the face completion image output by the face completion network or a real face image.

[0128] In combination with any of the embodiments provided by the present disclosure, the generative adversarial network further includes a second discriminator network, and the network loss for the training further includes: a third loss, which is used to indicate the difference between the discrimination result of the second discriminator network for the face completion image corresponding to the phoneme feature and the real corresponding result.

[0129] At least one embodiment of the present disclosure further provides an electronic device, as Figure 8 shown. The device includes a memory and a processor. The memory is used to store computer instructions that can run on the processor, and the processor is used to implement the speaking video generation method according to any embodiment of the present disclosure when executing the computer instructions.

[0130] At least one embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the speaking video generation method according to any embodiment of the present disclosure.

[0131] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0132] The embodiments in this specification are all described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the embodiment of the data processing device, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0133] The foregoing describes particular embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures need not be in the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0134] Embodiments of the subject matter and the functional operations described in this specification can be implemented in: digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or one or more of them in combination. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier to be executed by, or to control the operation of, a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, generated to encode and transmit information to a suitable receiver apparatus for execution by the data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0135] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform the corresponding functions by operating on input data and generating output. The processes and logical flows can also be performed by, or the apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0136] Computers suitable for executing computer programs include, for example, general and / or special purpose microprocessors, or any other type of central processing unit. Generally, the central processing unit will receive instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, etc., or the computer will be operably coupled to such mass storage devices to receive data therefrom or transfer data thereto, or both. However, a computer is not necessarily required to have such devices. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few.

[0137] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as including semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0138] Although this specification contains many specific implementation details, these should not be construed as limiting the scope of any invention or the scope of what is claimed, but rather as mainly describing the features of specific embodiments of a particular invention. Certain features described in multiple embodiments in this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may act in certain combinations as described above and even be claimed as such initially, one or more features from a claimed combination may in some cases be removed from that combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0139] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or sequentially, or that all illustrated operations be performed, to achieve a desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0140] Accordingly, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. In addition, the processes depicted in the figures are not necessarily in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0141] The foregoing is only a preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.

Claims

1. A method for generating a talking video, characterized in that, The method includes: Obtaining the phoneme features and acoustic features of the voice driving data, where the voice driving data includes at least one of audio and text; the phoneme features include features representing the pronunciation start and end times of each phoneme included in the audio corresponding to the voice driving data; the acoustic features are obtained by extracting features from the audio corresponding to the voice driving data; Obtaining at least one set of face key point information of the target object in the first image according to the phoneme features and the acoustic features, including: Obtaining a plurality of sub-phoneme features included in the phoneme features and the corresponding sub-acoustic features of the plurality of sub-phoneme features; inputting the sub-phoneme features and the corresponding sub-acoustic features into a face key point extraction network to obtain face key point information corresponding to the sub-phoneme features and the sub-acoustic features; wherein, the face key point information includes the position information of the key points corresponding to the facial features and the face contour in the first image; Obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and a second image including the face of the target object, where a set area of the mouth of the target object in the second image is occluded; Obtaining the speaking video of the target object according to the voice driving data and the at least one target face image.

2. The method according to claim 1, characterized in that, The obtaining the phoneme features and acoustic features of the voice driving data includes: Obtaining the phonemes included in the audio corresponding to the voice driving data and the time stamps corresponding to each phoneme to obtain the phoneme features of the voice driving data; Extracting features from the audio corresponding to the voice driving data to obtain the acoustic features of the voice driving data.

3. The method according to claim 1, wherein The face key point information includes 3D face key point information. Before obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and a second image including the face of the target object, the method further includes: Projecting the 3D face key point information onto a 2D plane to obtain 2D face key point information corresponding to the 3D face key point information; Updating the face key point information by using the 2D face key point information.

4. The method according to claim 1, wherein Before obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and a second image including the face of the target object, the method further includes: Performing filtering processing on multiple sets of face key point information so that the change amount between the face key point information of each image frame and the face key point information of the adjacent frame meets a set condition.

5. The method according to claim 1, wherein The obtaining at least one target face image corresponding to the voice driving data according to the at least one set of face key point information and a second image including the face of the target object includes: Inputting each set of face key point information and the second image into a face completion network to obtain a target face image corresponding to the face key point information, where the face completion network is used to complete the set area occluded in the second image according to the face key point information.

6. The method according to claim 1, characterized in that Obtaining the talking video of the target object according to the voice driving data and the at least one target face image includes: Fusing the at least one target face image with a set background image to obtain a first image sequence; Obtaining the talking video of the target object according to the first image sequence and the audio corresponding to the voice driving data.

7. The method according to any one of claims 3 to 6, characterized in that The face key point extraction network is trained using phoneme feature samples and corresponding acoustic feature samples, where the phoneme feature samples and the acoustic feature samples include the labeled face key point information of the target object.

8. The method according to claim 7, wherein The face key point extraction network is trained through the following method: Training an initial face key point extraction network according to the phoneme feature samples and the corresponding acoustic feature samples, and completing the training to obtain the face key point extraction network when the change of the network loss meets the convergence condition, where the network loss includes the difference between the face key point information predicted by the initial face key point extraction network and the labeled face key point information.

9. The method according to claim 7, wherein The phoneme feature samples and the acoustic feature samples are obtained by annotating the face key point information of an object to the phoneme features and acoustic features of the audio of the object.

10. The method according to claim 9, wherein The phoneme feature samples and the acoustic feature samples are obtained through the following method: Obtaining the talking video of the object; Obtaining multiple face images according to the talking video, and multiple audio frames corresponding to the face images; Obtaining the phoneme features and acoustic features of at least one audio frame corresponding to the face image; Obtaining face key point information according to the face image, and annotating the phoneme features and the acoustic features according to the face key point information to obtain the phoneme feature samples and the acoustic feature samples.

11. The method according to claim 5, characterized in that, The face completion network is trained using a generative adversarial network, the generative adversarial network includes the face completion network and a first discriminator network, and the network loss of the training includes: A first loss, which is used to indicate the difference between the face completion image output by the face completion network and the complete face image, where the complete face image is the face image corresponding to the face key point information; A second loss, which is used to indicate the difference between the classification result output by the first discriminator network for the input image and the annotation information of the input image, where the annotation information indicates that the input image is the face completion image output by the face completion network or a real face image.

12. The method according to claim 11, wherein The generative adversarial network further includes a second discriminator network, and the network loss of the training further includes: A third loss, which is used to indicate the difference between the discrimination result of the second discriminator network for the face completion image corresponding to the phoneme feature and the real corresponding result.

13. A speech video generation device, characterized in that, The device includes: A first acquisition unit, which is used to acquire the phoneme features and acoustic features of voice driving data, where the voice driving data includes at least one of audio and text; the phoneme features include features representing the pronunciation start and end times of each phoneme included in the audio corresponding to the voice driving data; the acoustic features are obtained by performing feature extraction on the audio corresponding to the voice driving data; A second acquisition unit, configured to acquire at least one set of facial key point information of a target object in the first image according to the phoneme feature and the acoustic feature; wherein, the facial key point information includes the position information of key points corresponding to facial features and the facial contour in the first image. A first obtaining unit, configured to obtain at least one target facial image corresponding to the voice driving data according to the at least one set of facial key point information and a second image including the face of the target object, wherein a set area of the mouth of the target object in the second image is occluded. A second obtaining unit, configured to obtain a speaking video of the target object according to the voice driving data and the at least one target facial image. Wherein, the second acquisition unit is specifically configured to: acquire a plurality of sub-phoneme features included in the phoneme feature and sub-acoustic features corresponding to the plurality of sub-phoneme features; input the sub-phoneme features and the corresponding sub-acoustic features into a facial key point extraction network to obtain facial key point information corresponding to the sub-phoneme features and the sub-acoustic features.

14. An electronic device, characterized in that, The device includes a memory and a processor, the memory is used to store computer instructions that can be run on the processor, and the processor is used to implement the method according to any one of claims 1 to 12 when executing the computer instructions.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Video generation method and device, electronic equipment and computer storage medium

    CN110677598A

  • Voice synthesis method and device based on virtual character, medium and equipment

    CN111369967A