Video generation method and apparatus, and device, medium and product

By acquiring key information from audio and video and combining it with a lip-sync rendering model, a video that matches the audio is generated, solving the problem of inconsistent video generation in existing technologies and improving the quality of video generation.

WO2025213845A1PCT designated stage Publication Date: 2025-10-16BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/140100
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-08
Filing Date
2024-12-17
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively generate videos that match audio sequences in scenarios involving language switching, sentence modification, and sentence replacement, especially when maintaining consistency in facial expressions and postures.

Method used

By acquiring reference video and audio sequences, and utilizing the facial key points corresponding to the audio, the lip-masking results of the target image, and the facial key points of the reference image, an image corresponding to the audio is generated. Then, a deformation fusion process is performed using a lip-shape rendering model to generate a video that matches the audio.

Benefits of technology

It improves the quality and accuracy of video generation while maintaining consistency of facial expressions and postures in audio and video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140100_16102025_PF_FP_ABST
    Figure CN2024140100_16102025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the embodiments of the present disclosure are a video generation method and apparatus, and a device, a medium and a product. The method comprises: after acquiring a reference video and an audio sequence, firstly, on the basis of a facial key point corresponding to an ith audio frame in the audio sequence, a lip mask result of a target image, which corresponds to the ith audio frame, in the reference video, at least two reference image frames selected from the reference video, and facial key points of the reference images, generating an image corresponding to the ith audio frame, wherein i is a positive integer, and i is less than or equal to the total number of frames in the audio sequence; and then, on the basis of images corresponding to the audio frames in the audio sequence, generating a video corresponding to the audio sequence.
Need to check novelty before this filing date? Find Prior Art

Description

A video generation method, device, apparatus, medium, and product

[0001] Cross-reference to Related Applications

[0002] The present application claims priority to Chinese Patent Application No. 202410418012.X, entitled "A Video Generation Method, Device, Apparatus, Medium, and Product", filed on April 8, 2024, which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] Embodiments of the present disclosure relate to the technical field of data processing, and in particular, to a video generation method, device, apparatus, medium, and product. BACKGROUND

[0004] Currently, some application scenarios are relatively common, such as language switching processing of a video, sentence modification processing in a video, or sentence replacement processing in a video, and the like. SUMMARY

[0005] The embodiments of the present disclosure provide a video generation method, device, apparatus, medium, and product, which are beneficial to improving the video generation effect.

[0006] To achieve the above-mentioned purpose, the technical scheme provided by the embodiments of the present disclosure is as follows:

[0007] The embodiments of the present disclosure provide a video generation method, which comprises:

[0008] obtaining a reference video and an audio sequence;

[0009] generating an image corresponding to the i-th frame of audio in the audio sequence according to the face key points corresponding to the i-th frame of audio, the mouth shape mask result of the target image corresponding to the i-th frame of audio in the reference video, at least two reference images selected from the reference video, and the face key points of each reference image; i is a positive integer, i≤total frame number in the audio sequence;

[0010] generating a video corresponding to the audio sequence according to the images corresponding to each frame of audio in the audio sequence.

[0011] In a possible implementation, the determination process of the at least two reference images comprises:

[0012] sorting the images in the reference video according to the mouth shape amplitude representation data of each image in the reference video to obtain an image sequence;

[0013] performing equal-interval sampling on the image sequence to obtain a sampling image;

[0014] According to the sampling image, the at least two reference images are determined.

[0015] In a possible implementation, the at least two reference images are determined according to the sampling image and at least one pose-similar image corresponding to the i-th audio in the reference video.

[0016] The similarity between the pose representation data of each pose-similar image and the pose representation data corresponding to the i-th audio reaches a preset similarity requirement.

[0017] In a possible implementation, the time corresponding to each pose-similar image in the reference video is different from the time corresponding to the i-th audio.

[0018] The time corresponding to the i-th audio is determined according to the time corresponding to the target image in the reference video.

[0019] In a possible implementation, the pose representation data corresponding to the i-th audio is determined according to the pose representation data of the target image.

[0020] In a possible implementation, the face key points corresponding to the i-th audio include face key point determination results of at least one audio in the audio sequence.

[0021] The at least one audio includes the i-th audio.

[0022] In a possible implementation, for any audio in the at least one audio, the face key point determination result of the audio is two-dimensional face key points obtained by projecting three-dimensional face key points corresponding to the audio to a two-dimensional plane.

[0023] The three-dimensional face key points corresponding to the audio are obtained by performing three-dimensional face key point determination processing on the audio according to part or all of the images in the reference video.

[0024] In a possible implementation, for any reference image in the at least two reference images, the face key points of the reference image are two-dimensional face key points obtained by projecting three-dimensional face key points of the reference image to a two-dimensional plane.

[0025] The three-dimensional face key points of the reference image are obtained by performing three-dimensional face key point determination processing on the reference image.

[0026] In a possible implementation, the determination of the image corresponding to the i-th audio includes:

[0027] For any reference image in the at least two reference images, pixel usage description information corresponding to the reference image is predicted according to the reference image, the face key points of the reference image, and the face key points corresponding to the i-th frame of audio; the pixel usage description information includes pixel adjustment description information and / or pixel fusion weight;

[0028] The at least two reference images are subjected to morphological fusion processing according to the pixel usage description information corresponding to the at least two reference images, to obtain a fused image;

[0029] The i-th frame of audio is subjected to image generation processing according to the fused image, the lip shape mask result of the target image, and the face key points corresponding to the i-th frame of audio, to obtain an image corresponding to the i-th frame of audio.

[0030] In a possible implementation, the pixel usage description information corresponding to each of the reference images is determined by an information prediction module in a lip shape rendering model;

[0031] The morphological fusion processing is implemented by a morphological fusion module in the lip shape rendering model;

[0032] The image generation processing is implemented by a face generation module in the lip shape rendering model.

[0033] In a possible implementation, the reference video and the audio sequence are determined according to a sample video;

[0034] After the image corresponding to the i-th frame of audio is generated, the method further includes:

[0035] The face generation module in the lip shape rendering model is updated according to difference representation data between the target image and the image corresponding to the i-th frame of audio, and the morphological fusion module and the information prediction module in the lip shape rendering model are updated according to the difference representation data between the target image and the image corresponding to the i-th frame of audio, and the difference representation data between the target image and the fused image.

[0036] Embodiments of the present disclosure provide a video generation apparatus, which includes:

[0037] A data acquisition unit is configured to acquire a reference video and an audio sequence;

[0038] An image generation unit is configured to generate an image corresponding to an i-th frame of audio according to face key points corresponding to the i-th frame of audio, a lip shape mask result of a target image corresponding to the i-th frame of audio in the reference video, at least two reference images selected from the reference video, and face key points of the reference images; i is a positive integer, and i is less than or equal to a total number of frames in the audio sequence.

[0039] a video generation unit configured to generate a video corresponding to the audio sequence according to images corresponding to each frame of audio in the audio sequence.

[0040] An electronic device is provided, and the device includes a processor and a memory.

[0041] The memory is configured to store instructions or a computer program.

[0042] The processor is configured to execute the instructions or the computer program in the memory, so that the electronic device performs a video generation method provided in the embodiments of the present disclosure.

[0043] A computer readable medium is provided, and the computer readable medium stores instructions or a computer program. When the instructions or the computer program are run on a device, the device performs a video generation method provided in the embodiments of the present disclosure.

[0044] A computer program product is provided, and the computer program product includes a computer program carried on a non-transitory computer readable medium. The computer program includes program code for performing a video generation method provided in the embodiments of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the related art, the drawings needed to be used in the embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present disclosure, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.

[0046] FIG. 1 is a flowchart of a video generation method provided in the embodiments of the present disclosure;

[0047] FIG. 2 is a schematic diagram of a video generation process provided in the embodiments of the present disclosure;

[0048] FIG. 3 is a schematic diagram of an image generation process provided in the embodiments of the present disclosure;

[0049] FIG. 4 is a structural schematic diagram of a video generation apparatus provided in the embodiments of the present disclosure;

[0050] FIG. 5 is a structural schematic diagram of an electronic device provided in the embodiments of the present disclosure. DETAILED DESCRIPTION

[0051] In order for those skilled in the art to better understand the solutions of the embodiments of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, not all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present disclosure.

[0052] As described above, for some application scenarios, such as language switching processing of a video, sentence modification processing in a video, or sentence replacement processing in a video, and the like, there can be the following needs: generating a video adapted to a certain audio sequence according to an existing video.

[0053] In order to better understand the technical solutions provided by the embodiments of the present disclosure, the video generation method provided by the embodiments of the present disclosure will be described below in conjunction with some drawings. As shown in FIG. 1, the video generation method provided by the embodiments of the present disclosure includes the following S1-S3. Wherein, FIG. 1 is a flowchart of a video generation method provided by the embodiments of the present disclosure.

[0054] S1: Obtain a reference video and an audio sequence.

[0055] Wherein, the reference video refers to a video required to be referred to when generating a video, such as the reference video shown in FIG. 2, so that the reference video is used to provide information other than the facial expression state, such as the mouth shape state, such as facial identifying information similar to facial contour features and facial feature distribution features, facial posture information, and the like.

[0056] In addition, the embodiments of the present disclosure do not limit the implementation of the above reference video. In order to facilitate understanding, two scenarios will be described below.

[0057] Scenario one, when the video generation method provided by the embodiments of the present disclosure is used to perform a certain generation task, such as a video generation task, the above reference video can be a video specified by a user through a certain means, such as a single-person video, so that the reference video meets some needs of the user, such as needs in terms of facial identifying information, facial posture, and the like. It should be noted that the present disclosure does not limit the means, for example, the reference video can be a video manually uploaded by the user, or a video selected by the user from some candidate videos, or a video downloaded by the user with the help of certain means.

[0058] Scene two, when the video generation method provided by the embodiment of the present disclosure is used to implement the model training process, the reference video mentioned above can be determined according to a sample video, so that the reference video includes part or all of the images in the sample video. Wherein, the sample video refers to a video required to be used in the model training process; and the embodiment of the present disclosure does not limit the implementation of the sample video.

[0059] In addition, the embodiment of the present disclosure does not limit the acquisition method of the reference video mentioned above.

[0060] The audio sequence refers to the audio required to be used in the video generation, such as the audio sequence shown in FIG. 2, so that the audio sequence is used to provide the facial expression state, such as the lip shape state and the like.

[0061] In addition, the embodiment of the present disclosure does not limit the implementation of the audio sequence mentioned above, in order to facilitate understanding, the following will be described in combination with two scenes.

[0062] Scene one, when the video generation method provided by the embodiment of the present disclosure is used to execute a certain generation task, such as a video generation task, the audio sequence mentioned above can be specified by a user with certain means; and the embodiment of the present disclosure does not limit the implementation of the audio sequence, for example, in some application scenarios, such as a video translation scenario, the audio sequence satisfies the following constraints: the semantic information of the sentence described by the audio sequence is consistent with the semantic information of the sentence described in the reference video, but the language of the sentence described by the audio sequence is different from the language of the sentence described in the reference video. For another example, in some application scenarios, such as a video sentence modification scenario, the audio sequence at least satisfies the following constraints: the sentence described by the audio sequence is partially the same as the sentence described in the reference video. For another example, in some application scenarios, such as a video sentence replacement scenario, the audio sequence at least satisfies the following constraints: the sentence described by the audio sequence is completely different from the sentence described in the reference video.

[0063] Scene two, when the video generation method provided by the embodiment of the present disclosure is used to implement the model training process, if the reference video mentioned above is determined according to a sample video, the audio sequence mentioned above can be determined according to the sample video, so that the audio sequence includes part or all of the audio in the sample video.

[0064] In addition, the embodiments of the present disclosure do not limit the association between the audio sequence and the reference video, for example, the two can satisfy the following constraint: for the i-th frame of audio in the audio sequence, there is a target image in the reference video corresponding to the i-th frame of audio. Wherein, the target image refers to an image in the reference video that has a corresponding relationship with the i-th frame of audio, so that the target image is used to provide information other than the facial expression state, such as facial features, facial posture, etc. for the face image generation process of the i-th frame of audio; and the embodiments of the present disclosure do not limit the implementation of the target image, for example, when the total number of frames in the audio sequence is equal to the total number of frames in the reference video, the target image can refer to the i-th frame of image in the reference video. Wherein, i is a positive integer, i≤I, I represents the total number of frames of the audio sequence.

[0065] It should be noted that the embodiments of the present disclosure do not limit the implementation of the corresponding relationship in the above paragraph, for example, when the total number of frames of the audio sequence is equal to the total number of frames of the reference video, the corresponding relationship can specifically include: the corresponding relationship between the i-th frame of audio in the audio sequence and the i-th frame of image in the reference video. For another example, when the total number of frames of the audio sequence is greater than the total number of frames of the reference video, the corresponding relationship can specifically include: the corresponding relationship between the i-th frame of audio in the audio sequence and the i-th frame of image in the reference video after frame increasing. Wherein, the reference video after frame increasing is obtained by performing a certain frame increasing processing on the original reference video, so that the total number of frames of the reference video after frame increasing is equal to the total number of frames of the audio sequence; and the embodiments of the present disclosure do not limit the implementation of the frame increasing processing, for example, it can be implemented by using any one of the existing or future appearing methods capable of increasing the total number of frames of a video, such as frame insertion processing on the original video or copy splicing processing on the original video. For another example, when the total number of frames of the audio sequence is less than the total number of frames of the reference video, the corresponding relationship can specifically include: the corresponding relationship between the i-th frame of audio in the audio sequence and the i-th frame of image in the reference video after frame decreasing. Wherein, the reference video after frame decreasing is obtained by performing a certain frame decreasing processing on the original reference video, so that the total number of frames of the reference video after frame decreasing is equal to the total number of frames of the audio sequence; and the embodiments of the present disclosure do not limit the implementation of the frame decreasing processing, for example, it can be implemented by using any one of the existing or future appearing methods capable of decreasing the total number of frames of a video, such as sampling processing on the original video or cropping processing on the original video. Wherein, i is a positive integer, i≤I, I represents the total number of frames of the audio sequence.

[0066] In addition, the embodiments of the present disclosure do not limit the acquisition method of the audio sequence.

[0067] S2: generating an image corresponding to the i-th frame of audio according to the facial key points corresponding to the i-th frame of audio in the audio sequence, the lip mask result of the target image corresponding to the i-th frame of audio in the reference video, the at least two reference images selected from the reference video, and the facial key points of each reference image; i is a positive integer, i≤total number of frames in the audio sequence.

[0068] The facial key points corresponding to the i-th frame of audio are used to represent the facial expression state, such as the lip shape state, of the object in the reference video under the i-th frame of audio. It should be noted that the embodiments of the object are not limited in the present disclosure, for example, the object can be implemented by any kind of animal or virtual image capable of presenting expressions.

[0069] In addition, the embodiments of the present disclosure do not limit the implementation of the facial key points corresponding to the i-th frame of audio described above, for example, any kind of key point capable of representing the facial expression state that exists at present or in the future can be implemented. For example, in some application scenarios, in order to better improve the generation effect, the facial key points corresponding to the i-th frame of audio can be implemented by two-dimensional facial key points. The two-dimensional facial key points are used to describe the facial expression state of an object in a two-dimensional space.

[0070] It can be seen that in a possible implementation, the facial key points corresponding to the i-th frame of audio described above can include two-dimensional facial key points of the i-th frame of audio. The two-dimensional facial key points of the i-th frame of audio are used to describe the two-dimensional facial expression state of the object in the reference video under the i-th frame of audio, and the two-dimensional facial key points of the i-th frame of audio can be determined according to the i-th frame of audio and the reference video.

[0071] In addition, the embodiments of the present disclosure do not limit the determination method of the two-dimensional facial key points of the i-th frame of audio described above, for example, any kind of method capable of generating two-dimensional facial key points according to a video and a frame of audio that exists at present or in the future can be implemented, such as a method realized by means of a pre-constructed two-dimensional key point generation model. The two-dimensional key point generation model is used to generate two-dimensional facial key points for the input data of the two-dimensional key point generation model, such as video+audio; and the embodiments of the present disclosure do not limit the implementation of the two-dimensional key point generation model.

[0072] Further, in order to better improve the generation effect, the embodiment of the present disclosure further provides a determination manner of the two-dimensional face key points of the i-th frame of audio, in which the two-dimensional face key points of the i-th frame of audio can be obtained by projecting the three-dimensional face key points corresponding to the i-th frame of audio to a two-dimensional plane. Wherein, the three-dimensional face key points corresponding to the i-th frame of audio are used to describe the three-dimensional facial expression state of the object in the reference video under the i-th frame of audio; and the three-dimensional face key points corresponding to the i-th frame of audio are obtained by performing three-dimensional face key point determination processing on the i-th frame of audio according to part or all of the images in the reference video. It should be noted that the embodiment of the present disclosure does not limit the implementation manner of the three-dimensional face key point determination processing, for example, it can be implemented by using any method that can perform three-dimensional face key point determination processing according to some images and a frame of audio, such as a method implemented by means of a pre-constructed three-dimensional key point generation model. Wherein, the three-dimensional key point generation model is used for three-dimensional face key point generation processing on the input data of the three-dimensional key point generation model, such as video+audio; and the embodiment of the present disclosure does not limit the implementation manner of the three-dimensional key point generation model.

[0073] For example, in some application scenarios, in order to better improve the generation effect, the determination process of the three-dimensional face key points corresponding to the i-th frame of audio can include the following steps 11-12.

[0074] Step 11: According to the reference video and the i-th frame of audio, predict the three-dimensional face parameters corresponding to the i-th frame of audio.

[0075] Wherein, the three-dimensional face parameters corresponding to the i-th frame of audio are used to describe the face state of the object in the reference video under the i-th frame of audio, such as facial features, facial expressions, facial poses, etc.

[0076] In addition, the embodiments of the present disclosure do not limit the implementation of the three-dimensional face parameters corresponding to the i-th frame of audio, for example, the three-dimensional face parameters corresponding to the i-th frame of audio can include face identity (Identity document, ID) parameters, face expression parameters, and face posture parameters. The face identity parameters are used to describe the identity characteristics of the face of the object in the reference video, such as the face contour, the distribution of facial features, and the like, so that the three-dimensional face model constructed based on the face identity parameters is in the state of having ID, no posture, and no expression. The face expression parameters are used to describe the face expression state under the i-th frame of audio, such as the mouth shape state, so that the three-dimensional face model constructed based on the face expression parameters is in the state of no ID, no posture, and having expression. The face posture parameters are used to describe the face posture under the i-th frame of audio, such as the side face, the front face, and the like, so that the three-dimensional face model constructed based on the face posture parameters is in the state of no ID, having posture, and no expression. It should be noted that the embodiments of the present disclosure do not limit the association relationship between the three parameters, for example, the three parameters are in a decoupled state.

[0077] In addition, the embodiments of the present disclosure do not limit the determination process of the three-dimensional face parameters corresponding to the i-th frame of audio described above, for example, any method capable of determining three-dimensional face parameters based on a video and a frame of audio can be used, such as a method implemented by means of a pre-constructed face parameter determination model. The face parameter determination model refers to a pre-constructed model having a three-dimensional face parameter determination function, such as a machine learning model, and the embodiments of the present disclosure do not limit the implementation of the face parameter determination model.

[0078] Also, in some application scenarios, such as the lip modification scenario, in order to better improve the generation effect, the three-dimensional face parameters corresponding to the i-th frame of audio satisfy the following constraints: ① the face identity parameters and the face expression parameters in the three-dimensional face parameters corresponding to the i-th frame of audio are determined according to the reference video and the i-th frame of audio, so that the face features described by the face identity parameters are consistent with the face features presented in the reference video, and the face expression state represented by the face expression parameters satisfies the face expression state requirement of the i-th frame of audio; ② the face posture parameters in the three-dimensional face parameters corresponding to the i-th frame of audio are determined according to the face posture parameters in the three-dimensional face parameters of the target image, so that the face posture parameters in the three-dimensional face parameters corresponding to the i-th frame of audio are consistent with the face posture parameters in the three-dimensional face parameters of the target image. The three-dimensional face parameters of the target image are obtained by performing three-dimensional face parameter determination processing on the target image, and the determination process of the three-dimensional face parameters of the target image is not limited in the embodiments of the present disclosure. For example, any method that can determine three-dimensional face parameters for an image can be used.

[0079] Step 12: determining the three-dimensional face key points corresponding to the i-th frame of audio according to the three-dimensional face parameters corresponding to the i-th frame of audio.

[0080] It should be noted that the embodiments of the present disclosure do not limit the implementation of step 12 above, for example, any method that can determine three-dimensional face key points based on three-dimensional face parameters can be used.

[0081] Based on the related content of steps 11 to 12 above, in some application scenarios, such as the scenario shown in FIG. 2, for the i-th frame of audio in the audio sequence, the face reconstruction processing can be performed according to the reference video to obtain the three-dimensional face parameters of each frame of image in the reference video; then, the three-dimensional face parameters corresponding to the i-th frame of audio are predicted according to the three-dimensional face parameters of part or all of the images in the reference video and the audio features of the i-th frame of audio; then, the three-dimensional face key points corresponding to the i-th frame of audio are determined according to the three-dimensional face parameters corresponding to the i-th frame of audio, so that the three-dimensional face key points can better represent the face state of the object in the reference video under the i-th frame of video, so that the face key points corresponding to the i-th frame of audio determined based on the three-dimensional face key points are more accurate, and thus the generation effect is improved.

[0082] In fact, in order to better improve the generation effect, the embodiment of the present disclosure also provides a possible implementation of the face key point corresponding to the i-th frame of audio, in which the face key point corresponding to the i-th frame of audio can include the face key point determination result of at least one frame of audio in the audio sequence, and the at least one frame of audio includes the i-th frame of audio. Wherein, the at least one frame of audio refers to the audio in the audio sequence that needs to be referenced when generating the face image of the i-th frame of audio; and the embodiment of the present disclosure does not limit the implementation of the at least one frame of audio, for example, the at least one frame of audio can include the i-T-th frame of audio, the i-T+1-th frame of audio, …, the i-th frame of audio, the i+1-th frame of audio, …, and the i+T-th frame of audio in the audio sequence, so that the at least one frame of audio can better represent the audio information carried by the i-th frame of audio and its context information, thereby making the image generated based on the at least one frame of audio more accurate.

[0083] In addition, for any audio in the at least one frame of audio, the face key point determination result of the audio refers to the face key point determined for the audio, so that the face key point determination result can represent the face state of the object in the reference video under the audio; and the embodiment of the present disclosure does not limit the implementation of the face key point determination result of the audio, for example, the face key point determination result of the audio can be implemented by three-dimensional face key points and / or two-dimensional face key points. It can be seen that in a possible implementation, the face key point determination result of the audio can be a two-dimensional face key point obtained by projecting the three-dimensional face key point corresponding to the audio onto a two-dimensional plane. Wherein, the three-dimensional face key point corresponding to the audio is obtained by performing three-dimensional face key point determination processing on the audio according to part or all of the images in the reference video; and the implementation of the three-dimensional face key point corresponding to the audio is similar to the implementation of the three-dimensional face key point corresponding to the i-th frame of audio, which will not be described here for brevity.

[0084] Based on the related content of the face key point corresponding to the i-th frame of audio, in a possible implementation, after obtaining the three-dimensional face key point corresponding to each frame of audio in the audio sequence, the face key point determination result of at least one frame of audio can be determined according to the three-dimensional face key point corresponding to the at least one frame of audio; and the face key point corresponding to the i-th frame of audio is determined according to the face key point determination result of the at least one frame of audio, so that the face key point corresponding to the i-th frame of audio includes the face key point determination result of the at least one frame of audio, thereby making the face key point corresponding to the i-th frame of audio better represent the face state of the object in the reference video under the i-th frame of audio, such as face features, facial expressions, facial poses, context information, etc., thereby making the image generated based on the face key point corresponding to the i-th frame of audio better.

[0085] In addition, for the target image corresponding to the i-th frame of audio in the reference video, the target image refers to an image existing in the reference video and used to provide information other than the facial expression state for the face image generation process of the i-th frame of audio, such as the i-th frame of image in the reference video; and the lip mask result of the target image is used to provide information other than the lip shape state in the target image; and the disclosure embodiments do not limit the determination process of the lip mask result, for example, the lip mask result of the target image can be obtained by performing mask processing on the lower half of the face in the target image, so that the lip mask result can represent the upper half of the face in the target image, so that the lip mask result can provide some other information in the target image, such as skin color, light, and the like, so that the generated image based on the lip mask result is better.

[0086] In addition, for the at least two frames of reference images, the at least two frames of reference images refer to images selected from the reference video, so that the at least two frames of reference images can provide some pixel information with use value to the face image generation process of the i-th frame of audio, such as pixel information related to the posture, lip shape, and the like.

[0087] In addition, the disclosure embodiments do not limit the determination process of the at least two frames of reference images, for example, when generating an image for the i-th frame of audio, the at least two frames of reference images can be randomly selected from the reference video, such as randomly selected for the i-th frame of audio. It can be seen that in some application scenarios, the at least two frames of reference images corresponding to different frames of audio in the above audio sequence refer to the same group of images randomly selected from the reference video. For example, in some application scenarios, the at least two frames of reference images corresponding to different frames of audio in the audio sequence refer to different groups of images randomly selected from the reference video.

[0088] In addition, in order to better improve the generation effect, the disclosure embodiments also provide a determination manner of the at least two frames of reference images, and in this manner, the determination process of the at least two frames of reference images can include the following steps 21-23.

[0089] Step 21: According to the lip shape amplitude representation data of each frame of image in the reference video, the images in the reference video are sorted to obtain an image sequence.

[0090] Wherein, the j-th frame of image in the reference video refers to an image existing in the reference video and located at the j-th arrangement position, j is a positive integer, j≤J, J is a positive integer, and J represents the total number of frames in the reference video.

[0091] In addition, for the jth image in the reference video, the mouth opening amplitude feature data of the jth image is used to represent the mouth opening amplitude of the object in the jth image, and the disclosure embodiments do not limit the determination manner of the mouth opening amplitude feature data of the jth image. For example, any existing or future method capable of determining the mouth opening amplitude of an image can be used, such as a method implemented by means of a pre-constructed mouth opening amplitude determination model. The mouth opening amplitude determination model refers to a model pre-constructed and capable of determining the mouth opening amplitude of an image, such as a machine learning model.

[0092] In addition, in order to better improve the generation effect, the disclosure embodiments further provide a determination manner of the mouth opening amplitude feature data of the jth image. In this manner, the determination process of the mouth opening amplitude feature data of the jth image can be as follows: the mouth opening amplitude feature data of the jth image is determined according to the mouth key point of the facial key point of the jth image, such as the two-dimensional facial key point and / or the three-dimensional facial key point, so that the mouth opening amplitude feature data can more accurately represent the mouth opening amplitude of the object in the jth image. The three-dimensional facial key point of the jth image is used to represent the state of the face state presented in the jth image in a three-dimensional space, and the three-dimensional facial key point of the jth image is obtained by performing three-dimensional facial key point determination processing on the jth image. It should be noted that the disclosure embodiments do not limit the implementation manner of the three-dimensional facial key point determination processing. The two-dimensional facial key point of the jth image is used to represent the face state of the object in the jth image in a two-dimensional space, and the disclosure embodiments do not limit the determination method of the two-dimensional facial key point of the jth image. For example, any existing or future method capable of determining the two-dimensional facial key point of an image can be used. For example, the two-dimensional facial key point of the jth image can be obtained by projecting the three-dimensional facial key point of the jth image to a two-dimensional plane, which is beneficial to improve the consistency of the key point related information, thereby improving the generation effect.

[0093] In addition, for the image sequence in step 21, the determination process of the image sequence can be as follows: according to the mouth opening amplitude feature data of each image in the reference video, the part or all images in the reference video are sorted to obtain the image sequence, so that all images in the image sequence are arranged in ascending order or descending order according to the mouth opening amplitude feature data, thereby making the image sequence better describe the mouth opening amplitude distribution presented in the reference video. It can be seen that in a possible implementation manner, the image sequence can include all images in the reference video, and all images in the image sequence are arranged in ascending order or descending order according to the mouth opening amplitude feature data.

[0094] According to the related content of step 21, for some application scenarios, after obtaining the reference video, the mouth shape amplitude representation data of each frame image in the reference video can be determined first; then the images are arranged in ascending order or descending order according to the mouth shape amplitude representation data, and an image sequence is obtained, so that the image sequence includes part or all of the images in the reference video, and the arrangement position of each image in the image sequence is positively or negatively correlated with the mouth shape amplitude representation data of the corresponding image in the reference video, so that the image sequence can better represent the mouth shape change in the reference video.

[0095] Step 22: equally interval sampling the image sequence to obtain a sampling image.

[0096] The sampling image refers to an image sampled from the image sequence, and the disclosure embodiments do not limit the sampling interval required for sampling, for example, the sampling interval can be determined according to actual needs, such as generation efficiency needs or generation accuracy needs.

[0097] According to the related content of step 22, for some application scenarios, after obtaining the image sequence, some images can be equally interval sampled from the image sequence as sampling images, so that these sampling images can represent different mouth shape states of the object in the reference video, thereby providing useful mouth shape information for the face image generation process of the i-th frame of audio, which is conducive to improving the generation effect.

[0098] Step 23: determining the at least two reference images according to the sampling image.

[0099] It should be noted that the disclosure embodiments do not limit the implementation of step 23, for example, step 23 can be specifically: determining the sampling image as the at least two reference images, so that the at least two reference images can provide some useful information such as mouth shape information for the face image generation process of each frame image in the audio sequence. It can be seen that in a possible implementation, the at least two reference images involved in the face image generation process of different frame images in the audio sequence are all the sampling images, so that the at least two reference images involved in the face image generation process of different frame images in the audio sequence remain the same.

[0100] In fact, in order to better improve the generation effect, the embodiment of the present disclosure further provides a possible implementation of the above step 23, in which the step 23 can be specifically: determining at least two reference images corresponding to the i-th frame of audio according to the above sampled image and at least one posture similar image corresponding to the i-th frame of audio in the reference video, so that the at least two reference images corresponding to the i-th frame of audio include the sampled image and the at least one posture similar image, so that the at least two reference images corresponding to the i-th frame of audio can provide as comprehensive useful information as possible for the face image generation process of the i-th frame of audio, such as lip shape information + posture information, etc., so as to better generate the image corresponding to the i-th frame of audio based on the at least two reference images subsequently. Wherein, the at least one posture similar image refers to some images in the reference video, which have the closest posture to the posture corresponding to the i-th frame of audio.

[0101] In addition, the embodiment of the present disclosure does not limit the implementation of the at least one posture similar image in the above paragraph, for example, the at least one posture similar image can satisfy the following constraint: the similarity between the posture representation data of each posture similar image and the posture representation data corresponding to the i-th frame of audio reaches a preset similarity requirement. It should be noted that the embodiment of the present disclosure does not limit the implementation of the preset similarity requirement, for example, the preset similarity requirement can be specifically: the similarity between the posture representation data of each posture similar image and the posture representation data corresponding to the i-th frame of audio all exceeds a preset similarity threshold. For another example, after the similarity between the posture representation data of each frame of image in the reference video and the posture representation data corresponding to the i-th frame of audio is determined, all images in the reference video are sorted in descending order according to the similarity to obtain a sorting result, the preset similarity requirement can be specifically: the arrangement position of each posture similar image in the sorting result is all ahead of a preset arrangement position threshold.

[0102] In addition, for any one of the at least one pose similar image, the pose representation data of the pose similar image is used to represent the face pose of the object in the pose similar image; and the disclosure embodiments do not limit the determination process of the pose representation data of the pose similar image, for example, any one of the existing or future methods capable of performing pose extraction processing on an image can be used to implement the determination process. For example, in order to better improve the generation effect, the determination process of the pose representation data of the pose similar image can be: determining the pose representation data of the pose similar image according to the face pose parameter in the three-dimensional face parameter of the pose similar image. The three-dimensional face parameter of the pose similar image is used to represent the face state of the object in the three-dimensional space in the pose similar image; and the implementation manner of the three-dimensional face parameter of the pose similar image is similar to the implementation manner of the three-dimensional face parameter corresponding to the i-th frame of audio. It should be noted that the disclosure embodiments do not limit the implementation manner of the step of "determining the pose representation data of the pose similar image according to the face pose parameter in the three-dimensional face parameter of the pose similar image", for example, it can be specifically: first, determining the three-dimensional face key point of the pose similar image by using the face pose parameter in the three-dimensional face parameter of the pose similar image, so that the three-dimensional face key point of the pose similar image can represent the face state in the case of no ID, no expression and pose; and then determining the three-dimensional face key point of the pose similar image as the pose representation data of the pose similar image.

[0103] In addition, for any one of the at least one pose similar image, the pose representation data of the pose similar image is used to represent the face pose of the object in the pose similar image; and the disclosure embodiments do not limit the determination process of the pose representation data of the pose similar image, for example, any one of the existing or future methods capable of performing pose extraction processing on an image can be used to implement the determination process. For example, in order to better improve the generation effect, the determination process of the pose representation data of the pose similar image can be: determining the pose representation data of the pose similar image according to the face pose parameter in the three-dimensional face parameter of the pose similar image. The three-dimensional face parameter of the pose similar image is used to represent the face state of the object in the three-dimensional space in the pose similar image; and the implementation manner of the three-dimensional face parameter of the pose similar image is similar to the implementation manner of the three-dimensional face parameter corresponding to the i-th frame of audio. It should be noted that the disclosure embodiments do not limit the implementation manner of the step of "determining the pose representation data of the pose similar image according to the face pose parameter in the three-dimensional face parameter of the pose similar image", for example, it can be specifically: first, determining the three-dimensional face key point of the pose similar image by using the face pose parameter in the three-dimensional face parameter of the pose similar image, so that the three-dimensional face key point of the pose similar image can represent the face state in the case of no ID, no expression and pose; and then determining the three-dimensional face key point of the pose similar image as the pose representation data of the pose similar image.

[0104] Further, in order to improve the generation effect, the embodiment of the present disclosure further provides a possible implementation of the at least one posture similar image, in which the at least one posture similar image satisfies the following constraint: the similarity between the posture representation data of each posture similar image and the posture representation data corresponding to the i-th frame of audio reaches the preset similarity requirement, and the time corresponding to each posture similar image in the reference video is different from the time corresponding to the i-th frame of audio. Wherein, the time corresponding to the i-th frame of audio is determined according to the time corresponding to the target image in the reference video; and the embodiment of the present disclosure does not limit the determination manner of the time corresponding to the i-th frame of audio, for example, it can be specifically: the time corresponding to the target image in the reference video is determined as the time corresponding to the i-th frame of audio.

[0105] Based on the above content, in a possible implementation, when the time corresponding to the i-th frame of audio is consistent with the time corresponding to the target image in the reference video, the at least one posture similar image corresponding to the i-th frame of audio can satisfy the following constraint: the similarity between the posture representation data of each posture similar image and the posture representation data corresponding to the i-th frame of audio reaches the preset similarity requirement, and the time corresponding to each posture similar image in the reference video is different from the time corresponding to the target image in the reference video, so that the at least one posture similar image corresponding to the i-th frame of audio can include multiple images existing in the reference video, in addition to the target image, which are closest to the posture corresponding to the i-th frame of audio, thereby effectively avoiding some interference caused by the target image, and further enabling the posture similar images to better provide useful posture information for the face image generation process of the i-th frame of audio, so as to improve the generation effect. Wherein, the posture corresponding to the i-th frame of audio refers to the face posture presented in the target image.

[0106] Based on the above related content of steps 21-23, the present embodiment provides two possible implementations of the above at least two reference images, one is full fixed, for example, by performing relevant processing on the reference video in terms of mouth shape to select some images with different mouth shape states from the reference video, such as the above sampling images, so as to be able to apply these images with different mouth shape states to the face image generation process of each frame of audio in the audio sequence in the subsequent process, so that the at least two reference images referred to by the face image generation process of different frames of audio in the audio sequence remain consistent. The other is partially fixed and partially dynamic, for example, by performing relevant processing on the reference video in terms of mouth shape to select some images with different mouth shape states from the reference video as a fixed part, and by performing relevant processing on the reference video in terms of posture to select some images with a posture close to the posture corresponding to the i-th frame of audio from the reference video as a dynamic part corresponding to the i-th frame of audio, so as to be able to apply the fixed part + the dynamic part corresponding to the i-th frame of audio to the face image generation process of the i-th frame of audio in the subsequent process, which is conducive to better improving the generation effect.

[0107] In addition, for any reference image in the above at least two reference images, the face key points of the reference image are used to describe the face state presented in the reference image, and the present embodiment does not limit the implementation of the face key points of the reference image, for example, the implementation of the face key points of the reference image is similar to the implementation of the face key points corresponding to the i-th frame of audio. As can be seen, in one possible implementation, if the face key points corresponding to the i-th frame of audio are implemented by two-dimensional face key points, the face key points of the reference image can also be implemented by two-dimensional face key points, and the face key points of the reference image and the face key points corresponding to the i-th frame of audio satisfy the following constraint: the face key points of the reference image and the face key points corresponding to the i-th frame of audio are corresponding in key point sequence number.

[0108] In addition, for any one of the at least two reference images, the face key points of the reference image are not limited in the embodiments of the present disclosure, for example, if the face key points corresponding to the i-th audio frame are obtained by projecting the three-dimensional face key points to a two-dimensional plane, the face key points of the reference image can be two-dimensional face key points obtained by projecting the three-dimensional face key points of the reference image to a two-dimensional plane, so that the face key points of the reference image are obtained in the same way as the face key points corresponding to the i-th audio frame, which can effectively avoid the defects caused by different two-dimensional key point acquisition mechanisms, thereby facilitating to improve the generation effect. The three-dimensional face key points of the reference image are used to describe the state of the face in the reference image in a three-dimensional space; and the three-dimensional face key points of the reference image are obtained by performing three-dimensional face key point determination processing on the reference image. It should be noted that the three-dimensional face key point determination processing is not limited in the embodiments of the present disclosure, for example, it can use any method that can perform three-dimensional face key point determination processing on an image, such as a method implemented by using a pre-constructed machine learning model with three-dimensional face key point determination processing function.

[0109] In addition, for the i-th audio in the audio sequence, the image corresponding to the i-th audio refers to an image generated for the i-th audio, so that the image corresponding to the i-th audio satisfies the following constraints: the facial expression state presented in the image corresponding to the i-th audio satisfies the facial expression state requirement of the i-th audio, and other information in the image corresponding to the i-th audio, except the facial expression state, is consistent with the corresponding information in the target image, so that the image corresponding to the i-th audio can represent the result of adjusting the facial expression state of the target image according to the i-th audio.

[0110] In fact, in order to better improve the generation effect, the embodiments of the present disclosure also provide a determination manner of the image corresponding to the i-th audio, in which the determination process of the image corresponding to the i-th audio can include the following steps 31-33.

[0111] Step 31: for any one of the at least two reference images, predicting pixel usage description information corresponding to the reference image according to the reference image, the face key points of the reference image, and the face key points corresponding to the i-th audio; the pixel usage description information includes pixel adjustment description information and / or pixel fusion weight.

[0112] The pixel usage description information corresponding to the reference image is used to describe how to use the pixels in the reference image in the face image generation process of the i-th frame of audio, so that the pixel usage description information can indicate how the pixels in the reference image affect the face image generation process of the i-th frame of audio.

[0113] In addition, the pixel usage description information corresponding to the reference image is not limited in the embodiments of the present disclosure, for example, the pixel usage description information corresponding to the reference image can include pixel adjustment description information corresponding to the reference image and / or pixel fusion weight corresponding to the reference image.

[0114] In addition, for the pixel adjustment description information corresponding to the reference image, the pixel adjustment description information is used to indicate how to adjust the pixels in the reference image in the face image generation process of the i-th frame of audio; and the embodiments of the present disclosure do not limit the implementation of the pixel adjustment description information, for example, the pixel adjustment description information can be implemented by using the offset of image pixel information. It can be seen that in a possible implementation, the pixel adjustment description information satisfies the following constraints: the size of the pixel adjustment description information is the same as the size of the reference image, and the position coordinates of each pixel point in the pixel adjustment description information are used to indicate the offset of the position coordinates of the corresponding pixel point in the reference image.

[0115] In addition, for the pixel fusion weight corresponding to the reference image, the pixel fusion weight is used to indicate how much the pixels in the reference image affect the face image generation process of the i-th frame of audio; and the embodiments of the present disclosure do not limit the implementation of the pixel fusion weight, for example, the pixel fusion weight can satisfy the following constraints: the size of the pixel fusion weight is the same as the size of the reference image, and the weight value of each pixel point in the pixel fusion weight is used to indicate the influence degree of the corresponding pixel point in the reference image.

[0116] Further, the embodiments of the present disclosure do not limit the implementation of step 31, for example, in order to better improve the generation effect, step 31 can be implemented by using an information prediction module in a lip shape rendering model. As can be seen, in a possible implementation, step 31 can be specifically: for any reference image in the at least two reference images, the information prediction module predicts and outputs the pixel usage description information corresponding to the reference image according to the reference image, the face key points of the reference image, and the face key points corresponding to the i-th frame of audio. Wherein, the lip shape rendering model is used for lip shape rendering processing for input data of the lip shape rendering model, such as the lip shape rendering processing shown in FIG. 2 or FIG. 3; and the embodiments of the present disclosure do not limit the implementation of the lip shape rendering model, for example, the lip shape rendering model can at least include the information prediction module, such as the information prediction module shown in FIG. 3. Wherein, the information prediction module is used for predicting the pixel usage description information corresponding to a reference image according to the reference image, the face key points of the reference image, and the face key points corresponding to the i-th frame of audio; and the embodiments of the present disclosure do not limit the implementation of the information prediction module, for example, in order to better improve the generation effect, the information prediction module can include a convolutional neural network (CNN) and a feature injection module, so that the information prediction module has the function of fusing the reference image, the face key points of the reference image, and the face key points corresponding to the i-th frame of audio. It should be noted that the embodiments of the present disclosure do not limit the implementation of the feature injection module, for example, the feature injection module can be implemented by using AdaIN (Adaptive Instance Normalization) or SPADE.

[0117] Based on the related content of step 31, after obtaining the at least two reference images, the face key points of each reference image, and the face key points corresponding to the i-th frame of audio, these data can be input into the lip shape rendering model, so that the information prediction module in the lip shape rendering model can predict and output the pixel usage description information corresponding to the n-th reference image according to the n-th reference image, the face key points of the n-th reference image, and the face key points corresponding to the i-th frame of audio, so that the pixel usage description information can indicate how to use the pixels in the n-th reference image in the face image generation process of the i-th frame of audio, such as position offset + influence degree, n is a positive integer, n≤N, N is a positive integer, N represents the number of images in the at least two reference images, so that subsequent image generation processing can be performed on the i-th frame of audio based on the pixel usage description information corresponding to the reference images.

[0118] Step 32: performing morphing fusion processing on the at least two reference images according to the pixel usage description information corresponding to the at least two reference images, to obtain a fused image.

[0119] The fused image refers to an image obtained by performing morphing fusion processing on the at least two reference images according to the pixel usage description information corresponding to the at least two reference images, so that the fused image is adapted to the i-th audio.

[0120] In addition, the embodiments of the present disclosure do not limit the implementation of step 32 above. For example, in order to better improve the generation effect, step 32 can be specifically: after the information prediction module in the lip shape rendering model outputs the pixel usage description information corresponding to each reference image, the morphing fusion module in the lip shape rendering model performs morphing fusion processing on all reference images according to the pixel usage description information corresponding to all reference images, to obtain and output a fused image, so that the fused image can represent the pixel integration usage result of the reference images under the i-th audio, so that the fused image is adapted to the i-th audio, and thus the fused image can meet the following constraint: the facial expression state presented in the fused image meets the facial expression state requirement of the i-th audio, and the other parts or all information in the fused image except the facial expression state are consistent with the corresponding information in the target image. The morphing fusion module is used to perform pixel integration usage processing on the reference images according to the pixel usage description information corresponding to each reference image, and the embodiments of the present disclosure do not limit the implementation of the morphing fusion module. For example, the morphing fusion module can be implemented by using any information integration network, such as grid_sample.

[0121] Based on the related content of step 32 above, for the lip shape rendering model, after the information prediction module in the lip shape rendering model predicts and outputs the pixel usage description information corresponding to the n-th reference image according to the n-th reference image, the facial key points of the n-th reference image, and the facial key points corresponding to the i-th audio, n is a positive integer, n≤N, the morphing fusion module in the lip shape rendering model, such as the morphing fusion module shown in FIG. 3, can perform morphing fusion processing on all reference images according to the pixel usage description information corresponding to all reference images, to obtain and output a fused image, so that the fused image can represent the facial image generation result of the i-th audio, so that the image corresponding to the i-th audio can be determined based on the facial image generation result subsequently.

[0122] Step 33: performing image generation processing on the fused image, the lip shape mask result of the target image, and the facial key points corresponding to the i-th audio, to obtain an image corresponding to the i-th audio.

[0123] It should be noted that the embodiments of the present disclosure do not limit the implementation of step 33 above, for example, in order to better improve the generation effect, step 33 can be specifically: the face generation module in the lip shape rendering model performs image generation processing on the basis of the lip shape mask results of the fusion image and the target image, and the face key points corresponding to the i-th frame of audio, to obtain and output the image corresponding to the i-th frame of audio, so that the image quality of the image corresponding to the i-th frame of audio is better than that of the fusion image, which is beneficial to improve the generation effect. Wherein, the face generation module is used to perform optimization generation processing on the fusion image on the basis of the lip shape mask result of the target image and the face key points corresponding to the i-th frame of audio, so that the image output by the face generation module is better than the fusion image, so that the image output by the face generation module is more suitable for the i-th frame of audio; and the embodiments of the present disclosure do not limit the implementation of the face generation module, for example, the face generation module can be implemented by using any kind of image generation network existing or appearing in the future.

[0124] Based on the related content of steps 31 to 33 above, for some application scenarios, after obtaining the at least two reference images, the face key points of each reference image, the lip shape mask result of the target image corresponding to the i-th frame of audio, and the face key points corresponding to the i-th frame of audio, these data can be input into the lip shape rendering model, so that the lip shape rendering model can generate and output the image corresponding to the i-th frame of audio according to these data, such as the image corresponding to the i-th frame of audio shown in FIG. 3. Wherein, because the lip shape rendering model has good performance, the image corresponding to the i-th frame of audio generated by means of the lip shape rendering model is better, which is beneficial to improve the generation effect.

[0125] Based on the related content of S2 above, in some application scenarios, for the i-th frame of audio in the audio sequence, after obtaining the face key points corresponding to the i-th frame of audio, the lip shape mask result of the target image corresponding to the i-th frame of audio in the reference video, at least two reference images corresponding to the i-th frame of audio and the face key points of each reference image, the lip shape rendering model can generate and output the image corresponding to the i-th frame of audio according to these data, so that the image corresponding to the i-th frame of audio can represent the face state of the object in the target image under the i-th frame of audio, which can realize the expression adjustment processing of the target image based on the i-th frame of audio.

[0126] S3: generating a video corresponding to the audio sequence according to the images corresponding to each frame of audio in the audio sequence.

[0127] It should be noted that the embodiments of the present disclosure do not limit the implementation of S3 above, for example, in order to better improve the generation effect, S3 can be specifically: after obtaining the image corresponding to each frame of audio in the audio sequence, the video corresponding to the audio sequence can be generated according to the audio sequence and the image corresponding to each frame of audio in the audio sequence, so that the video corresponding to the audio sequence includes the audio sequence and the image corresponding to each frame of audio in the audio sequence, so that the object in the video corresponding to the audio sequence is consistent with the object in the reference video, and the facial expression state presented by the video corresponding to the audio sequence is consistent with the facial expression state required by the audio sequence, thereby enabling the video corresponding to the audio sequence to represent the facial state change of the object in the reference video under the audio sequence, so as to realize the adjustment of the facial expression state of the object in a video based on the audio sequence.

[0128] Based on the related content of S1 to S3 above, for the video generation method provided by the embodiments of the present disclosure, after obtaining the reference video and the audio sequence, first, the image corresponding to the i-th frame of audio in the audio sequence is generated according to the facial key points corresponding to the i-th frame of audio in the audio sequence, the lip mask result of the target image corresponding to the i-th frame of audio in the reference video, at least two reference images selected from the reference video, and the facial key points of each reference image, so that the facial expression state presented in the image corresponding to the i-th frame of audio meets the expression requirement of the i-th frame of audio, such as the lip requirement, and the information other than the facial expression state presented in the image corresponding to the i-th frame of audio, such as the facial features, the facial posture and the like, is consistent with the corresponding information presented in the target image; i is a positive integer, i≤total number of frames in the audio sequence; then, the video corresponding to the audio sequence is generated according to the image corresponding to each frame of audio in the audio sequence, so that the object presented in the video corresponding to the audio sequence is consistent with the object presented in the reference video, and the video corresponding to the audio sequence can represent the facial state change of the object under the audio sequence, so as to realize the generation of the video adapted to the audio sequence. Wherein, since the at least two reference images can as much as possible comprehensively present the information in the reference video, which has a use value in the process of generating the facial image of the i-th frame of audio, such as lip, facial posture and the like, so that the image generated based on these reference images for the i-th frame of audio can better represent the facial state of the object under the i-th frame of audio, thereby enabling the video generated based on these images to better represent the facial state change of the object under the audio sequence, and further enabling the finally generated video to be more adapted to the audio sequence, so as to facilitate improving the video generation effect.

[0129] In addition, the embodiments of the present disclosure do not limit the execution subject of the video generation method provided by the embodiments of the present disclosure. For example, the video generation method provided by the embodiments of the present disclosure can be applied to a terminal device or a server. For another example, the video generation method provided by the embodiments of the present disclosure can also be implemented by means of a data interaction process between a terminal device and a server. The terminal device can be a smartphone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a standalone server, a cluster server or a cloud server.

[0130] In addition, the embodiments of the present disclosure do not limit the application scenario of the video generation method provided by the embodiments of the present disclosure. For example, the video generation method can be used to complete a certain video generation task, such as a language switching processing task of a video, a sentence modification processing task of a video, or a sentence replacement processing task of a video. Moreover, the data processing logic used when the video generation method is used to complete the video generation task is similar to the data processing logic shown in S1-S3 above. For the sake of brevity, the details are not repeated here.

[0131] For another example, the video generation method provided by the embodiments of the present disclosure can be applied to a model training scenario. Based on this, the embodiments of the present disclosure also provide a model training process, which can specifically include at least the following steps 41-43.

[0132] Step 41: Obtain a reference video and an audio sequence, wherein the reference video and the audio sequence are both determined according to a sample video.

[0133] It should be noted that the related content of step 41 can be referred to the related content of S1 above.

[0134] It can be seen that in some application scenarios, for the current round, the reference video and the audio sequence are determined from a sample video, so that the reference video includes part or all of the images in the sample video, and the audio sequence includes part or all of the audio in the sample video, and there is a corresponding relationship between the target image in the reference video and the i th frame of audio in the audio sequence, such as the corresponding time of the target image in the sample video being the same as the corresponding time of the i th frame of audio in the sample video, so that the subsequent training process of the current round can be completed by means of the reference video and the audio sequence.

[0135] Step 42: generating, by a lip-sync model, an image corresponding to the i th frame of audio in the audio sequence according to the face key points corresponding to the i th frame of audio, the lip-sync mask result of the target image corresponding to the i th frame of audio in the reference video, at least two reference images selected from the reference video, and the face key points of each reference image; i is a positive integer, i≤total number of frames in the audio sequence.

[0136] The lip shape rendering model is used for lip shape rendering processing on input data of the lip shape rendering model, as shown in FIG. 2 or FIG. 3.

[0137] In addition, the embodiments of the present disclosure do not limit the implementation of the lip shape rendering model. For example, in order to better improve the generation effect, the lip shape rendering model can include an information prediction module, a deformation fusion module and a face generation module, so that the working principle of the lip shape rendering model can be: after inputting the face key points corresponding to the i th frame of audio in the audio sequence, the lip shape mask result of the target image corresponding to the i th frame of audio in the reference video, at least two reference images corresponding to the i th frame of audio, and the face key points of each reference image into the lip shape rendering model, first, the information prediction module in the lip shape rendering model predicts and outputs the pixel usage description information corresponding to the reference images according to the reference images, the face key points of the reference images, and the face key points corresponding to the i th frame of audio; then, the deformation fusion module in the lip shape rendering model performs deformation fusion processing on the reference images according to the pixel usage description information to obtain and output a fusion image; and then, the face generation module in the lip shape rendering model generates an image corresponding to the i th frame of audio according to the fusion image, the face key points corresponding to the i th frame of audio and the lip shape mask result of the target image, so that the performance of the lip shape rendering model can be measured based on the difference between the image corresponding to the i th frame of audio and the ground truth (GT) corresponding to the i th frame of audio. The ground truth corresponding to the i th frame of audio is used to guide the face image generation processing of the i th frame of audio. Moreover, the embodiments of the present disclosure do not limit the implementation of the ground truth corresponding to the i th frame of audio. For example, the ground truth corresponding to the i th frame of audio can be implemented by using the target image.

[0138] Step 43: updating the face generation module in the lip shape rendering model according to the difference representation data between the target image and the image corresponding to the i th frame of audio, and updating the deformation fusion module and the information prediction module in the lip shape rendering model according to the difference representation data between the target image and the image corresponding to the i th frame of audio, and the difference representation data between the target image and the fusion image, i is a positive integer, i≤total number of frames in the audio sequence, and returning to continue executing the above step 41 and subsequent steps until a preset stopping condition is reached.

[0139] The difference representation data between the target image and the image corresponding to the i-th frame of audio is used to represent the difference between the ground truth corresponding to the i-th frame of audio and the image corresponding to the i-th frame of audio, and the embodiments of the present disclosure do not limit the implementation of the difference representation data. For example, the difference representation data can be implemented by using any existing or future method capable of measuring the difference between two images. For another example, in order to better improve the model training effect, the difference representation data can be determined according to the similarity, such as the Euclidean distance or the cosine distance, between the image features of the target image and the image features of the image corresponding to the i-th frame of audio.

[0140] In addition, for the difference representation data between the target image and the image corresponding to the i-th frame of audio, the difference representation data can participate in the update process of all modules in the lip-synch model, such as the update process of the information prediction module, the update process of the deformation fusion module, and the update process of the face generation module, by means of gradient backpropagation.

[0141] In addition, the difference representation data between the target image and the fused image is used to represent the difference between the ground truth corresponding to the i-th frame of audio and the fused image, and the embodiments of the present disclosure do not limit the implementation of the difference representation data. For example, the difference representation data can be implemented by using any existing or future method capable of measuring the difference between two images. For another example, in order to better improve the model training effect, the difference representation data can be determined according to the similarity, such as the Euclidean distance or the cosine distance, between the image features of the target image and the image features of the fused image.

[0142] In addition, for the difference representation data between the target image and the fused image, the difference representation data can participate in the update process of other modules in the lip-synch model except the face generation module, such as the update process of the information prediction module and the update process of the deformation fusion module, by means of gradient backpropagation.

[0143] Further, for the preset stop condition described above, the preset stop condition refers to a condition required for stopping training of the lip-sync model, and the preset stop condition is not limited in the embodiments of the present disclosure. For example, the preset stop condition can specifically include that the model loss of the lip-sync model is lower than a pre-set loss threshold. For another example, the preset stop condition can include that the change rate of the model loss of the lip-sync model is lower than a pre-set change rate threshold. For another example, the preset stop condition can include that the number of updates of the lip-sync model reaches a pre-set number threshold. The model loss of the lip-sync model is used to represent the performance of the lip-sync model, and the loss of the lip-sync model can be determined according to the difference representation data between the target image and the image corresponding to the i-th frame of audio, and the difference representation data between the target image and the fused image. It should be noted that the calculation manner of the loss of the lip-sync model is not limited in the embodiments of the present disclosure.

[0144] Based on the related content of steps 41 to 43 described above, in some application scenarios, for the lip-sync model, the information prediction module, the deformation fusion module and the face generation module in the lip-sync model can be trained at the same time, and the output result of the face generation module is supervised by the GT, and the output result of the deformation fusion module is also weakly supervised by the GT, such as perception loss, which is beneficial to improve the model training effect.

[0145] Based on the video generation method provided in the embodiments of the present disclosure, the embodiments of the present disclosure further provide a video generation device, which will be explained and described below in combination with FIG. 4. FIG. 4 is a structural schematic diagram of a video generation device provided in the embodiments of the present disclosure. It should be noted that the technical details of the video generation device provided in the embodiments of the present disclosure are described above in combination with the video generation method.

[0146] As shown in FIG. 4, the video generation device 400 provided in the embodiments of the present disclosure includes:

[0147] The data acquisition unit 401 is configured to acquire a reference video and an audio sequence.

[0148] The image generation unit 402 is configured to generate an image corresponding to the i-th frame of audio in the audio sequence according to the face key points corresponding to the i-th frame of audio, the lip-sync mask result of the target image corresponding to the i-th frame of audio in the reference video, at least two reference images selected from the reference video, and the face key points of each reference image; i is a positive integer, i≤total number of frames in the audio sequence.

[0149] The video generation unit 403 is configured to generate a video corresponding to the audio sequence according to the images corresponding to each frame of audio in the audio sequence.

[0150] In a possible implementation, the determining of the at least two reference images comprises: sorting images in the reference video according to the mouth shape amplitude feature data of the images, to obtain an image sequence; sampling the image sequence at equal intervals to obtain sampled images; and determining the at least two reference images according to the sampled images.

[0151] In a possible implementation, the at least two reference images are determined according to the sampled images and at least one posture-similar image corresponding to the i-th audio in the reference video, a similarity between posture feature data of each of the posture-similar images and the posture feature data corresponding to the i-th audio reaching a preset similarity requirement.

[0152] In a possible implementation, a time corresponding to each of the posture-similar images in the reference video is different from a time corresponding to the i-th audio, and the time corresponding to the i-th audio is determined according to a time corresponding to the target image in the reference video.

[0153] In a possible implementation, the posture feature data corresponding to the i-th audio is determined according to the posture feature data of the target image.

[0154] In a possible implementation, the face key points corresponding to the i-th audio include a face key point determination result of at least one audio in the audio sequence, and the at least one audio includes the i-th audio.

[0155] In a possible implementation, for any audio in the at least one audio, the face key point determination result of the audio is two-dimensional face key points obtained by projecting three-dimensional face key points corresponding to the audio to a two-dimensional plane, and the three-dimensional face key points corresponding to the audio are obtained by performing three-dimensional face key point determination processing on the audio according to part or all of the images in the reference video.

[0156] In a possible implementation, for any reference image in the at least two reference images, the face key points of the reference image are two-dimensional face key points obtained by projecting three-dimensional face key points of the reference image to a two-dimensional plane, and the three-dimensional face key points of the reference image are obtained by performing three-dimensional face key point determination processing on the reference image.

[0157] In a possible implementation, the image generation unit 402 is specifically configured to: for any one of the at least two reference images, predict pixel usage description information of the reference image according to the reference image, face key points of the reference image, and face key points corresponding to the i-th frame of audio; the pixel usage description information includes pixel adjustment description information and / or pixel fusion weight; perform morphological fusion processing on the at least two reference images according to the pixel usage description information of the at least two reference images, to obtain a fused image; and perform image generation processing according to the fused image, a mouth shape mask result of the target image, and the face key points corresponding to the i-th frame of audio, to obtain an image corresponding to the i-th frame of audio.

[0158] In a possible implementation, the pixel usage description information of each of the reference images is determined by an information prediction module in the mouth shape rendering model; the morphological fusion processing is implemented by a morphological fusion module in the mouth shape rendering model; and the image generation processing is implemented by a face generation module in the mouth shape rendering model.

[0159] In a possible implementation, the reference video and the audio sequence are determined according to a sample video.

[0160] The video generation apparatus 400 further includes:

[0161] The model updating unit is configured to update a face generation module in the mouth shape rendering model according to difference representation data between the target image and the image corresponding to the i-th frame of audio, and update a morphological fusion module and an information prediction module in the mouth shape rendering model according to the difference representation data between the target image and the image corresponding to the i-th frame of audio, and difference representation data between the target image and the fused image.

[0162] Based on the above related content of the video generation apparatus 400, it can be known that the working principle of the video generation apparatus 400 provided in the embodiments of the present disclosure is as follows: after the reference video and the audio sequence are acquired, first, the image corresponding to the i-th frame of audio in the audio sequence is generated according to the face key points corresponding to the i-th frame of audio, the lip mask result of the target image corresponding to the i-th frame of audio in the reference video, the at least two reference images selected from the reference video, and the face key points of each reference image, so that the facial expression state presented in the image corresponding to the i-th frame of audio meets the expression requirement of the i-th frame of audio, such as the lip requirement, and the information other than the facial expression state presented in the image corresponding to the i-th frame of audio, such as the facial features, the facial posture and the like, is consistent with the corresponding information presented in the target image; i is a positive integer, i≤the total number of frames in the audio sequence; then, the video corresponding to the audio sequence is generated according to the images corresponding to each frame of audio in the audio sequence, so that the object presented in the video corresponding to the audio sequence is consistent with the object presented in the reference video, and the video corresponding to the audio sequence can present the facial state change of the object under the audio sequence, so that the video adapted to the audio sequence can be generated. Wherein, since the at least two reference images can present as much as possible the information in the reference video, such as the lip, the facial posture and the like, which has a use value in the process of generating the facial image of the i-th frame of audio, so that the image generated based on these reference images for the i-th frame of audio can better present the facial state of the object under the i-th frame of audio, thereby making the video finally generated based on these images better present the facial state change of the object under the audio sequence, and further making the finally generated video more adapted to the audio sequence, so as to be conducive to improving the video generation effect.

[0163] In addition, the embodiments of the present disclosure further provide an electronic device, which comprises a processor and a memory: the memory is used for storing instructions or computer programs; and the processor is used for executing the instructions or computer programs in the memory, so that the electronic device executes any implementation manner of the video generation method provided in the embodiments of the present disclosure.

[0164] Referring to FIG. 5, a structural schematic diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers and the like. The electronic device shown in FIG. 5 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.

[0165] As shown in FIG. 5, the electronic device 500 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or loaded into a random access memory (RAM) 503 from a storage device 508. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0166] In general, the following devices can be connected to the I / O interface 505: input devices 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 508 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 509. The communication devices 509 can allow the electronic device 500 to communicate wirelessly or wired with other devices to exchange data. Although FIG. 5 shows the electronic device 500 with various devices, it should be understood that all of the shown devices are not required to be implemented or possessed. More or less devices can be alternatively implemented or possessed.

[0167] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 509, or installed from the storage devices 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0168] The electronic device provided by the embodiments of the present disclosure and the method provided by the above-mentioned embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiments can be referred to the above-mentioned embodiments, and the present embodiments have the same beneficial effects as the above-mentioned embodiments.

[0169] The embodiments of the present disclosure also provide a computer readable medium, wherein instructions or computer programs are stored in the computer readable medium, and when the instructions or computer programs are run on a device, the device is caused to perform any of the embodiments of the video generation method provided by the embodiments of the present disclosure.

[0170] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as a carrier wave in a propagated data signal, which bears computer-readable program code. Such a propagated data signal can take many forms, including but not limited to electro-magnetic, optical or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer a program for use by or in connection with an instruction execution system, apparatus or device. Program code contained in the computer-readable medium can be transmitted using any suitable medium, including but not limited to wire, cable, optical fiber, RF (radio frequency), etc., or any suitable combination of the foregoing.

[0171] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.

[0172] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and is not assembled into the electronic device.

[0173] The computer-readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device can execute the method described above.

[0174] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0175] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0176] The units involved in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit / module does not constitute a limitation on the unit itself.

[0177] The functions described in the foregoing description can be implemented in part or in whole in software, which can be included as part of an operating system or a specific application, program or process. Furthermore, the functions can be implemented in software executed by one or more processors, microcontrollers or control units.

[0178] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0179] It should be noted that the various embodiments described in the specification are progressive and each embodiment focuses on the differences from other embodiments. The same and similar parts between embodiments can be mutually referred to. For the system or device disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.

[0180] It should be understood that in the embodiments of the present disclosure, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship between the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0181] It is also to be noted that, as used in the specification and the appended claims, the singular forms "a," "an" and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a component" can include a plurality of such components, and so forth. In this document, the terms "include" and "comprise" and their conjugates mean "including but not limited to," and the terms "consist of and "consist essentially of mean "including only those things listed and not including other things." The term "consisting of" in reference to a chemical composition refers only to the specified materials or steps. The term "consisting essentially of" means that the composition can include additional steps or materials, but only if the additional steps or materials do not materially alter the basic and novel characteristics of the claimed composition. Further, the use of "first," "second," "third," etc., to describe a process, method, or article of manufacture does not imply that these elements must be in a given sequence, even where the words "after," "before," "top," "bottom," etc., are used.

[0182] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM, flash memory, ROM, electrically programmable ROM (EPROM or EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC.

[0183] The foregoing description of the disclosed embodiments enables a person skilled in the art to implement or use the disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the disclosure. Thus, the present disclosure is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video generation method, wherein the method comprises: Obtain reference video and audio sequences; Generate an image corresponding to the i-th audio frame based on the facial key points corresponding to the i-th audio frame in the audio sequence, a lip mask result of a target image corresponding to the i-th audio frame in the reference video, at least two reference images selected from the reference video, and the facial key points of each of the reference images; i is a positive integer, i≤the total number of frames in the audio sequence; A video corresponding to the audio sequence is generated based on the image corresponding to each audio frame in the audio sequence.

2. The method according to claim 1, wherein the process of determining the at least two reference images comprises: Sorting the images in the reference video according to the lip amplitude representation data of each frame image in the reference video to obtain an image sequence; Performing equal-interval sampling on the image sequence to obtain sampled images; The at least two reference images are determined based on the sampled image.

3. The method according to claim 2, wherein the at least two reference image frames are determined based on the sampled image and at least one similar posture image corresponding to the i-th frame of audio in the reference video; The similarity between the posture representation data of each of the posture similar images and the posture representation data corresponding to the i-th frame of audio meets a preset similarity requirement.

4. The method according to claim 3, wherein the time corresponding to each of the posture-similar images in the reference video is different from the time corresponding to the i-th frame of audio; The time corresponding to the i-th frame of audio is determined based on the time corresponding to the target image in the reference video. 5 . The method according to claim 3 , wherein the posture representation data corresponding to the i-th frame of audio is determined based on the posture representation data of the target image.

6. The method according to claim 1, wherein the facial key points corresponding to the i-th audio frame include facial key point determination results of at least one audio frame in the audio sequence; The at least one frame of audio includes the i-th frame of audio.

7. The method according to claim 6, wherein for any audio in the at least one frame of audio, the facial key point determination result of the audio is a two-dimensional facial key point obtained by projecting the three-dimensional facial key point corresponding to the audio onto a two-dimensional plane; The three-dimensional facial key points corresponding to the audio are obtained by performing three-dimensional facial key point determination processing on the audio based on part or all of the images in the reference video.

8. The method according to claim 7, wherein for any reference image of the at least two reference images, the facial key points of the reference image are two-dimensional facial key points obtained by projecting the three-dimensional facial key points of the reference image onto a two-dimensional plane; The three-dimensional facial key points of the reference image are obtained by performing a three-dimensional facial key point determination process on the reference image.

9. The method according to claim 1, wherein the process of determining the image corresponding to the i-th audio frame comprises: For any reference image of the at least two reference images, predicting pixel usage description information corresponding to the reference image based on the reference image, facial key points of the reference image, and facial key points corresponding to the i-th frame of audio; the pixel usage description information includes pixel adjustment description information and / or pixel fusion weights; performing deformation fusion processing on the at least two reference image frames according to pixel usage description information corresponding to the at least two reference image frames to obtain a fused image; An image generation process is performed based on the fused image, the lip mask result of the target image, and the facial key points corresponding to the i-th frame of audio to obtain an image corresponding to the i-th frame of audio.

10. The method according to claim 9, wherein the pixel usage description information corresponding to each reference image is determined by using an information prediction module in a lip rendering model; The deformation fusion processing is achieved by using the deformation fusion module in the lip rendering model; The image generation process is implemented by utilizing the face generation module in the lip rendering model.

11. The method according to claim 10, wherein the reference video and the audio sequence are both determined based on a sample video; After generating the image corresponding to the i-th frame of audio, the method further includes: Based on the difference representation data between the target image and the image corresponding to the i-th frame of audio, the face generation module in the lip rendering model is updated, and based on the difference representation data between the target image and the image corresponding to the i-th frame of audio, and the difference representation data between the target image and the fused image, the deformation fusion module and the information prediction module in the lip rendering model are updated.

12. A video generation device, wherein the device comprises: a data acquisition unit, configured to acquire reference video and audio sequences; an image generation unit, configured to generate an image corresponding to the i-th audio frame based on facial key points corresponding to the i-th audio frame in the audio sequence, a lip mask result of a target image corresponding to the i-th audio frame in the reference video, at least two reference images selected from the reference video, and facial key points of each of the reference images; i is a positive integer, i≤the total number of frames in the audio sequence; The video generation unit is used to generate a video corresponding to the audio sequence based on the image corresponding to each audio frame in the audio sequence.

13. An electronic device, wherein the device comprises: processor and memory; The memory is used to store instructions or computer programs; The processor is configured to execute the instructions or computer program in the memory, so that the electronic device executes the method according to any one of claims 1 to 11.

14. A computer-readable medium, wherein the computer-readable medium stores instructions or a computer program, and when the instructions or the computer program are executed on a device, the device is caused to execute the method according to any one of claims 1 to 11.

15. A computer program product, wherein the program product comprises a computer program carried on a non-transitory computer-readable medium, the computer program comprising program code for executing the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Video generation method and device and electronic equipment

    CN112927712A

  • Video generation and model training method and device, equipment and storage medium

    CN116385604A

  • Speaking face generation method and device, electronic equipment and storage medium

    CN116844215A

  • Video generation method and related equipment

    CN117640994A

  • Text and audio-based real-time face reenactment

    US20200234690A1