Video generation method and apparatus, device, medium, and product

By acquiring target video and audio sequences, predicting facial key points in each frame of the audio sequence, and generating corresponding images, the problem of modifying and replacing sentences in videos is solved, and the facial expressions of the speaker are accurately adjusted.

WO2026036905A1PCT designated stage Publication Date: 2026-02-19BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/102406
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-16
Filing Date
2025-06-20
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle the modification and replacement of statements in videos, especially when maintaining consistency in the facial expressions of the speakers.

Method used

By acquiring the target video and audio sequence, the facial key points corresponding to each frame of the audio sequence are predicted, and the corresponding images are generated using these key points. The final video is then generated by combining the reference image and the masking results.

Benefits of technology

It enables accurate modification or replacement of sentences in the video while maintaining the consistency of the speaker's facial expressions, thus improving the quality and consistency of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025102406_19022026_PF_FP_ABST
    Figure CN2025102406_19022026_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a video generation method and apparatus, a device, a medium, and a product. The method comprises: acquiring a target video and an audio sequence, wherein the target video is used for describing the state of a speaker in a first statement, and there is a difference between a second statement described by the audio sequence and the first statement; on the basis of the target video and the audio sequence, determining a facial key point corresponding to each audio frame in the audio sequence; on the basis of a facial key point corresponding to an i-th audio frame in the audio sequence, at least two reference image frames corresponding to the i-th audio frame in a first region, at least two reference image frames corresponding to the i-th audio frame in a second region, and a lower half face masking result of a target image corresponding to the i-th audio frame, generating an image corresponding to the i-th audio frame, i being less than or equal to the number of audio frames in the audio sequence; and on the basis of the images corresponding to the audio frames in the audio sequence, generating a video corresponding to the audio sequence.
Need to check novelty before this filing date? Find Prior Art

Description

A video generation method, device, apparatus, medium, and product

[0001] The present application claims priority to the Chinese patent application No. 202411132560.2, filed on August 16, 2024, and entitled "A video generation method, device, apparatus, medium, and product", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the technical field of data processing, and in particular to a video generation method, device, apparatus, medium, and product. BACKGROUND

[0003] For some application scenarios, such as language switching processing of a video, sentence modification processing in a video, or sentence replacement processing in a video, the following requirements may exist: modifying the lip shape of an existing video according to a certain audio sequence. SUMMARY

[0004] The present application provides a video generation method, device, apparatus, medium, and product.

[0005] To achieve the above-mentioned purpose, the technical scheme provided by the present application is as follows:

[0006] The present application provides a video generation method, which comprises: obtaining a target video and an audio sequence, the target video being used to describe the state of a speaking object under first speaking content, the lower half of the face of the speaking object comprising a first area and a second area, the first area comprising a mouth, and the second area comprising other areas in the lower half of the face except the first area, the audio sequence being used to describe second speaking content, and the second speaking content being different from the first speaking content; predicting face key points corresponding to each frame of audio in the audio sequence according to the target video and the audio sequence; generating an image corresponding to the i-th frame of audio according to the face key points corresponding to the i-th frame of audio, at least two reference images corresponding to the i-th frame of audio in the first area, at least two reference images corresponding to the i-th frame of audio in the second area, and a lower half face mask result of a target image corresponding to the i-th frame of audio, the target video comprising the target image and each reference image, i being a positive integer, and i being less than or equal to the number of audio frames in the audio sequence; and generating a video corresponding to the audio sequence according to the images corresponding to each frame of audio in the audio sequence.

[0007] In a possible implementation, the face key points corresponding to the i-th frame of audio include the key points corresponding to the i-th frame of audio in the first region and the key points corresponding to the i-th frame of audio in the second region.

[0008] In a possible implementation, the generating the image corresponding to the i-th frame of audio according to the face key points corresponding to the i-th frame of audio, the at least two reference images corresponding to the i-th frame of audio in the first region, the at least two reference images corresponding to the i-th frame of audio in the second region, and the lower half face mask result of the target image corresponding to the i-th frame of audio includes: generating the image corresponding to the i-th frame of audio according to the key points corresponding to the i-th frame of audio in the first region, the at least two reference images corresponding to the i-th frame of audio in the first region, the key points corresponding to the i-th frame of audio in the second region, the at least two reference images corresponding to the i-th frame of audio in the second region, and the lower half face mask result of the target image corresponding to the i-th frame of audio.

[0009] In a possible implementation, the determining process of the image corresponding to the i-th frame of audio includes: determining the generation result of the first region corresponding to the i-th frame of audio according to the face key points corresponding to the i-th frame of audio and the at least two reference images corresponding to the i-th frame of audio in the first region; determining the generation result of the second region corresponding to the i-th frame of audio according to the face key points corresponding to the i-th frame of audio and the at least two reference images corresponding to the i-th frame of audio in the second region; and generating the image corresponding to the i-th frame of audio according to the generation result of the first region corresponding to the i-th frame of audio, the generation result of the second region corresponding to the i-th frame of audio, and the lower half face mask result of the target image corresponding to the i-th frame of audio.

[0010] In a possible implementation, the face key points corresponding to the i-th frame of audio include the key points corresponding to the i-th frame of audio in the first region and the key points corresponding to the i-th frame of audio in the second region.

[0011] The determining process of the generation result of the first region corresponding to the i-th frame of audio according to the face key points corresponding to the i-th frame of audio and the at least two reference images corresponding to the i-th frame of audio in the first region includes: determining the generation result of the first region corresponding to the i-th frame of audio according to the key points corresponding to the i-th frame of audio in the first region and the at least two reference images corresponding to the i-th frame of audio in the first region.

[0012] The generation result of the second region corresponding to the i-th frame of audio is determined according to the face key points corresponding to the i-th frame of audio and at least two frames of reference images corresponding to the i-th frame of audio in the second region.

[0013] In a possible implementation, the determination of the key points corresponding to the i-th frame of audio in the first region comprises: searching for a key point used for describing the first region from the face key points corresponding to the i-th frame of audio; and determining the key points corresponding to the i-th frame of audio in the first region according to the key point used for describing the first region.

[0014] In a possible implementation, the determination of the key points corresponding to the i-th frame of audio in the second region comprises: searching for a key point used for describing the second region from the face key points corresponding to the i-th frame of audio; and determining the key points corresponding to the i-th frame of audio in the second region according to the key point used for describing the second region.

[0015] In a possible implementation, the determination of the key points corresponding to the i-th frame of audio in the second region comprises: searching for a key point used for describing the lower half face from the face key points corresponding to the i-th frame of audio; and determining the key points corresponding to the i-th frame of audio in the second region according to the key point used for describing the lower half face.

[0016] In a possible implementation, the generation result of the first region corresponding to the i-th frame of audio is further determined according to the key point determination result of each frame of reference image in the at least two frames of reference images corresponding to the i-th frame of audio in the first region; for any reference image in the at least two frames of reference images corresponding to the i-th frame of audio in the first region, the key point determination result of the reference image is used to describe the state of the first region in the reference image.

[0017] In a possible implementation, the generation result of the second region corresponding to the i-th frame of audio is further determined according to the key point determination result of each frame of reference image in the at least two frames of reference images corresponding to the i-th frame of audio in the second region; for any reference image in the at least two frames of reference images corresponding to the i-th frame of audio in the second region, the key point determination result of the reference image is used to describe the state of the second region in the reference image, or the key point determination result of the reference image is used to describe the state of the lower half face in the reference image.

[0018] In a possible implementation, the at least two frames of reference images corresponding to the i-th frame of audio in the first region are different from the at least two frames of reference images corresponding to the i-th frame of audio in the second region.

[0019] In a possible implementation, the determining of the at least two frames of reference images corresponding to the i-th frame of audio in the first region comprises: sorting images in the target video according to mouth shape amplitude representation data of the images, to obtain an image sequence; performing equal-interval sampling on the image sequence to obtain sampled images; and determining the at least two frames of reference images corresponding to the i-th frame of audio in the first region according to the sampled images.

[0020] In a possible implementation, for any one of the at least two frames of reference images corresponding to the i-th frame of audio in the second region, a distance between a sequence position of the reference image in the target video and a sequence position of a target image corresponding to the i-th frame of audio in the target video is not greater than a preset distance threshold.

[0021] In a possible implementation, the audio sequence comprises M sub-sequences, where M is a positive integer.

[0022] The determining of the video corresponding to the audio sequence comprises: performing adjustment processing on an image corresponding to an n-th frame of audio in an m-th sub-sequence according to images corresponding to other audios in the m-th sub-sequence, to obtain an adjustment result of the image corresponding to the n-th frame of audio in the m-th sub-sequence, where n is a positive integer and n is less than a number of frames of audio in the m-th sub-sequence, a time sequence continuity presented by the adjustment result of each frame of audio in the m-th sub-sequence is higher than a time sequence continuity presented by images corresponding to the frames of audio in the m-th sub-sequence, m is a positive integer and m is less than M; and generating the video corresponding to the audio sequence according to the adjustment result of each frame of audio in the at least one sub-sequence.

[0023] In a possible implementation, the adjustment result of each frame of audio in the m-th sub-sequence is determined by a decoder; input data of a time sequence module in the decoder comprises a plurality of images; and the time sequence module is configured to: perform integration processing on the plurality of images to obtain overall data, perform self-attention processing on the overall data to obtain a self-attention processing result, and perform splitting processing on the self-attention processing result to obtain an adjustment result of each image in the plurality of images.

[0024] In a possible implementation, the decoder comprises at least one processing unit arranged in sequence, the processing unit comprises a processing module and a time sequence module, and input data of the time sequence module comprises output data of the processing module.

[0025] The application provides a video generation apparatus, comprising: a data acquisition unit, configured to acquire a target video and an audio sequence, the target video being used for describing a state of a speaking object under first speaking content, a lower half of a face of the speaking object comprising a first area and a second area, the first area comprising a mouth, and the second area comprising other areas of the lower half of the face except the first area, and the audio sequence being used for describing second speaking content, the second speaking content being different from the first speaking content; a data prediction unit, configured to predict face key points corresponding to each frame of audio in the audio sequence according to the target video and the audio sequence; a first generation unit, configured to generate an image corresponding to the i-th frame of audio according to the face key points corresponding to the i-th frame of audio, at least two reference images corresponding to the i-th frame of audio under the first area, at least two reference images corresponding to the i-th frame of audio under the second area, and a lower half face mask result of a target image corresponding to the i-th frame of audio, the target video comprising the target image and each reference image, i being a positive integer, and i being less than or equal to a number of audio frames in the audio sequence; and a second generation unit, configured to generate a video corresponding to the audio sequence according to images corresponding to each frame of audio in the audio sequence.

[0026] The application provides an electronic device, comprising: a processor and a memory; the memory is configured to store instructions or a computer program; and the processor is configured to execute the instructions or the computer program in the memory, so that the electronic device executes the video generation method provided in the application.

[0027] The application provides a computer readable medium, which stores instructions or a computer program, and when the instructions or the computer program are executed on a device, the device executes the video generation method provided in the application.

[0028] The application provides a computer program product, which comprises a computer program carried on a non-transitory computer readable medium, and the computer program comprises program codes for executing the video generation method provided in the application. BRIEF DESCRIPTION OF DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the application or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0030] FIG. 1 is a flowchart of a video generation method according to an embodiment of the present application;

[0031] FIG. 2 is a schematic diagram of a lower half face according to an embodiment of the present application;

[0032] FIG. 3 is a schematic diagram of another lower half face according to an embodiment of the present application;

[0033] FIG. 4 is a schematic diagram of a video generation process according to an embodiment of the present application;

[0034] FIG. 5 is a schematic diagram of a structure of a decoder according to an embodiment of the present application;

[0035] FIG. 6 is a schematic diagram of a structure of a video generation apparatus according to an embodiment of the present application;

[0036] FIG. 7 is a schematic diagram of a structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0037] In order to make the person skilled in the art better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor fall within the scope of protection of the present application.

[0038] In order to better understand the technical scheme provided by the present application, the video generation method provided by the present application will be described below in combination with some drawings. As shown in FIG. 1, the video generation method provided by the embodiment of the present application includes the following S1-S4. Wherein, FIG. 1 is a flowchart of a video generation method according to an embodiment of the present application.

[0039] S1: obtaining a target video and an audio sequence, the target video being used to describe a state of a speaking object under a first speaking content, a lower half face of the speaking object including a first area and a second area, the first area including a mouth, and the second area including other areas in the lower half face except the first area, and the audio sequence being used to describe a second speaking content, the second speaking content being different from the first speaking content.

[0040] Wherein, the target video refers to an existing video which needs to be processed for lip modification, so that the target video is used to provide other information except the facial expression state such as the lip state, such as the facial posture information, the facial identification information similar to the facial contour characteristics and the facial feature distribution characteristics, etc.

[0041] In addition, the target video can at least satisfy the following constraint: the target video is used to describe the state of the speaking object under the first speaking content, such as the state of the mouth shape, the facial posture, the body movement, and the like, so that the target video can at least represent the speaking content and the corresponding mouth shape presented before the mouth shape modification processing.

[0042] The first speaking content refers to the sentence described by the target video, such as the sentence "today the weather is very good, suitable for sunning the quilt", so that the first speaking content can represent the speaking content presented by the speaking object in the target video, thereby enabling the first speaking content to represent the semantic information expressed by the target video.

[0043] The speaking object refers to the object described by the target video, such as a person, a virtual image, an animal, and the like, and the present application does not limit the implementation of the speaking object, for example, the speaking object can be implemented by any kind of animal or virtual image capable of presenting an expression.

[0044] In addition, the present application does not limit the presentation mode of the speaking object in the target video, for example, the target video can at least be used to describe the lower half of the face of the speaking object. It can be seen that in a possible implementation, the target video can be used to describe the state of the face (such as the upper half of the face and the lower half of the face) of the speaking object under the first speaking content. For example, in a possible implementation, the target video can be used to describe the state of the lower half of the face of the speaking object under the first speaking content.

[0045] It should be noted that for any object, the face of the object is used to describe the facial features of the object, and the present application does not limit the implementation of the face of the object, for example, it can be specifically implemented by the face shown in FIG. 2 or FIG. 3. In addition, the lower half of the face of the object refers to the lower half of the face of the object, so that the lower half of the face can represent the part of the face of the object that changes greatly with different speaking contents, thereby enabling the lower half of the face to at least describe the mouth of the object, and the present application does not limit the implementation of the lower half of the face of the object, for example, it can be specifically implemented by the lower half of the face blocked by the blocking object in FIG. 2 or FIG. 3. In addition, the upper half of the face of the object refers to the upper half of the face of the object, so that the upper half of the face is used to describe the other area of the face except the lower half of the face, thereby enabling the upper half of the face to represent the part of the face of the object that changes little or even not with different speaking contents.

[0046] It should be further noted that the lower half face can be set according to actual needs in different application scenarios, for example, in some scenarios, in order to better improve the efficiency, the lower half face can be implemented by the lower half face occluded by the occlusion in FIG. 2, so that the lower half face is only used to describe the mouth and the surrounding area, so that subsequent mouth shape adjustment processing can be realized faster with some processing for the lower half face. For example, in some scenarios, in order to better improve the accuracy, the lower half face can be implemented by the lower half face occluded by the occlusion in FIG. 3, so that the lower half face is not only used to describe the mouth and the surrounding area, but also used to describe part of other parts (such as the nose), so that subsequent mouth shape adjustment processing can be more accurately realized with some processing for the lower half face.

[0047] Based on the above three paragraphs, in some scenarios, such as full-body scenarios, half-body scenarios, or head close-up scenarios, for the speaking object described by the target video, the speaking object can at least meet the following constraints: the lower half face of the speaking object includes a first area and a second area, the first area includes the mouth, and the second area includes other areas in the lower half face except the first area. In order to facilitate understanding, the two areas are introduced as follows.

[0048] For the first area described above, such as the mouth area shown in FIG. 2 or FIG. 3, the first area refers to an area that presents a large change with the change of the speaking content, so that the first area can at least include the mouth, so that the first area can be used to describe the mouth. In addition, the present application does not limit the implementation of the first area, for example, the first area can be implemented by the circumscribed rectangle of the mouth, such as the mouth area shown in FIG. 2 or FIG. 3, so that subsequent mouth shape adjustment processing can be realized by processing the first area.

[0049] For the second area described above, the second area refers to an area in the lower half face of the speaking object that presents a small change (even no change) with the change of the speaking content, so that the second area can include other areas in the lower half face except the first area, so that the second area can be used to describe other areas in the lower half face except the mouth area, and further, the second area can represent the surrounding area of the mouth area, so that subsequent adaptive adjustment of the surrounding area during mouth shape adjustment can be realized by processing the second area, so that the adaptation degree between the mouth area and the surrounding area during mouth shape adjustment can be effectively improved, thereby facilitating the improvement of the generation effect.

[0050] In addition, the present application does not limit the implementation of the target video, and in order to facilitate understanding, two scenarios are described as follows.

[0051] In a first scenario, when the video generation method is used to perform a video generation task, such as a lip modification task, the target video can be a video specified by a user through a certain means, such as a single-person video, so that the target video meets some needs of the user, such as needs in facial identification information, facial posture, and the like. It should be noted that the present application does not limit the means, for example, the target video can be a video manually uploaded by the user, or a video selected by the user from some candidate videos, or a video downloaded by the user through certain means.

[0052] In a second scenario, when the video generation method is used to implement a model training process, the target video can be determined according to a sample video, so that the target video includes part or all of the sample video. The sample video refers to a video required in the model training process, and the present application does not limit the implementation of the sample video.

[0053] In addition, the present application does not limit the acquisition method of the target video. For ease of understanding, the following will be described in conjunction with two scenarios.

[0054] In a first scenario, when the video generation method is used to perform a video generation task, such as a lip modification task, the acquisition process of the target video can be to receive a video provided by a user or an upstream task as the target video, such as the target video shown in FIG. 4. The upstream task is used to provide some data, such as the target video, for the video generation task, so that the video generation task can be completed based on the data subsequently.

[0055] In a second scenario, when the video generation method is used to implement a model training process, the acquisition process of the target video can be to determine the target video from a sample video, so that the target video includes part or all of the sample video, so that the model training process can be completed by means of the target video subsequently.

[0056] The audio sequence refers to an audio required in the lip modification processing of the target video, such as the audio shown in FIG. 4, so that the audio sequence can describe the facial expression state, such as the lip state, expected to be achieved through the lip modification processing to a certain extent.

[0057] In addition, the audio sequence can at least satisfy the following constraint: the audio sequence is used to describe second speech content, and the second speech content is different from the first speech content, so that the sentence described by the audio sequence is different from the sentence described by the target video, so that subsequent lip modification can be performed on the target video based on the audio sequence. The second speech content refers to the sentence described by the audio sequence, such as the sentence “today is a good day, suitable for going out for a walk”, so that the second speech content can represent the sentence that is expected to be changed by the lip modification.

[0058] In addition, the present application does not limit the implementation of the above-mentioned audio sequence, and the following two scenarios are described below for ease of understanding.

[0059] Scenario one, when the video generation method provided by the present application is used to perform a certain video generation task, such as a lip modification task, the above-mentioned audio sequence can be specified by a user with certain means; and the present application does not limit the implementation of the audio sequence, for example, in some application scenarios, such as a video translation scenario, the audio sequence satisfies the following constraint: the semantic information of the sentence described by the audio sequence is consistent with the semantic information of the sentence described in the target video, but the language of the sentence described by the audio sequence is different from the language of the sentence described in the target video. For another example, in some application scenarios, such as a video sentence modification scenario, the audio sequence at least satisfies the following constraint: the sentence described by the audio sequence is partially the same as the sentence described in the target video. For another example, in some application scenarios, such as a video sentence replacement scenario, the audio sequence at least satisfies the following constraint: the sentence described by the audio sequence is completely different from the sentence described in the target video.

[0060] Scenario two, when the video generation method provided by the present application is used to implement a model training process, if the above-mentioned target video is determined according to a sample video, the above-mentioned audio sequence can be determined according to the sample video, so that the audio sequence includes part or all of the audio in the sample video.

[0061] In addition, the present application does not limit the implementation of the above-mentioned audio sequence, and the following two scenarios are described below for ease of understanding.

[0062] Scenario one, when the video generation method provided by the present application is used to perform a certain video generation task, such as a lip modification task, the above-mentioned audio sequence is obtained by receiving a piece of audio provided by a user or an upstream task as an audio sequence.

[0063] In the second scenario, when the video generation method provided in the present application is used to implement the model training process, the above-mentioned audio sequence acquisition process is: determining the audio sequence from the sample video, so that the audio sequence includes part or all of the audio in the sample video, so as to subsequently complete the model training process by means of the audio sequence.

[0064] Furthermore, the present application does not limit the association between the above-mentioned audio sequence and the above-mentioned target video, for example, both can satisfy the following constraint: for the i-th frame of audio in the audio sequence, there is a target image corresponding to the i-th frame of audio in the target video. Wherein, the i-th frame of audio refers to the audio in the audio sequence that is in the i-th arrangement position. i is a positive integer, i≤I, I represents the number of audio frames of the audio sequence.

[0065] For the i-th frame of audio in the above paragraph, the target image corresponding to the i-th frame of audio refers to the image in the target video that has a corresponding relationship with the i-th frame of audio, so that the target image is used to provide other information in addition to the facial expression state for the face image generation process of the i-th frame of audio, such as facial features, facial posture and the like. In addition, the present application does not limit the implementation of the target image, for example, when the number of audio frames in the audio sequence is the same as the number of image frames in the target video, the target image can refer to the i-th frame of image in the target video.

[0066] It should be noted that the present application does not limit the implementation of the corresponding relationship in the above paragraph, for example, when the number of audio frames in the audio sequence is equal to the number of image frames in the target video, the corresponding relationship can specifically include: the corresponding relationship between the i-th frame of audio in the audio sequence and the i-th frame of image in the target video. For another example, when the number of audio frames in the audio sequence is greater than the number of image frames in the target video, the corresponding relationship can specifically include: the corresponding relationship between the i-th frame of audio in the audio sequence and the i-th frame of image in the target video after frame increasing. The target video after frame increasing is obtained by performing a certain frame increasing processing on the original target video, so that the number of image frames in the target video after frame increasing is equal to the number of audio frames in the audio sequence; and the present application does not limit the implementation of the frame increasing processing, for example, it can be implemented by using any method capable of increasing the number of image frames in a video, such as frame insertion processing or copy splicing processing on the original video. For another example, when the number of audio frames in the audio sequence is less than the number of image frames in the target video, the corresponding relationship can specifically include: the corresponding relationship between the i-th frame of audio in the audio sequence and the i-th frame of image in the target video after frame decreasing. The target video after frame decreasing is obtained by performing a certain frame decreasing processing on the original target video, so that the number of image frames in the target video after frame decreasing is equal to the number of audio frames in the audio sequence; and the present application does not limit the implementation of the frame decreasing processing, for example, it can be implemented by using any method capable of decreasing the number of image frames in a video, such as sampling processing or cropping processing on the original video. Wherein, i is a positive integer, i≤I, I represents the number of audio frames in the audio sequence.

[0067] S2: According to the target video and the audio sequence, predict the face key points corresponding to each frame of audio in the audio sequence.

[0068] Wherein, for the i-th frame of audio in the audio sequence, the face key point corresponding to the i-th frame of audio is used to represent the facial expression state, such as the mouth shape state, of the speaker described by the target video at the i-th frame of audio.

[0069] In addition, the present application does not limit the implementation of the face key point corresponding to the i-th frame of audio in the above, for example, it can be implemented by using any key point capable of representing the facial expression state. For another example, in some application scenarios, in order to better improve the generation effect, the face key point corresponding to the i-th frame of audio can be implemented by using two-dimensional face key point. Wherein, the two-dimensional face key point is used to describe the facial expression state of an object in a two-dimensional space.

[0070] It can be seen that, in a possible implementation, the facial key points corresponding to the i-th frame of audio can include two-dimensional facial key points of the i-th frame of audio. The two-dimensional facial key points of the i-th frame of audio are used to represent a two-dimensional facial expression state of the speaking object described by the target video at the i-th frame of audio. The two-dimensional facial key points of the i-th frame of audio can be determined according to the i-th frame of audio and the target video.

[0071] In addition, the present application does not limit the determination manner of the two-dimensional facial key points of the i-th frame of audio. For example, any existing or future method capable of generating two-dimensional facial key points according to a video and a frame of audio can be used, such as a method implemented by means of a pre-constructed two-dimensional key point generation model. The two-dimensional key point generation model is used to generate two-dimensional facial key points for input data of the two-dimensional key point generation model, such as a video+audio. The present application does not limit the implementation manner of the two-dimensional key point generation model.

[0072] In addition, in order to better improve the generation effect, the present application further provides a determination manner of the two-dimensional facial key points of the i-th frame of audio. In this manner, the two-dimensional facial key points of the i-th frame of audio can be obtained by projecting three-dimensional facial key points corresponding to the i-th frame of audio to a two-dimensional plane. The three-dimensional facial key points corresponding to the i-th frame of audio are used to represent a three-dimensional facial expression state of the speaking object described by the target video at the i-th frame of audio. The three-dimensional facial key points corresponding to the i-th frame of audio are obtained by performing three-dimensional facial key point determination processing on the i-th frame of audio according to part or all of the images in the target video. It should be noted that the present application does not limit the implementation manner of the three-dimensional facial key point determination processing. For example, any existing or future method capable of determining three-dimensional facial key points according to some images and a frame of audio can be used, such as a method implemented by means of a pre-constructed three-dimensional key point generation model. The three-dimensional key point generation model is used to generate three-dimensional facial key points for input data of the three-dimensional key point generation model, such as a video+audio. The present application does not limit the implementation manner of the three-dimensional key point generation model.

[0073] For example, in some application scenarios, in order to better improve the generation effect, the determination process of the three-dimensional facial key points corresponding to the i-th frame of audio can include the following steps 11-12.

[0074] Step 11: According to the target video and the i-th frame of audio, the three-dimensional facial parameters corresponding to the i-th frame of audio are predicted.

[0075] The third-dimensional face parameter corresponding to the i-th frame of audio is used to represent the face state of the speaker described by the target video at the i-th frame of audio, such as face features, facial expressions, face postures, etc.

[0076] In addition, the present application does not limit the implementation of the third-dimensional face parameter corresponding to the i-th frame of audio. For example, the third-dimensional face parameter corresponding to the i-th frame of audio can include an identity document (ID) parameter, a facial expression parameter, and a face posture parameter. The identity document parameter is used to represent the identity features of the face of the speaker described by the target video, such as the face contour and the distribution of facial features, so that the third-dimensional face model constructed based on the identity document parameter is in the state of having ID, no posture, and no expression. The facial expression parameter is used to describe the facial expression state at the i-th frame of audio, such as the mouth shape state, so that the third-dimensional face model constructed based on the facial expression parameter is in the state of no ID, no posture, and having expression. The face posture parameter is used to describe the face posture at the i-th frame of audio, such as the side face and the front face, so that the third-dimensional face model constructed based on the face posture parameter is in the state of no ID, having posture, and no expression. It should be noted that the present application does not limit the association between the three parameters, such as the mutual decoupling state between the three parameters.

[0077] In addition, the present application does not limit the determination process of the third-dimensional face parameter corresponding to the i-th frame of audio. For example, it can use any existing or future method that can determine the third-dimensional face parameter based on a video and a frame of audio, such as a method implemented by means of a pre-constructed face parameter determination model. The face parameter determination model refers to a pre-constructed model with the function of determining the third-dimensional face parameter, such as a machine learning model, and the present application does not limit the implementation of the face parameter determination model.

[0078] Also, in some application scenarios, such as the lip modification scenario, in order to better improve the generation effect, the three-dimensional face parameters corresponding to the i-th frame of audio satisfy the following constraints: ① the face identity parameters and the face expression parameters in the three-dimensional face parameters corresponding to the i-th frame of audio are determined according to the target video and the i-th frame of audio, so that the face features described by the face identity parameters are consistent with the face features presented in the target video, and the face expression state represented by the face expression parameters satisfies the face expression state requirement of the i-th frame of audio; ② the face posture parameters in the three-dimensional face parameters corresponding to the i-th frame of audio are determined according to the face posture parameters in the three-dimensional face parameters of the target image corresponding to the i-th frame of audio, so that the face posture parameters in the three-dimensional face parameters corresponding to the i-th frame of audio are consistent with the face posture parameters in the three-dimensional face parameters of the target image. The three-dimensional face parameters of the target image are obtained by performing three-dimensional face parameter determination processing on the target image; and the present application does not limit the determination process of the three-dimensional face parameters of the target image, for example, it can be implemented by using any existing or future method capable of performing three-dimensional face parameter determination processing on an image.

[0079] Step 12: determining the three-dimensional face key points corresponding to the i-th frame of audio according to the three-dimensional face parameters corresponding to the i-th frame of audio.

[0080] It should be noted that the present application does not limit the implementation of step 12 above, for example, it can be implemented by using any existing or future method capable of determining three-dimensional face key points based on three-dimensional face parameters.

[0081] Based on the related content of steps 11 to 12 above, in some application scenarios, for the i-th frame of audio in the audio sequence, the face reconstruction processing can be performed according to the target video to obtain the three-dimensional face parameters of each frame of image in the target video; then, the three-dimensional face parameters corresponding to the i-th frame of audio are predicted according to the three-dimensional face parameters of part or all of the images in the target video and the audio features of the i-th frame of audio; then, the three-dimensional face key points corresponding to the i-th frame of audio are determined according to the three-dimensional face parameters corresponding to the i-th frame of audio, so that the three-dimensional face key points can better represent the face state of the speaking object described by the target video at the i-th frame of video, so that the face key points corresponding to the i-th frame of audio determined based on the three-dimensional face key points are more accurate, and thus the generation effect is improved.

[0082] It is found through research that when the same prediction process is used to simultaneously implement prediction processing for different facial regions of the speaker, the face of the speaker is generated with poor quality due to interference between different facial regions. Therefore, in order to improve the generation effect, different and independent prediction processes can be used for different facial regions to overcome the defects caused by the interference.

[0083] Based on the above research, the present application also provides a possible implementation of the face key point corresponding to the i-th frame of audio, in which the face key point corresponding to the i-th frame of audio at least satisfies the following constraints: the face key point corresponding to the i-th frame of audio includes a key point corresponding to the i-th frame of audio in the first region and a key point corresponding to the i-th frame of audio in the second region. In order to facilitate understanding, the two key points are introduced below.

[0084] The above "key point corresponding to the i-th frame of audio in the first region" is used to describe the state of the first region of the speaker under the i-th frame of audio, so that the "key point corresponding to the i-th frame of audio in the first region" is used to affect the prediction process of the first region, so that the finally generated first region for the i-th frame of audio at least satisfies the constraints described by the "key point corresponding to the i-th frame of audio in the first region", such as mouth shape state constraints and the like.

[0085] In addition, the present application does not limit the determination process of the above "key point corresponding to the i-th frame of audio in the first region", for example, it can be specifically: first, find a key point for describing the first region from the face key point corresponding to the i-th frame of audio, so that the key point for describing the first region can represent the state of the first region under the i-th frame of audio; then, determine the "key point corresponding to the i-th frame of audio in the first region" according to the key point for describing the first region, so that the "key point corresponding to the i-th frame of audio in the first region" includes the key point for describing the first region, so that the "key point corresponding to the i-th frame of audio in the first region" can more accurately describe the state of the first region under the i-th frame of audio.

[0086] The above "key point corresponding to the i-th frame of audio in the second region" is used to describe the state of the second region of the speaker under the i-th frame of audio, so that the "key point corresponding to the i-th frame of audio in the second region" is used to affect the prediction process of the second region, so that the finally generated second region for the i-th frame of audio at least satisfies the constraints described by the "key point corresponding to the i-th frame of audio in the second region".

[0087] In addition, the application does not limit the determination process of the key point corresponding to the i-th frame of audio in the second region, which can be specifically: first, searching for a key point for describing the second region from the face key points corresponding to the i-th frame of audio, so that the key point for describing the second region can represent the state of the second region in the i-th frame of audio; then, determining the key point corresponding to the i-th frame of audio in the second region according to the key point for describing the second region, so that the key point corresponding to the i-th frame of audio in the second region includes the key point for describing the second region, so that the key point corresponding to the i-th frame of audio in the second region can more accurately describe the state of the second region in the i-th frame of audio.

[0088] In addition, since the first region is used to describe the mouth region, and the second region is used to describe the surrounding region of the mouth region, so there is some correlation between the boundary of the second region and the boundary of the first region, in order to better meet the constraints of these correlations, the application also provides a determination process of the key point corresponding to the i-th frame of audio in the second region, which can be specifically: searching for a key point for describing the lower half face from the face key points corresponding to the i-th frame of audio, so that the key point for describing the lower half face can not only represent the state of the second region in the i-th frame of audio, but also represent the state of the correlation between the second region and the first region in the i-th frame of audio; then, determining the key point corresponding to the i-th frame of audio in the second region according to the key point for describing the lower half face, so that the key point corresponding to the i-th frame of audio in the second region includes the key point for describing the lower half face, so that the second region generated based on the key point corresponding to the i-th frame of audio in the second region meets the constraints of these correlations, and the second region generated based on the key point corresponding to the i-th frame of audio in the second region is more accurate, which is beneficial to improve the generation effect.

[0089] Based on the above-mentioned related content of S2, for some scenes, after obtaining the target video and the audio sequence, the face key points corresponding to each frame of audio in the audio sequence can be predicted based on the target video and the audio sequence, so that the face corresponding to each frame of audio in the audio sequence can be generated based on the face key points subsequently.

[0090] S3: generating an image corresponding to the i-th frame of audio according to the face key point corresponding to the i-th frame of audio, at least two reference images corresponding to the i-th frame of audio in the first region, at least two reference images corresponding to the i-th frame of audio in the second region, and the lower half face mask result of the target image corresponding to the i-th frame of audio, the target video including the target image and the reference images, i being a positive integer, i≤the number of audio frames in the audio sequence.

[0091] wherein, for the i-th frame audio in the audio sequence, the target image corresponding to the i-th frame audio refers to an image existing in the target video and used to provide other information than the facial expression state for the i-th frame audio in the target video, such as the i-th frame image in the target video, etc.

[0092] In addition, for the target image corresponding to the i-th frame audio, the lower half face mask result of the target image is used to provide other information than the lower half face in the target image; and the application does not limit the determination process of the lower half face mask result, for example, it can be specifically: performing mask processing on the lower half face in the target image to obtain the lower half face mask result of the target image, so that the lower half face mask result can at least represent the upper half face in the target image, thereby making the lower half face mask result be able to provide some other information in the target image, such as the upper half face, the background, the light, etc., and further making the image generated based on the lower half face mask result consistent with the target image in these other information, which is thus beneficial to improve the generation effect.

[0093] For the i-th frame audio in the audio sequence, the at least two reference images corresponding to the i-th frame audio under the first region refer to images existing in the target video and required for reference when performing prediction processing on the first region, such as the plurality of images 1 shown in FIG. 4, so that the at least two reference images corresponding to the i-th frame audio under the first region are used to affect the prediction processing of the first region, thereby making the finally generated first region for the i-th frame audio at least satisfy some constraints described by the at least two reference images corresponding to the i-th frame audio under the first region.

[0094] In addition, the application does not limit the implementation of the at least two reference images corresponding to the i-th frame audio under the first region, for example, it can include some images randomly extracted from the target video.

[0095] In addition, in order to better improve the generation effect, the application further provides a determination manner of the at least two reference images corresponding to the i-th frame audio under the first region, which can specifically include the following steps 21-23.

[0096] Step 21: sorting the images in the target video according to the mouth shape amplitude representation data of each frame image in the target video to obtain an image sequence.

[0097] wherein, the j-th frame image in the target video refers to an image existing in the target video and located at the j-th arrangement position, j is a positive integer, j≤J, J is a positive integer, and J represents the number of image frames in the target video.

[0098] In addition, for the jth image in the target video, the mouth opening amplitude feature data of the jth image is used to represent the mouth opening amplitude of the speaker in the jth image, and the application does not limit the determination method of the mouth opening amplitude feature data of the jth image. For example, it can be implemented by using any existing or future method capable of determining the mouth opening amplitude of an image, such as a method implemented by using a pre-constructed mouth opening amplitude determination model. The mouth opening amplitude determination model refers to a model pre-constructed and capable of determining the mouth opening amplitude of an image, such as a machine learning model.

[0099] In addition, in order to better improve the generation effect, the application further provides a determination method of the mouth opening amplitude feature data of the jth image. In this method, the determination process of the mouth opening amplitude feature data of the jth image can be: determining the mouth opening amplitude feature data of the jth image according to the mouth key point of the facial key point of the jth image, such as the two-dimensional facial key point and / or the three-dimensional facial key point, so that the mouth opening amplitude feature data can more accurately represent the mouth opening amplitude of the speaker in the jth image. The three-dimensional facial key point of the jth image is used to represent the state of the face state presented in the jth image in the three-dimensional space; and the three-dimensional facial key point of the jth image is obtained by performing three-dimensional facial key point determination processing on the jth image. It should be noted that the application does not limit the implementation of the three-dimensional facial key point determination processing. The two-dimensional facial key point of the jth image is used to represent the face state of the object in the jth image in the two-dimensional space; and the application does not limit the determination method of the two-dimensional facial key point of the jth image. For example, it can be implemented by using any existing or future method capable of determining the two-dimensional facial key point of an image. For example, the two-dimensional facial key point of the jth image can be obtained by projecting the three-dimensional facial key point of the jth image to a two-dimensional plane, which is beneficial to improve the consistency of the key point related information, thereby improving the generation effect.

[0100] In addition, for the image sequence in step 21, the determination process of the image sequence can be: sorting part or all of the images in the target video according to the mouth opening amplitude feature data of each image in the target video to obtain an image sequence, so that all the images in the image sequence are arranged in ascending or descending order according to the mouth opening amplitude feature data, thereby making the image sequence better describe the mouth opening amplitude distribution presented in the target video.

[0101] It can be seen that, in a possible implementation, the image sequence can include all images in the target video; and all images in the image sequence are arranged in ascending order or descending order according to the lip amplitude feature data.

[0102] Based on the related content of step 21, for some application scenarios, after obtaining the target video, the lip amplitude feature data of each frame image in the target video can be determined first; then the images are arranged in ascending order or descending order according to the lip amplitude feature data, to obtain an image sequence, so that the image sequence includes part or all images in the target video, and the arrangement position of each image in the image sequence is positively or negatively correlated with the lip amplitude feature data of the corresponding image in the target video, so that the image sequence can better represent the lip shape change in the target video.

[0103] Step 22: equally interval sampling the image sequence to obtain a sampling image.

[0104] The sampling image refers to an image sampled from the image sequence; and the application does not limit the sampling interval required for the sampling, for example, the sampling interval can be determined according to actual needs, such as generation efficiency needs or generation accuracy needs.

[0105] Based on the related content of step 22, for some application scenarios, after obtaining the image sequence, some images can be equally interval sampled from the image sequence as sampling images, so that the sampling images can represent different lip shape states of the speaker in the target video, thereby the sampling images can provide useful lip shape information for the generation process of the first region corresponding to the i th frame of audio, which is conducive to improving the generation effect.

[0106] Step 23: determining at least two frames of reference images corresponding to the i th frame of audio in the first region according to the sampling image.

[0107] It should be noted that the application does not limit the implementation of step 23, for example, step 23 can be specifically: determining the sampling image as the at least two frames of reference images corresponding to the i th frame of audio in the first region, so that the at least two frames of reference images corresponding to the i th frame of audio in the first region can provide some reference information such as lip shape information for the prediction processing of the first region corresponding to the i th frame of audio.

[0108] Based on the related content of steps 21-23 above, in some scenarios, the present application selects some images with different mouth shape states from the target video through related processing on the target video, such as the sampling images above, so as to apply these images with different mouth shape states to the prediction processing of the first region corresponding to each frame of audio in the audio sequence, so that the at least two reference images referred to by the prediction processing of the first region corresponding to different frames of audio in the audio sequence remain consistent.

[0109] For the i-th frame of audio in the audio sequence, the at least two reference images corresponding to the second region of the i-th frame of audio refer to images existing in the target video and required to be referred to when performing prediction processing on the second region, such as the multiple images 2 shown in FIG. 4, so that the at least two reference images corresponding to the second region of the i-th frame of audio are used to affect the prediction processing of the second region, so that the finally generated second region for the i-th frame of audio at least meets some constraints described by the at least two reference images corresponding to the second region of the i-th frame of audio.

[0110] In addition, the at least two reference images corresponding to the second region of the i-th frame of audio can at least meet the following constraint: the at least two reference images corresponding to the second region of the i-th frame of audio are different from the at least two reference images corresponding to the first region of the i-th frame of audio, so that the present application can realize prediction of different face regions by referring to different images, thereby facilitating improvement of the generated effect.

[0111] In addition, the present application does not limit the implementation of the at least two reference images corresponding to the second region of the i-th frame of audio, for example, it can include some images randomly extracted from the target video.

[0112] It is found through research that for any frame of image in a video, although the mouth shape described by the image can be different from the mouth shape described by other images located near the image, other information described by the image, such as background, face posture, etc., is very close to, or even the same as, the corresponding information described by other images located near the image.

[0113] Based on the above research, in order to better improve the generation effect, the present application also provides a possible implementation of the above "at least two frames of reference images corresponding to the i-th frame of audio in the second region", in which the "at least two frames of reference images corresponding to the i-th frame of audio in the second region" can at least satisfy the following constraint: for any reference image in the "at least two frames of reference images corresponding to the i-th frame of audio in the second region", the distance between the arrangement position of the reference image in the target video and the arrangement position of the target image corresponding to the i-th frame of audio in the target video is not greater than a preset distance threshold, so that the "at least two frames of reference images corresponding to the i-th frame of audio in the second region" not only includes the target image corresponding to the i-th frame of audio in the target video, but also includes images near the target image in the target video, so that the "at least two frames of reference images corresponding to the i-th frame of audio in the second region" can as comprehensively as possible represent the state of the i-th frame of audio in other aspects except for the lip shape.

[0114] Wherein, the preset distance threshold can be pre-specified, or can be determined according to the video generation requirements of actual application scenarios in advance, which is not limited in the present application. For example, the preset distance threshold can be 3.

[0115] Based on the above two paragraphs, in a possible implementation, if the target video and the audio sequence are in a time alignment state, the target image corresponding to the i-th frame of audio in the audio sequence is the i-th frame of image in the target video, and when the above preset distance threshold is D, the "at least two frames of reference images corresponding to the i-th frame of audio in the second region" can include the i-D-th frame of image in the target video, the i-D+1-th frame of image in the target video, and so on, and the i+D-th frame of image in the target video, so that the arrangement positions of each reference image in the "at least two frames of reference images corresponding to the i-th frame of audio in the second region" in the target video belong to the interval [i-D, i+D], and D is a positive integer, such as D=3.

[0116] It can be seen that, in a possible implementation, the determination process of the above "at least two frames of reference images corresponding to the i-th frame of audio in the second region" can be: first, obtaining the arrangement position of the target image corresponding to the i-th frame of audio in the target video, such as the arrangement position i; then, according to the preset distance threshold and the arrangement position, calculating the position interval, such as the interval [i-D, i+D], so that the difference between each position in the position interval and the arrangement position of the target image in the target video does not exceed the preset distance threshold; and then, searching for images with arrangement positions belonging to the position interval from the target video as the "at least two frames of reference images corresponding to the i-th frame of audio in the second region".

[0117] For the i-th frame of audio in the audio sequence, the image corresponding to the i-th frame of audio refers to an image generated or predicted for the i-th frame of audio, such that the image corresponding to the i-th frame of audio satisfies the following constraints: the facial expression state presented in the image corresponding to the i-th frame of audio satisfies the facial expression state requirement of the i-th frame of audio, and other information in the image corresponding to the i-th frame of audio, except for the facial expression state, is consistent with the corresponding information in the target image corresponding to the i-th frame of audio, so that the image corresponding to the i-th frame of audio can represent the result of performing facial expression state adjustment processing on the target image corresponding to the i-th frame of audio according to the i-th frame of audio.

[0118] In addition, the present application does not limit the implementation of S3 above, for example, it can be implemented by using any machine learning model for image generation processing, such as a Gan model or a diffusion model. It can be seen that in one possible implementation, S3 can be specifically: performing image generation processing by the machine learning model according to the facial key points corresponding to the i-th frame of audio, the at least two frames of reference images corresponding to the i-th frame of audio in the first region, the at least two frames of reference images corresponding to the i-th frame of audio in the second region, and the lower half face mask result of the target image corresponding to the i-th frame of audio, to obtain the image corresponding to the i-th frame of audio, so that the image generation processing satisfies the following constraints: the at least two frames of reference images corresponding to the i-th frame of audio in the first region are used to affect the generation processing of the first region in the image corresponding to the i-th frame of audio, and the at least two frames of reference images corresponding to the i-th frame of audio in the second region are used to affect the generation processing of the second region in the image corresponding to the i-th frame of audio, so as to realize the generation of different regions according to different reference images, thereby facilitating the improvement of the generation effect.

[0119] In addition, in order to better improve the generation effect, the present application also provides one possible implementation of S3 above, in which when the facial key points corresponding to the i-th frame of audio in the audio sequence above include the key points corresponding to the i-th frame of audio in the first region and the key points corresponding to the i-th frame of audio in the second region, S3 can be specifically: generating the image corresponding to the i-th frame of audio according to the key points corresponding to the i-th frame of audio in the first region, the at least two frames of reference images corresponding to the i-th frame of audio in the first region, the key points corresponding to the i-th frame of audio in the second region, the at least two frames of reference images corresponding to the i-th frame of audio in the second region, and the lower half face mask result of the target image corresponding to the i-th frame of audio.

[0120] Also, the application does not limit the implementation of S3 shown in the above paragraph, for example, it can be implemented by using any machine learning model for implementing the generation processing, such as a Gan model or a diffusion model. As can be seen, in one possible implementation, S3 can specifically be: performing image generation processing on the lower half face mask result of the target image corresponding to the i-th frame of audio by the machine learning model according to the key points corresponding to the i-th frame of audio in the first region, the at least two frames of reference images corresponding to the i-th frame of audio in the first region, the key points corresponding to the i-th frame of audio in the second region, the at least two frames of reference images corresponding to the i-th frame of audio in the second region, and the target image corresponding to the i-th frame of audio, to obtain the image corresponding to the i-th frame of audio, so that the image generation processing satisfies the following constraints: the key points corresponding to the i-th frame of audio in the first region and the at least two frames of reference images corresponding to the i-th frame of audio in the first region are both used to affect the generation processing of the first region in the image corresponding to the i-th frame of audio, and the key points corresponding to the i-th frame of audio in the second region and the at least two frames of reference images corresponding to the i-th frame of audio in the second region are both used to affect the generation processing of the second region in the image corresponding to the i-th frame of audio, so that different regions can be generated according to different reference images and different key points, thereby facilitating the improvement of the generation effect.

[0121] It has been found through research that, in order to better improve the generation effect, different and independent processing processes can be used for different regions to realize prediction. Based on this, the application also provides one possible implementation of S3 above, in which S3 can specifically include the following steps 31-33.

[0122] Step 31: determining the generation result of the first region corresponding to the i-th frame of audio according to the face key points corresponding to the i-th frame of audio and the at least two frames of reference images corresponding to the i-th frame of audio in the first region.

[0123] The generation result of the first region corresponding to the i-th frame of audio is used to describe the state of the first region of the speaker under the i-th frame of audio, such as the lip shape state.

[0124] In addition, the application does not limit the implementation of step 31 above, for example, it can be implemented by using any machine learning model for implementing the generation processing, such as a Gan model or a diffusion model. As can be seen, in one possible implementation, step 31 can specifically be: performing generation processing, such as mouth region generation processing, by the machine learning model according to the face key points corresponding to the i-th frame of audio and the at least two frames of reference images corresponding to the i-th frame of audio in the first region, to obtain the generation result of the first region corresponding to the i-th frame of audio, so that the generation result can represent the state of the first region of the speaker under the i-th frame of audio.

[0125] Further, in order to improve the generation effect, the present application further provides a possible implementation of the step 31, in which when the face key points corresponding to the i-th frame of audio include at least the key points corresponding to the i-th frame of audio in the first region, the step 31 can be specifically: determining the generation result of the first region corresponding to the i-th frame of audio according to the key points corresponding to the i-th frame of audio in the first region and at least two reference images corresponding to the i-th frame of audio in the first region, so as to effectively avoid the interference caused by the key points of the regions other than the first region, thereby facilitating the improvement of the generation effect.

[0126] In addition, the present application does not limit the implementation of the step 31 in the above paragraph, for example, it can be implemented by using any machine learning model for implementing the generation process, such as Gan model or diffusion model. It can be seen that in a possible implementation, the step 31 can be specifically: performing generation processing, such as mouth region generation processing, etc., by the machine learning model according to the key points corresponding to the i-th frame of audio in the first region and at least two reference images corresponding to the i-th frame of audio in the first region, to obtain the generation result of the first region corresponding to the i-th frame of audio, so that the generation result can more accurately represent the state of the first region of the speaker in the i-th frame of audio.

[0127] Further, in order to improve the generation effect, the present application further provides a possible implementation of the step 31, in which the step 31 can be specifically: determining the generation result of the first region corresponding to the i-th frame of audio according to the face key points corresponding to the i-th frame of audio (or the key points corresponding to the i-th frame of audio in the first region), at least two reference images corresponding to the i-th frame of audio in the first region, and the key point determination results of each of the at least two reference images corresponding to the i-th frame of audio in the first region.

[0128] In which, for any reference image in the at least two reference images corresponding to the i-th frame of audio in the first region, the key point determination result of the reference image is used to describe the state of the first region in the reference image.

[0129] It can be seen that in a possible implementation, for the k-th reference image in the at least two reference images corresponding to the i-th frame of audio in the first region, the key point determination result of the k-th reference image is used to describe the state of the first region in the k-th reference image, k is a positive integer, k≤K, K is a positive integer, and K represents the number of the at least two reference images.

[0130] In addition, the application does not limit the determination process of the key point determination result of the kth reference image in the above paragraph, for example, any existing or future key point detection method for a first region in an image can be used, such as a machine learning model with a pre-constructed mouth region key point detection function.

[0131] In addition, in order to better improve the generation effect, the determination process of the key point determination result of the kth reference image in the above paragraph can be: first, determining the face key point of the kth reference image according to the kth reference image, so that the face key point is used to describe the face state presented in the kth reference image; then, searching for a key point for describing the first region from the face key point of the kth reference image; and then, determining the key point determination result of the kth reference image according to the key point for describing the first region, so that the key point determination result is used to describe the state of the first region in the kth reference image.

[0132] The face key point of the kth reference image in the above paragraph is used to describe the face state presented in the kth reference image, and the application does not limit the implementation of the face key point of the kth reference image, for example, the implementation of the face key point of the kth reference image is similar to the implementation of the face key point corresponding to the ith frame of audio in the above paragraph.

[0133] It can be seen that in a possible implementation, if the face key point corresponding to the ith frame of audio in the above paragraph is implemented by using a two-dimensional face key point, the face key point of the kth reference image can also be implemented by using a two-dimensional face key point, and the face key point of the kth reference image and the face key point corresponding to the ith frame of audio satisfy the following constraint: the face key point of the kth reference image and the face key point corresponding to the ith frame of audio are corresponding in key point serial number.

[0134] In addition, the application does not limit the implementation of the face key point of the kth reference image in the above paragraph, for example, the face key point of the kth reference image can be implemented by using any existing or future face key point recognition method, such as a machine learning model with a face key point recognition function.

[0135] For example, if the face key point corresponding to the i-th frame of audio above is obtained by projecting the three-dimensional face key point to a two-dimensional plane, the face key point of the k-th reference image above can be a two-dimensional face key point obtained by projecting the three-dimensional face key point of the k-th reference image to a two-dimensional plane, so that the face key point of the k-th reference image is obtained in the same way as the face key point corresponding to the i-th frame of audio, which can effectively avoid defects caused by different two-dimensional key point obtaining mechanisms, thereby improving the generation effect. The three-dimensional face key point of the k-th reference image is used to describe the state of the face presented in the k-th reference image in three-dimensional space; and the three-dimensional face key point of the k-th reference image is obtained by three-dimensional face key point determination processing of the k-th reference image. It should be noted that the present application does not limit the implementation of the three-dimensional face key point determination processing, for example, it can use any method that can determine the three-dimensional face key point of an image, such as a method implemented by a pre-constructed machine learning model with three-dimensional face key point determination processing function.

[0136] In addition, the present application does not limit the implementation of the above step of "determining the generation result of the first region corresponding to the i-th frame of audio according to the face key point corresponding to the i-th frame of audio (or the key point corresponding to the i-th frame of audio in the first region), the at least two frames of reference images corresponding to the i-th frame of audio in the first region, and the key point determination result of each frame of reference image in the at least two frames of reference images corresponding to the i-th frame of audio in the first region". For example, it can be implemented by using any machine learning model for generation processing, such as Gan model or diffusion model.

[0137] In addition, in order to better improve the generation effect, the present application also provides a possible implementation of the above step 31, in which the step 31 can specifically include the following steps 311-312.

[0138] Step 311: For any reference image in the at least two frames of reference images corresponding to the i-th frame of audio in the first region, predict the pixel usage description information corresponding to the reference image according to the reference image, the key point determination result of the reference image, and the face key point corresponding to the i-th frame of audio (or the key point corresponding to the i-th frame of audio in the first region); the pixel usage description information includes pixel adjustment description information and / or pixel fusion weight.

[0139] The pixel usage description information of the kth reference image corresponds to pixels of the kth reference image, and is used to describe how the pixels of the kth reference image are used in the generation of the first region corresponding to the ith frame of audio, so that the pixel usage description information can indicate how the part or all of the pixels of the kth reference image affect the generation of the first region corresponding to the ith frame of audio.

[0140] In addition, the present application does not limit the implementation of the pixel usage description information of the kth reference image, for example, the pixel usage description information of the kth reference image can include pixel adjustment description information of the kth reference image and / or pixel fusion weight of the kth reference image.

[0141] In addition, for the pixel adjustment description information of the kth reference image, the pixel adjustment description information is used to indicate how to adjust the part or all of the pixels of the kth reference image in the generation of the first region corresponding to the ith frame of audio; and the present application does not limit the implementation of the pixel adjustment description information, for example, the pixel adjustment description information can be implemented by using the offset of the image pixel information. It can be seen that, in a possible implementation, the pixel adjustment description information satisfies the following constraints: the size of the pixel adjustment description information is the same as the size of the kth reference image, and the position coordinates of each pixel point in the pixel adjustment description information are used to indicate the offset of the position coordinates of the corresponding pixel point in the kth reference image.

[0142] In addition, for the pixel fusion weight of the kth reference image, the pixel fusion weight is used to indicate how the pixels of the kth reference image affect the generation of the first region corresponding to the ith frame of audio; and the present application does not limit the implementation of the pixel fusion weight, for example, the pixel fusion weight can satisfy the following constraints: the size of the pixel fusion weight is the same as the size of the kth reference image, and the weight value of each pixel point in the pixel fusion weight is used to indicate the influence degree of the corresponding pixel point in the kth reference image.

[0143] In addition, the present application does not limit the implementation of the step 311, for example, in order to better improve the generation effect, the step 311 can be implemented by using the prediction module in the mouth region rendering model (such as the prediction model of the mouth region shown in FIG. 4).

[0144] It can be seen that in a possible implementation, step 311 can be specifically: for any one of the at least two reference images corresponding to the first region of the i-th frame of audio, the prediction module in the mouth region rendering model performs prediction processing according to the reference image, the key point determination result of the reference image, and the face key point corresponding to the i-th frame of audio (or the key point corresponding to the first region of the i-th frame of audio), such as the offset prediction processing 1 shown in FIG. 4, to obtain and output the pixel usage description information corresponding to the reference image, such as the processing result 1 shown in FIG. 4.

[0145] For the mouth region rendering model shown in the above two paragraphs, the prediction model of the mouth region shown in FIG. 4 can be used for mouth region rendering processing of the input data of the mouth region prediction model; and the present application does not limit the implementation of the mouth region prediction model, for example, the prediction module in the mouth region rendering model is used to perform prediction processing according to a reference image, a key point determination result of the reference image, and a face key point corresponding to the i-th frame of audio (or a key point corresponding to the first region of the i-th frame of audio), to obtain and output pixel usage description information corresponding to the reference image; and the present application does not limit the implementation of the prediction module, for example, in order to better improve the generation effect, the prediction module can include a convolutional neural network (CNN) and a feature injection module, so that the prediction module has the function of fusing the reference image, the key point determination result of the reference image, and the face key point corresponding to the i-th frame of audio (or the key point corresponding to the first region of the i-th frame of audio). It should be noted that the present application does not limit the implementation of the feature injection module, for example, the feature injection module can be implemented by using AdaIN (Adaptive Instance Normalization) or SPADE.

[0146] Based on the related content of step 311, after obtaining any one of the at least two reference images corresponding to the first area of the i-th frame of audio, the key point determination result of each reference image, and the key point corresponding to the first area of the i-th frame of audio, these data can be input into the mouth region rendering model, so that the prediction module in the mouth region rendering model can predict and output the pixel usage description information corresponding to the k-th reference image according to the k-th reference image, the key point determination result of the k-th reference image, and the key point corresponding to the first area of the i-th frame of audio, so that the pixel usage description information can indicate how to use the pixels in the k-th reference image in the generation process of the first area corresponding to the i-th frame of audio, such as position offset and influence degree, k is a positive integer, k≤K, K is a positive integer, so that subsequent generation processing of the first area of the i-th frame of audio can be based on the pixel usage description information corresponding to the reference image.

[0147] Step 312: deforming and fusing the at least two reference images according to the pixel usage description information corresponding to the at least two reference images to obtain the generation result of the first area corresponding to the i-th frame of audio, so that the generation result can represent the image obtained by deforming and fusing the at least two reference images according to the pixel usage description information corresponding to the at least two reference images, so that the first area described by the generation result is adapted to the i-th frame of audio.

[0148] It should be noted that the present application does not limit the implementation of step 312, for example, in order to better improve the generation effect, step 312 can be specifically: after the prediction module in the mouth region rendering model outputs the pixel usage description information corresponding to each reference image, the deforming and fusing module in the mouth region rendering model deforms and fuses the reference images according to the pixel usage description information corresponding to the reference images, such as the generation processing 1 shown in FIG. 4, to obtain and output the generation result of the first area corresponding to the i-th frame of audio, such as the generation result 1 shown in FIG. 4, so that the fused image can represent the pixel integration result of the reference images under the i-th frame of audio, so that the first area described by the generation result is adapted to the i-th frame of audio, and thus the generation result can meet the following constraint: the first area represented in the generation result meets the mouth region state requirement of the i-th frame of audio. Wherein, the deforming and fusing module is used to integrate and use the pixels in the reference images according to the pixel usage description information corresponding to each reference image; and the present application does not limit the implementation of the deforming and fusing module, for example, the deforming and fusing module can adopt any kind of information integration network, such as grid_sample.

[0149] Based on the related content of the above steps 311 to 312, for the above mouth region rendering model, first, the prediction module in the mouth region rendering model predicts and outputs the pixel usage description information corresponding to the kth reference image among the at least two reference images corresponding to the first region of the ith audio under the first region, the key point determination result of the kth reference image, and the key point corresponding to the first region of the ith audio under the first region according to the ith audio under the first region, k is a positive integer, k≤K; then, the deformation fusion module in the mouth region rendering model performs deformation fusion processing on the reference images according to the pixel usage description information corresponding to the reference images to obtain and output the generation result of the first region corresponding to the ith audio, so that the image corresponding to the ith audio can be determined based on the generation result subsequently.

[0150] It can be seen that for some application scenarios, after obtaining the at least two reference images corresponding to the first region of the ith audio, the key point determination result of each reference image, and the face key point corresponding to the ith audio (or the key point corresponding to the first region of the ith audio), these data can be input into the mouth region rendering model, such as the prediction model of the mouth region shown in FIG. 4, so that the mouth region rendering model can generate and output the generation result of the first region corresponding to the ith audio according to these data, such as the generation result 1 shown in FIG. 4. Wherein, because the mouth region rendering model has good performance, the first region generated by means of the mouth region rendering model is better, which is conducive to improving the generation effect.

[0151] Based on the related content of the above step 31, for some scenarios, after obtaining the face key point corresponding to the ith audio and the at least two reference images corresponding to the first region of the ith audio, the generation result of the first region corresponding to the ith audio can be determined by means of these data, so that the generation result can represent the state of the first region of the speaking object under the ith audio.

[0152] Step 32: determining the generation result of the second region corresponding to the ith audio according to the face key point corresponding to the ith audio and the at least two reference images corresponding to the second region of the ith audio under the second region.

[0153] Wherein, the generation result of the second region corresponding to the ith audio is used to describe the state of the second region of the speaking object under the ith audio, such as the lip shape state and the like.

[0154] In addition, the present application does not limit the implementation of the above step 32, for example, the implementation of the implementation of the step 32 is similar to the implementation of the above step 31, and for the sake of brevity, it will not be repeated here.

[0155] For the convenience of understanding, some possible implementations of step 32 are taken as examples for illustration below.

[0156] In example 1, in a possible implementation, step 32 above can be specifically: determining the generation result of the second region corresponding to the i-th frame of audio according to the key points corresponding to the i-th frame of audio under the second region and at least two frames of reference images corresponding to the i-th frame of audio under the second region, so as to effectively avoid the interference caused by the key points of the regions other than the lower half of the face, thereby facilitating to improve the generation effect.

[0157] In example 2, in a possible implementation, step 32 above can be specifically: determining the generation result of the second region corresponding to the i-th frame of audio according to the face key points corresponding to the i-th frame of audio (or the key points corresponding to the i-th frame of audio under the second region), at least two frames of reference images corresponding to the i-th frame of audio under the second region, and the key point determination result of each frame of reference image in the at least two frames of reference images corresponding to the i-th frame of audio under the second region.

[0158] For any reference image in the at least two frames of reference images corresponding to the i-th frame of audio under the second region, the key point determination result of the reference image is used to describe the state of the second region or the lower half of the face in the reference image.

[0159] It can be seen that, in a possible implementation, for the q-th reference image in the at least two frames of reference images corresponding to the i-th frame of audio under the second region, the key point determination result of the q-th reference image is used to describe the state of the second region in the q-th reference image, q is a positive integer, q≤Q, Q is a positive integer, and Q represents the number of the at least two frames of reference images.

[0160] In another possible implementation, for the q-th reference image in the at least two frames of reference images corresponding to the i-th frame of audio under the second region, the key point determination result of the q-th reference image is used to describe the state of the lower half of the face in the q-th reference image, so that the key point determination result is not only used to describe the characteristics of the second region, but also can describe the characteristics of the related content between the second region and the first region, q is a positive integer, q≤Q, Q is a positive integer, so as to facilitate to improve the generation effect.

[0161] In Example 3, in a possible implementation, step 32 above can be implemented by means of a non-mouth region rendering model for predicting a region in the lower half of the face other than the mouth region (also referred to as a non-mouth region), and for the non-mouth region rendering model, first, a prediction module in the non-mouth region rendering model performs prediction processing on the qth reference image among the at least two reference images corresponding to the second region under the ith frame of audio, the key point determination result of the qth reference image, and the face key point corresponding to the ith frame of audio (or the key point corresponding to the second region under the ith frame of audio), as shown in offset prediction processing 2 in FIG. 4, to obtain and output pixel usage description information corresponding to the qth reference image, as shown in processing result 2 in FIG. 4, q is a positive integer and q≤Q; then a morphing fusion module in the non-mouth region rendering model performs morphing fusion processing on the reference images according to the pixel usage description information corresponding to the reference images, as shown in generation processing 2 in FIG. 4, to obtain and output a generation result of the second region corresponding to the ith frame of audio, as shown in generation result 2 in FIG. 4, so that the generation result can be used to determine the image corresponding to the ith frame of audio.

[0162] It should be noted that, for the pixel usage description information corresponding to the qth reference image in the above paragraph, the pixel usage description information corresponding to the qth reference image is used to describe how to use the pixels in the qth reference image in the generation process of the second region corresponding to the ith frame of audio, so that the pixel usage description information can indicate the influence of part or all of the pixels in the qth reference image on the generation process of the second region corresponding to the ith frame of audio; and the implementation of the pixel usage description information corresponding to the qth reference image is similar to the implementation of the pixel usage description information corresponding to the kth reference image above, which will not be described here for brevity.

[0163] Based on the related content of step 32 above, for some scenarios, after obtaining the face key point corresponding to the ith frame of audio and the at least two reference images corresponding to the second region under the ith frame of audio, the generation result of the second region corresponding to the ith frame of audio can be determined by means of these data, so that the generation result can indicate the state of the second region of the speaker under the ith frame of audio.

[0164] Step 33: generating the image corresponding to the ith frame of audio according to the generation result of the first region corresponding to the ith frame of audio, the generation result of the second region corresponding to the ith frame of audio, and the lower half of the face mask result of the target image corresponding to the ith frame of audio.

[0165] It should be noted that the present application does not limit the implementation of the above step 33, for example, it can be implemented by means of any machine learning model for realizing the generation processing, such as a Gan model or a diffusion model. It can be seen that in one possible implementation, the step 33 can be specifically: generating and outputting the image corresponding to the i-th frame of audio by the machine learning model according to the generation result of the first region corresponding to the i-th frame of audio, the generation result of the second region corresponding to the i-th frame of audio, and the lower half face mask result of the target image corresponding to the i-th frame of audio.

[0166] For example, in order to better improve efficiency, the above step 33 can be specifically: performing a Blend processing, such as a splicing processing or a Blend processing shown by the Blend module in FIG. 4, on the generation result of the first region corresponding to the i-th frame of audio, the generation result of the second region corresponding to the i-th frame of audio, and the lower half face mask result of the target image corresponding to the i-th frame of audio, to obtain the image corresponding to the i-th frame of audio.

[0167] Based on the related content of the above steps 31 to 33, in some scenarios, the prediction processing of the mouth region and the prediction processing of the non-mouth region in the lower half face can be split into two completely independent processing processes, such as the processing processes shown in FIG. 4, to ensure that the prediction processing of the two regions does not interfere with each other, which is beneficial to improve the generation effect.

[0168] It can be seen that in one possible implementation, for the i-th frame of audio in the above audio sequence, the image corresponding to the i-th frame of audio satisfies the following constraints: the first region in the image corresponding to the i-th frame of audio is generated according to at least two frames of reference images corresponding to the i-th frame of audio in the first region and part or all of the face key points corresponding to the i-th frame of audio, the second region in the image corresponding to the i-th frame of audio is generated according to at least two frames of reference images corresponding to the i-th frame of audio in the second region and part or all of the face key points corresponding to the i-th frame of audio, and the other regions in the image corresponding to the i-th frame of audio except the first region and the second region are generated according to the lower half face mask result of the target image corresponding to the i-th frame of audio.

[0169] Based on the related content of S3 above, in some scenarios, for the i-th frame of audio in the audio sequence, the generation result of the first region corresponding to the i-th frame of audio is determined according to the key points corresponding to the i-th frame of audio in the first region and at least two frames of reference images corresponding to the i-th frame of audio in the first region, and the generation result of the second region corresponding to the i-th frame of audio is determined according to the key points corresponding to the i-th frame of audio in the second region and at least two frames of reference images corresponding to the i-th frame of audio in the second region; then, the generation result of the first region, the generation result of the second region, and the lower half face mask result of the target image corresponding to the i-th frame of audio are spliced to obtain the image corresponding to the i-th frame of audio, i is a positive integer, i ≤ the number of audio frames in the audio sequence.

[0170] S4: generating a video corresponding to the audio sequence according to the images corresponding to each frame of audio in the audio sequence.

[0171] It should be noted that the present application does not limit the implementation of S4 above, for example, it can specifically be: generating a video corresponding to the audio sequence according to the images corresponding to each frame of audio in the audio sequence, so that the video includes the images corresponding to each frame of audio in the audio sequence.

[0172] For example, S4 above can be: generating a video corresponding to the audio sequence according to the audio sequence and the images corresponding to each frame of audio in the audio sequence, so that the video includes the audio sequence and the images corresponding to each frame of audio in the audio sequence.

[0173] For example, S4 above can be implemented by any machine learning model for realizing the generation process, such as a Gan model or a diffusion model. It can be seen that in one possible implementation, S4 can specifically be: performing video generation processing by the machine learning model according to the images corresponding to each frame of audio in the audio sequence to obtain and output the video corresponding to the audio sequence.

[0174] It is found through research that for the images corresponding to each frame of audio in the audio sequence, since these images are generated one by one, the effect presented in the time sequence continuity between these images may be relatively poor, so in order to better improve the generation effect, the present application also provides one possible implementation of S4 above, in which when the audio sequence includes M sub-sequences, M is a positive integer, S4 can specifically include the following steps 41-42.

[0175] Step 41: adjusting the image corresponding to the nth frame of audio in the mth sub-sequence according to the images corresponding to the audios other than the nth frame of audio in the mth sub-sequence, to obtain an adjustment result of the nth frame of audio in the mth sub-sequence, where n is a positive integer, n≤the number of audio frames in the mth sub-sequence, the temporal continuity between the adjustment result of each frame of audio in the mth sub-sequence is higher than the temporal continuity between the images corresponding to the audios in the mth sub-sequence, m is a positive integer, m≤M.

[0176] wherein the audio sequence above includes M sub-sequences; and the M sub-sequences satisfy the following constraint: there is no intersection between any two sub-sequences, and the union of the M sub-sequences includes all audios in the audio sequence.

[0177] The mth sub-sequence refers to a sub-sequence determined from the audio sequence above and arranged in the mth position; and the mth sub-sequence can include multiple frames of audio, such as L frames of audio as shown in FIG. 4, where L is a positive integer.

[0178] In addition, the mth sub-sequence can at least satisfy the following constraints: the arrangement position of each frame of audio in the mth sub-sequence in the audio sequence is earlier than the arrangement position of each frame of audio in the m+1th sub-sequence in the audio sequence, and / or the arrangement position of each frame of audio in the mth sub-sequence in the audio sequence is later than the arrangement position of each frame of audio in the m-1th sub-sequence in the audio sequence.

[0179] In addition, for the mth sub-sequence, the nth frame of audio in the mth sub-sequence refers to the audio arranged in the nth position in the mth sub-sequence; and the adjustment result corresponding to the nth frame of audio is obtained by adjusting the image corresponding to the nth frame of audio in the mth sub-sequence according to the images corresponding to the audios other than the nth frame of audio in the mth sub-sequence, so that the temporal continuity between the adjustment result corresponding to the nth frame of audio and the adjustment result corresponding to the adjacent audio is higher than the temporal continuity between the image corresponding to the nth frame of audio and the image corresponding to the adjacent audio, where n is a positive integer, n≤the number of audio frames in the mth sub-sequence. The temporal continuity is used to describe the content continuity between multiple frames of images arranged close to each other, such as the continuity of background changes, the continuity of expression changes, etc.; and the application does not limit the calculation method of the temporal continuity, which can be implemented by using any method capable of calculating the temporal continuity between multiple frames of images, such as existing or future methods.

[0180] Also, the present application does not limit the implementation of the above step 41, for example, it can be implemented by means of any method that can be optimized in terms of temporal continuity, such as a machine learning model with temporal continuity optimization function.

[0181] Furthermore, in order to better improve the generation effect, the present application also provides a possible implementation of the above step 41, in which the step 41 can be implemented by means of a decoder (Decoder) for realizing temporal continuity optimization processing, such as the decoder shown in FIG. 4 or FIG. 5. Wherein, the decoder is used for performing temporal continuity optimization processing on the input data (such as multi-frame images) of the decoder; and the present application does not limit the implementation of the decoder, for example, it can be implemented by means of Gan model or diffusion model.

[0182] In addition, in order to better improve the optimization effect, the present application provides a possible implementation of the above decoder, in which the decoder can at least meet the following constraints: the input data of the temporal module in the decoder includes multiple images, and the temporal module is used for: integrating the multiple images to obtain overall data, performing self-attention processing on the overall data to obtain self-attention processing result, and splitting the self-attention processing result to obtain the adjustment result of the multiple images.

[0183] Wherein, the temporal module is used for performing temporal continuity optimization processing on the input data of the temporal module; and the temporal module can at least meet the following constraints: the input data of the temporal module includes multiple images; and the temporal module is used for performing temporal continuity optimization processing on the multiple images to obtain the adjustment result of the multiple images.

[0184] In addition, in order to better improve the optimization effect, the present application also provides a possible implementation of the temporal module, in which when the input data of the temporal module includes multiple images, the temporal module can at least meet the following constraints: integrating the multiple images to obtain overall data, performing self-attention processing on the overall data, such as temporal self-attention processing, to obtain self-attention processing result, and splitting the self-attention processing result to obtain the adjustment result of the multiple images.

[0185] It should be noted that the application does not limit the implementation of the integration processing in the above paragraph, for example, it can be implemented by using any existing or future method capable of integrating multiple images. For another example, in order to improve efficiency, the integration processing can be implemented by using reshape. In addition, the application does not limit the implementation of the split processing in the above paragraph, for example, it can be implemented by using any existing or future method capable of splitting one data into multiple images. For another example, in order to improve efficiency, the split processing can be implemented by using reshape.

[0186] Based on the above two paragraphs, in a possible implementation, the timing module can include an integration network, a self-attention network, and a split network. In order to facilitate understanding, the following will introduce these three networks respectively.

[0187] For the integration network described above, the input data of the integration network includes multiple images, so that the integration network is used for integration processing on the multiple images. In addition, the application does not limit the implementation of the integration network, for example, it can be implemented by using reshape. As an example, when the size of the multiple images is [B, L, C, W, H], B represents batch, L represents the number of images in the multiple images, C represents image channel number, W represents image width, and H represents image height, the integration network can be specifically used for reshaping the multiple images to obtain data with a size of [BxWxH, C, L] as the overall data corresponding to the multiple objects.

[0188] For the self-attention network described above, the input data of the self-attention network includes the output data of the integration network, such as the overall data with a size of [BxWxH, C, L], so that the self-attention network is used for self-attention processing on the overall data, such as self-attention processing on L, to obtain a self-attention processing result, so that the size of the self-attention processing result is consistent with the size of the overall data. In addition, the application does not limit the implementation of the self-attention network, for example, it can be implemented by using any self-attention network.

[0189] For the split network described above, the input data of the split network includes the output data of the self-attention network, such as the self-attention processing result with a size of [BxWxH, C, L], so that the split network is used for split processing on the self-attention processing result to obtain a split result, so that the size of the split result is consistent with the size of the input data of the integration network. In addition, the application does not limit the implementation of the split network, for example, it can be implemented by using reshape.

[0190] Based on the related content of the above timing module, in some scenarios, the timing module can be implemented by the timing module shown in FIG. 5, so that the working principle of the timing module can be: first, reshape the input data of the timing module to obtain the overall data; then, perform self-attention processing on the timing of the overall data to obtain the self-attention processing result; then, reshape the self-attention processing result to obtain the output data of the timing module.

[0191] In addition, in order to better improve the generation effect, the present application also provides a possible implementation manner of the above decoder, in which manner, the decoder can include at least one processing unit arranged in sequence, the processing unit includes a processing module and a timing module, and the input data of the timing module includes the output data of the processing module, so that the decoder can better realize the optimization processing of the input data of the decoder by means of multiple processing. For the convenience of understanding, the following will be described in conjunction with examples.

[0192] As an example, as shown in FIG. 5, when the above decoder includes 5 processing units and 1 up-sampling module, if the input data of the decoder includes multiple images, the working principle of the decoder is:

[0193] First, the processing module in the first processing unit performs the first down-sampling processing on the multiple images; then, the timing module in the first processing unit performs the timing continuity optimization processing on the output data of the processing module in the first processing unit;

[0194] Then, the processing module in the second processing unit performs the second down-sampling processing on the output data of the timing module in the first processing unit; then, the timing module in the second processing unit performs the timing continuity optimization processing on the output data of the processing module in the second processing unit;

[0195] Then, the processing module in the third processing unit performs the third down-sampling processing on the output data of the timing module in the second processing unit; then, the timing module in the third processing unit performs the timing continuity optimization processing on the output data of the processing module in the third processing unit;

[0196] Then, the processing module in the fourth processing unit performs the first up-sampling processing on the output data of the timing module in the third processing unit; then, the timing module in the fourth processing unit performs the timing continuity optimization processing on the output data of the processing module in the fourth processing unit;

[0197] Then, the processing module in the fifth processing unit performs a second upsampling processing on the output data of the timing module in the fourth processing unit; then, the timing module in the fifth processing unit performs a timing continuity optimization processing on the output data of the processing module in the fifth processing unit.

[0198] Finally, the upsampling module performs a third upsampling processing on the output data of the timing module in the fifth processing unit, to obtain the output data of the decoder.

[0199] It can be seen that, in a possible implementation, the above decoder can include T down-sampling modules and T upsampling modules arranged in sequence, and a timing module is added behind each of the modules except the last upsampling module, so that the decoder can better realize the timing optimization processing.

[0200] It should be noted that the present application does not limit the implementation of the "T down-sampling modules and T upsampling modules" in the above paragraph, for example, it can be implemented by using a plurality of down-sampling modules and their corresponding upsampling modules involved in any UNet network (such as the UNet network in the Gan model or the UNet network in the diffusion model), so that the "T down-sampling modules and T upsampling modules" have the characteristics of the UNet network itself, such as being presented in the U-shaped manner.

[0201] It should also be noted that the present application does not limit the implementation of the down-sampling module, for example, it can be implemented by using CNN. In addition, the present application also does not limit the implementation of the upsampling module, for example, it can be implemented by using CNN.

[0202] Based on the related content of the above step 41, in some scenarios, for the mth sub-sequence in the above audio sequence, after obtaining the images corresponding to each frame of audio in the mth sub-sequence, the images can be input into the decoder together, so that the decoder can perform self-attention processing on the images as a whole in terms of timing, so that the decoder can adjust the image corresponding to the nth frame of audio in the mth sub-sequence according to the images corresponding to the other audios in the mth sub-sequence except the nth frame of audio, to obtain the adjustment result of the nth frame of audio in the mth sub-sequence, n is a positive integer, n≤the number of audio frames in the mth sub-sequence, and then the timing continuity presented by the adjustment results corresponding to each frame of audio in the mth sub-sequence is higher than the timing continuity presented by the images corresponding to each frame of audio in the mth sub-sequence, so that the timing continuity optimization processing can be realized, m is a positive integer, m≤M.

[0203] Step 42: generating a video corresponding to the audio sequence according to the adjustment result corresponding to each frame of audio in the at least one sub-sequence.

[0204] It should be noted that the present application does not limit the implementation of step 42, for example, it can be specifically: generating a video corresponding to the audio sequence according to the adjustment result corresponding to each frame of audio in the at least one sub-sequence, so that the video includes the adjustment result corresponding to each frame of audio in the sub-sequence.

[0205] Based on the related content of steps 41-42 above, in some scenarios, after obtaining the images corresponding to each frame of audio in the audio sequence, the images can be divided into multiple parts, so that each part includes multiple images arranged in continuous positions, such as the images corresponding to L frames of audio shown in FIG. 4; then each part is processed by a decoder for temporal continuity optimization. Wherein, since the decoder is used for performing self-attention processing on the multiple images input to the decoder as a whole in terms of time, so that the decoder can make full use of the inter-frame relationship between the multiple images, such as the relationship in the image space, the relationship in various image content aspects (such as facial expressions, etc.), etc., so that the decoder can optimize the multiple images by referring to the inter-frame continuity between the multiple images, which is beneficial to generate a video that is continuous in time.

[0206] Based on the related content of S1-S4 above, for the video generation method provided by the embodiments of the present application, first, the target video and the audio sequence are obtained, so that the target video is used to describe the state of the speaker under the first speaking content, and the second speaking content described by the audio sequence is different from the first speaking content; then, according to the target video and the audio sequence, the facial key points corresponding to each frame of audio in the audio sequence are determined, so that the facial key points are used to describe the facial state of the speaker under the corresponding audio; then, according to the facial key points corresponding to the i-th frame of audio in the audio sequence, the at least two frames of reference images corresponding to the first region of the i-th frame of audio, the at least two frames of reference images corresponding to the second region of the i-th frame of audio, and the lower half face mask result of the target image corresponding to the i-th frame of audio, the image corresponding to the i-th frame of audio is generated, i is a positive integer, i≤ the number of audio frames in the audio sequence; then, according to the images corresponding to each frame of audio in the audio sequence, a video corresponding to the audio sequence is generated, so that the video is used to describe the state of the speaker under the second speaking content, such as the lip shape state, etc., so as to meet the demand for modifying the lip shape of the existing video.

[0207] The first region comprises the mouth of the speaker, so that at least two frames of reference images corresponding to the i-th frame of audio under the first region can represent images required for generating the mouth, so that the mouth generated based on the reference images is more suitable for the i-th frame of audio, thereby facilitating improved generation effect.

[0208] In addition, the second region comprises the other region of the lower half of the face of the speaker except the first region, so that at least two frames of reference images corresponding to the i-th frame of audio under the second region can represent images required for generating the other region, so that the other region generated based on the reference images is more suitable for the i-th frame of audio, thereby facilitating improved generation effect.

[0209] In addition, the lower half of the face mask result of the target image corresponding to the i-th frame of audio is used to represent the upper half of the face suitable for the i-th frame of audio, so that the upper half of the face generated based on the lower half of the face mask result is more suitable for the i-th frame of audio, thereby facilitating improved generation effect.

[0210] In addition, the at least two frames of reference images corresponding to the i-th frame of audio under the first region, the at least two frames of reference images corresponding to the i-th frame of audio under the second region, and the target image corresponding to the i-th frame of audio are all determined from the target video, so that the images generated based on the three kinds of images satisfy the following constraints with the corresponding images in the target video: the lip shape is different, and other contents except the lip shape are consistent, so as to better meet the demand for lip shape modification of the existing video.

[0211] Further, the present application does not limit the execution subject of the video generation method provided by the embodiments of the present application. For example, the video generation method provided by the embodiments of the present application can be applied to a terminal device or a server. For another example, the video generation method provided by the embodiments of the present application can also be implemented by means of data interaction process between the terminal device and the server. The terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a stand-alone server, a cluster server or a cloud server.

[0212] It should be noted that the present application does not limit the application scenario of the above-mentioned video generation method. For example, when the video generation method is applied to model training, any existing or future model training method can be used for implementation, and the present application does not make specific limitation thereon.

[0213] Based on the related content of the above-mentioned video generation method, the technical solutions provided by the present application have the advantages shown in the following ①-③.

[0214] ① This application separates the prediction process of the mouth region (e.g., the first region mentioned above) from the prediction process of other regions of the lower face besides the mouth region (e.g., the second region mentioned above) into two independent processes to ensure that the prediction processes of these two regions do not interfere with each other, which is beneficial to improving the generation effect.

[0215] ②This application improves the generation quality of different regions by using different reference image selection strategies for the mouth region and other regions of the lower face other than the mouth region, which is beneficial to improving the generation effect.

[0216] ③This application enhances timing stability by utilizing a timing module, which helps improve the generation effect.

[0217] Based on the video generation method provided in the embodiments of this application, this application also provides a video generation device, which will be explained and described below with reference to FIG6. FIG6 is a schematic diagram of the structure of a video generation device provided in the embodiments of this application. It should be noted that for technical details of the video generation device provided in the embodiments of this application, please refer to the relevant content of the video generation method above.

[0218] As shown in Figure 6, the video generation apparatus 600 provided in this embodiment includes:

[0219] The data acquisition unit 601 is used to acquire a target video and an audio sequence. The target video is used to describe the state of the speaker under the first speech content. The lower half of the speaker's face includes a first region and a second region. The first region includes the mouth. The second region includes other regions of the lower half of the face besides the first region. The audio sequence is used to describe the second speech content. The second speech content is different from the first speech content.

[0220] The data prediction unit 602 is used to predict facial key points corresponding to each frame of audio in the audio sequence based on the target video and the audio sequence.

[0221] The first generation unit 603 is used to generate an image corresponding to the i-th frame of audio based on the facial key points corresponding to the i-th frame of audio in the audio sequence, at least two reference images corresponding to the i-th frame of audio in the first region, at least two reference images corresponding to the i-th frame of audio in the second region, and the lower half face mask result of the target image corresponding to the i-th frame of audio. The target video includes the target image and each of the reference images, where i is a positive integer and i ≤ the number of audio frames in the audio sequence.

[0222] The second generation unit 604 is used to generate a video corresponding to the audio sequence based on the images corresponding to each frame of audio in the audio sequence.

[0223] In a possible implementation, the face key points corresponding to the i-th frame of audio include key points corresponding to the i-th frame of audio in the first region and key points corresponding to the i-th frame of audio in the second region.

[0224] The first generation unit 603 is specifically configured to generate an image corresponding to the i-th frame of audio according to the key points corresponding to the i-th frame of audio in the first region, at least two frames of reference images corresponding to the i-th frame of audio in the first region, the key points corresponding to the i-th frame of audio in the second region, at least two frames of reference images corresponding to the i-th frame of audio in the second region, and a lower half face mask result of a target image corresponding to the i-th frame of audio.

[0225] In a possible implementation, the first generation unit 603 is specifically configured to determine a generation result of the first region corresponding to the i-th frame of audio according to the face key points corresponding to the i-th frame of audio and at least two frames of reference images corresponding to the i-th frame of audio in the first region; determine a generation result of the second region corresponding to the i-th frame of audio according to the face key points corresponding to the i-th frame of audio and at least two frames of reference images corresponding to the i-th frame of audio in the second region; and generate an image corresponding to the i-th frame of audio according to the generation result of the first region corresponding to the i-th frame of audio, the generation result of the second region corresponding to the i-th frame of audio, and a lower half face mask result of a target image corresponding to the i-th frame of audio.

[0226] In a possible implementation, the face key points corresponding to the i-th frame of audio include key points corresponding to the i-th frame of audio in the first region and key points corresponding to the i-th frame of audio in the second region.

[0227] The first generation unit 603 is specifically configured to determine a generation result of the first region corresponding to the i-th frame of audio according to the key points corresponding to the i-th frame of audio in the first region and at least two frames of reference images corresponding to the i-th frame of audio in the first region; and determine a generation result of the second region corresponding to the i-th frame of audio according to the key points corresponding to the i-th frame of audio in the second region and at least two frames of reference images corresponding to the i-th frame of audio in the second region.

[0228] In a possible implementation, the determination of the key points corresponding to the i-th frame of audio in the first region includes: searching for a key point used to describe the first region from the face key points corresponding to the i-th frame of audio; and determining the key points corresponding to the i-th frame of audio in the first region according to the key point used to describe the first region.

[0229] In a possible implementation, the determining of the key point corresponding to the i-th frame of audio under the second region includes: searching for a key point used for describing the second region from the face key points corresponding to the i-th frame of audio; and determining the key point corresponding to the i-th frame of audio under the second region according to the key point used for describing the second region.

[0230] In a possible implementation, the determining of the key point corresponding to the i-th frame of audio under the second region includes: searching for a key point used for describing the lower half of the face from the face key points corresponding to the i-th frame of audio; and determining the key point corresponding to the i-th frame of audio under the second region according to the key point used for describing the lower half of the face.

[0231] In a possible implementation, the generation result of the first region corresponding to the i-th frame of audio is further determined according to key point determination results of each of the at least two frames of reference images corresponding to the i-th frame of audio under the first region; for any one of the at least two frames of reference images corresponding to the i-th frame of audio under the first region, the key point determination result of the reference image is used to describe a state of the first region in the reference image.

[0232] In a possible implementation, the generation result of the second region corresponding to the i-th frame of audio is further determined according to key point determination results of each of the at least two frames of reference images corresponding to the i-th frame of audio under the second region; for any one of the at least two frames of reference images corresponding to the i-th frame of audio under the second region, the key point determination result of the reference image is used to describe a state of the second region in the reference image, or the key point determination result of the reference image is used to describe a state of the lower half of the face in the reference image.

[0233] In a possible implementation, there is a difference between the at least two frames of reference images corresponding to the i-th frame of audio under the first region and the at least two frames of reference images corresponding to the i-th frame of audio under the second region.

[0234] In a possible implementation, the determining of the at least two frames of reference images corresponding to the i-th frame of audio under the first region includes: sorting the images in the target video according to the lip shape amplitude representation data of the images, to obtain an image sequence; performing equal-interval sampling on the image sequence to obtain a sampling image; and determining the at least two frames of reference images corresponding to the i-th frame of audio under the first region according to the sampling image.

[0235] In a possible implementation, for any reference image in the at least two reference images corresponding to the i-th frame of audio under the second region, a distance between a position of the reference image in the target video and a position of a target image corresponding to the i-th frame of audio in the target video is not greater than a preset distance threshold.

[0236] In a possible implementation, the audio sequence includes M sub-sequences, where M is a positive integer.

[0237] The determination of the video corresponding to the audio sequence includes: performing adjustment processing on an image corresponding to an n-th frame of audio in an m-th sub-sequence according to images corresponding to other audio frames in the m-th sub-sequence except the n-th frame of audio, to obtain an adjustment result of the image corresponding to the n-th frame of audio in the m-th sub-sequence, where n is a positive integer and n is less than a number of audio frames in the m-th sub-sequence, a time sequence continuity presented by the adjustment result of each frame of audio in the m-th sub-sequence is higher than a time sequence continuity presented by images corresponding to the audio frames in the m-th sub-sequence, and m is a positive integer and m is less than M; and generating the video corresponding to the audio sequence according to the adjustment result of each frame of audio in the at least one sub-sequence.

[0238] In a possible implementation, the adjustment result of each frame of audio in the m-th sub-sequence is determined by a decoder; input data of a time sequence module in the decoder includes a plurality of images; and the time sequence module is configured to: perform integration processing on the plurality of images to obtain overall data, perform self-attention processing on the overall data to obtain a self-attention processing result, and perform splitting processing on the self-attention processing result to obtain an adjustment result of each image in the plurality of images.

[0239] In a possible implementation, the decoder includes at least one processing unit arranged in sequence, the processing unit includes a processing module and a time sequence module, and input data of the time sequence module includes output data of the processing module.

[0240] Based on the related content of the video generation apparatus 600, the working principle of the video generation apparatus 600 provided in the application is that: first, a target video and an audio sequence are acquired, so that the target video is used to describe the state of a speaker under first speech content, and the second speech content described by the audio sequence is different from the first speech content; then, according to the target video and the audio sequence, face key points corresponding to each frame of audio in the audio sequence are determined, so that the face key points are used to describe the face state of the speaker under the corresponding audio; then, according to the face key points corresponding to the i-th frame of audio in the audio sequence, at least two frames of reference images corresponding to the i-th frame of audio under a first region, at least two frames of reference images corresponding to the i-th frame of audio under a second region, and a lower face mask result of a target image corresponding to the i-th frame of audio, an image corresponding to the i-th frame of audio is generated, i is a positive integer, i≤ the number of audio frames in the audio sequence; then, according to the images corresponding to each frame of audio in the audio sequence, a video corresponding to the audio sequence is generated, so that the video is used to describe the state of the speaker under the second speech content, such as lip shape state, and thus the demand for modifying the lip shape of the existing video can be met.

[0241] In addition, an electronic device is also provided in the embodiment of the application, and the device includes a processor and a memory: the memory is used to store instructions or computer programs; and the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation manner of the video generation method provided in the embodiment of the application.

[0242] Referring to FIG. 7, a structural schematic diagram of an electronic device 700 suitable for implementing the embodiments of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. The electronic device shown in FIG. 7 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.

[0243] As shown in FIG. 7, the electronic device 700 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 702 or loaded into a random access memory (RAM) 703 from a storage device 708. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0244] Generally, the following devices can be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 can allow the electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG. 7 shows the electronic device 700 with various devices, it should be understood that all of the illustrated devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.

[0245] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.

[0246] The electronic device provided by the embodiments of the present disclosure and the method provided by the above-described embodiments belong to the same inventive concept, and technical details not described in detail in the present embodiments can be referred to the above-described embodiments, and the present embodiments have the same beneficial effects as the above-described embodiments.

[0247] The embodiments of the present disclosure also provide a computer readable medium, wherein instructions or a computer program are stored in the computer readable medium, and when the instructions or the computer program are executed on a device, the device is caused to perform any of the embodiments of the video generation method provided by the embodiments of the present disclosure.

[0248] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can send, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.

[0249] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.

[0250] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and is not assembled into the electronic device.

[0251] The computer-readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device can execute the method described above.

[0252] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0253] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0254] The units involved in the embodiments of the present disclosure can be implemented by software, or can be implemented by hardware. In some cases, the name of the unit / module does not constitute a limitation on the unit itself.

[0255] The functions described in the above description herein can be performed at least in part by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0256] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0257] It should be noted that the various embodiments described in the specification are progressive and each embodiment focuses on the differences from other embodiments. The same and similar parts between embodiments can be mutually referred to. For the system or device disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.

[0258] It should be understood that in this application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship between the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0259] It is also to be noted that, as used in the specification and the appended claims, the singular forms "a," "an" and "the" include plural referents unless otherwise indicated. Furthermore, to the extent that the terms "including," "includes," "having," "has," "with," or "contains" are used in either the detailed description and the claims, such terms are intended to be inclusive in a manner similar to the term "comprising" as an open transition term without precluding any additional or other elements.

[0260] The embodiments disclosed herein can each be implemented as a method, apparatus, or article of manufacture using programming instructions. The embodiments disclosed herein can be implemented using software, firmware, hardware, or a combination thereof. The various elements of the disclosed embodiments, as well as the procedural aspects of the disclosed embodiments, can be implemented using a variety of programming instructions, software, firmware, or the like. In addition, one or more of the disclosed embodiments can be implemented by a computer system having a processor and a memory, where the memory stores programming instructions for execution by the processor. The programming instructions can be stored in a variety of ways, including on a computer diskette, on a computer hard drive, on a computer memory, on a computer tape, or on other computer storage devices.

[0261] The above description of disclosed embodiments provides enough information to enable one of ordinary skill in the art to practice the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Accordingly, the application is not to be restricted based on the embodiments shown but is to be accorded the widest scope of the principles and novel features disclosed herein.

Claims

1. A method for generating a video, comprising: obtaining a target video and an audio sequence, the target video being used to describe a state of a speaking object under a first speaking content, a lower half face of the speaking object comprising a first region and a second region, the first region comprising a mouth, and the second region comprising other regions of the lower half face except the first region, the audio sequence being used to describe a second speaking content, the second speaking content being different from the first speaking content; predicting, according to the target video and the audio sequence, facial landmarks corresponding to each frame of audio in the audio sequence; generating, according to facial landmarks corresponding to an i th frame of audio in the audio sequence, at least two reference images corresponding to the first region under the i th frame of audio, at least two reference images corresponding to the second region under the i th frame of audio, and a lower half face mask result of a target image corresponding to the i th frame of audio, an image corresponding to the i th frame of audio, the target video comprising the target image and each of the reference images, i being a positive integer, i ≤ a number of audio frames in the audio sequence; and generating, according to images corresponding to each frame of audio in the audio sequence, a video corresponding to the audio sequence. 2.The method of claim 1, wherein the facial landmarks corresponding to the i th frame of audio comprise facial landmarks corresponding to the first region under the i th frame of audio, and facial landmarks corresponding to the second region under the i th frame of audio; and the generating, according to the facial landmarks corresponding to the i th frame of audio, the at least two reference images corresponding to the first region under the i th frame of audio, the at least two reference images corresponding to the second region under the i th frame of audio, and the lower half face mask result of the target image corresponding to the i th frame of audio, the image corresponding to the i th frame of audio, comprises: generating, according to the facial landmarks corresponding to the first region under the i th frame of audio, the at least two reference images corresponding to the first region under the i th frame of audio, the facial landmarks corresponding to the second region under the i th frame of audio, the at least two reference images corresponding to the second region under the i th frame of audio, and the lower half face mask result of the target image corresponding to the i th frame of audio, the image corresponding to the i th frame of audio. 3.The method of claim 1, wherein the determining the image corresponding to the i th frame of audio comprises: determining, according to the facial landmarks corresponding to the i th frame of audio, and the at least two reference images corresponding to the first region under the i th frame of audio, a generation result of the first region corresponding to the i th frame of audio; determining, according to the facial landmarks corresponding to the i th frame of audio, and the at least two reference images corresponding to the second region under the i th frame of audio, a generation result of the second region corresponding to the i th frame of audio; and generating, according to the generation result of the first region corresponding to the i th frame of audio, the generation result of the second region corresponding to the i th frame of audio, and the lower half face mask result of the target image corresponding to the i th frame of audio, the image corresponding to the i th frame of audio. ​ ​ ​ ​ ​ ​ ​ ​ ​ 4. The method of claim 3, wherein the facial landmarks corresponding to the i-th frame of audio comprise landmarks corresponding to the i-th frame of audio under the first region and landmarks corresponding to the i-th frame of audio under the second region; determining the generation result of the first region corresponding to the i-th frame of audio according to the facial landmarks corresponding to the i-th frame of audio and at least two reference images corresponding to the i-th frame of audio under the first region comprises: determining the generation result of the first region corresponding to the i-th frame of audio according to the landmarks corresponding to the i-th frame of audio under the first region and at least two reference images corresponding to the i-th frame of audio under the first region; determining the generation result of the second region corresponding to the i-th frame of audio according to the facial landmarks corresponding to the i-th frame of audio and at least two reference images corresponding to the i-th frame of audio under the second region comprises: determining the generation result of the second region corresponding to the i-th frame of audio according to the landmarks corresponding to the i-th frame of audio under the second region and at least two reference images corresponding to the i-th frame of audio under the second region.

5. The method of claim 2 or 4, wherein the process of determining the landmarks corresponding to the i-th frame of audio under the first region comprises: finding landmarks used to describe the first region from the facial landmarks corresponding to the i-th frame of audio; and determining the landmarks corresponding to the i-th frame of audio under the first region according to the landmarks used to describe the first region.

6. The method of claim 2 or 4, wherein the process of determining the landmarks corresponding to the i-th frame of audio under the second region comprises: finding landmarks used to describe the second region from the facial landmarks corresponding to the i-th frame of audio, and determining the landmarks corresponding to the i-th frame of audio under the second region according to the landmarks used to describe the second region; or the process of determining the landmarks corresponding to the i-th frame of audio under the second region comprises: finding landmarks used to describe the lower half of the face from the facial landmarks corresponding to the i-th frame of audio, and determining the landmarks corresponding to the i-th frame of audio under the second region according to the landmarks used to describe the lower half of the face.

7. The method of claim 3, wherein the generation result of the first region corresponding to the i-th frame of audio is further determined according to the landmark determination result of each of the at least two reference images corresponding to the i-th frame of audio under the first region; for any one of the at least two reference images corresponding to the i-th frame of audio under the first region, the landmark determination result of the reference image is used to describe the state of the first region in the reference image.

8. The method of claim 3, wherein the generation result of the second region corresponding to the i-th frame of audio is further determined according to the landmark determination result of each of the at least two reference images corresponding to the i-th frame of audio under the second region; ​ For any one of the at least two reference images corresponding to the i-th frame of audio under the second region, a key point determination result of the reference image is used to describe a state of the second region in the reference image, or a key point determination result of the reference image is used to describe a state of the lower half face in the reference image.

9. The method of claim 1, wherein there is a difference between the at least two reference images corresponding to the i-th frame of audio under the first region and the at least two reference images corresponding to the i-th frame of audio under the second region.

10. The method of claim 1, wherein the determination process of the at least two reference images corresponding to the i-th frame of audio under the first region comprises: sorting images in the target video according to mouth shape amplitude representation data of the images to obtain an image sequence; equidistantly sampling the image sequence to obtain sampled images; determining the at least two reference images corresponding to the i-th frame of audio under the first region according to the sampled images.

11. The method of claim 1, wherein for any one of the at least two reference images corresponding to the i-th frame of audio under the second region, a distance between an arrangement position of the reference image in the target video and an arrangement position of a target image corresponding to the i-th frame of audio in the target video is not greater than a preset distance threshold.

12. The method of claim 1, wherein the audio sequence comprises M sub-sequences, M being a positive integer; the determination process of the video corresponding to the audio sequence comprises: adjusting a target image corresponding to an n-th frame of audio in an m-th sub-sequence according to images corresponding to other audio frames in the m-th sub-sequence except the n-th frame of audio to obtain an adjustment result of the target image, n being a positive integer, n≤ a number of audio frames in the m-th sub-sequence, a time sequence continuity presented by the adjustment results of the audio frames in the m-th sub-sequence being higher than a time sequence continuity presented by the images corresponding to the audio frames in the m-th sub-sequence, m being a positive integer, m≤M; generating the video corresponding to the audio sequence according to the adjustment results of the audio frames in the at least one sub-sequence.

13. The method of claim 12, wherein the adjustment results of the audio frames in the m-th sub-sequence are determined by a decoder; input data of a time sequence module in the decoder comprises a plurality of images; the time sequence module is configured to: integrate the plurality of images to obtain overall data, perform self-attention processing on the overall data to obtain a self-attention processing result, and split the self-attention processing result to obtain the adjustment results of the images in the plurality of images.

14. The method of claim 13, wherein the decoder comprises at least one processing unit arranged in sequence, the processing unit comprises a processing module and a time sequence module, and input data of the time sequence module comprises output data of the processing module.

15. A video generation apparatus, comprising: The data acquisition unit is configured to acquire a target video and an audio sequence, the target video being used to describe a state of a speaking object under first speaking content, a lower half face of the speaking object including a first region and a second region, the first region including a mouth, and the second region including other regions of the lower half face except the first region, and the audio sequence being used to describe second speaking content, the second speaking content being different from the first speaking content. The data prediction unit is configured to predict, according to the target video and the audio sequence, face key points corresponding to each frame of audio in the audio sequence. The first generation unit is configured to generate an image corresponding to an i-th frame of audio in the audio sequence according to the face key points corresponding to the i-th frame of audio, at least two reference images corresponding to the i-th frame of audio in the first region, at least two reference images corresponding to the i-th frame of audio in the second region, and a lower half face mask result of a target image corresponding to the i-th frame of audio, the target video including the target image and each reference image, i being a positive integer, i≤ the number of audio frames in the audio sequence. The second generation unit is configured to generate a video corresponding to the audio sequence according to the images corresponding to each frame of audio in the audio sequence.

16. An electronic device, the device comprising: A processor and a memory; The memory is configured to store instructions or a computer program. The processor is configured to execute the instructions or the computer program in the memory, so that the electronic device executes the method in any one of claims 1-14.

17. A computer readable medium having stored therein instructions or a computer program, which when executed on a device, cause the device to execute the method in any one of claims 1-14.

18. A computer program product, characterised in that, It includes a computer program carried on a non-transitory computer readable medium, the computer program containing program codes for executing the method in any one of claims 1-14.

Citation Information

Patent Citations

  • Digital human video generation method and device, electronic equipment and storage medium

    CN113886643A

  • Training method and device of image generation model and image generation method and device

    CN116797877A

  • Video generation method and device and electronic equipment

    CN117880606A

  • Video generation method and device, computer equipment and storage medium

    CN117939253A

  • Electric vehicle height restriction opening / closing facility

    KR102632687B1