Video generation method, model training method, device and computer program product

By obtaining the target audio and reference images, combining global visual features and lip features, and using the audio alignment model to generate video frames synchronized with the audio, the problem of poor synchronization between lip movements and audio in the existing technology is solved, achieving more vivid and natural video generation and improving the user experience.

CN120658921APending Publication Date: 2025-09-16BEIJING AUTONAVI YUNMAP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510538730.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In videos generated by existing technologies, the lip movements of characters and the audio are poorly synchronized, resulting in a lack of naturalness in the video content and a poor visual experience for users.

Method used

By obtaining the target audio and reference images, the global visual features and lip features of the audio clip are determined, and the pre-trained audio alignment model is used to generate video frames synchronized with the audio. By combining facial expressions, body movements and background features, more vivid and natural videos are generated.

Benefits of technology

It improves the naturalness of character expressions in the video and the synchronization of lip movements with audio, enhancing the user's visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120658921A_ABST
    Figure CN120658921A_ABST
Patent Text Reader

Abstract

The invention discloses a video generation method and device, a model training method and device and a computer program product, and the video generation method comprises the steps: obtaining a target audio and a reference picture used for generating a video, and the reference picture comprises a sound production object; determining a global visual feature of each to-be-generated video frame corresponding to the audio clip according to the clip features of one or more audio clips corresponding to the target audio and the reference image; according to the pronunciation feature of each audio frame of the target audio and the lip feature of the sounding object in the reference picture, determining the lip feature of the sounding object in a to-be-generated video frame corresponding to the audio frame; and generating each video frame according to the lip feature and the global visual feature corresponding to the to-be-generated video frame. Through the scheme provided by the invention, the expression of characters in the generated video is more vivid and natural, the lip action and the audio can be accurately synchronized, and the visual experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a video generation method, a model training method, a device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the rapid development of computer and artificial intelligence technologies, video synthesis technology has been increasingly used in various scenarios. Video synthesis usually involves first obtaining a piece of audio and then generating a corresponding video based on the audio. The generated video can be used to emit the audio through a specific character. Video synthesis can be applied to various scenarios such as voice broadcasting, film and television clip generation, animation generation, and user interaction.

[0003] When generating a video corresponding to audio, the related technology usually determines the lip movements of the character in each frame based on the pronunciation characteristics of the audio, thereby generating the corresponding video.

[0004] However, the videos generated by the lip and audio matching scheme in the related art usually show the characters opening their mouths stiffly to express the audio content. The generated video content lacks naturalness and provides a poor visual experience to users. Summary of the Invention

[0005] This application provides a video generation method, model training method, device, electronic device, computer-readable storage medium, and computer program product that can not only make the expressions of characters in the generated videos more vivid and natural, but also accurately synchronize lip movements with audio, thereby improving the user's visual experience. The specific solution is as follows:

[0006] In a first aspect, the present application provides a video generation method, the method comprising:

[0007] Obtaining target audio and a reference picture for generating a video, wherein the reference picture includes a sound-emitting object;

[0008] determining, based on the segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment, the global visual features being used to indicate overall visual features of the video frame based on the reference image;

[0009] Determining, based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame;

[0010] According to the lip features and the global visual features corresponding to the video frame to be generated, each video frame synchronized with the target audio and based on the reference picture and produced by the sound object is generated.

[0011] Optionally, before determining the global visual features of each to-be-generated video frame corresponding to the audio segment based on the segment features of one or more audio segments corresponding to the target audio and the reference image, the method further includes:

[0012] Acquiring an action intensity parameter for controlling the action intensity of the sound-emitting object;

[0013] The determining, based on the segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment includes:

[0014] Determining, based on the segment features of each audio segment of the target audio and the reference image and the motion intensity parameter, global visual features of each to-be-generated video frame corresponding to the audio segment;

[0015] The determining, based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame includes:

[0016] According to the pronunciation features of each audio frame of the target audio, the action intensity parameter and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the video frame to be generated corresponding to the audio frame are determined.

[0017] Optionally, determining, based on the segment features of each audio segment of the target audio and the reference image and the motion intensity parameter, global visual features of each to-be-generated video frame corresponding to the audio segment includes:

[0018] Determining the difference in position change of key points of the sound-emitting object between the video frames to be generated according to the action intensity parameter;

[0019] Determining, based on the segment features of each audio segment of the target audio and the difference in position changes between the reference image and the key points, global visual features of each to-be-generated video frame corresponding to the audio segment;

[0020] The determining, based on the pronunciation features of each audio frame of the target audio, the action intensity parameter, and the lip features of the sound-making object in the reference picture, of the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame includes:

[0021] According to the pronunciation features of each audio frame of the target audio, the difference in the position changes of the key points, and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the video frame to be generated corresponding to the audio frame are determined.

[0022] Optionally, before generating, based on the lip features and the global visual features corresponding to the video frames to be generated, the video frames of the sound produced by the sound-producing object that are synchronized with the target audio and based on the reference picture, the method further includes:

[0023] Extracting facial features of the sound-making object from the reference image;

[0024] The step of generating, based on the lip features and the global visual features corresponding to the video frame to be generated, video frames synchronized with the target audio and produced by the sound object and based on the reference picture, includes:

[0025] According to the lip features and the global visual features corresponding to the video frame to be generated, and based on the facial features of the sound-emitting object, each video frame synchronized with the target audio and produced by the sound-emitting object based on the reference picture is generated.

[0026] Optionally, before determining the global visual features of each to-be-generated video frame corresponding to the audio segment based on the segment features of one or more audio segments corresponding to the target audio and the reference image, the method further includes:

[0027] Obtaining a description text for describing the content displayed by the reference image;

[0028] The determining, based on the segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment includes:

[0029] determining, based on the segment features of one or more audio segments corresponding to the target audio, the reference image, and the description text, global visual features of each to-be-generated video frame corresponding to the audio segment;

[0030] The determining, based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame includes:

[0031] According to the pronunciation features of each audio frame of the target audio, the reference picture and the description text, the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame are determined.

[0032] Optionally, the video generation method generates each video frame based on a pre-trained video generation model, wherein the video generation model includes an audio alignment model;

[0033] The global visual features and the lip features are determined in the following manner:

[0034] Determining first input information of the audio alignment model according to the target audio and the reference image;

[0035] The first input information is input into the audio alignment model so that the audio alignment model outputs the global visual features of each to-be-generated video frame corresponding to the audio segment based on the segment features of one or more audio segments corresponding to the target audio and the reference image, and outputs the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference image.

[0036] Optionally, the segment features include: at least one of semantic features, tone features, emotional features, background sound features, pitch features, sound intensity features, pause features, and speech speed features;

[0037] The global visual features include at least one of facial expression features, limb movement features, background features, and lip features.

[0038] In a second aspect, the present application provides a method for training a video generation model, wherein the video generation model includes an audio alignment model, and the method includes:

[0039] Acquire a training sample, where the training sample includes a sample video and a sample audio that matches the sample video;

[0040] Determining a correspondence between each sample video segment of the sample video and each sample audio segment of the sample audio;

[0041] Determining first sample input information based on the corresponding sample video clips and sample audio clips, and inputting the first sample input information into the audio alignment model to be trained, so that the audio alignment model to be trained learns the association between the global visual features of the sample video clips and the corresponding sample audio clips, thereby obtaining a first-stage audio alignment model;

[0042] Determining second sample input information based on the sample video frames and sample audio frames corresponding to the sample video and the sample audio, and inputting the second sample input information into the first-stage audio alignment model so that the first-stage audio alignment model learns the association between the lip features of the sample video frames and the corresponding sample audio frames, thereby obtaining a second-stage audio alignment model;

[0043] The trained audio alignment model is obtained based on the second-stage audio alignment model.

[0044] In a third aspect, the present application provides a video generation device, comprising:

[0045] an acquisition unit, configured to acquire target audio and a reference picture for generating a video, wherein the reference picture includes a sound-emitting object;

[0046] a determination unit configured to determine, based on segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment, the global visual features being used to indicate overall visual features of the video frame based on the reference image; and to determine, based on pronunciation features of each audio frame of the target audio and lip features of the sound-making subject in the to-be-generated video frame corresponding to the audio frame, lip features of the sound-making subject in the reference image;

[0047] A generating unit is configured to generate, based on the lip features and the global visual features corresponding to the video frames to be generated, video frames synchronized with the target audio and produced by the sound-producing object and based on the reference picture.

[0048] In a fourth aspect, the present application also provides an electronic device comprising: a processor, a memory, and computer program instructions stored on the memory and executable on the processor; when the processor executes the computer program instructions, the method as described in any one of the first to second aspects is implemented.

[0049] In a fifth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement any one of the methods in the first to second aspects.

[0050] In a sixth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method as described in any one of the first to second aspects.

[0051] Compared with the prior art, this application has the following advantages:

[0052] The video generation method provided in an embodiment of the present application obtains target audio and a reference image for generating a video, wherein the reference image includes a sound-emitting object and provides a basis for generating the video to be generated; then, based on the segment features of one or more audio segments corresponding to the target audio and the reference image, the global visual features of each video frame to be generated corresponding to the audio segment are determined, wherein the global visual features are used to indicate the overall visual features of the video frame based on the reference image. That is, the present application determines the global visual features of each video frame that is consistent with the entire audio segment based on each audio segment of the target audio. Since the segment features of an audio segment can usually reflect the semantics, tone, attitude and other contextual content of the audio segment, when generating a video, different audio contexts usually correspond to different visual features such as expressions, body movements, and background changes. The global visual features of each video frame determined by the present application based on the segment features of the audio segment can match the context reflected by the audio segment, so that the video subsequently generated based on the global visual features can reflect the overall context of the target audio as a whole, thereby improving the overall naturalness of the generated video. The present application also determines the lip features of the sound-making object in the video frame to be generated corresponding to the audio frame based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference picture. The determined lip features can accurately match the pronunciation lip shape of each audio frame of the target audio, so that the subsequent video frames determined according to the lip features can be more accurately synchronized with the pronunciation of the target audio; and then, based on the lip features corresponding to the video frame to be generated and the global visual features, each video frame that is synchronized with the target audio and is pronounced by the sound-making object based on the reference picture can be generated.

[0053] It can be seen that the solution provided in the present application generates a video based on the target audio and reference image, which can not only accurately synchronize the lip movements of the sound-producing object with the pronunciation of the target audio, but also make the generated video match the overall context of the target audio in terms of non-strictly time-aligned visual elements such as facial expressions, body movements or background changes, rather than rigidly expressing the video content only through the lips, making the expression of the characters in the generated video more vivid and natural, thereby improving the user's visual experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is a schematic diagram of an application scenario of the video generation method provided by this application;

[0055] Figure 2 This is a flowchart of an example of a video generation method provided in an embodiment of the present application;

[0056] Figure 3This is a flow chart of an example of generating a video using a video generation model provided in an embodiment of the present application;

[0057] Figure 4 Schematic diagram of training the first-stage audio alignment model in an embodiment of the present application;

[0058] Figure 5 Schematic diagram of training the second-stage audio alignment model in an embodiment of the present application;

[0059] Figure 6 This is a flowchart of an example of a video generation model training process according to an embodiment of the present application;

[0060] Figure 7 This is a structural block diagram of the electronic device provided in this application. DETAILED DESCRIPTION

[0061] In order to enable those skilled in the art to better understand the technical solutions of this application, the following clearly and completely describes this application in conjunction with the drawings in the embodiments of this application. However, this application can be implemented in many other ways different from the following description. Therefore, based on the embodiments provided in this application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of this application.

[0062] It should be noted that the terms "first", "source domain", "third", etc. in the claims, description and drawings of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. The data used in this way are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including", "having" and their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0063] In order to facilitate understanding of the various embodiments of the present application, the application background of the embodiments is described.

[0064] With the rapid development of computer and artificial intelligence technologies, video synthesis technology has been increasingly used in various scenarios. Video synthesis usually involves first obtaining a segment of audio and then generating a corresponding video based on the audio. The generated video can be used to produce the audio through a specific character. Video synthesis can be applied to various scenarios such as voice broadcasting, film and television drama clip generation, animation generation, and user interaction. When generating a video corresponding to audio, the related technology usually determines the lip movements of the character in each frame based on the pronunciation characteristics of the audio, thereby generating the corresponding video. However, the videos generated by this lip and audio matching scheme in the related technology usually show that the character opens his mouth stiffly to express the audio content, and the generated video content lacks naturalness, which provides a poor visual experience for users.

[0065] To address the above issues, the present invention provides a video generation method, model training method, apparatus, electronic device, computer-readable storage medium, and computer program product. These methods aim to make the expressions of characters in the generated videos more vivid and natural, while also accurately synchronizing lip movements with audio, thereby enhancing the user's visual experience.

[0066] The video generation method provided in this application can be applied to video generation in various professional fields. Specifically, it can be applied to voice broadcasting, film and television clip generation, animation generation, video creation, user interaction, virtual assistants, etc., but is not limited to these.

[0067] In order to facilitate understanding of the method embodiment of this application, its application scenario is introduced. Figure 1 , Figure 1 The following is a schematic diagram of an application scenario of the solution provided in the embodiment of the present application. This application scenario is a schematic illustration and is not intended to be a specific description of its application scenario. Figure 1 As shown, in this application scenario, a server 102 and a client 101 are provided. In this embodiment, a connection is established between the client 101 and the server 102 via network communication, thereby performing data transmission.

[0068] The client 101 can be an electronic device with display and data processing functions, such as a mobile phone, a tablet computer (pad), a smart watch, a desktop computer, a smart TV, a VR device, a vehicle-mounted device, a wearable device, a laptop computer, etc. The client 101 is used to obtain the target audio and reference image input by the user, and send the target audio and reference image to the server 102, so that the server 102 generates each video frame that is synchronized with the target audio and is sounded by the sound-emitting object based on the reference image, and obtains each video frame from the server 102 to display the generated video to the user. The client 101 can also be used to send access requests, interactive information, etc. to the server 102, so that the server 102 sends the corresponding request data to the client 101 for display.

[0069] The server 102 has high computing power. The server 102 can be a server, which has high-speed processor (central processing unit, CPU) computing power, long-term reliable operation, powerful input / output (input / output, I / O) external data throughput capacity and better scalability. The server 102 can be a single server or a server cluster. The server 102 is used to obtain the target audio and reference picture input by the user from the client 101, and generate each video frame synchronized with the target audio and based on the reference picture through the sound object, and send the generated video frames to the client 101. The server 102 can also provide other specific services for the client 101, such as user information access, website access, application access, etc., which are not specifically limited in this application.

[0070] The client 101 and the server 102 may communicate with each other using various communication systems, such as a wired communication system or a wireless communication system. Examples of the wireless communication system include a global system for mobile communications (GSM), a code division multiple access (CDMA), a wideband code division multiple access (WCDMA), a general packet radio service (GPRS), a long term evolution (LTE), an LTE frequency division duplex (FDD), an LTE time division duplex (TDD), a universal mobile telecommunication system (UMTS), a worldwide interoperability for microwave access (WiMAX), a future fifth generation (5G) system or new radio (NR), a satellite communication system, and the like.

[0071] In other application scenarios, only a client may be provided, and the user may select target audio and reference images on the client, and the client may generate a video based on the target audio and reference images. The solution provided in this application may also be applied to other application scenarios, and is not specifically limited in this application.

[0072] Example 1

[0073] The first embodiment of the present application provides a video generation method, which can be applied to electronic devices, such as servers, desktop computers, laptops, mobile phones, tablet computers (pads), smart watches, smart TVs, VR devices, vehicle-mounted devices, wearable devices, and other electronic devices with data processing functions.

[0074] like Figure 2 As shown, the video generation method provided in the first embodiment of the present application includes the following steps S110 to S140.

[0075] Step S110: Obtain target audio and reference pictures for generating a video, wherein the reference pictures include a sound-emitting object.

[0076] The target audio may be audio selected by the user or input by the user, or may be audio downloaded from the Internet by the electronic device, or may be audio selected from pre-stored audio by the electronic device, etc. The target audio may include audio of speech or singing, etc., and this application does not specifically limit this. It is understood that the target audio includes voice information such as speech and singing, rather than audio that does not include voice information, such as light music.

[0077] The reference image can be an image selected or input by the user, or an image downloaded from the Internet or selected from pre-stored images by the electronic device. The reference image can be a photograph, a composite image, a video screenshot, an image of a cartoon character, etc. The sounding object included in the reference image can be a person, a cartoon character, an animal, a virtual character in a game, etc., but is not limited to these. Those skilled in the art can flexibly set the specific form of the target audio and reference image, and this application does not specifically limit it.

[0078] Step S120: Determine the global visual features of each to-be-generated video frame corresponding to the audio segment based on the segment features of one or more audio segments corresponding to the target audio and the reference image. The global visual features are used to indicate the overall visual features of the video frame based on the reference image.

[0079] In step S120, the target audio can be divided into one or more audio segments. Specifically, the target audio can be divided into multiple audio segments according to a preset division unit. The preset division unit can be any division unit from 5 frames to 20 frames, or other more or fewer division units. This application does not specifically limit it. When the division unit is relatively long, it can better capture the overall characteristics of the longer audio segment, so that the determined global visual features can better reflect the contextual characteristics of the audio as a whole. When the division unit is relatively short, it can more accurately reflect the characteristics of each audio segment, so that the determined global visual features can more accurately correspond to the audio segment of each small unit. Optionally, the target audio can also be divided into only one audio segment, and the division unit is the entire target audio, that is, the target audio is not divided. Those skilled in the art can set the length of the audio segment according to actual needs.

[0080] The segment features of an audio clip may include at least one of the semantic features, tone features, emotional features, background sound features, pitch features, sound intensity features, pause features, and speech speed features of the audio clip, but are not limited thereto. These audio clip features can reflect the overall visual features of the video image as a whole.

[0081] The above-mentioned global visual features may include at least one of expression features, limb movement features, background features, and lip features, but are not limited thereto. These visual features can indicate the picture style of the video frame as a whole.

[0082] Since the audio segment is composed of multiple audio frames, the video frames to be generated corresponding to the audio segment, that is, the audio frames included in the audio segment respectively correspond to the video frames to be generated, that is, each audio frame corresponds to a video frame to be generated.

[0083] Exemplarily, in step S120, the global visual features of each to-be-generated video frame corresponding to the audio segment may be determined specifically according to steps 1 to 4 below.

[0084] Step 1: Process the audio segment to extract segment features.

[0085] For example, the segment features may be extracted from the audio segment by using a pre-trained audio feature extraction module.

[0086] Step 2: Analyze the reference image to obtain visual element features of the reference image.

[0087] The visual elements of the reference image may include, but are not limited to, color distribution, object shape, background setting, and person appearance. Specifically, the visual element features of the reference image may be extracted using a pre-trained image processing model such as a convolutional neural network.

[0088] Step 3: Combine the segment features extracted from the audio segment with the visual element features obtained from the reference image to obtain a combined feature.

[0089] Specifically, multimodal fusion can be used to combine features. For example, the visual element features and the fragment features can be vector-splicing to obtain a spliced ​​feature vector, and the spliced ​​vector can be determined as the combined feature. Alternatively, a cross-attention algorithm or a cross-modal transformation algorithm can be used to combine the visual element features and the fragment features to obtain a combined feature, so that the audio features can effectively guide the generation of visual features.

[0090] Step 4: Based on the above combined features, use a generative adversarial network, variational autoencoder, or other pre-trained generative model to predict and generate global visual features for each video frame to be generated.

[0091] Step S130: Determine the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame according to the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference image.

[0092] Since the global visual features determined in step S120 are the overall visual features of the video frame, the features of the lip details may not be accurate enough. Therefore, the lip features determined in step S130 through the pronunciation features of each audio frame can more precisely synchronize the lip shape of the object with the pronunciation of the target audio.

[0093] The pronunciation features of the audio frame may include phonemes, pronunciation position and method, pitch, timbre and other pronunciation-related features.

[0094] Exemplarily, in step S130 , the following steps a to c may be followed to determine the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame.

[0095] Step a: Extract the pronunciation features of each audio frame from the target audio.

[0096] For example, speech processing techniques in related technologies (such as Mel-frequency cepstral coefficients, linear predictive coding, etc.) can be used to extract pronunciation features of each audio frame from the target audio. The pronunciation features include but are not limited to phoneme category, fundamental frequency, resonance peak position, energy distribution, etc.

[0097] Step b: Identify the lip features in the reference image.

[0098] Specifically, a pre-trained image recognition model can be used to identify and extract lip features of the speaker in a reference image. The image recognition model can be a facial landmark detection model or a lip region recognition model. Lip features may include, but are not limited to, lip shape (such as opening size, width, and degree of lip opening), color, texture, vermilion border, and the relative position of the upper and lower lips.

[0099] Step c: Determine the lip features corresponding to the pronunciation features of the audio frame through a pre-trained lip feature mapping model.

[0100] The lip feature mapping model can be trained using a supervised training algorithm, an unsupervised training algorithm, or other training algorithms in related technologies. The lip feature mapping model can be a deep neural network, a convolutional neural network, a recurrent neural network, or other models.

[0101] In one embodiment, the video generation method can generate each video frame based on a pre-trained video generation model, such as Figure 3As shown, the video generation model includes an audio alignment model, which is used to output the global visual features of each to-be-generated video frame corresponding to the audio segment based on the segment features of one or more audio segments corresponding to the target audio and the reference image, and output the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference image. In other words, the audio alignment model can be used to obtain both the above-mentioned global visual features and the lip features of the sound-making object in the to-be-generated video frame. The training of the video generation model will be described in detail later. In this case, in one embodiment, the global visual features in step S120 and the lip features of the sound-making object in the to-be-generated video frame in step S130 can be determined according to the following steps S10 to S20.

[0102] Step S10: Determine first input information of the audio alignment model according to the target audio and the reference image.

[0103] Specifically, the target audio and reference image can be determined as the first input information of the audio alignment model, or the target audio and reference image can be further processed or other information can be added to obtain the first input information, so that the audio alignment model can perform data inference more efficiently and accurately.

[0104] Exemplarily, the target audio can be divided into one or more audio segments, and the segment features of the audio segments are extracted, and the pronunciation features of each audio frame of the target audio are extracted. Based on the segment features of the audio segments, the pronunciation features of each audio frame, and the reference image, the first input information of the audio alignment model is determined. Specifically, the segment features of the audio segments, the pronunciation features of each audio frame, and the reference image can be determined as the first input information of the audio alignment model, or the segment features of the audio segments, the pronunciation features of each audio frame, and the reference image can be further processed, or other information can be added to obtain the first input information of the audio alignment model.

[0105] In one embodiment, the video generation model may be a diffusion model. Since diffusion models can be used to generate high-resolution and realistic images and are easier to train, when the video generation model is a diffusion model, the resolution and realism of the generated video frames can be improved, and the model training process can also be made easier. The audio alignment model can be understood as a denoising layer of the video generation model. The video generation model may also be other models, such as a neural network model, a generative adversarial network, a variational autoencoder, etc., which are not specifically limited in this application.

[0106] Step S10 may determine the first input information of the audio alignment model according to the following steps S11 to S12.

[0107] Step S11: determining a first number of padding frames according to the number of target audio frames, where the first number is 1 less than the number of target audio frames or is the same as the number of target audio frames.

[0108] Exemplarily, if the number of frames of the target audio is 100, the number of filler frames is 99, or 100. The filler frame can be a frame of any picture, for example, the filler frame can be a solid color frame, a frame with the same content as the reference picture, a random picture frame, etc. The filler frame is used to enable the video generation model to restore the content in the filler frame to a video frame that meets the requirements based on the target audio and the reference image. When the number of filler frames is 1 less than the number of frames of the target audio, the reference image can be used as the first frame of the video frame to be generated. When the number of filler frames is the same as the number of frames of the target audio, the video generation model can restore the content in each filler frame to each video frame that meets the requirements based on the target audio and the reference image. Those skilled in the art can set it according to actual needs, wherein using the reference image as the first frame of the video frame to be generated can more efficiently obtain each video frame. Therefore, the following example uses the reference image as the first frame of the video to be generated as an example for explanation.

[0109] In practical applications, the sampling frequencies of audio and video may be different. This application can determine the video frame corresponding to the audio frame of the target audio according to the sampling frequencies of audio and video. For example, if four sampling units of audio are one audio frame and five sampling units of video are one video frame, then the frame correspondence between audio and video can be determined according to the different sampling frequencies of audio and video, such as Figure 5 As shown, an audio frame corresponding to four sampling units in the audio corresponds to a video frame corresponding to five sampling units in the video. The number of target audio frames can also be determined according to the number of sampling units included in a single audio frame.

[0110] Step S12: determining first input information of the audio alignment model according to an image sequence formed by the reference image and each of the padded frames and the target audio.

[0111] Specifically, the image sequence formed by the reference image and each of the filling frames and the target audio can be determined as the first input information of the audio alignment model, or the image sequence and the target audio can be further processed or other data can be added to obtain the first input information.

[0112] Since the diffusion model is a process of obtaining an image that meets the requirements by adding noise to the image and then gradually denoising it, this embodiment enables the audio alignment model to perform denoising and recovery inference based on the padded image by setting a padding frame, thereby obtaining accurate video frames after denoising.

[0113] In a specific embodiment, step S12 may determine the first input information of the audio alignment model according to the following steps S12a to S12c.

[0114] Step S12a: Feature encoding is performed on the target audio to obtain audio encoding features, where the audio encoding features include segment features of each audio segment of the target audio and pronunciation features of each audio frame of the target audio.

[0115] Specifically, such as Figure 3 As shown, the target audio can be encoded into audio coding features using a pre-trained audio coding model. Specifically, the target audio can be input into the audio coding model to obtain the audio coding features. The audio coding model is used to encode the target audio to obtain segment features of each audio segment of the target audio and pronunciation features of each audio frame of the target audio. The video generation model includes the audio coding model.

[0116] Step S12b: converting the audio coding features into converted audio coding features having the same dimension as that corresponding to the audio alignment model.

[0117] Since different models usually have different corresponding data dimensions when performing data processing, the audio coding features can be converted into converted audio coding features with the same dimension as the dimension corresponding to the audio alignment model, so that the audio alignment model can process the converted audio coding features.

[0118] Specifically, such as Figure 3 As shown, the audio coding features can be converted into converted audio coding features with the same dimension as the dimension corresponding to the audio alignment model through a pre-trained first dimension alignment model. Specifically, the audio coding features can be input into the first dimension alignment model to obtain converted audio coding features with the same dimension as the dimension corresponding to the audio alignment model.

[0119] Step S12c: determining first input information of the audio alignment model according to an image sequence formed by the reference image and each of the padded frames and the converted audio coding features.

[0120] Specifically, the above-mentioned image sequence and the converted audio coding features can be determined as the first input information of the audio alignment model. The image sequence can also be further processed or other information can be added to generate the first input information of the audio alignment model, so that the process of generating video frames can refer to more information, or it can be more convenient to generate images and improve the accuracy of the generated video frames. The relevant content of generating the first input information after adding other information will be introduced later.

[0121] Step S20: Input the first input information into the audio alignment model, so that the audio alignment model outputs the global visual features of each to-be-generated video frame corresponding to the audio segment based on the segment features of one or more audio segments corresponding to the target audio and the reference image, and outputs the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference image.

[0122] This embodiment uses a pre-trained audio alignment model to efficiently and accurately obtain the global visual features of each video frame to be generated and the lip features of the sound-making object in the video frame to be generated, making video generation more efficient and accurate.

[0123] Step S140: Generate video frames synchronized with the target audio and produced by the sound-producing object based on the reference image according to the lip features and the global visual features corresponding to the video frame to be generated.

[0124] Specifically, the lip features and the global visual features corresponding to the video frame to be generated can be combined to obtain a video frame with fused features, and each video frame with fused features can be determined as a video frame synchronized with the target audio and produced by the sound-producing subject based on the reference image. Specifically, the lip features and the global visual features corresponding to the video frame to be generated can be input into a pre-trained feature fusion model to obtain a video frame with fused features.

[0125] Since the determined global visual features and lip features correspond to the target audio and are based on the reference image, the lip features of the object when speaking in each generated video frame can be synchronously corresponded to the pronunciation features of each frame of the target audio, and the global visual features of each video frame can also be consistent with the overall context of each audio clip of the target audio, and the generated video frames are consistent with the overall content of the reference image.

[0126] The video generation method provided in an embodiment of the present application obtains target audio and a reference image for generating a video, wherein the reference image includes a sound-emitting object and provides a basis for generating the video to be generated; then, based on the segment features of one or more audio segments corresponding to the target audio and the reference image, the global visual features of each video frame to be generated corresponding to the audio segment are determined, wherein the global visual features are used to indicate the overall visual features of the video frame based on the reference image. That is, the present application determines the global visual features of each video frame that is consistent with the entire audio segment based on each audio segment of the target audio. Since the segment features of an audio segment can usually reflect the semantics, tone, attitude and other contextual content of the audio segment, when generating a video, different audio contexts usually correspond to different visual features such as expressions, body movements, and background changes. The global visual features of each video frame determined by the present application based on the segment features of the audio segment can match the context reflected by the audio segment, so that the video subsequently generated based on the global visual features can reflect the overall context of the target audio as a whole, thereby improving the overall naturalness of the generated video. The present application also determines the lip features of the sound-making object in the video frame to be generated corresponding to the audio frame based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference picture. The determined lip features can accurately match the pronunciation lip shape of each audio frame of the target audio, so that the subsequent video frames determined according to the lip features can be more accurately synchronized with the pronunciation of the target audio; and then, based on the lip features corresponding to the video frame to be generated and the global visual features, each video frame that is synchronized with the target audio and is pronounced by the sound-making object based on the reference picture can be generated.

[0127] It can be seen that the solution provided by the present application generates a video based on the target audio and reference image, which can not only accurately synchronize the lip movements of the sound-producing object with the pronunciation of the target audio, but also make the generated video match the overall context of the target audio in terms of non-strictly time-aligned visual elements such as facial expressions, body movements or background changes, rather than rigidly expressing the video content only through the lips, making the expression of the characters in the generated video more vivid and natural, and improving the user's visual experience.

[0128] In one embodiment, before step S120, the following step S120a may be further included.

[0129] Step S120a: Acquire an action intensity parameter for controlling the action intensity of the sound-emitting object.

[0130] The above-mentioned action intensity parameter can be data input by the user. For example, the action intensity parameter can be represented by a normalized numerical value between 0 and 1. The larger the numerical value, the greater the action intensity of the sound-emitting object, and the smaller the numerical value, the smaller the action intensity of the sound-emitting object. Action intensity can also be represented by numerical values ​​in other ranges, such as numerical values ​​in the range of 1 to 100, 1 to 10, etc., or the action intensity can also be represented by characters, for example, it can include four levels of intensity from a to d, or the action intensity parameter can also be represented by text, for example, the action intensity parameter can include five levels such as weak, relatively weak, medium, relatively strong, and strong. Those skilled in the art can flexibly set the specific form of the action intensity parameter according to actual needs.

[0131] The motion intensity of the sound-generating object is used to represent the magnitude of the motion amplitude of the sound-generating object. The greater the motion intensity, the greater the motion amplitude of the sound-generating object in the generated video, and vice versa.

[0132] The above-mentioned motion intensity parameters may include: facial motion intensity parameters and body motion intensity parameters. The facial motion intensity parameters are used to control the motion intensity of the face of the sound-generating subject, such as controlling the mouth opening amplitude intensity, frowning amplitude, facial skin movement intensity, etc. The body motion intensity parameters are used to control the motion intensity of the limbs of the sound-generating subject, such as controlling the arm swing amplitude, leg step amplitude, hand joint movement intensity, etc. Since the motion characteristics of limb and facial movements are different, this embodiment divides the motion intensity parameters into facial motion intensity parameters and body motion intensity parameters, which can more finely control the movement of different types of movements and improve the user's flexibility in controlling the motion intensity.

[0133] Correspondingly, step S120 can be implemented according to the following step S121.

[0134] Step S121: determining global visual features of each to-be-generated video frame corresponding to the audio segment according to the segment features of each audio segment of the target audio, and based on the reference image and the motion intensity parameter.

[0135] Specifically, the audio clip can be processed to extract clip features, the reference image can be analyzed to obtain visual element features of the reference image, and the clip features can be combined with the visual element features of the reference image based on the above-mentioned action intensity parameters to obtain combined features. Based on the above-mentioned combined features, a generative adversarial network, a variational autoencoder or other pre-trained generative model can be used to predict and generate global visual features of each video frame to be generated.

[0136] In a specific embodiment, step S121 may determine the global visual features of each to-be-generated video frame corresponding to the audio segment according to the following steps S121a to S121b.

[0137] Step S121a: determining the key point position change difference between each to-be-generated video frame according to the action intensity parameter.

[0138] The difference in the position changes of the key points between the video frames to be generated can reflect the movement intensity of the sound-emitting object. For example, if the neck key point and the waist key point are at positions 1 and 2 respectively in the first frame, and the neck key point and the waist key point are at positions 2 and 3 respectively in the second frame, then the changes of the two key points are both 1 pixel to the right, and there is no difference in the changes of the two key points, indicating that the neck key point and the waist key point move synchronously between the two frames, and the movement amplitude is very small. If the neck key point and the waist key point are at positions 2 and 5 respectively in the second frame, then the change of the neck key point is 1 pixel to the right, and the change of the waist key point is 3 pixels to the right. The difference in the position changes of the two key points is large, indicating that the neck key point and the waist key point do not move synchronously between the two frames, and the movement amplitude is large.

[0139] Specifically, the video generation model may further include an intensity control model. Step S121a may specifically determine the key point position change differences of the sound-producing object between each to-be-generated video frame by inputting the action intensity parameter into the intensity control model to obtain the key point position change differences of the sound-producing object between each to-be-generated video frame. Specifically, the intensity control model is used to convert the action intensity parameter into key point position change differences, thereby controlling the action intensity of the sound-producing object between different video frames when subsequently generating video frames.

[0140] Alternatively, the difference in key point position changes can also be obtained in the following manner: a mapping relationship is established between the action intensity parameter and the key point change amplitude, which mapping relationship may include a linear or nonlinear mapping. For example, the linear relationship may include the action intensity parameter being proportional to the key point change amplitude, and the difference in key point position changes of the sound-emitting object between each video frame to be generated corresponding to the action intensity parameter is obtained by the established mapping relationship.

[0141] Step S121b: determining the global visual features of each to-be-generated video frame corresponding to the audio segment according to the segment features of each audio segment of the target audio and based on the difference in position changes between the reference image and the key points.

[0142] Step S130 can be implemented according to the following step S131.

[0143] Step S131: Determine the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame according to the pronunciation features of each audio frame of the target audio, the action intensity parameter, and the lip features of the sound-making object in the reference image.

[0144] Specifically, the pronunciation features of each audio frame can be extracted from the target audio, the lip features in the reference image can be determined, and the lip features corresponding to the pronunciation features of the audio frame can be determined based on the above-mentioned action intensity parameters and through a pre-trained lip feature mapping model.

[0145] In a specific embodiment, step S131 can be implemented according to the following steps S131a.

[0146] Step S131a: Determine the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame based on the pronunciation features of each audio frame of the target audio, the key point position change differences, and the lip features of the sound-making object in the reference image.

[0147] The process of determining global visual features and lip features in steps S121 and S131 can be referred to the description of steps S120 and S130 above and will not be detailed here. In this embodiment, after converting the action intensity parameter into key point position change differences, the key point position change differences can more accurately and detailedly represent the action intensity between different frames, thereby better guiding the generation of global visual features and lip features.

[0148] By setting the motion intensity parameters, this embodiment allows users to flexibly set the required motion intensity of the sound object according to their own needs, so that the generated global visual features and lip features are consistent with the motion intensity parameters, meeting the user's customized video output needs, improving the flexibility of video generation, and providing a better user experience.

[0149] In one embodiment, before step S140, the video generation method may further include the following step S140a.

[0150] Step S140a: extracting facial features of the sound-making object from the reference image.

[0151] Specifically, such as Figure 3 As shown, the facial region of the sounding subject can be determined from a reference image. For example, the facial region of the sounding subject can be intercepted or cropped from the reference image, and the facial features of the sounding subject can be extracted from the facial region. Specifically, the facial features of the sounding subject can be extracted from the facial region using a pre-trained facial feature extraction model. The facial feature extraction model can adopt a model that has been trained in related technologies. The facial features of the sounding subject may include, but are not limited to, key point features, face shape features, facial structure features, skin features, etc.

[0152] Step S140 can be implemented according to the following step S141.

[0153] Step S141: Based on the lip features and the global visual features corresponding to the video frame to be generated and based on the facial features of the sound-emitting object, generate each video frame synchronized with the target audio and produced by the sound-emitting object based on the reference picture.

[0154] The implementation of step S141 can refer to the description of the implementation process of step S140 above. This embodiment refers to the facial features of the sound-emitting subject when generating each video frame, so that the face of the sound-emitting subject in the generated video frame can better maintain the facial features of the sound-emitting subject and is less likely to suffer from facial deformation, facial distortion, etc., thereby enabling the generated video to better achieve a balance between identity consistency and movement diversity, and better maintain the identity characteristics of the character in long video sequences and large-scale movements.

[0155] In one embodiment, before step S120, the following step S120b may be further included.

[0156] Step S120b: Obtain description text for describing the content displayed by the reference image.

[0157] The description text may be a pre-set text for the reference image, or may be a description text for the reference image input by the user.

[0158] Correspondingly, step S120 can be implemented according to the following step S122.

[0159] Step S122: determining global visual features of each to-be-generated video frame corresponding to the audio segment based on the segment features of one or more audio segments corresponding to the target audio, the reference image, and the description text.

[0160] Step S130 can be implemented according to the following step S132.

[0161] Step S132: determining the lip features of the sound-producing object in the to-be-generated video frame corresponding to the audio frame according to the pronunciation features of each audio frame of the target audio, the reference picture, and the description text.

[0162] In the case where the description text is added, the above step S121 can be specifically implemented according to the following step S121c, and the step S131 can be implemented according to the following step S131b.

[0163] Step S121c: determining the global visual features of each to-be-generated video frame corresponding to the audio segment according to the segment features of each audio segment of the target audio, and based on the reference image, the action intensity parameter, and the description text.

[0164] Step S131b: Determine the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame based on the pronunciation features of each audio frame of the target audio, the action intensity parameter, the lip features of the sound-making object in the reference picture, and the description text.

[0165] This embodiment guides the generation of global visual features and lip features by adding descriptive text, further adding guiding conditions to the generated video, and can match the generated video frames with the descriptive text, thereby making the generated video more realistic and vivid.

[0166] In a specific embodiment, when step S120 is implemented according to step S121 and step S130 is implemented according to step S131, step S12c can be implemented according to the following steps S12c1 to S12c2 when determining the first input information.

[0167] Step S12c1: combining the image sequence formed by the reference image and each of the filling frames with the action intensity parameter to obtain an intensity fusion image sequence.

[0168] Specifically, the video generation model may further include an intensity control model. Correspondingly, step S12c1 may be implemented according to the following steps A to B.

[0169] Step A: Inputting the action intensity parameter into the intensity control model to obtain the position change difference of the key points of the sound object between each video frame to be generated.

[0170] Step B: combining the image sequence formed by the reference image and each of the filling frames with the key point position change differences of the sound-emitting object between each of the video frames to be generated to obtain an intensity fusion image sequence.

[0171] This embodiment uses the intensity control model to efficiently and accurately obtain the differences in the position changes of the key points of the sound-emitting object between each video frame to be generated, so as to better reflect the visual characteristics between each frame and obtain a more accurate intensity fusion image sequence, thereby more efficiently and accurately generating a video that matches the motion intensity reflected by the action intensity parameter.

[0172] In a specific embodiment, when the motion intensity parameter includes a facial motion intensity parameter and a body motion intensity parameter, as Figure 3 As shown, the intensity control model may include a first-dimensional transformation module, a second-dimensional transformation module, and a variation difference generation module. The variation difference generation module may include a residual network and a pooling layer, each for performing multi-layer data processing to improve data processing accuracy. Accordingly, step A may be implemented according to steps A1 to A4.

[0173] Step A1: input the facial movement intensity parameter into the first dimensional transformation module to obtain N-dimensional facial movement intensity parameters, where N is a positive integer.

[0174] Step A2: inputting the limb movement intensity parameter into the second dimensional transformation module to obtain the N-dimensional limb movement intensity parameter.

[0175] The above-mentioned first dimensional transformation module and the second dimensional transformation module are respectively used to convert the facial movement intensity parameters and the body movement intensity parameters into N-dimensional vector parameters, that is, the two dimensional transformation modules are used to convert the facial movement intensity parameters and the body movement intensity parameters into vector spaces of the same dimension to facilitate subsequent feature combination.

[0176] Step A3: combining the N-dimensional facial movement intensity parameters and the N-dimensional limb movement intensity parameters to obtain a combined intensity parameter.

[0177] Specifically, the N-dimensional facial movement intensity parameters and the N-dimensional limb movement intensity parameters may be vector-joined.

[0178] Step A4: inputting the combined dimension parameters into the change difference generation module to obtain the key point position change difference of the sound object between each video frame to be generated.

[0179] The variation difference generation module is used to generate the key point position variation differences of the sound-emitting object between each video frame to be generated. This embodiment can conveniently fuse the two action intensity parameters and generate the key point position variation differences through the variation difference generation module. Through the reasonable model structure design, the efficiency and accuracy of the key point position variation difference generation are improved.

[0180] In a specific embodiment, when the video generation model is a diffusion model, step B can obtain an intensity fusion image sequence according to the following steps B1 to B2.

[0181] Step B1: adding noise to each image in an image sequence formed by the reference image and each of the padded frames to obtain a noisy image sequence.

[0182] The added noise may be random noise to improve the robustness and accuracy of the prediction. The added noise may also be other noise, such as preset noise, which is not specifically limited in this application.

[0183] B2: Combining the key point position change difference with the noisy image sequence to obtain an intensity fused noisy image sequence.

[0184] In this embodiment, the video generation model is used to denoise the intensity-fused noisy image sequence to obtain each video frame. Specifically, the video generation model denoises the intensity-fused noisy image sequence through each denoising layer. By using a diffusion model, this embodiment can accurately and quickly obtain clear and accurate video frames during the denoising process.

[0185] Step S12c2: Determine first input information of the audio alignment model according to the intensity fusion image sequence and the converted audio coding features.

[0186] This embodiment uses the action intensity parameter as reference information of the first input information of the audio alignment model, so that the audio alignment model is guided by the action intensity parameter when inferring global visual features and lip features. Through steps S12c1 to S12c2 and step S20, the generation and determination of global visual features and lip features in steps S121 and S131 can be realized through the audio alignment model.

[0187] In a specific embodiment, when the present application obtains the descriptive text of the reference image through step S120b to guide video generation, that is, when the global visual features are determined according to step S121c and the lip features are determined according to step S131b, step S12c2 can determine the first input information of the audio alignment model according to the following steps C to D.

[0188] Step C: combining the description text with the intensity fusion image sequence to obtain a text intensity fusion image sequence.

[0189] Specifically, such as Figure 3 As shown, the description text can be encoded to obtain a text code, and the text code is combined with the intensity fusion image sequence to obtain a text intensity fusion image sequence. When the text code is combined with the intensity fusion image sequence, as shown in FIG. Figure 3 As shown, the above-mentioned video generation model may also include a picture-text fusion model, and a pre-trained picture-text fusion model may be used to combine the text encoding with the intensity fusion image sequence to obtain a text intensity fusion image sequence. The picture-text fusion model may be a cross-attention model or other types of models, which is not specifically limited in this application.

[0190] Step D: Determine first input information of the audio alignment model based on the text intensity fusion image sequence and the converted audio coding features.

[0191] Specifically, the text intensity fused image sequence and the converted audio coding features can be determined as the first input information of the audio alignment model, or the text intensity fused image sequence and the converted audio coding features can be further processed or other guiding information can be added to obtain the first input information, so that the audio alignment model can more accurately generate global visual features and lip features.

[0192] This embodiment uses the description text as reference information for the first input information of the audio alignment model, so that the audio alignment model is guided by the description text when inferring global visual features and lip features. Through steps C to D and step S20, the generation and determination of global visual features and lip features in steps S121c and S131b can be realized through the audio alignment model.

[0193] In one embodiment, Figure 3 As shown, the video generation model may further include a face alignment model. The above step S141 may generate video frames synchronized with the target audio and based on the reference picture and produced by the sound object according to the following steps S141a to S141c.

[0194] Step S141a: Determine second input information of the face alignment model based on the facial features of the sound-producing object, the lip features corresponding to the video frame to be generated, and the global visual features.

[0195] Specifically, the facial features of the sound-emitting object, the lip features corresponding to the video frame to be generated, and the global visual features can be determined as the second input information of the face alignment model. The facial features of the sound-emitting object, the lip features corresponding to the video frame to be generated, or the global visual features can be further processed or other features can be added to obtain the second input information.

[0196] Step S141b: inputting the second input information into the face alignment model, so that the face alignment model combines the facial features with the lip features and the global visual features to obtain facial fusion video features corresponding to the video frame to be generated.

[0197] The face alignment model is used to combine the facial features with the lip features and the global visual features to obtain facial fusion video features corresponding to the video frame to be generated.

[0198] Step S141c: Generate a video of the sound produced by the sound-producing object based on the reference picture and synchronized with the target audio according to the facial fusion video features.

[0199] This embodiment can efficiently and accurately combine facial features with lip features and global visual features through the face alignment model, thereby obtaining facial fusion video features under the guidance of facial features.

[0200] In one embodiment, the video generation model may further include a pre-processing model, which is used to perform image feature processing during the denoising process of the images in the intensity fusion noisy image sequence to obtain a processed image sequence. Figure 3 As shown, the pre-processing model may include a first normalization layer, a self-attention layer, a second normalization layer, etc. The first normalization layer is used to normalize the images in the intensity fusion denoised image sequence during the denoising process to obtain first normalized image information. The self-attention layer is used to add image feature information according to the reference image on the basis of the first normalized image information to obtain feature-added image information. For example, image composition information, image layout information, image texture information and other image feature information can be added. The second normalization layer can normalize the feature-added image information to obtain second normalized image information.

[0201] Correspondingly, step S12c2 can be implemented according to the following steps E to F.

[0202] Step E: Inputting the intensity fused and noisy image sequence into the pre-processing model, so that the pre-processing model performs image feature processing on the intensity fused and noisy image sequence during the denoising process to obtain a processed image sequence.

[0203] Image feature processing may include the above-mentioned normalization processing, image feature addition processing, etc., which is not specifically limited in this application.

[0204] Step F: Determine first input information of the audio alignment model based on the processed image sequence and the converted audio coding features.

[0205] For example, the processed image sequence and the converted audio coding features can be determined as the first input information of the audio alignment model, or other features can be added or the processed image sequence and the converted audio coding features can be further feature processed to obtain the first input information.

[0206] By setting a pre-processing model, this embodiment enables the video generation model to perform image processing in the denoising process of more layers. Since each denoising layer can restore data of some dimensions of the noisy image, the pre-processing model can increase the reliability and accuracy of the video generation model in generating video frames.

[0207] In a specific embodiment, when the video generation model is a diffusion model, the above-mentioned step B2 can be implemented according to the following steps: adding a time step to the key point position change difference to obtain the key point position change difference in the time dimension; combining the key point position change difference in the time dimension with the noisy image sequence to obtain an intensity fusion noisy image sequence. Specifically, a time step can be assigned to each key point position change difference, and the time step can represent the corresponding position of the key point position change difference in the video frame sequence, and the time step can be represented by a position encoding or a time-embeddable vector. By adding a time step to the key point position change difference, it is possible to better ensure that the diffusion model understands and generates the temporal continuity in the video, so that the action intensity change between video frames has a natural and smooth transition.

[0208] The following is an illustrative introduction to the video generation method provided by this application through a specific example. Figure 3 As shown, in this example, each video frame is generated by a video generation model, the video generation model is a diffusion model, and the video generation model includes a denoising model and a non-denoising model. Figure 3 In the denoising model, the intensity image combination layer, the first normalization layer, the self-attention layer, the second normalization layer, the image-text fusion model, the audio alignment model, the face alignment model, the third normalization layer, and the feedforward neural network layer are arranged in sequence. The output of the previous layer in the denoising model serves as the input of the next layer. The non-denoising model includes the intensity control model and can also include other non-denoising models for processing data. These non-denoising models will be introduced in the following steps. Figure 3 As shown, the video generation method in this example includes the following steps (1) to (10).

[0209] Step (1): Obtain input information of the video generation model.

[0210] The input information of the video generation model includes: an image sequence formed by a reference image and a first number of filling frames, a description text for describing the content displayed by the reference image, a target audio, a facial action intensity parameter, and a body action intensity parameter.

[0211] Step (2): Encode the image sequence formed by the reference image and the first number of filling frames through a video coding model and convert them into a hidden domain to obtain an encoded image sequence, and add noise to the encoded image sequence to obtain a noisy image sequence.

[0212] Step (3): input the facial action intensity parameters and the body action intensity parameters into the intensity control model to obtain the difference in the position change of the key points of the sound-producing object between each video frame to be generated.

[0213] Step (4): adding a time step to the key point position change difference to obtain the key point position change difference in the time dimension.

[0214] Step (5): Input the position change difference of the key points in the time dimension and the noisy image sequence into the denoising model, combine the position change difference of the key points in the time dimension with the noisy image sequence through the intensity image combination layer of the denoising model to obtain an intensity fused noisy image sequence, and sequentially subject the intensity fused noisy image sequence to denoising by the first normalization layer, the self-attention layer, and the second normalization layer to obtain a processed image sequence, wherein the output result of the previous layer in the denoising model is the input of the next layer.

[0215] Step (6): Encode the description text through the text encoding model to obtain text encoding, input the text encoding into the image-text fusion model, and the image-text fusion model fuses the output data of the previous layer with the text encoding to obtain output data.

[0216] Step (7): The target audio is passed through the audio coding model to obtain audio coding, and the audio coding is converted into the same dimension as the audio alignment model through the first dimensional transformation model to obtain the converted audio coding features.

[0217] Step (8): Input the converted audio coding features and the data output by the image-text fusion model into the audio alignment model, and input the data output by the audio alignment model into the face alignment model.

[0218] Step (9): A facial region is captured from the reference image, and facial features are extracted from the facial region through a facial feature extraction model. The facial features are converted into the same dimension as the dimension corresponding to the face alignment model through a second dimensional transformation model to obtain a converted dimension facial feature, and the converted dimension facial feature is also input into the face alignment model.

[0219] Step (10): The data output by the face alignment model is denoised after passing through the third normalization layer and the nonlinear transformation layer to obtain a denoised image sequence, and the denoised image sequence is decoded by the image decoding model to obtain each video frame.

[0220] The detailed process of video generation in this example can be found in the relevant description above and will not be described in detail in this example.

[0221] Example 2

[0222] The second embodiment of the present application also provides a training method for a video generation model, which is applied to electronic devices, which may be servers, desktop computers, laptops, mobile phones, tablet computers (pads), smart watches, smart TVs, VR devices, vehicle-mounted devices, wearable devices, and other electronic devices with data processing functions. The video generation model includes an audio alignment model and may also include other models.

[0223] The training method includes the following steps S210 to S250.

[0224] Step S210: Acquire a training sample, where the training sample includes a sample video and a sample audio that matches the sample video.

[0225] The training samples may be pre-stored samples, samples downloaded from the network, or samples received from other devices, etc., which is not specifically limited in this application.

[0226] Step S220: Determine the correspondence between each sample video segment of the sample video and each sample audio segment of the sample audio.

[0227] Specifically, the sample audio can be divided into sample audio segments. The specific division method can refer to the audio segment division method in the first embodiment, and is not specifically limited in this application. The division of the sample video segments is the same as the division of the sample audio segments, so that the sample audio segments and the sample video segments can be synchronized. The sample video segments and sample audio segments having a corresponding relationship mean that the two are synchronized in playback time.

[0228] Step S230: Determine first sample input information based on the corresponding sample video clips and sample audio clips, and input the first sample input information into the audio alignment model to be trained, so that the audio alignment model to be trained learns the association between the global visual features of the sample video clips and the corresponding sample audio clips, and obtains a first-stage audio alignment model.

[0229] Specifically, the corresponding sample video clips and sample audio clips may be determined as the first sample input information, or the corresponding sample video clips and sample audio clips may be further processed or other guiding information may be added to generate the first sample input information.

[0230] In step S230, model training methods such as supervised learning, semi-supervised learning, unsupervised learning, and adversarial learning in related technologies can be used to enable the audio alignment model to be trained to learn the association between the global visual features of the sample video clips and the corresponding sample audio clips.

[0231] For example, the first frame of the sample audio clip and the sample video can be input into the audio alignment model to be trained, or the first frame of the sample audio clip and the sample video can be further processed or added with other guidance information and then input into the audio alignment model to be trained to obtain the first output video frames output by the audio alignment model to be trained. According to the differences between each first output video frame and each frame in the corresponding sample video clip, the model parameters of the audio alignment model to be trained are adjusted. By iterating multiple training times with a large number of samples, when the difference between each first output video frame and each frame in the corresponding sample video clip meets the convergence condition or is less than the preset difference threshold, the first stage audio alignment model is obtained.

[0232] For example, Figure 4 As shown, Figure 4 are the encoding information corresponding to the corresponding sample audio clips and sample video clips, Figure 4 The upper half is the video encoding information corresponding to the sample video clip, and the lower half is the audio encoding information corresponding to the sample audio clip. Step S230 can encode the sample audio clip and the sample video clip separately to obtain the sample audio clip encoding and the sample video clip encoding, and combine the sample audio clip encoding and the sample video clip encoding to obtain the first sample input information. Since the first sample input information combines the corresponding sample audio clip encoding and the sample video clip encoding as a whole, after the first sample input information is input into the audio alignment model to be trained, the audio alignment model to be trained can learn the correlation between the global visual features of the sample video clip and the corresponding sample audio clip through learning algorithms such as cross-attention learning, such as Figure 4 As shown, the horizontal coding blocks represent the coding of the sample audio segment, and the vertical coding blocks represent the coding of the sample video segment. The horizontal coding blocks and the vertical coding blocks are combined as a whole for learning to learn the association between the global visual features of the sample video segment and the corresponding sample audio segment.

[0233] In a specific embodiment, in step S230, when the video generation model is a diffusion model, the first sample input information can be determined according to the following steps S231 to S232.

[0234] Step S231: obtaining a first preset noise, and adding noise to each frame of the sample video using the first preset noise to obtain a first noisy sample video frame.

[0235] Step S232: determining first sample input information according to the corresponding first noisy sample video segment and sample audio segment, where the first noisy sample video segment is a segment consisting of the first noisy sample video frames.

[0236] Specifically, the corresponding first noisy sample video clip and sample audio clip may be determined as the first sample input information, or the first sample input information may be obtained by further processing the corresponding first noisy sample video clip and sample audio clip or adding other information.

[0237] In step SS230, the first-stage audio alignment model can be trained according to the following steps S233 to S234.

[0238] Step S233: Input the first sample input information into the audio alignment model to be trained, so that the audio alignment model to be trained denoises the first noisy sample video frame according to the sample audio segment to obtain a first denoising result, and obtains a first predicted noise corresponding to the sample video frame according to the first denoising result.

[0239] Step S234: According to the difference between the first predicted noise and the first preset noise, the model parameters of the audio alignment model to be trained are adjusted to obtain a first-stage audio alignment model.

[0240] Specifically, the model parameters of the audio alignment model to be trained can be adjusted based on the principle of reducing the difference between the first predicted noise and the first preset noise until the difference between the first predicted noise and the first preset noise is less than the preset difference.

[0241] Step S240: Determine the second sample input information based on the corresponding sample video frames and sample audio frames in the sample video and the sample audio, and input the second sample input information into the first-stage audio alignment model, so that the first-stage audio alignment model learns the association between the lip features of the sample video frame and the corresponding sample audio frame, and obtains the second-stage audio alignment model.

[0242] Specifically, the corresponding sample video frames and sample audio frames may be determined as the second sample input information, or the corresponding sample video frames and sample audio frames may be further processed or other guiding information may be added to generate the first sample input information.

[0243] In step S240, model training methods such as supervised learning, semi-supervised learning, unsupervised learning, and adversarial learning in related technologies can also be used to enable the first-stage audio alignment model to learn the association between the lip features of the sample video frame and the corresponding sample audio frame.

[0244] For example, the sample audio frame and the first frame of the sample video can be input into the audio alignment model to be trained, or the sample audio frame and the first frame of the sample video can be further processed or added with other guidance information and then input into the first-stage audio alignment model to obtain the second output video frames output by the first-stage audio alignment model. According to the difference between each second output video frame and the corresponding sample video frame, the model parameters of the first-stage audio alignment model are adjusted. By iterating multiple training cycles with a large number of samples, when the difference between each second output video frame and the corresponding sample video frame meets the convergence condition or is less than the preset difference threshold, the second-stage audio alignment model is obtained.

[0245] For example, Figure 5 As shown, Figure 5 are the encoding information corresponding to the corresponding sample audio frame and sample video frame respectively, Figure 5 The upper half is the video encoding information corresponding to each frame of the sample video, and the lower half is the audio encoding information corresponding to each frame of the sample audio. Step S240 can encode the sample audio frame and the sample video frame separately to obtain the sample audio frame encoding and the sample video frame encoding, determine the sample audio frame encoding corresponding to the sample video frame encoding, and combine the corresponding sample audio frame encoding and the sample video frame encoding to obtain the second sample input information. Since the second sample input information aligns and combines the corresponding sample audio frame encoding and the sample video frame encoding at the frame level, after the second sample input information is input into the first-stage audio alignment model, the first-stage audio alignment model can learn the correlation between the lip features of the sample video frame and the corresponding sample audio frame through learning algorithms such as attention learning. Figure 5 As shown, in the sample audio frame encoding and the sample video frame encoding, the encoding blocks of the same color represent the corresponding video frames and audio frames, and the respective encoding blocks of the corresponding video frames and audio frames are combined for overall learning to learn the association between the lip features of the sample video frame and the corresponding sample audio frame.

[0246] In a specific embodiment, in step S240, when the video generation model is a diffusion model, the second sample input information can be determined according to the following steps S241 to S242.

[0247] Step S241: obtaining a second preset noise, and adding noise to each frame of the sample video with the second preset noise to obtain a second noisy sample video frame.

[0248] Step S242: Determine second sample input information according to the corresponding second noisy sample video frame and sample audio frame.

[0249] Specifically, the corresponding second noisy sample video frame and sample audio frame may be determined as the second sample input information, or the corresponding second noisy sample video frame and sample audio frame may be further processed or other information may be added to obtain the second sample input information.

[0250] In step SS240, the second-stage audio alignment model can be trained according to the following steps S243 to S244.

[0251] Step S243: Input the second sample input information into the first-stage audio alignment model, so that the first-stage audio alignment model denoises the second noisy sample video frame according to the sample audio frame to obtain a second denoising result, and obtains a second predicted noise corresponding to the sample video frame according to the second denoising result.

[0252] Step S244: According to the difference between the second predicted noise and the second preset noise, the model parameters of the first-stage audio alignment model are adjusted to obtain the second-stage audio alignment model.

[0253] Specifically, the model parameters of the first-stage audio alignment model can be adjusted based on the principle of reducing the difference between the second predicted noise and the second preset noise until the difference between the second predicted noise and the second preset noise is less than the preset difference.

[0254] In a specific embodiment, step S244 may adjust the model parameters of the first-stage audio alignment model according to the following step S244a.

[0255] Step S244a: Adjust the model parameters of the first-stage audio alignment model according to the difference between the predicted noise corresponding to the lip area in the second predicted noise and the preset noise corresponding to the lip area in the second preset noise.

[0256] This embodiment focuses on the noise prediction accuracy of the lip area, and can concentrate training attention on the lip area, thereby improving the model's learning effect on lip features in the second stage of training and improving the accuracy of model training.

[0257] In a specific embodiment, step S244a can be implemented according to the following steps G to H.

[0258] Step G: Adjust the model parameters of the first-stage audio alignment model in a first manner with a first preset probability, wherein the first manner is: adjusting the model parameters of the first-stage audio alignment model according to the difference between the predicted noise corresponding to the lip area in the second predicted noise and the preset noise corresponding to the lip area in the second preset noise.

[0259] Step H: When the first method is not hit, the model parameters of the first-stage audio alignment model are adjusted by the second method, and the second method is: adjusting the model parameters of the first-stage audio alignment model according to the difference between the first second predicted noise and the second preset noise.

[0260] In order to avoid over-regularization and improve the naturalness of head movement and background dynamics, this application adopts a first preset probability η to control the application of lip constraints, so that the model can strike a balance between paying attention to lip movement and maintaining the naturalness of the overall movement.

[0261] Step S250: Obtain a trained audio alignment model based on the second-stage audio alignment model.

[0262] Specifically, the second-stage audio alignment model can be determined as the trained audio alignment model, or the second-stage audio alignment model can be further optimized to obtain the trained audio alignment model, which is not specifically limited in this application.

[0263] The model training method provided in this application is divided into two stages to respectively learn the association between the global visual features of the sample video clips and the corresponding sample audio clips, and the association between the lip features of the sample video frames and the corresponding sample audio frames. That is, this application performs dual-stage audio-visual alignment (DAVA) training. First, the weak correlation between audio and video elements is learned at the clip level, including but not limited to facial expressions, body movements, and background changes; then, accurate lip synchronization is achieved at the frame level. This staged learning method can not only capture lip movements directly related to audio, but also notice movements that are not strictly time-aligned, such as eyebrow movements and shoulder swaying, so that the naturalness and realism of each video frame generated by the trained audio alignment model are better.

[0264] In one embodiment, the video generation model may further include an intensity control model. Before step S230, the training method may further include the following steps S230a to S230b.

[0265] Step S230a: Determine the sample action intensity parameter of the sample video.

[0266] Specifically, the sample action intensity parameter may be determined according to the following steps S230a1 to S230a2.

[0267] Step S230a1: Determine a key point position sequence consisting of key point positions corresponding to each sample video frame.

[0268] For example, a sample video frame may include multiple key points. For example, the sample video frame includes 3 key points. When 4 sample video frames are included, the key point position sequence is a sequence corresponding to the positions of the 3 key points in the 4 sample video frames.

[0269] Step S230a2: determining the key point position change difference corresponding to the key point position sequence, and determining the sample action intensity parameter of the sample video according to the key point position change difference.

[0270] The key point position variation differences corresponding to the key point position sequence are similar to the key point position variation differences between the video frames to be generated in the first embodiment and will not be described in detail here. Specifically, the key point position variance corresponding to the key point position sequence can be determined as the key point position variation differences corresponding to the key point position sequence. The variance can also be normalized through normalization, and the normalized variance can be determined as the key point position variation differences corresponding to the key point position sequence. This embodiment can effectively reflect the action intensity between the sample video frames by calculating the variance.

[0271] Step S230b: inputting the sample action intensity parameters into the intensity control model to be trained, so that the intensity control model to be trained converts the sample action intensity parameters into the position change difference of the sample key points of the sound-generating object in the sample video.

[0272] Correspondingly, step S230 may determine the first sample input information according to the following step S235.

[0273] Step S235: determining first sample input information according to the corresponding sample video clips and sample audio clips and the position change difference of the sample key points.

[0274] For example, the corresponding sample video clips and sample audio clips, and the difference in position changes of the sample key points can be determined as the first sample input information, or the corresponding sample video clips and sample audio clips, and the difference in position changes of the sample key points can be further processed or other information can be added to obtain the first sample input information.

[0275] The step S240 may determine the second sample input information according to the following step S245 .

[0276] Step S245: determining second sample input information according to the corresponding sample video frames and sample audio frames in the sample video and the sample audio, and the position change difference of the sample key points.

[0277] For example, the corresponding sample video frames and sample audio frames in the sample video and the sample audio, and the difference in position changes of the sample key points can be determined as the second sample input information. The corresponding sample video frames and sample audio frames in the sample video and the sample audio, and the difference in position changes of the sample key points can also be further processed or other information can be added to obtain the second sample input information.

[0278] The training method may further include the following steps S250 to S260.

[0279] Step S250: adjusting the model parameters of the intensity control model to be trained according to the difference between the first predicted noise and the first preset noise to obtain a first-stage intensity control model.

[0280] Step S260: adjusting the model parameters of the first-stage intensity control model according to the difference between the second predicted noise and the second preset noise to obtain a trained intensity control model.

[0281] The intensity control model's parameter adjustment method is similar to that of the audio alignment model and will not be detailed here. This embodiment adds an intensity control model, allowing the trained video generation model to control the motion intensity of the generated video through the intensity control model, improving user flexibility and video generation flexibility.

[0282] In a specific embodiment, the video generation model may further include a face alignment model, and step S233 may determine the first predicted noise corresponding to the sample video frame according to the following steps S233a to S233c.

[0283] Step S233a: extracting sample facial features of the sample sound-making object from the video frames of the sample video.

[0284] The method for extracting the facial features of the sample may refer to the method for extracting the facial features in the first embodiment, and will not be described in detail here.

[0285] Step S233b: Determine third input information according to the sample facial features and the first denoising result.

[0286] For example, the sample facial features and the first denoising result may be determined as the third input information, or the sample facial features and the first denoising result may be further processed or added with other information to obtain the third input information.

[0287] Step S233c: Inputting the third input information into the face alignment model to be trained, so that the face alignment model to be trained denoises the first denoising result according to the sample facial features to obtain a third denoising result, and obtaining a first predicted noise corresponding to the sample video frame according to the third denoising result.

[0288] In a specific embodiment, step S243 may determine the second predicted noise corresponding to the sample video frame according to the following steps S243a to S243b.

[0289] Step S243a: Determine fourth input information according to the sample facial features and the second denoising result.

[0290] For example, the sample facial features and the second denoising result may be determined as the fourth input information, or the sample facial features and the second denoising result may be further processed or added with other information to obtain the fourth input information.

[0291] Step S243b: Inputting the fourth input information into the face alignment model to be trained, so that the face alignment model to be trained denoises the second denoising result according to the sample facial features to obtain a fourth denoising result, and obtaining a second predicted noise corresponding to the sample video frame according to the fourth denoising result.

[0292] The model training method may further include the following steps S270 to S280.

[0293] Step S270: According to the difference between the first predicted noise and the first preset noise, the model parameters of the face alignment model to be trained are adjusted to obtain a first-stage face alignment model:.

[0294] Step S280: adjusting the model parameters of the first-stage face alignment model according to the difference between the second predicted noise and the second preset noise to obtain a trained face alignment model.

[0295] The model parameter adjustment method of the face alignment model is similar to the parameter adjustment method of the audio alignment model and will not be detailed here. This embodiment adds a face alignment model so that the trained video generation model can use the face alignment model to constrain the motion of the generated video to better ensure that the generated video maintains its identity.

[0296] The following is an illustrative introduction to the model training method provided by this application through a specific example. Figure 6 As shown, the video generation model to be trained in this example is a diffusion model. Similar to the first embodiment, the video generation model includes a denoising model and a non-denoising model. Figure 6In the denoising model, the intensity image combination layer, the first normalization layer, the self-attention layer, the second normalization layer, the image-text fusion model, the audio alignment model, the face alignment model, the third normalization layer, and the feedforward neural network layer are arranged in sequence. The output of the previous layer in the denoising model serves as the input of the next layer. The non-denoising model includes an intensity control model and may also include other non-denoising models for processing data. These non-denoising models will be introduced in the following steps. Among them, some models in the video generation model to be trained can be pre-trained models, and some models are models to be trained. Figure 6 In the example, the ones marked with flames are models to be trained, and the ones without flames are pre-trained models. By using some pre-trained models, the efficiency of model training can be improved, so that targeted and efficient model training can be carried out. The training process includes the first training process and the second training process, such as Figure 6 As shown, the first training process includes the following steps (1) to (12).

[0297] Step (1): Obtain training samples.

[0298] The training samples include: a sample video, a sample description text for describing the content displayed in the first frame of the sample video, and a sample audio corresponding to the sample video.

[0299] Step (2): extracting sample facial motion intensity parameters and sample body motion intensity parameters from the sample video.

[0300] Step (3): The sample video is converted into a latent domain by video encoding using a pre-trained video encoding model to obtain a sample encoded image sequence, and a first preset noise is added to the sample encoded image sequence to obtain a sample noisy image sequence.

[0301] Step (4): Input the sample facial action intensity parameters and the sample body action intensity parameters into the intensity control model to be trained to obtain the output key point position change difference.

[0302] Step (5): adding a time step to the position change difference of the sample key points to obtain the position change difference of the sample key points in the time dimension.

[0303] Step (six): Input the difference in position change of the key points in the sample time dimension and the sample noisy image sequence into the denoising model, combine the difference in position change of the key points in the sample time dimension with the sample noisy image sequence through the pre-trained intensity image combination layer in the denoising model to obtain a sample intensity fusion noisy image sequence, and subject the sample intensity fusion noisy image sequence to denoising processing of the pre-trained first normalization layer, self-attention layer, and second normalization layer in sequence to obtain a sample processed image sequence, wherein the output result of the previous layer in the denoising model is the input of the next layer.

[0304] Step (seven): Encode the sample description text through the pre-trained text encoding model to obtain the sample text encoding, and input the sample text encoding into the image-text fusion model to be trained. The image-text fusion model to be trained fuses the output data of the previous layer with the sample text encoding to obtain the output data.

[0305] Step (eight): The sample audio is passed through the pre-trained audio coding model to obtain the sample audio code, and the sample audio code is converted into the same dimension as the audio alignment model to be trained through the first dimension transformation model to be trained to obtain the sample converted audio code features.

[0306] Step (9): Input the corresponding sample converted audio segment encoding features and sample output segment data into the audio alignment model to be trained, the sample converted audio segment encoding features are the encoding features corresponding to the audio segment of the sample audio, and the sample output segment data are the data output by the graphic fusion model corresponding to the video segment corresponding to the audio segment of the sample audio, so that the audio alignment model to be trained learns the association between the global visual features of the sample video segment and the corresponding sample audio segment, and inputs the input output by the audio alignment model to be trained into the face alignment model to be trained.

[0307] Step (10): A sample facial area is captured from the first sample video frame, and sample facial features are extracted from the sample facial area using a pre-trained facial feature extraction model. The sample facial features are converted into the same dimension as the dimension corresponding to the face alignment model to be trained using the second dimension transformation model to be trained, thereby obtaining sample converted dimension facial features. The sample converted dimension facial features are also input into the face alignment model to be trained.

[0308] Step (11): The data output by the face alignment model to be trained is denoised after passing through a pre-trained third normalization layer and a pre-trained nonlinear transformation layer to obtain a denoised video frame and a first predicted noise.

[0309] Step (12): Adjust the model parameters of each model to be trained according to the difference between the first predicted noise and the first preset noise to obtain a first-stage audio alignment model.

[0310] The second training process is similar to the training process of the first training process, except that the model to be trained in the first training process becomes the first-stage model in the second stage, that is, the model that has undergone the first-stage training process, and the second training process changes step (nine) in the first training process into the following step (thirteen).

[0311] Step (thirteen): Input the corresponding sample converted audio frame coding features and sample output frame data into the first stage audio alignment model, the sample converted audio frame coding features are the coding features corresponding to the audio frame of the sample audio, and the sample output frame data are the data output by the graphic fusion model corresponding to the video frame corresponding to the audio frame of the sample audio, so that the first stage audio alignment model learns the association between the lip features of the sample video frame and the corresponding sample audio frame, and inputs the data output by the first stage audio alignment model into the first stage face alignment model.

[0312] The specific training process in this example can be referred to the detailed description above and will not be described in detail here.

[0313] The model training process in this embodiment is similar to the relevant steps in the video generation process in the first embodiment. Similar content has been described in detail in the first embodiment. For the details of the relevant technical features and the effects achieved, please refer to the corresponding description of the video generation method embodiment provided in the above first embodiment.

[0314] The following explains the encoding and decoding process in this application through the first repair encoding method described in the first embodiment.

[0315] Divide the data to be transmitted into n source data packets, and divide the block size of the data to be transmitted into n;

[0316] Sequentially or in other ways, perform XOR calculations on the contents of n complete source data packets to obtain the encoded content in the repair data packet. Note that the number of data packets in the block now becomes n+1, that is, n source data packets and one repair data packet;

[0317] Writing the value of n as a data frame in the repair packet into a repair data packet, transmitting the repair data packet and each source data packet to the client, or transmitting the data packet to the client in another manner, so that when the client detects packet loss, if it detects the presence of the repair data packet, it will perform decoding and repair;

[0318] When the client receives a redundant data packet, it starts decoding. The decoding process is the same as the encoding process. The contents of the data packet are XORed. The obtained content is the content of the lost packet, and the recovery process ends.

[0319] Example 3

[0320] The third embodiment of the present application further provides a video generation device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For details of the relevant technical features and the effects achieved, please refer to the corresponding description of the video generation method provided above. The video generation device provided in this embodiment includes:

[0321] an acquisition unit, configured to acquire target audio and a reference picture for generating a video, wherein the reference picture includes a sound-emitting object;

[0322] a determination unit configured to determine, based on segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment, the global visual features being used to indicate overall visual features of the video frame based on the reference image; and to determine, based on pronunciation features of each audio frame of the target audio and lip features of the sound-making subject in the to-be-generated video frame corresponding to the audio frame, lip features of the sound-making subject in the reference image;

[0323] A generating unit is configured to generate, based on the lip features and the global visual features corresponding to the video frames to be generated, video frames synchronized with the target audio and produced by the sound-producing object and based on the reference picture.

[0324] Example 4

[0325] The fourth embodiment of the present application also provides an electronic device embodiment corresponding to the video generation method provided in the first embodiment. The following description of the electronic device embodiment is only illustrative. The electronic device embodiment is as follows:

[0326] Please refer to Figure 7 Understand the above electronic devices, Figure 7 Schematic diagram of an electronic device. The electronic device provided in this embodiment includes: a processor 1001, a memory 1002, a communication bus 1003, and a communication interface 1004;

[0327] The memory 1002 is used to store computer instructions for data processing. When the computer instructions are read and executed by the processor 1001, the following steps are performed:

[0328] Obtaining target audio and a reference picture for generating a video, wherein the reference picture includes a sound-emitting object;

[0329] determining, based on the segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment, the global visual features being used to indicate overall visual features of the video frame based on the reference image;

[0330] Determining, based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame;

[0331] According to the lip features and the global visual features corresponding to the video frame to be generated, each video frame synchronized with the target audio and based on the reference picture and produced by the sound object is generated.

[0332] The fifth embodiment of the present application also provides an electronic device embodiment corresponding to the model training method provided in the second embodiment. The following description of the electronic device embodiment is only illustrative. The electronic device embodiment is as follows:

[0333] The electronic device provided in this embodiment includes: a processor, a memory, a communication bus, and a communication interface;

[0334] The memory is used to store computer instructions for data processing, which, when read and executed by the processor, perform the following steps:

[0335] Acquire a training sample, where the training sample includes a sample video and a sample audio that matches the sample video;

[0336] Determining a correspondence between each sample video segment of the sample video and each sample audio segment of the sample audio;

[0337] Determining first sample input information based on the corresponding sample video clips and sample audio clips, and inputting the first sample input information into the audio alignment model to be trained, so that the audio alignment model to be trained learns the association between the global visual features of the sample video clips and the corresponding sample audio clips, thereby obtaining a first-stage audio alignment model;

[0338] Determining second sample input information based on the sample video frames and sample audio frames corresponding to the sample video and the sample audio, and inputting the second sample input information into the first-stage audio alignment model so that the first-stage audio alignment model learns the association between the lip features of the sample video frames and the corresponding sample audio frames, thereby obtaining a second-stage audio alignment model;

[0339] The trained audio alignment model is obtained based on the second-stage audio alignment model.

[0340] The sixth embodiment of the present application also provides a computer-readable storage medium for implementing the method described in the first embodiment. The computer-readable storage medium embodiment provided in this application is described in a relatively simple manner. For relevant parts, please refer to the corresponding description of the above method embodiment. The embodiment described below is merely illustrative.

[0341] The computer-readable storage medium provided in this embodiment stores computer instructions, which, when executed by a processor, implement the following steps:

[0342] Obtaining target audio and a reference picture for generating a video, wherein the reference picture includes a sound-emitting object;

[0343] determining, based on the segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment, the global visual features being used to indicate overall visual features of the video frame based on the reference image;

[0344] Determining, based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame;

[0345] According to the lip features and the global visual features corresponding to the video frame to be generated, each video frame synchronized with the target audio and based on the reference picture and produced by the sound object is generated.

[0346] The seventh embodiment of the present application also provides a computer-readable storage medium for implementing the method described in the second embodiment. The computer-readable storage medium embodiment provided in this application is described in a relatively simple manner. For relevant parts, please refer to the corresponding description of the above method embodiment. The embodiment described below is merely illustrative.

[0347] The computer-readable storage medium provided in this embodiment stores computer instructions, which, when executed by a processor, implement the following steps:

[0348] Acquire a training sample, where the training sample includes a sample video and a sample audio that matches the sample video;

[0349] Determining a correspondence between each sample video segment of the sample video and each sample audio segment of the sample audio;

[0350] Determining first sample input information based on the corresponding sample video clips and sample audio clips, and inputting the first sample input information into the audio alignment model to be trained, so that the audio alignment model to be trained learns the association between the global visual features of the sample video clips and the corresponding sample audio clips, thereby obtaining a first-stage audio alignment model;

[0351] Determining second sample input information based on the sample video frames and sample audio frames corresponding to the sample video and the sample audio, and inputting the second sample input information into the first-stage audio alignment model so that the first-stage audio alignment model learns the association between the lip features of the sample video frames and the corresponding sample audio frames, thereby obtaining a second-stage audio alignment model;

[0352] The trained audio alignment model is obtained based on the second-stage audio alignment model.

[0353] The eighth embodiment of the present application also provides a computer program product for implementing the video generation method described in the first embodiment. The computer program product embodiment provided in this application is described in a relatively simple manner. For relevant parts, please refer to the corresponding description of the above method embodiment. The embodiment described below is merely illustrative.

[0354] The computer program product provided in this embodiment includes a computer program, which, when executed by a processor, implements the following steps:

[0355] Obtaining target audio and a reference picture for generating a video, wherein the reference picture includes a sound-emitting object;

[0356] determining, based on the segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment, the global visual features being used to indicate overall visual features of the video frame based on the reference image;

[0357] Determining, based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame;

[0358] According to the lip features and the global visual features corresponding to the video frame to be generated, each video frame synchronized with the target audio and based on the reference picture and produced by the sound object is generated.

[0359] The ninth embodiment of the present application also provides a computer program product for implementing the training method described in the second embodiment. The computer program product embodiment provided in this application is described in a relatively simple manner. For relevant parts, please refer to the corresponding description of the above method embodiment. The embodiment described below is merely illustrative.

[0360] The computer program product provided in this embodiment includes a computer program, which, when executed by a processor, implements the following steps:

[0361] Acquire a training sample, where the training sample includes a sample video and a sample audio that matches the sample video;

[0362] Determining a correspondence between each sample video segment of the sample video and each sample audio segment of the sample audio;

[0363] Determining first sample input information based on the corresponding sample video clips and sample audio clips, and inputting the first sample input information into the audio alignment model to be trained, so that the audio alignment model to be trained learns the association between the global visual features of the sample video clips and the corresponding sample audio clips, thereby obtaining a first-stage audio alignment model;

[0364] Determining second sample input information based on the sample video frames and sample audio frames corresponding to the sample video and the sample audio, and inputting the second sample input information into the first-stage audio alignment model so that the first-stage audio alignment model learns the association between the lip features of the sample video frames and the corresponding sample audio frames, thereby obtaining a second-stage audio alignment model;

[0365] The trained audio alignment model is obtained based on the second-stage audio alignment model.

[0366] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0367] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0368] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.

[0369] 2. Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0370] Although the present application is disclosed as above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.

Claims

1. A video generation method, characterized in that: The method comprises: Obtaining target audio and a reference picture for generating a video, wherein the reference picture includes a sound-emitting object; determining, based on the segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment, the global visual features being used to indicate overall visual features of the video frame based on the reference image; Determining, based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame; According to the lip features and the global visual features corresponding to the video frame to be generated, each video frame synchronized with the target audio and based on the reference picture and produced by the sound object is generated.

2. The video generation method according to claim 1, characterized in that Before determining the global visual features of each to-be-generated video frame corresponding to the audio segment based on the segment features of one or more audio segments corresponding to the target audio and the reference image, the method further includes: Acquiring an action intensity parameter for controlling the action intensity of the sound-emitting object; The determining, based on the segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment includes: Determining, based on the segment features of each audio segment of the target audio and the reference image and the motion intensity parameter, global visual features of each to-be-generated video frame corresponding to the audio segment; The determining, based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame includes: According to the pronunciation features of each audio frame of the target audio, the action intensity parameter and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the video frame to be generated corresponding to the audio frame are determined.

3. The video generation method according to claim 2, characterized in that The determining, based on the segment features of each audio segment of the target audio and the reference image and the motion intensity parameter, global visual features of each to-be-generated video frame corresponding to the audio segment includes: Determining the difference in position change of key points of the sound-emitting object between the video frames to be generated according to the action intensity parameter; Determining, based on the segment features of each audio segment of the target audio and the difference in position changes between the reference image and the key points, global visual features of each to-be-generated video frame corresponding to the audio segment; The determining, based on the pronunciation features of each audio frame of the target audio, the action intensity parameter, and the lip features of the sound-making object in the reference picture, of the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame includes: According to the pronunciation features of each audio frame of the target audio, the difference in the position changes of the key points, and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the video frame to be generated corresponding to the audio frame are determined.

4. The video generation method according to claim 1, wherein: Before generating, based on the lip features and the global visual features corresponding to the video frames to be generated, the video frames of the sound produced by the sound-producing object, which are synchronized with the target audio and based on the reference picture, the method further includes: Extracting facial features of the sound-making object from the reference image; The step of generating, based on the lip features and the global visual features corresponding to the video frame to be generated, video frames synchronized with the target audio and produced by the sound object and based on the reference picture, includes: According to the lip features and the global visual features corresponding to the video frame to be generated, and based on the facial features of the sound-emitting object, each video frame synchronized with the target audio and produced by the sound-emitting object based on the reference picture is generated.

5. The video generation method according to claim 1, wherein: Before determining the global visual features of each to-be-generated video frame corresponding to the audio segment based on the segment features of one or more audio segments corresponding to the target audio and the reference image, the method further includes: Obtaining a description text for describing the content displayed by the reference image; The determining, based on the segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment includes: determining, based on the segment features of one or more audio segments corresponding to the target audio, the reference image, and the description text, global visual features of each to-be-generated video frame corresponding to the audio segment; The determining, based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference picture, the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame includes: According to the pronunciation features of each audio frame of the target audio, the reference picture and the description text, the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame are determined.

6. The video generation method according to claim 1, wherein: The video generation method generates each video frame based on a pre-trained video generation model, wherein the video generation model includes an audio alignment model; The global visual features and the lip features are determined in the following manner: Determining first input information of the audio alignment model according to the target audio and the reference image; The first input information is input into the audio alignment model so that the audio alignment model outputs the global visual features of each to-be-generated video frame corresponding to the audio segment based on the segment features of one or more audio segments corresponding to the target audio and the reference image, and outputs the lip features of the sound-making object in the to-be-generated video frame corresponding to the audio frame based on the pronunciation features of each audio frame of the target audio and the lip features of the sound-making object in the reference image.

7. The video generation method according to any one of claims 1 to 6, characterized in that: The segment features include: at least one of semantic features, tone features, emotional features, background sound features, pitch features, sound intensity features, pause features, and speech speed features; The global visual features include at least one of facial expression features, limb movement features, background features, and lip features.

8. A method for training a video generation model, characterized in that: The video generation model includes an audio alignment model, and the method includes: Acquire a training sample, where the training sample includes a sample video and a sample audio that matches the sample video; Determining a correspondence between each sample video segment of the sample video and each sample audio segment of the sample audio; Determining first sample input information based on the corresponding sample video clips and sample audio clips, and inputting the first sample input information into the audio alignment model to be trained, so that the audio alignment model to be trained learns the association between the global visual features of the sample video clips and the corresponding sample audio clips, thereby obtaining a first-stage audio alignment model; Determining second sample input information based on the sample video frames and sample audio frames corresponding to the sample video and the sample audio, and inputting the second sample input information into the first-stage audio alignment model so that the first-stage audio alignment model learns the association between the lip features of the sample video frames and the corresponding sample audio frames, thereby obtaining a second-stage audio alignment model; The trained audio alignment model is obtained based on the second-stage audio alignment model.

9. A video generating device, characterized in that: The device comprises: an acquisition unit, configured to acquire target audio and a reference picture for generating a video, wherein the reference picture includes a sound-emitting object; a determination unit configured to determine, based on segment features of one or more audio segments corresponding to the target audio and the reference image, global visual features of each to-be-generated video frame corresponding to the audio segment, the global visual features being used to indicate overall visual features of the video frame based on the reference image; and to determine, based on pronunciation features of each audio frame of the target audio and lip features of the sound-making subject in the to-be-generated video frame corresponding to the audio frame, lip features of the sound-making subject in the reference image; A generating unit is configured to generate, based on the lip features and the global visual features corresponding to the video frames to be generated, video frames synchronized with the target audio and produced by the sound-producing object and based on the reference picture.

10. A computer program product, characterized in that include: The method comprises a computer program, which implements the method according to any one of claims 1 to 8 when executed by a processor.