Training method of video generation model, video generation method and device, electronic equipment, medium and program product
By integrating audio and facial data, the video generation model is trained, and the problem of insufficient consistency and authenticity of video generation in the prior art is solved, and higher quality video generation is achieved.
Patent Information
- Application Number
- CN202510220571.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has shortcomings in the consistency, authenticity and image details of video generation, and cannot fully meet users' expectations and experience in the video generation process.
By obtaining the sample training data set, including video frame data, audio data and facial data, fusion of audio and facial data, generating control data, and training the video generation model to generate target video data under the control of the control data.
The video generation quality based on audio-driven is improved, the video consistency and authenticity are enhanced, and the image details are natural and dynamic changes are improved.
Smart Images

Figure CN119996770A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of large language models, generative artificial intelligence, virtual digital humans, and specifically to a training method for a video generation model, a video generation method, a device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.
[0003] In recent years, artificial intelligence technology, especially technology related to video generation, has gradually shown its application potential in digital media, human-computer interaction, etc. With the rapid development of multimedia technology, the fusion of audio and visual information can be applied in more scenarios to meet the growing demand for digital content.
[0004] The methods described in this section are not necessarily methods that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any method described in this section is considered to be prior art simply because it is included in this section. Similarly, unless otherwise indicated, the issues mentioned in this section should not be considered to have been recognized in any prior art. Summary of the invention
[0005] The present disclosure provides a video generation model training method, a video generation method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of the present disclosure, a method for training a video generation model is provided, comprising: obtaining a sample training data set, the sample training data set comprising sample video frame data, sample audio data, and sample facial data associated with sample facial features; fusing the sample audio data and the sample facial data to obtain sample control data; based on the sample video frame data, training a video generation model to generate target video data under the control of the sample control data.
[0007] According to another aspect of the present disclosure, a video generation method is provided, comprising: acquiring image data and audio data for generating video data, the image data comprising a single video frame; providing the image data and the audio data to a video generation model to generate the video data, wherein the video generation model is trained according to the training method of the video generation model as described above.
[0008] According to another aspect of the present disclosure, a training device for a video generation model is provided, comprising: a sample data acquisition module, configured to acquire a sample training data set, the sample training data set comprising sample video frame data, sample audio data, and sample facial data associated with sample facial features; a data fusion module, configured to fuse the sample audio data and the sample facial data to obtain sample control data; and a generation training module, configured to train a video generation model to generate target video data under the control of the sample control data based on the sample video frame data.
[0009] According to another aspect of the present disclosure, a video generation device is provided, comprising: a data acquisition module, configured to acquire image data and audio data for generating video data, the image data comprising a single video frame; and a video generation module, configured to provide the image data and the audio data to a video generation model to generate video data, wherein the video generation model is trained according to the training device of the video generation model as described above.
[0010] According to another aspect of the present disclosure, an electronic device is provided, comprising at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the video generation model training method and / or video generation method as described above in the present disclosure.
[0011] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to enable a computer to execute the video generation model training method and / or video generation method as described above in the present disclosure.
[0012] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the video generation model training method and / or video generation method as described above in the present disclosure.
[0013] According to one or more embodiments of the present disclosure, the quality of audio-driven video generation can be improved.
[0014] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0016] Figure 1 A schematic diagram showing an exemplary system in which various methods described herein may be implemented according to an embodiment of the present disclosure;
[0017] Figure 2 A flowchart of a method for training a video generation model according to an embodiment of the present disclosure is shown;
[0018] Figure 3 A schematic diagram showing the generation of target video data based on training a video generation model based on sample video frame data according to an embodiment of the present disclosure is shown;
[0019] Figure 4 A schematic diagram showing the fusion of sample audio data and sample facial data according to an embodiment of the present disclosure is shown;
[0020] Figure 5 A schematic diagram showing obtaining second fused feature information by calculation based on first fused feature information according to an embodiment of the present disclosure is shown;
[0021] Figure 6 A schematic diagram showing generation of a subsequent video segment among multiple video segments based on temporal attention between the subsequent video segment and the previous video segment according to an embodiment of the present disclosure is shown;
[0022] Figure 7 A schematic diagram showing a first-stage training process of a video generation model according to an embodiment of the present disclosure is shown;
[0023] Figure 8 A schematic diagram showing a second stage training process of a video generation model according to an embodiment of the present disclosure;
[0024] Fig. 9 A flow chart of a video generation method according to an embodiment of the present disclosure is shown;
[0025] Fig.10 A structural block diagram of a training device for a video generation model according to an embodiment of the present disclosure is shown;
[0026] Fig.11A structural block diagram of a video generating device according to an embodiment of the present disclosure is shown;
[0027] Fig.12 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0028] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0029] In the present disclosure, unless otherwise specified, the use of the terms "first", "second", etc. to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements, and such terms are only used to distinguish one element from another element. In some examples, the first element and the second element may refer to the same instance of the element, and in some cases, based on the description of the context, they may also refer to different instances.
[0030] The terms used in the description of various examples in this disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element can be one or more. In addition, the term "and / or" used in this disclosure covers any one of the listed items and all possible combinations.
[0031] Among related technologies, the technology of generating two-dimensional video from a single facial image through audio drive has been initially applied in various fields such as digital media, games, and film and television creation. However, the current technology still has deficiencies in terms of the coherence, authenticity, and image details of video generation, which cannot fully meet the expectations and experience of users in the process of video generation.
[0032] To this end, the embodiments of the present disclosure provide a more effective audio-driven video generation technology.
[0033] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0034] Figure 1 FIG. 1 is a schematic diagram of an exemplary system 100 in which various methods and apparatuses described herein may be implemented according to an embodiment of the present disclosure. Figure 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more applications.
[0035] In an embodiment of the present disclosure, the server 120 may run one or more services or software applications that enable execution of the video generation model training method and / or video generation method described in the embodiment of the present disclosure.
[0036] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtualized environments and virtualized environments. In some embodiments, these services may be provided as web-based services or cloud services, such as provided to users of client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0037] exist Figure 1 In the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may differ from the system 100. Therefore, Figure 1 is one example of a system for implementing the various methods described herein and is not intended to be limiting.
[0038] The user may use client devices 101, 102, 103, 104, 105 and / or 106 to provide audio data, image data, etc. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via the interface. Figure 1 Only six client devices are depicted, but one skilled in the art will appreciate that the present disclosure may support any number of client devices.
[0039] Client devices 101, 102, 103, 104, 105 and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, game systems, thin clients, various messaging devices, sensors or other sensing devices, etc. These computer devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smart phones, tablet computers, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Game systems may include various handheld game devices, Internet-enabled game devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.
[0040] The network 110 may be any type of network known to those skilled in the art that may support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. By way of example only, the one or more networks 110 may be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0041] Server 120 may include one or more general purpose computers, dedicated server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0042] The computing units in the server 120 may run one or more operating systems including any of the above operating systems and any commercially available server operating systems. The server 120 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0043] In some implementations, server 120 may include one or more applications to analyze and consolidate data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.
[0044] In some embodiments, the server 120 may be a server of a distributed system, or a server combined with a blockchain. The server 120 may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in a cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and virtual private servers (VPS) services.
[0045] The system 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. The databases 130 may reside in various locations. For example, the database used by the server 120 may be local to the server 120, or may be remote from the server 120 and may communicate with the server 120 via a network-based or dedicated connection. The databases 130 may be of different types. In some embodiments, the databases used by the server 120 may be, for example, relational databases. One or more of these databases may store, update, and retrieve data to and from the databases in response to commands.
[0046] In some embodiments, one or more of the databases 130 may also be used by applications to store application data. The databases used by the applications may be different types of databases, such as a key-value store, an object store, or a conventional store backed by a file system.
[0047] Figure 1The system 100 may be configured and operated in various ways to enable the application of various methods and apparatuses described in the present disclosure.
[0048] The following describes in detail various aspects of the video generation model training method and the video generation method according to the embodiments of the present disclosure.
[0049] It is understandable that the video generation model mentioned in the embodiments of the present disclosure may involve a generative artificial intelligence model, such as a large language model. In general, the video generation model can generate a video that meets the user's intention based on the prompt provided by the user, such as images, audio, and video. Since the basic principles and architecture of the video generation model are technologies known in the art, the present disclosure does not go into too much detail about these aspects to avoid confusing the main purpose of the present disclosure.
[0050] In some embodiments, the training process of the video generation model may involve two stages, namely, a static facial image generation stage and a dynamic video generation stage. The static facial image generation stage may involve the video generation model learning how to generate an image consistent with its identity based on a reference image. For example, if the reference image is for a specific person, the video generation model will also generate an image corresponding to the person. The dynamic video generation stage may involve the video generation model learning how to generate a coherent video based on a reference image or video, so that the actions of the specific person in the video are coherent.
[0051] Figure 2 A flowchart of a method 200 for training a video generation model according to an embodiment of the present disclosure is shown.
[0052] In some embodiments, method 200 may involve the second stage of the training process as described above, ie, the dynamic video generation stage.
[0053] like Figure 2 As shown, method 200 includes step S201, step S202 and step S203.
[0054] In step S201, a sample training data set is obtained, where the sample training data set includes sample video frame data, sample audio data, and sample facial data associated with sample facial features.
[0055] In the example, obtaining a sample training data set can be a basic step in the video generation model training process, and the sample training data set can include a large amount or even a massive amount of training data. The training data collected by this step can be used to provide the effectiveness and accuracy of subsequent model training. In an embodiment of the present disclosure, in order to train a video generation model based on audio-driven video generation, the sample training data set includes training data related to the two dimensions of video and audio, namely sample video data and sample audio data. At the same time, in order to enable the video generation model to better learn how to truly generate facial details adapted to the audio during the training process (such as human lips or facial movements should conform to the content of the audio itself), the sample training data set also includes sample facial data associated with sample facial features. The face mentioned in the embodiment of the present disclosure may include a human face, such as a human face, and may also include a face of a virtual image, such as an animated or cartoon character, or a digital human character.
[0056] In an example, the sample video frame data may include a single frame or multiple frames of images, which may be extracted from a video. A single frame image may be, for example, a static facial image, and multiple frames of images may be, for example, a continuous video frame sequence reflecting dynamic facial changes. For example, the sample video frame data may include a continuous video frame sequence of a virtual digital human character saying the sentence "Welcome to my live studio".
[0057] In the example, the sample audio data may correspond to the sample video frame data described above. For example, while extracting sample video data from a video, the corresponding audio data may be separately extracted as the sample audio data. For example, if the sample video frame data includes a continuous sequence of video frames in which a virtual digital human character says the sentence "Welcome to my live studio", the sample audio data may include the voice of the sentence "Welcome to my live studio". As previously mentioned, in an embodiment of the present disclosure, the sample training data set includes training data related to the two dimensions of video and audio, so as to be used for the video generation model to learn how to generate videos based on audio drive. Since the sample audio data may include non-content information such as intonation and rhythm in addition to content information related to the voice content, it is convenient for the model to learn the correlation between facial dynamic changes and voice content, and then learn how to generate videos based on audio drive.
[0058] In the example, the sample facial features may involve facial features and attributes of one or more people or virtual images when speaking, which may be quantified in one or more dimensions to facilitate the video generation model's understanding of facial details such as micro-expressions. For example, the sample facial features may be quantified in dimensions such as head or facial movement (such as rotation, tilt, etc.) and facial expressions (such as smile, frown, etc.). Therefore, the sample facial data associated with the sample facial features may carry the above-mentioned quantified information to facilitate the video generation model to better learn how to truly generate facial details adapted to the audio during the training process. Taking the aforementioned virtual digital human character saying "Welcome to my live studio" as an example, the sample facial data may be related to the virtual digital human character, or may also be related to other virtual digital human characters. For example, the sample facial features may be related to different ages, genders or personalities, such as for young people with cheerful personalities, the head movement involved in the sample facial features may be relatively larger, and the facial expressions may be relatively richer. Accordingly, the sample facial features may be extracted based on a certain number of people, or may be manually set to conform to a certain virtual image.
[0059] In step S202, the sample audio data and the sample facial data are fused to obtain sample control data.
[0060] In the example, the sample control data can be configured as a control constraint on the video generation process in the training of the video generation model, and the control constraint includes explicit control such as facial key point control due to the introduction of sample facial data. Therefore, in addition to the sample audio data portion as a driving factor, the sample control data also includes a sample facial data portion for further reflecting detailed facial features and attributes (such as micro-expressions). The fusion operation may, for example, include a weighted sum operation so that the information carried by both the sample audio data and the sample facial data can be expressed in a unified sample control data. Accordingly, in an embodiment of the present disclosure, the sample facial data also serves as a driving factor and participates in the training process of the video generation model in the form of sample control data together with the sample audio data.
[0061] In step S203, based on the sample video frame data, the training video generation model generates target video data under the control of the sample control data.
[0062] In the example, the video generation model can be based on the architecture of the known diffusion model (StableDiffusion) in deep learning, and the sample control data can be introduced as a control constraint in the video generation process of the diffusion model. In the video generation process, the diffusion model can utilize spatial attention mechanisms, cross attention mechanisms, and audio attention mechanisms. As a result, the video generation model can learn how to generate videos that are consistent with the sample video frame data, while taking into account the correspondence between audio and video and the naturalness and coherence of the facial details of the characters. For example, the video generation model can learn how to generate a video with natural and coherent mouth shape, expression, posture, etc. when a virtual digital human character says the sentence "Welcome to my live studio."
[0063] Therefore, in the training method 200 of the video generation model according to the embodiment of the present disclosure, by using sample facial data associated with sample facial features in the sample training data set, and using the fused sample audio data and sample facial data as control constraints of the video generation model in the video generation process, explicit control can be conveniently and cleverly introduced in the training process of the video generation model. The explicit control can guide the video generation model to learn detailed facial features and attributes, and then guide the video generation model to learn how to apply such facial details on the basis of taking into account the correspondence between audio and video, so as to have the ability to finely simulate the micro-expressions of characters in the video, thereby improving the authenticity of the video generation.
[0064] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0065] In some embodiments, the sample video frame data may include a plurality of sample motion frames. Figure 2 The step S203, based on the sample video frame data, training the video generation model to generate the target video data under the control of the sample control data, may include: extracting features from multiple sample motion frames to obtain sample motion feature information associated with the dynamic changes between the multiple sample motion frames; providing the sample motion feature information to the video generation model, so that the video generation model generates the target video data based on the sample motion feature information under the control of the sample control data.
[0066] In the example, the above-mentioned multiple sample motion frames may include a sequence of key frames that reflect the dynamic changes of the character's face. These key frames can capture the characteristics of the character's facial expressions, mouth shape changes, head posture, etc. when speaking. Therefore, the selection of the multiple sample motion frames can reflect the coherence of the character's facial movements when speaking, so that the video generation model can learn how to generate videos with motion coherence. In addition, these sample motion frames can also be enhanced, such as random cropping, rotation, brightness adjustment, etc., to increase the complexity of the data, so that the model can better cope with various practical application scenarios.
[0067] In the example, the feature extraction operation can be performed by a convolutional neural network (CNN). For example, channel compression can be performed by a convolutional neural network, thereby not only obtaining sample motion feature information associated with the dynamic changes between multiple sample motion frames through feature extraction operations, but also facilitating the reduction of computational complexity of subsequent processes. The dynamic changes between the above-mentioned multiple sample motion frames may, for example, include changes in facial expressions, mouth shape changes, head postures, etc. when a person is speaking, so the sample motion feature information can reflect the movement coherence of the person when speaking, thereby helping the video generation model to understand this movement coherence through a cross-attention mechanism during the video generation process, thereby facilitating the generation of videos with natural and smooth facial movements. In addition, the convolutional neural network can be pre-trained in the first stage of training as described above, so that the learned weights can help the second stage of training to converge quickly.
[0068] Therefore, by extracting features from multiple sample motion frames and providing the obtained sample motion feature information to the video generation model, the model can learn the dynamic change characteristics of facial movements, thereby improving the motion coherence of the generated video.
[0069] Figure 3 A schematic diagram of generating target video data based on training a video generation model based on sample video frame data according to an embodiment of the present disclosure is shown.
[0070] like Figure 3 As shown, the sample video frame data may include multiple sample motion frames 301. Feature extraction may be performed on the multiple sample motion frames 301, for example, by a convolutional neural network, to obtain sample motion feature information 302. The extracted sample motion feature information 302 can reflect the dynamic change information of the character in terms of mouth shape, expression, head posture, etc. The sample motion feature information 302 is provided to the video generation model 310, so that the video generation model 310 can refer to the sample motion feature information 302 in the process of generating the target video data 304, thereby generating the target video data 304 with good motion coherence under the control of the sample control data 303.
[0071] In some embodiments, the above-mentioned step of extracting features from multiple sample motion frames to obtain sample motion feature information associated with dynamic changes between the multiple sample motion frames may include: extracting spatial feature information and temporal feature information of the multiple sample motion frames, wherein the sample motion feature information includes spatial feature information and temporal feature information.
[0072] In the example, the spatial feature information may be related to the spatial attributes (e.g., geometric shape, positional relationship) of local areas in the face (e.g., lips, eyes, eyebrows, etc.). The spatial feature information of certain local areas may be captured during feature extraction using a convolutional neural network. The temporal feature information may describe the characteristics of facial or head motion changes between different frames, such as the trajectory of head posture changes, the law of expression transitions, etc. When extracting temporal feature information, the TSM (Temporal Shift Module) algorithm may be applied to capture the law of motion changes over time, wherein the available sample motion frames from multiple sample motion frames may be calculated so as to cover as large a range of sample motion frames as possible.
[0073] In the example, the video generation model can use the cross-attention mechanism to understand the spatial feature information and the temporal feature information in the process of learning video generation through the training process. The video generation model can understand the characteristics, features and attributes of facial structure through spatial feature information, and can understand the temporal continuity and motion laws between the previous and next frames through temporal feature information, thereby learning how to generate target video data with natural facial motion. For example, when generating a mouth opening action, since the spatial feature information can reflect the static features of the lip shape, and the temporal feature information can reflect the dynamic features in the continuous change process from closing to opening, the sample motion feature information after the combination of the two can accurately describe the evolution law of the mouth opening action.
[0074] Therefore, the spatial feature information and temporal feature information are extracted from the sample motion frames. These feature information not only contains static details, but also reflects the dynamic change rules between multiple frames, providing rich data support for the subsequent video generation model. Using these feature information in the training of the video generation model can enable the model to better generate dynamic videos that conform to the real motion rules.
[0075] In some embodiments, the sample facial data may include sample movement data and sample expression data, the sample movement data may be associated with facial movement features in the sample facial features, and the sample expression data may be associated with facial expression features in the sample facial features.
[0076] In the example, the sample facial data may include at least two key parts: sample movement data and sample expression data. The sample movement data may be used to describe the dynamic change characteristics of the character during the speech process, such as the rotation, tilt and overall position movement of the head. The sample expression data may be used to describe the expression change characteristics of the character's face driven by different emotions or voices. Expression changes may involve, for example, the movements of the mouth, eyes and eyebrows, such as smiling, frowning, opening the mouth, etc.
[0077] Therefore, by comprehensively considering the sample movement data and the sample expression data, the target video data generated in the subsequent model training process can accurately reflect the facial expressions and movement movements of the characters.
[0078] In some embodiments, such as in combination Figure 2 The step S202 of fusing the sample audio data and the sample facial data to obtain the sample control data may include: processing the sample audio data to obtain sample audio feature information, processing the sample movement data to obtain sample movement feature information, and processing the sample expression data to obtain sample expression feature information; fusing the sample audio feature information, the sample movement feature information and the sample expression feature information with a predetermined weight to obtain first fused feature information, wherein the sample control data includes the first fused feature information.
[0079] In the example, before data fusion, the sample audio data, sample motion data and sample expression data can be processed. For the sample audio data, a pre-trained audio embedding model (such as a Wav2Vec model) can be used to convert it into a feature vector with a predetermined dimension. For the sample motion data and sample expression data, feature extraction can be performed through a fully connected layer network.
[0080] In the example, after completing the above-mentioned feature extraction, the extracted sample audio feature information, sample movement feature information, and sample expression feature information can be fused to generate sample control data. For example, a predetermined weight can be assigned to each of the three feature information (for example, the weight ratio of sample audio feature information, sample movement feature information, and sample expression feature information can be 5:2:3), and a weighted average or other fusion strategy can be used to obtain the first fused feature information. The first fused feature information can comprehensively consider the audio content and changes in facial movements, expressions, etc., so that the generated target video data can accurately reflect the voice content and facial movements of the person when speaking. For example, when simulating the speaker's mouth shape, the audio feature information can determine the content of the voice, while the movement feature information and the expression feature information can respectively control the changes in head movement and facial expressions.
[0081] Therefore, by processing and fusing the sample audio data, sample motion data and sample expression data into the first fused feature information as a whole, and including it in the sample control data for use in the training of subsequent video generation models, the video generation model can be guided to more finely simulate the facial details of the characters in the generated video, such as micro-expressions.
[0082] Figure 4 A schematic diagram of fusing sample audio data and sample facial data according to an embodiment of the present disclosure is shown.
[0083] like Figure 4 As shown, the sample facial data 410 may include sample movement data 410a and sample expression data 410b. The sample movement data 410a and the sample expression data 410b may be feature extracted through a fully connected layer network, respectively, so as to obtain sample movement feature information 411 and sample expression feature information 412, respectively. For the sample audio data 420, the pre-trained audio embedding model Wav2Vec may be used to extract the sample audio feature information 421. Finally, the predetermined weights of the three feature information, the sample movement feature information 411, the sample expression feature information 412 and the sample audio feature information 421, may be set, and the three feature information may be fused by weighted average or other fusion strategies to obtain the first fused feature information 430.
[0084] In some embodiments, such as in combination Figure 2 The method 200 may further include: encoding the first fused feature information to obtain first encoded fused feature information; performing Hadamard products on the first encoded fused feature information with the first mask, the second mask and the third mask respectively to obtain corresponding first mask feature information, second mask feature information and third mask feature information; fusing the first mask feature information, the second mask feature information and the third mask feature information to obtain second fused feature information, wherein the sample control data includes the second fused feature information.
[0085] In the example, in order to further enhance the training effect of the video generation model, the first fused feature information may be encoded to convert the first fused feature information into a form more suitable for subsequent operations. The encoding process may be performed using a vector encoder.
[0086] In the example, the mask can be used to finely control the local area of the image. After obtaining the first coded fusion feature information, the first coded fusion feature information can be respectively subjected to Hadamard product operation with the first mask, the second mask and the third mask to obtain the first mask feature information, the second mask feature information and the third mask feature information, which can respectively represent the adjustment range of the mouth area, the facial expression area and the head movement area.
[0087] In the example, the first mask feature information, the second mask feature information, and the third mask feature information can be fused, for example, the three mask feature information can be given respective weights. Since the first mask feature information can represent the characteristics of the mouth area, and the lip movement is strongly associated with the audio, the weight can be set relatively high, and the weight ratio with the second mask feature information representing the facial expression characteristics and the third mask feature information representing the head movement characteristics can be 5:3:2. Then, weighted averaging or other fusion strategies can be used to fuse the three mask feature information into the second fused feature information. The second fused feature information can be used as part of the sample control data to guide the video generation model to learn how to more accurately generate the facial movements and expression changes of the characters when speaking.
[0088] Therefore, this region-based feature adjustment method provides accurate and detailed control constraints for the video generation model during the training process, allowing the model to more effectively learn how to generate target video data that can truly and naturally reflect the relationship between audio and facial dynamics.
[0089] Figure 5 A schematic diagram of obtaining second fused feature information through calculation based on first fused feature information according to an embodiment of the present disclosure is shown.
[0090] like Figure 5 As shown, the first fused feature information 501 can be encoded to obtain the first encoded fused feature information 510. The first encoded fused feature information 510 can be respectively subjected to Hadamard product operations with the first mask 511 representing the mouth area, the second mask 512 representing the facial expression area, and the third mask 513 representing the head movement to obtain the corresponding first mask feature information 521 representing the mouth area feature, the second mask feature information 522 representing the facial expression area feature, and the third mask feature information 523 representing the head movement feature. Finally, the predetermined weights of the first mask feature information 521, the second mask feature information 522, and the third mask feature information 523 can be set, and the three feature information can be fused by weighted average or other fusion strategies to obtain the second fused feature information 530.
[0091] In some embodiments, the first mask, the second mask, and the third mask may be used to adjust features of different areas of the face that are associated with the audio.
[0092] In the example, the first mask can focus on the mouth area, for example, it can cover the lips and surrounding areas to enhance the synchronization of mouth shape and audio; the second mask can cover the facial expression area, such as eyes, eyebrows and cheeks, to control things like the amplitude of expression; the third mask can target posture, such as movement and rotation of the head, to adjust the overall movement of the head.
[0093] Therefore, by using masks to adjust the features of the mouth area, facial expression area and posture, the model can accurately control the dynamic changes of the mouth, face and posture, ensuring that the generated mouth shape, facial expression movements, etc. are highly matched with the audio content.
[0094] In some embodiments, the target video data may include a plurality of video segments generated one by one, and a subsequent video segment of the plurality of video segments may be generated based on a temporal attention with respect to a previous video segment.
[0095] In the example, the generation of the target video data may not be completed all at once, but may be generated one by one in sequence in units of video segments. Each video segment contains a fixed number of continuous frames, for example, the number of continuous frames may be set to 14 frames.
[0096] In the example, the temporal attention mechanism can be used in the generation process of the target video data. In order to ensure the coherence between the video segments, the video generation model can consider the dynamic changes in one or more previously generated video segments when generating the current video segment, so as to ensure that the generated current video segment is coherent and smooth with the previous video segment. That is, the video frame in the previous video segment can serve as a reference frame to help the video generation model understand the coherence of the character's movements. For example, when simulating the changes in the character's mouth shape, the video generation model can learn how to dynamically adjust the mouth shape in the subsequent video segment based on the opening and closing state of the mouth in the previous video segment based on the temporal attention mechanism, thereby avoiding the phenomenon of incoherent movements.
[0097] Therefore, by enabling the video generation model to utilize the temporal attention mechanism in the process of generating multiple video segments one by one, it can ensure that the details in each video segment, such as mouth shape, eye movement, etc., are correctly continued in time, thereby making the transition between the generated multiple video segments in terms of movements and expressions smoother.
[0098] Figure 6 A schematic diagram showing generation of a subsequent video segment among multiple video segments based on temporal attention between the subsequent video segment and the previous video segment according to an embodiment of the present disclosure is shown.
[0099] like Figure 6As shown, multiple video segments 601, 602, and 603 are generated one by one in chronological order. Among them, video segment 601 is a previous video segment relative to video segments 602 and 603, and video segments 602 and 603 are subsequent videos relative to video segment 601. Similarly, video segment 602 is a previous video segment relative to video segment 603, and is a subsequent video segment relative to video segment 601. Accordingly, video segment 602 can be generated based on the temporal attention between it and video segment 601, and similarly, video segment 603 can be generated based on the temporal attention between it and at least one of video segments 601 and 602.
[0100] In some embodiments, the video generation model may include a denoising network, wherein the denoising network may be pre-trained to generate a target video frame from noise based on a reference video frame, the above-mentioned target video data includes the target video frame, and the target video frame and the reference video frame are processed and generated for the same target object.
[0101] In the example, as mentioned above, the video generation model can be based on the architecture of the diffusion model known in deep learning, and the denoising network is usually used in the architecture of the diffusion model. In the embodiment of the present disclosure, as mentioned above, the training process of the video generation model can involve two stages, namely, the static facial image generation stage and the dynamic video generation stage, so the denoising network can be pre-trained in the first stage to help the video generation model learn how to generate an image consistent with its identity based on the reference image.
[0102] In the example, the reference video frame can be a single facial image provided to the denoising network. The reference video frame can provide various information related to the face of the person, such as facial key point information, texture information, etc. During the pre-training process of the denoising network, the denoising network can learn how to recover the target video frame with the same target object as the reference video frame from the noise. For example, the denoising network can gradually recover the image content that matches the eye details in the reference video frame during the denoising process based on the eye details of the person in the reference video frame, such as the shape and texture features of the eyes. In the dynamic video generation stage, the target video data can include multiple consecutive target video frames, each of which can be generated by the denoising network to form the target video data.
[0103] Therefore, by pre-training the denoising network in the first stage of training of the video generation model, the training difficulty of the video generation model can be simplified. The second stage of training on dynamic video generation can be carried out on the basis that the video generation model first has the ability to generate images with consistent identity based on the reference image. The effectiveness of the video generation model can be improved by deploying training strategies.
[0104] In some embodiments, the reference video frame may be represented in the form of a latent vector in a latent space.
[0105] In the example, the latent space is also often referred to as the latent space or latent space, and similarly, the latent vector is also often referred to as the latent vector or latent vector. The reference video frame can be converted into a latent vector after being processed by a variational autoencoder (VAE), so that the subsequent model can understand and process the data.
[0106] Therefore, by converting the reference video frame into a latent vector, not only can we provide the video generation model with an easy-to-process data form, but we can also reduce the amount of computation for subsequent operations and improve computational efficiency.
[0107] In some embodiments, the video generation model may further include a convolutional neural network for channel compression, wherein the convolutional neural network may be pre-trained to provide first information associated with a facial region in a reference video frame to affect a generation process of a target video frame.
[0108] In the example, the convolutional neural network can extract feature information related to the facial area in the reference video frame through channel compression operation, such as position features and shape features of key parts such as eyes, nose, and mouth, that is, the first information.
[0109] In the example, in the process of generating a target video frame from noise based on a reference video frame by the denoising network, the first information about the facial region extracted by the convolutional neural network can be applied to the denoising process by means of a cross-attention mechanism. For example, as mentioned above, when the reference video frame to be generated is for a specific person, the denoising network will also generate image content corresponding to the person in the process of generating a target video frame from noise based on the reference video frame.
[0110] Therefore, by applying the facial features extracted by the convolutional neural network to the denoising network to generate the target video frame, the generated target video frame can be made consistent with the reference video frame in terms of facial features, which can improve the consistency of portrait identity compared to the traditional method of using a face encoder.
[0111] In some embodiments, the video generation model may further include a reference network, which may be pre-trained to provide second information associated with the visual texture in the reference video frame to affect the generation process of the target video frame.
[0112] In the example, the reference network may be, for example, a neural network for image feature extraction, which may extract feature information related to visual texture, i.e., second information, from the reference video frame. Feature information related to visual texture may involve details in the image, such as facial skin texture, hair details, light and shadow changes, etc. These feature information may affect the realism and delicacy of the video frame.
[0113] In the example, in the process of generating a target video frame from noise based on a reference video frame by a denoising network, the second information about visual texture extracted by the reference network can be applied to the denoising process by means of a cross-attention mechanism. For example, the reference video frame may contain skin highlights under specific lighting, and the reference network can extract the brightness distribution characteristics of the area with skin highlights and guide it in the denoising process to present the same lighting effect.
[0114] Therefore, by applying the visual texture features extracted by the reference network to the process of generating target video frames by the denoising network, the generated target video frames can be made more natural, realistic and rich in image details.
[0115] In some embodiments, the reference network may have fixed network parameters after being pre-trained.
[0116] In the example, after the first phase of training of the video generation model is completed, the pre-training of the reference network can also be completed, so that the network parameters obtained through pre-training can be fixed. That is, in the subsequent second phase of training of the video generation model, the fixed network parameters can be directly used without adjusting them.
[0117] Therefore, from the perspective of training strategy, deploying the pre-training of the reference network in the first stage and fixing its network parameters in the second stage can enable the training of the second stage to focus on the level of dynamic video generation rather than the level of static image generation, thereby simplifying the training process as a whole to improve training efficiency.
[0118] In some embodiments, the training method of the video generation model (such as combining Figure 2 The method 200) may include a first training phase for static facial image generation and a second training phase for dynamic video generation.
[0119] In the example, the modules involved in the first stage of training may include a VAE encoder module, a convolutional neural network module, a reference network module, and a denoising network module. The VAE encoder module can encode the reference video frame, and the convolutional neural network module can perform channel compression on the encoded data to obtain feature information associated with the facial area; at the same time, the reference network module can extract feature information associated with the visual texture of the encoded data. The features extracted by the reference network module and the convolutional neural network module are then applied to the denoising process in the form of a cross-attention mechanism, so that the model has the ability to generate static facial images. After the first stage of training is completed, the network parameters of the reference network are fixed and used in the second stage of training.
[0120] In the example, the modules involved in the second stage of training may include a convolutional neural network module, a motion feature extraction module, an audio feature extraction module, an expression feature extraction module, a reference network module, and a denoising network module. The convolutional neural network module can extract motion features from sample motion frames to obtain spatial and temporal features associated with dynamic changes in the face; the motion features, audio features, and expression features extracted by the motion feature extraction module, the audio feature extraction module, and the expression feature extraction module are fused with certain weights to obtain fused features; at the same time, the reference network can extract visual texture features from one or more sample motion frames as reference video frames. Finally, the visual texture features, spatial and temporal features, and fused features are applied to the denoising process in the form of a cross-attention mechanism, so that the model has the ability to generate dynamic videos consistent with the audio content.
[0121] Therefore, by training the video generation model in two stages: static image generation and dynamic video generation, the final generated video can have higher facial rendering accuracy and generation stability, and be more realistic and natural in terms of micro-expressions and facial movements.
[0122] Figure 7 A schematic diagram of the first stage training process of the video generation model according to an embodiment of the present disclosure is shown.
[0123] As mentioned above, the first stage of the training process refers to the static facial image generation stage. Figure 7As shown, the reference video frame 701 can be converted into a latent vector 702 after being processed by a VAE encoder, and the latent vector 702 can contain key information in the reference video frame 701. The convolutional neural network can be used to extract feature information from the latent vector 702 through channel compression to obtain first information 703 associated with the facial area. At the same time, the reference network can be used to extract feature information from the latent vector 702 to obtain second information 704 associated with the visual texture. Then, the first information 703 and the second information 704 can be provided to the denoising network 705, so that the denoising network 705 can generate a target video frame 706 that is the same as the target object of the reference video frame 701 during the denoising process using a cross-attention mechanism. After the first stage of training is completed, the pre-training of the reference network is immediately completed, so the network parameters of the reference network can be fixed for use in the second stage of training. The first stage of training can involve spatial attention and cross-attention.
[0124] Figure 8 A schematic diagram of the second stage training process of the video generation model according to an embodiment of the present disclosure is shown.
[0125] like Figure 8 As shown, in the second stage of training, the reference network with fixed parameters after the pre-training in the first stage can be used to extract features from one or more sample motion frames as reference video frames in the multiple sample motion frames 801 included in the sample video frame data to obtain visual texture features 802 associated with visual texture. At the same time, the multiple sample motion frames 801 can be feature extracted by a convolutional neural network to obtain spatial and temporal features 803 associated with dynamic changes in the face. On the other hand, feature extraction operations can be performed on sample motion data, sample expression data, and sample audio data to obtain motion features 804, audio features 805, and expression features 806, respectively, and then these three features are fused with certain weights to obtain fused features 807. Finally, the visual texture features 802, spatial and temporal features 803, and fused features 807 can be provided to the denoising network 808 to act on the denoising process in the process of generating video by the denoising network 808 using a cross-attention mechanism, thereby obtaining target video data 809. The training process of the second stage can involve spatial attention, cross-attention, audio attention, and temporal attention.
[0126] Fig. 9 A flowchart of a video generating method 900 according to an embodiment of the present disclosure is shown.
[0127] like Fig. 9 As shown, method 900 includes step S901 and step S902.
[0128] In step S901 , image data and audio data for generating video data are acquired, where the image data includes a single video frame.
[0129] In the example, the user may expect to use the video generation model to generate a video of interest, and at the same time expect the video to match a predetermined voice, for example, expecting to generate a video of a virtual digital person saying "Welcome to my live broadcast room! Today I will introduce the latest new products to everyone, welcome to pay attention and ask questions!" Accordingly, the image data can be provided by the user when using the video generation model, for example, the user can provide a picture or photo of the virtual digital person or a picture or photo containing the face of the virtual digital person. At the same time, the user can also provide audio data related to the voice of the sentence "Welcome to my live broadcast room! Today I will introduce the latest new products to everyone, welcome to pay attention and ask questions!"
[0130] In step S902, the image data and the audio data are provided to a video generation model to generate video data, wherein the video generation model is trained according to the video generation model training method as described above (for example, combined with Figure 2 The video generation model is obtained by training using the training method 200).
[0131] In the example, after the image data and the audio data are input into the video generation model, the model can generate a video that matches the audio data based on the image data. For example, a user can provide a picture of a person's smiling face and an audio of a song to the video generation model. After receiving these two types of data, the video generation model can generate a video in which the person can adjust his mouth shape and change his expression to the rhythm of the song, simulating the effect of humming the song.
[0132] Therefore, in the video generation method 900 according to the embodiment of the present disclosure, since the video generation model used can be trained to simulate the micro-expressions of characters in the video more finely while taking into account the correspondence between audio and video, the authenticity of the video generation can be improved, thereby improving the user experience.
[0133] Fig.10 A structural block diagram of a video generation model training device 1000 according to an embodiment of the present disclosure is shown.
[0134] like Fig.10As shown, the apparatus 1000 includes a sample data acquisition module 1001, a data fusion module 1002, and a generation training module 1003. The sample data acquisition module 1001 is configured to acquire a sample training data set, which includes sample video frame data, sample audio data, and sample facial data associated with sample facial features. The data fusion module 1002 is configured to fuse the sample audio data and the sample facial data to obtain sample control data. The generation training module 1003 is configured to generate target video data based on the sample video frame data and train the video generation model under the control of the sample control data.
[0135] The operations of the sample data acquisition module 1001, the data fusion module 1002 and the training generation module 1003 may correspond to the following: Figure 2 The operations of steps S201, S202 and S203 are shown in FIG. Therefore, the details of each aspect thereof will not be repeated here.
[0136] According to an embodiment of the present disclosure, a video generating device is also provided.
[0137] Fig.11 A structural block diagram of a video generating device 1100 according to an embodiment of the present disclosure is shown.
[0138] like Fig.11 As shown, the apparatus 1100 includes a data acquisition module 1101 and a video generation module 1102. The data acquisition module 1101 may be configured to acquire image data and audio data for generating video data, wherein the image data includes a single video frame. The video generation module 1102 may be configured to provide the image data and audio data to a video generation model to generate video data, wherein the video generation model is a training device for the video generation model as described above (e.g., in combination with Fig.10 The device 1000) is trained.
[0139] The operations of the data acquisition module 1101 and the video generation module 1102 may correspond to the following: Fig. 9 The operations of steps S901 and S902 are shown in FIG. Therefore, the details of each aspect thereof will not be repeated here.
[0140] According to an embodiment of the present disclosure, an electronic device, a readable storage medium and a computer program product are also provided.
[0141] According to an embodiment of the present disclosure, an electronic device is also provided, comprising at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0142] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to enable a computer to execute the method as described above.
[0143] According to an embodiment of the present disclosure, a computer program product is also provided, including a computer program, wherein the computer program implements the method as described above when executed by a processor.
[0144] refer to Fig.12 , a block diagram of an electronic device 1200 that can be used as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0145] like Fig.12 As shown, the electronic device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the electronic device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0146] Multiple components in the electronic device 1200 are connected to the I / O interface 1205, including: an input unit 1206, an output unit 1207, a storage unit 1208, and a communication unit 1209. The input unit 1206 can be any type of device that can input information to the electronic device 1200. The input unit 1206 can receive input digital or character information, and generate key signal input related to user settings and / or function control of the electronic device, and can include but is not limited to a mouse, a keyboard, a touch screen, a track pad, a track ball, a joystick, a microphone, and / or a remote controller. The output unit 1207 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1208 can include but is not limited to a disk and an optical disk. The communication unit 1209 allows the electronic device 1200 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device and / or the like.
[0147] The computing unit 1201 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1201 performs the various methods and processes described above. For example, in some embodiments, the method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1208. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1200 via ROM 1202 and / or a communication unit 1209. When the computer program is loaded into RAM 1203 and executed by the computing unit 1201, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to perform the method described above in any other appropriate manner (e.g., by means of firmware).
[0148] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0149] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0150] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0151] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0152] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0153] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0154] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0155] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but only by the claims after authorization and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, each step can be performed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. It is important that with the evolution of technology, many elements described herein can be replaced by equivalent elements that appear after the present disclosure.
Claims
1. A method for training a video generation model, comprising: Acquire a sample training data set, wherein the sample training data set includes sample video frame data, sample audio data, and sample facial data associated with sample facial features; fusing the sample audio data and the sample facial data to obtain sample control data; as well as Based on the sample video frame data, the training video generation model generates target video data under the control of the sample control data.
2. The method according to claim 1, wherein: The sample video frame data includes a plurality of sample motion frames, and the training video generation model generates target video data under the control of the sample control data based on the sample video frame data, including: performing feature extraction on the plurality of sample motion frames to obtain sample motion feature information associated with dynamic changes between the plurality of sample motion frames; and The sample motion feature information is provided to the video generation model, so that the video generation model generates the target video data under the control of the sample control data based on the sample motion feature information.
3. The method according to claim 2, wherein: The extracting features from the plurality of sample motion frames to obtain sample motion feature information associated with dynamic changes between the plurality of sample motion frames includes: The spatial feature information and the temporal feature information of the plurality of sample motion frames are extracted, wherein the sample motion feature information includes the spatial feature information and the temporal feature information.
4. The method according to any one of claims 1 to 3, wherein: The sample facial data includes sample movement data and sample expression data, wherein the sample movement data is associated with facial movement features in the sample facial features, and the sample expression data is associated with facial expression features in the sample facial features.
5. The method according to claim 4, wherein: The fusing the sample audio data and the sample facial data to obtain sample control data includes: Processing the sample audio data to obtain sample audio feature information, processing the sample movement data to obtain sample movement feature information, and processing the sample expression data to obtain sample expression feature information; and The sample audio feature information, the sample movement feature information and the sample expression feature information are fused with a predetermined weight to obtain first fused feature information, wherein the sample control data includes the first fused feature information.
6. The method according to claim 5, wherein: The method further comprises: Encoding the first fused feature information to obtain first encoded fused feature information; Performing Hadamard products on the first coded fusion feature information and the first mask, the second mask and the third mask respectively to obtain corresponding first mask feature information, second mask feature information and third mask feature information; and The first mask feature information, the second mask feature information and the third mask feature information are fused to obtain second fused feature information, wherein the sample control data includes the second fused feature information.
7. The method according to claim 6, wherein: The first mask, the second mask, and the third mask are respectively used to adjust features of different areas of the face that are associated with the audio.
8. The method according to any one of claims 1 to 7, wherein: The target video data includes a plurality of video segments generated one by one, wherein a subsequent video segment of the plurality of video segments is generated based on temporal attention with respect to a preceding video segment.
9. The method according to any one of claims 1 to 8, wherein: The video generation model includes a denoising network, wherein the denoising network is pre-trained to generate a target video frame from noise based on a reference video frame, the target video data includes the target video frame, and the target video frame and the reference video frame are generated by processing for the same target object.
10. The method according to claim 9, wherein: The reference video frame is represented in the form of a latent vector in a latent space.
11. The method according to claim 9 or 10, wherein: The video generation model also includes a convolutional neural network for channel compression, wherein the convolutional neural network is pre-trained to provide first information associated with the facial area in the reference video frame to act on the generation process of the target video frame.
12. The method according to claim 11, wherein: The video generation model also includes a reference network that is pre-trained to provide second information associated with visual texture in the reference video frame to affect the generation process of the target video frame.
13. The method according to claim 12, wherein: The reference network has fixed network parameters after being pre-trained.
14. The method according to any one of claims 1 to 13, wherein: The method comprises a first training phase for static facial image generation and a second training phase for dynamic video generation.
15. A video generation method, comprising: Acquire image data and audio data for generating video data, wherein the image data includes a single video frame; as well as The image data and the audio data are provided to a video generation model to generate the video data, wherein the video generation model is trained according to the method according to any one of claims 1 to 14.
16. A training device for a video generation model, comprising: A sample data acquisition module is configured to acquire a sample training data set, wherein the sample training data set includes sample video frame data, sample audio data, and sample facial data associated with sample facial features; a data fusion module, configured to fuse the sample audio data and the sample facial data to obtain sample control data; as well as A training module is generated, which is configured to generate target video data based on the sample video frame data and the training video generation model under the control of the sample control data.
17. A video generating device, comprising: a data acquisition module configured to acquire image data and audio data for generating video data, wherein the image data includes a single video frame; as well as A video generation module is configured to provide the image data and the audio data to a video generation model to generate the video data, wherein the video generation model is trained according to the device according to claim 16.
18. An electronic device, comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-15.
19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1-15.
20. A computer program product comprising a computer program, wherein: The computer program implements the method according to any one of claims 1-15 when executed by a processor.