Method and apparatus for training video generation model, device, storage medium, and product

By training the video generation model, the problem of the one-to-one correspondence between phonemes and visemes in voice-driven 3D facial expression generation was solved, and accurate matching of facial expressions and voice content was achieved, improving the naturalness and generation efficiency of virtual objects.

WO2025209111A1PCT designated stage Publication Date: 2025-10-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/081595
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-02
Filing Date
2025-03-10
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

In the prior art, when voice drives the generation of 3D facial expressions, the correspondence between phonemes and visemes is not one-to-one, resulting in a low matching degree between facial expressions and voice.

Method used

By obtaining sample pairs of speech and video, the video generation model is trained to enable it to generate facial expressions that accurately match the speech content based on the speech. The model is continuously trained using sample videos as targets to improve the accuracy of generating facial expressions.

Benefits of technology

The generated facial expressions accurately match the speech content, which improves the naturalness of virtual objects and user experience, and increases the convenience and efficiency of generating facial expressions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025081595_09102025_PF_FP_ABST
    Figure CN2025081595_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A method and an apparatus for training a video generation model, a device, a storage medium, and a product, relating to the technical field of artificial intelligence. The method comprises: obtaining at least two first sample pairs, the first sample pairs each comprising a first sample speech and a first sample video corresponding to the first sample speech, and a facial expression of a virtual object in the first sample video matching the speech content of the first sample speech (201); by means of a video generation model and on the basis of the first sample speech of each of the at least two first sample pairs, determining a predicted video corresponding to the first sample speech of each of the at least two first sample pairs (202); and on the basis of the predicted video and the first sample video corresponding to each of the at least two first sample pairs, training the video generation model (203). The video generation model trained via the method can generate, on the basis of speech, facial expressions for a virtual object that accurately match the speech content.
Need to check novelty before this filing date? Find Prior Art

Description

Training method, device, equipment, storage medium and product for video generation model

[0001] This application claims priority to Chinese patent application No. 202410395181.6 filed on April 2, 2024, entitled “Training method, apparatus, device, storage medium and product for video generation model”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of artificial intelligence technology, and in particular to a training method, apparatus, device, storage medium and product for a video generation model. Background Art

[0003] In recent years, voice-driven 3D (Three Dimensions) facial technology has been increasingly used in various fields. This technology converts speech into 3D facial expressions that are relevant to the spoken content. This technology can be applied to voice assistants, virtual objects, and other fields, for example, synchronizing facial expressions with speech.

[0004] In related technologies, when generating facial expressions based on speech, the phoneme-viseme method is usually used. For example, a correspondence between phonemes and visemes in speech is established, and facial expressions are generated based on this correspondence. Phonemes are the smallest pronunciation units with distinctive meanings in human language, while visemes are phonemes presented in a visual way, used to depict the mouth posture when pronouncing a word. However, there is not a one-to-one correspondence between phonemes and visemes. For example, different phonemes correspond to the same viseme, and the same facial expressions will be generated based on different speech, resulting in a low degree of matching between facial expressions and speech. Summary of the Invention

[0005] The embodiments of the present application provide a method, apparatus, device, storage medium, and product for training a video generation model, which can generate facial expressions for virtual objects based on speech that accurately match the speech content. The technical solutions provided by this application are as follows.

[0006] In one aspect, a method for training a video generation model is provided, the method being performed by a computer device, the method comprising:

[0007] Obtain at least two first sample pairs, each first sample pair including a first sample voice and a first sample video corresponding to the first sample voice, wherein the facial expression of the virtual object in the first sample video matches the voice content of the first sample voice;

[0008] Determining, by a video generation model, based on the first sample speech of each of the at least two first sample pairs, a predicted video corresponding to the first sample speech of each of the at least two first sample pairs, wherein the video generation model is used to generate a video based on speech;

[0009] The video generation model is trained based on the predicted video and the first sample video corresponding to the first sample speech of each of the at least two first sample pairs.

[0010] In another aspect, a training apparatus for a video generation model is provided, the apparatus comprising:

[0011] An acquisition module is configured to acquire at least two first sample pairs, each first sample pair comprising a first sample voice and a first sample video corresponding to the first sample voice, wherein the facial expression of the virtual object in the first sample video matches the voice content of the first sample voice;

[0012] a processing module, configured to determine, by means of a video generation model, a predicted video corresponding to the first sample speech of each of the at least two first sample pairs based on the first sample speech of each of the at least two first sample pairs, wherein the video generation model is configured to generate a video based on speech;

[0013] A training module is used to train the video generation model based on the predicted video and the first sample video corresponding to the respective first sample voices of the at least two first sample pairs.

[0014] On the other hand, a computer device is provided, which includes a processor and a memory, wherein the memory is used to store at least one program, and the at least one program is loaded and executed by the processor to implement the training method of the video generation model in the embodiment of the present application.

[0015] On the other hand, a computer-readable storage medium is provided, in which at least one program is stored. The at least one program is loaded and executed by a processor to implement the training method of the video generation model in the embodiment of the present application.

[0016] On the other hand, a computer program product is provided, which includes at least one program segment, wherein the at least one program segment is stored in a computer-readable storage medium, and a processor of a computer device reads the at least one program segment from the computer-readable storage medium, and the processor executes the at least one program segment, so that the computer device executes the training method of the video generation model described in any of the above implementation methods.

[0017] The embodiment of the present application provides a method for training a video generation model, which trains the video generation model based on sample speech and sample videos corresponding to the sample speech, wherein the facial expressions of the virtual object in the sample video match the speech content of the sample speech, and then continuously trains the video generation model with the sample video as the target, so that the video generation model can gradually learn how to more accurately generate facial expressions that accurately match the speech content based on the speech. The video generation model trained in this way can generate a video synchronized with the speech, and in the video, the facial expressions of the virtual object accurately match the speech content, that is, using the trained video generation model, it is possible to generate facial expressions that accurately match the speech content for the virtual object based on the speech. For example, when the trained video generation model is applied to a virtual human live broadcast scene, it can generate more natural and accurate facial expressions for the virtual human, improving the user experience. In addition, generating facial expressions based on the video generation model can improve the convenience and efficiency of generating facial expressions. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIG1 is a schematic diagram of an implementation environment provided by an embodiment of the present application.

[0019] FIG2 is a flowchart of a method for training a video generation model provided in an embodiment of the present application.

[0020] FIG3 is a flowchart of another method for training a video generation model provided in an embodiment of the present application.

[0021] FIG4 is a flowchart of another method for training a video generation model provided in an embodiment of the present application.

[0022] FIG5 is a schematic diagram of an expansion of model training data provided in an embodiment of the present application.

[0023] FIG6 is a training flowchart of a video generation model provided in an embodiment of the present application.

[0024] FIG7 is a schematic diagram of an application scenario of a video generation model provided in an embodiment of the present application.

[0025] FIG8 is a flowchart of a virtual human live broadcast scenario provided by an embodiment of the present application.

[0026] FIG9 is a schematic diagram of a game interface provided in an embodiment of the present application.

[0027] FIG10 is a schematic diagram of enhancing a sample video frame provided in an embodiment of the present application.

[0028] FIG11 is a training flowchart of a sample video processing model provided in an embodiment of the present application.

[0029] FIG12 is a flowchart of obtaining training data provided in an embodiment of the present application.

[0030] Figure 13 is a schematic diagram of a training device for a video generation model provided in an embodiment of the present application.

[0031] FIG14 is a block diagram of a terminal provided in an embodiment of the present application.

[0032] FIG15 is a block diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the sample pairs involved in this application were all obtained with full authorization.

[0034] The following is an introduction to the professional terms involved in this application:

[0035] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI domains. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0036] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning. Pretrained models are the latest development in deep learning, integrating these techniques.

[0037] Key technologies in speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction. Large model technology is revolutionizing the development of speech technology. For example, pre-trained models based on the Transformer architecture offer strong generalization and versatility, making them highly capable of handling a wide range of speech processing tasks.

[0038] Natural language processing (NLP) is a key area of ​​research in computer science and artificial intelligence. It studies theories and methods that enable effective communication between humans and computers using natural language. Natural language processing involves natural language, the language we use daily, and is closely related to linguistics. It also involves computer science and mathematics. Pre-trained models, a key technology for model training in artificial intelligence, are derived from large language models (LLMs) in the NLP field. After fine-tuning, LLMs can be widely applied to downstream tasks. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question-answering, knowledge graphs, and other technologies.

[0039] The following is an introduction to the implementation environment involved in this application.

[0040] The video generation model training method provided in the embodiment of the present application can be executed by a computer device, which can be provided as a server or a terminal. The following is a schematic diagram of the implementation environment of the video generation model training method provided in the embodiment of the present application.

[0041] Referring to Figure 1, Figure 1 is a schematic diagram of an implementation environment of a training method for a video generation model provided in an embodiment of the present application, wherein the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected via wired or wireless communication, which is not limited in this application. In some embodiments, the server 102 is used to train the video generation model, and the trained video generation model is used to generate a video based on speech, and the facial expression of the virtual object in the video matches the speech content of the speech. An application is installed on the terminal 101, which is used to generate a video based on speech and then display the video. In some embodiments, the terminal 101 is embedded with a trained video generation model, and the terminal 101 generates a video based on speech through the video generation model. In other embodiments, the terminal 101 generates a video based on speech through the video generation model on the server 102. In this application, a video refers to a series of continuous image frames. For example, a video is a playable video obtained by recording a real object. For another example, a video is a 3D animation that can be rendered to output a series of continuous image frames and played in sequence. This application does not limit this.

[0042] In some embodiments, the terminal 101 is a smart phone, a tablet computer, a laptop computer, a desktop computer, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a VR (Virtual Reality) device, an AR (Augmented Reality) device, etc., but is not limited thereto. In some embodiments, the server 102 is an independent server or a server cluster or distributed system composed of at least two servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the server 102 is primarily responsible for computing tasks, and the terminal 101 is responsible for secondary computing tasks; alternatively, the server 102 is responsible for secondary computing services, and the terminal 101 is responsible for primary computing tasks; alternatively, the server 102 and the terminal 101 use a distributed computing architecture for collaborative computing.

[0043] Refer to Figure 2, which is a flowchart of a training method for a video generation model provided in an embodiment of the present application. The method is executed by a computer device and includes the following steps.

[0044] 201. A computer device obtains at least two first sample pairs, each of which includes a first sample voice and a first sample video corresponding to the first sample voice, wherein the facial expression of a virtual object in the first sample video matches the voice content of the first sample voice.

[0045] In an embodiment of the present application, the first sample video includes a virtual object, and the virtual object emits a first sample voice and has a facial expression that matches the voice content of the first sample voice. A virtual object refers to an movable object in a virtual environment, and the movable object can be a virtual character, a virtual animal, an animated character, etc., but is not limited thereto. A virtual object has its own shape and volume in a virtual environment and occupies a part of the space in the virtual environment. A virtual environment refers to an environment provided (or displayed) when an application is running on a terminal device, and the virtual environment can be a two-dimensional virtual environment, a 2.5-dimensional virtual environment, or a three-dimensional virtual environment, etc. Exemplarily, when the virtual environment is a three-dimensional virtual environment, the virtual object is a three-dimensional model created based on animation skeleton technology.

[0046] In the embodiment of the present application, matching facial expressions with speech content means that facial expressions match at least one of the pronunciation, tone, speech speed, and emotional state of the speech content. For example, the pronunciation of the speech content matches the lip shape of the facial expression. For another example, a serious tone matches a serious facial expression, and a humorous tone matches a humorous facial expression. For another example, the slower the speech speed, the smoother the facial expression. The emotional state refers to the emotional state expressed by the speech content, and the emotional state can be determined by at least one of the modal particles, tone, and speech speed in the speech content; for example, if the emotional state is happy, the matching facial expression is happy, and if the emotional state is sad, the matching facial expression is sad.

[0047] In an embodiment of the present application, a computer device can render based on a first sample video to display the first sample video on the computer device. Optionally, the first sample video includes at least two sample video frames, and rendering is performed based on the at least two sample video frames to display the first sample video including the at least two sample video frames.

[0048] In some embodiments, the facial expressions of the virtual object are controlled by the parameters of at least two expression controllers of the virtual object's face. Accordingly, the first sample video includes the parameters of at least two expression controllers of the virtual object's face. The first sample video includes at least two sample video frames. For example, the data form of the first sample video is a matrix, each vector in the matrix is ​​a sample video frame, the dimension of each vector is the same as the number of expression controllers, and the element value of each dimension in the vector represents the parameter of an expression controller. In other words, the parameters of the expression controller represented by the element value of each dimension in the vector together constitute the sample video frame. In some embodiments, the virtual object is a 3D virtual object, and the facial expression is a 3D facial expression. Accordingly, the first sample video is a 3D facial animation video.

[0049] 202. The computer device determines, through a video generation model, predicted videos corresponding to the first sample speech of each of the at least two first sample pairs based on the first sample speech of each of the at least two first sample pairs, where the video generation model is used to generate video based on speech.

[0050] In an embodiment of the present application, for each first sample pair, the computer device processes the first sample speech in the first sample pair using a video generation model to obtain a predicted video corresponding to the first sample speech, where the predicted video includes a virtual object. The training goal of the video generation model is to generate a video based on the speech, with the facial expression of the virtual object in the video matching the speech content.

[0051] 203. The computer device trains a video generation model based on the predicted video and the first sample video corresponding to the first sample speech of each of the at least two first sample pairs.

[0052] In an embodiment of the present application, for each first sample pair, the model parameters of the video generation model are adjusted based on the loss value between the predicted video of the first sample pair and the first sample video.

[0053] In an embodiment of the present application, a computer device iteratively trains a video generation model based on at least two first samples for respective predicted videos and first sample videos until a preset requirement is met. Meeting the preset requirement may mean that the loss value between the predicted video and the first sample video converges, or that the loss value reaches a preset threshold, or that the number of iterations reaches a preset number, which are not specifically limited herein.

[0054] The embodiment of the present application provides a method for training a video generation model, which trains a video generation model based on sample speech and sample videos corresponding to the sample speech, wherein the facial expressions of the virtual objects in the sample videos match the speech content of the sample speech, and then continuously trains the video generation model with the sample videos as the target, so that the video generation model can gradually learn how to more accurately generate facial expressions that accurately match the speech content based on the speech. The video generation model trained in this way can generate a video synchronized with the speech, and the facial expressions of the virtual objects in the video accurately match the speech content. That is, using the trained video generation model, facial expressions that accurately match the speech content can be generated for the virtual objects based on the speech. For example, when the trained video generation model is applied to a virtual human live broadcast scene, more natural and accurate facial expressions can be generated for the virtual human, improving the user experience. In addition, generating facial expressions based on the video generation model can improve the convenience and efficiency of generating facial expressions.

[0055] Figure 2 above illustrates the basic process of a video generation model training method. The following further describes the video generation model training method based on Figure 3. Referring to Figure 3, Figure 3 is a flowchart of another video generation model training method provided in an embodiment of the present application. The method is executed by a computer device and includes the following steps.

[0056] 301. A computer device obtains at least two second sample pairs, each second sample pair including a second sample speech and a second sample video corresponding to the second sample speech, wherein the facial expression of the subject in the second sample video matches the speech content of the second sample speech.

[0057] In an embodiment of the present application, the second sample video includes an object. Optionally, the second sample video corresponding to the second sample voice may be a video recorded when the object speaks, and the video includes the object's facial expression, and the second sample voice is the voice recorded by the object.

[0058] The object in the second sample video can be a real object, which facilitates the acquisition of the second sample video. It should be noted that the acquisition of the second sample voice and the second sample video of the object is authorized by the object. Optionally, an authorization interface is displayed on the terminal used by the object, and the authorization interface displays prompt information, an agreement control, and a disagreement control. The prompt information is used to prompt the acquisition of the object's voice and video, and the agreement control is used to indicate that the object agrees that the terminal obtains its voice and video. In response to the triggering operation of the agreement control, the object's voice and video are acquired.

[0059] The objects in the second sample videos of at least two second sample pairs can be the same object, thereby improving the convenience of obtaining at least two second sample pairs. The objects in the second sample videos of at least two second sample pairs can be different objects, thereby improving the diversity of samples. The second sample voices of at least two second sample pairs can be different, and accordingly, the second sample videos of at least two second sample pairs can be different, thereby improving the diversity of samples.

[0060] 302. The computer device determines, through a sample video processing model, the processed second sample videos corresponding to at least two second sample pairs based on the respective second sample videos of at least two second sample pairs, wherein the sample video processing model is used to generate a sample video including a virtual object based on a sample video including an object.

[0061] In an embodiment of the present application, for each second sample pair, the second sample video in the second sample pair is processed by the sample video processing model to obtain a processed second sample video. The facial expression of the virtual object in the processed second sample video matches the facial expression of the object in the second sample video before processing. That is, the processed second sample video contains a virtual object, and the virtual object not only corresponds to the object in the second sample video, but also has the same facial expression as the object in the second sample video.

[0062] The training process of the sample video processing model is shown in the embodiment of FIG4 , and will not be described in detail here.

[0063] In some embodiments, taking the sample video frame in the second sample video as the first sample video frame as an example (other naming methods may also be used, which is not limited to this), accordingly, the computer device executes step 302, including the following steps:

[0064] For each second sample pair, the computer device determines, through the sample video processing model, the processed first sample video frames corresponding to each of the at least two first sample video frames in the second sample video in the second sample pair; and determines the processed second sample videos corresponding to each of the at least two second sample pairs based on the processed first sample video frames corresponding to each of the at least two second sample pairs.

[0065] This process is that, for any second sample video in a second sample pair, the computer device processes each sample video frame in the second sample video separately through the sample video processing model to obtain a processed sample video frame corresponding to each sample video frame, thereby converting each sample video frame including an object into a sample video frame including a virtual object.

[0066] The sample video processing model is configured to generate sample video frames including a virtual object based on sample video frames including an object. In the sample video frames including the virtual object, the virtual object not only corresponds to the object in the sample video frames including the object, but also has the same facial expression as the object. For example, if the object in the second sample video is a real object, the sample video processing model can generate the facial expression of the virtual object based on the facial expression of the real object, thereby achieving the simulation of the facial expression of the real object by the virtual object.

[0067] In some embodiments, a process in which a computer device determines processed second sample videos corresponding to at least two second sample pairs based on processed first sample video frames corresponding to the at least two second sample pairs includes the following steps:

[0068] The computer device arranges the processed first sample video frames in ascending order of sampling time based on the sampling time sequence of at least two first sample video frames in the second sample video of the second sample pair, thereby obtaining a processed second sample video. Furthermore, the data format of each processed first sample video frame is a vector, and the data format of the processed second sample video is a matrix. The at least two vectors are arranged in ascending order of sampling time of the at least two first sample video frames, thereby obtaining a processed second sample video.

[0069] In an embodiment of the present application, each sample video frame in the second sample video is processed separately by a sample video processing model to obtain a sample video frame containing a virtual object. This achieves refined processing so that each processed sample video frame is accurate, thereby improving the accuracy of the sample video. This process is to convert the second sample video containing an object into a second sample video containing a virtual object through the sample video processing model to achieve sample expansion. For example, the virtual object is a 3D virtual object, and the object in the second sample video is a real object. Through the sample video processing model, the sample video containing a real object can be converted into a sample video containing a 3D virtual object, so that the sample video containing the 3D virtual object can be used in the training process of the video generation model to achieve sample expansion.

[0070] For example, see Figure 5, which is a schematic diagram of expanding model training data provided by an embodiment of the present application. In this embodiment, the second sample video includes a real object, and the computer device processes the sample video frames in the second sample video using the sample video processing model to obtain a processed sample video frame corresponding to each sample video frame, thereby obtaining a processed second sample video. This second sample video includes a 3D virtual object, thereby expanding the training data of the video generation model.

[0071] 303. The computer device determines at least two first sample pairs based on the second sample voices of the at least two second sample pairs and the processed second sample videos corresponding to the at least two second samples.

[0072] In an embodiment of the present application, each first sample pair includes a first sample voice and a first sample video corresponding to the first sample voice, and the facial expression of the virtual object in the first sample video matches the voice content of the first sample voice.

[0073] In some embodiments, the computer device enhances the second sample speech to expand the sample. The computer device enhances the second sample speech of each of at least two second sample pairs to obtain at least two enhanced second sample speech, each of which has the same speech content as the second sample speech before enhancement. At least two first sample pairs are determined based on the second sample speech of each of the at least two second sample pairs, the at least two enhanced second sample speech, and the processed second sample videos corresponding to each of the at least two second sample pairs. For any first sample pair, the first sample speech in the first sample pair is either the second sample speech before enhancement or the second sample speech after enhancement, and each of the enhanced second sample speech has the same second sample video corresponding to the second sample speech before enhancement.

[0074] This process means that after obtaining the second sample pair, the computer device can enhance the second sample speech in the second sample pair, and use the enhanced second sample speech as the first sample speech to participate in the training process of the video generation model, or use the second sample speech before enhancement (that is, the original second sample speech) as the first sample speech to participate in the training process of the video generation model.

[0075] In an embodiment of the present application, by enhancing the second sample speech in each second sample pair and using the enhanced second sample speech as the first sample speech to participate in the training process of the video generation model, the diversity of samples is increased, and training the video generation model based on these multiple types of samples can improve the generalization of the video generation model.

[0076] In some embodiments, the computer device enhances the second sample voices of at least two second sample pairs to obtain at least two enhanced second sample voices, including at least one of the following: the computer device adjusts the pitch of the second sample voices of at least two second sample pairs to obtain at least two enhanced second sample voices; the computer device adds reverberation to the second sample voices of at least two second sample pairs to obtain at least two enhanced second sample voices; the computer device adds noise to the second sample voices of at least two second sample pairs to obtain at least two enhanced second sample voices.

[0077] Pitch is primarily determined by the frequency of a sound, rising and falling with the frequency. Furthermore, pitch is also related to the volume of a sound. Therefore, adjusting the pitch of a sample speech refers to adjusting at least one of the frequency and volume of the sample speech, such as increasing or decreasing the frequency of the sample speech, increasing or decreasing the volume of the sample speech, etc. In some embodiments, a computer device adjusts the pitch of the sample speech using a pitch adjustment device, which is a device used to adjust pitch.

[0078] Adding reverberation to the sample speech refers to adding an acoustic effect that simulates multiple reflections of sound within a closed or semi-enclosed space. For example, reverberation in at least two different scenarios can be added to the sample speech to further increase the diversity of the sample. The at least two different scenarios can be indoors, outdoors, etc.

[0079] Adding noise to the sample speech refers to adding different types of chaotic sound signals to the sample speech. For example, at least two different types of noise are added to the sample speech to further increase the diversity of the sample. The different types of noise can be Gaussian noise, white noise, impulse noise, etc.

[0080] It should be noted that the aforementioned methods for enhancing the sample speech can be freely combined, such as adding at least one of reverberation and noise to the pitch-adjusted sample speech to further increase sample diversity. The aforementioned methods are merely optional implementations of enhancing the sample speech. The computer device can also enhance the sample speech through other optional implementations, which will not be detailed here.

[0081] In an embodiment of the present application, by adjusting the pitch of the sample speech and adding reverberation or noise to the sample speech, the sample is expanded without changing the speech content of the sample speech. The above-mentioned enhancement methods are highly convenient, thereby improving the efficiency of enhancing the sample speech.

[0082] In an embodiment of the present application, the process of obtaining at least two first sample pairs is implemented through the above steps 301-303. In this embodiment, the sample video in the sample speech-sample video including the object (i.e., the second sample pair) is processed by the sample video processing model to obtain a sample video including the virtual object (i.e., the processed sample video). Since the sample speech is synchronized with the sample video including the object, the sample video including the virtual object synchronized with the sample speech is obtained, and then the parallel data of the sample speech-sample video including the virtual object is obtained. By obtaining the parallel data of the sample speech-sample video including the virtual object through the sample video processing model, it is no longer necessary for the animator to construct the facial expressions of the virtual objects in the sample video frame by frame, thereby improving the efficiency of obtaining the sample video including the facial expressions of the virtual objects, and these data are used to train the video generation model for generating videos based on speech, thereby improving the efficiency of obtaining the training data of the video generation model.

[0083] It should be noted that the above steps 301-303 are merely optional implementation methods for obtaining at least two first sample pairs. The computer device may also implement this process through other optional implementation methods, which will not be described in detail here.

[0084] 304. Determine predicted videos corresponding to the first sample speech of each of the at least two first sample pairs based on the first sample speech of each of the at least two first sample pairs using a video generation model, where the video generation model is used to generate video based on speech.

[0085] In some embodiments, the computer device processes the first sample speech in each first sample pair using a video generation model to obtain a predicted video corresponding to the first sample speech, including the following steps:

[0086] For each first sample pair, the computer device processes the first sample speech in the first sample pair using a feature extraction module in the video generation model to obtain speech features of the first sample speech. The feature extraction module is used to extract the speech features. The computer device processes the speech features using an attention module in the video generation model to obtain attention features of the speech features. The attention module is used to extract the attention features. The computer device processes the attention features using a regression module in the video generation model to obtain a predicted video. The regression module is used to generate the predicted video based on the attention features.

[0087] The following describes the implementation of step 304 by taking the first sample speech as the enhanced second sample speech as an example with reference to formula (1) to formula (4).

[0088] Schematically, the second sample speech is represented by A={a 1 ,…,aT}, the predicted video is represented as Y = {y 1 ,…,y M a represents a frame of speech signal, and y represents a frame of video signal. T represents the number of samples of the second sample speech. For example, at 16 kHz, the number of samples per second is 16,000. M represents the number of video frames in the predicted video. For example, the predicted video contains 50 frames per second.

[0089] The computer device enhances the second sample speech through the data enhancement module to obtain the enhanced second sample speech, that is, the first sample speech. This process can be expressed by the following formula (1). a (A) (1).

[0090] Wherein, A′ represents the second sample speech after enhancement, that is, the first sample speech, and A represents the second sample speech before enhancement, that is, the original second sample speech. a (A) indicates that the second sample speech A is enhanced.

[0091] Optionally, the feature extraction module in the video generation model is a wav2vec2 module (a speech feature extraction module). The first sample speech is input into the feature extraction module to obtain speech features. This process can be expressed by the following formula (2). c =wav2vec2(A′) (2).

[0092] Among them, H c Represents the voice features, M represents the number of video frames in the predicted video, the speech feature includes at least two speech sub-features, and the at least two speech sub-features correspond one-to-one to at least two video frames. represents the speech sub-feature corresponding to the Mth video frame, and wav2vec2(A′) represents the speech feature extracted from A′.

[0093] It should be noted that extracting speech features through the wav2vec2 module is only an optional implementation method. The computer device can also extract speech features through other models, which will not be repeated here.

[0094] The attention features in the video generation model can be single-head attention features or multi-head attention features, and can be self-attention features or cross-attention features. Optionally, the attention module is a multi-layer FFT (Feed Forward Transformer) module. The speech features are input into the attention module to extract new deep features of the speech features and obtain attention features. This process can be expressed by the following formula (3). H d =FFT(Hc ) (3).

[0095] Among them, H d represents the attention feature, FFT(H c ) represents the extraction of speech features H c attention characteristics.

[0096] Optionally, the regression module in the video generation model is a fully connected layer, that is, a linear prediction layer, which is used to perform nonlinear transformation on the attention features to obtain the predicted video. This process can be expressed by the following formula (4). d ) (4).

[0097] Among them, Y′ represents the predicted video, Linear(H d ) represents the attention feature H d Perform nonlinear transformation.

[0098] For example, see Figure 6, which is a training flow chart of a video generation model provided by an embodiment of the present application. Among them, the original sample speech (that is, the second sample speech) is first enhanced to obtain an enhanced sample speech (that is, the first sample speech), and the enhanced sample speech is processed by the wav2vec2 module in the video generation model to obtain speech features, and the speech features are processed by the N-layer FFT module in the video generation model to obtain attention features, where N is an integer greater than 1. The attention features are processed by the linear prediction layer in the video generation model to obtain a predicted video. Among them, the FFT module includes a multi-head attention unit and a convolution layer, and the speech features are processed by the multi-head attention unit and the convolution layer in turn to obtain attention features.

[0099] 305. The computer device trains a video generation model based on the predicted video and the first sample video corresponding to the first sample speech of each of the at least two first sample pairs.

[0100] In some embodiments, the computer device determines, for each first sample pair, a loss value between the predicted video of the first sample pair and the first sample video, and adjusts model parameters of the video generation model based on the loss value of each first sample pair.

[0101] In some embodiments, the predicted video includes at least two predicted video frames, the first sample video includes at least two sample video frames, the at least two predicted video frames correspond one-to-one to the at least two sample video frames, and the computer device trains the video generation model based on the predicted video and the first sample video of the at least two first samples, including the following steps:

[0102] For each first sample pair, the computer device determines a loss value based on the difference between at least two predicted video frames and at least two sample video frames respectively; and adjusts the model parameters of the video generation model based on the loss values ​​corresponding to each of the at least two first sample pairs.

[0103] The number of the at least two predicted video frames is the same as the number of the at least two sample video frames, and the at least two predicted video frames and the at least two sample video frames correspond one to one according to sampling time, that is, one predicted video frame corresponds to one sample video frame.

[0104] In an embodiment of the present application, a loss value is comprehensively determined based on the difference between at least two predicted video frames in the predicted video and at least two sample video frames in the first sample video, so that the loss value is more accurate and comprehensive, and then the model parameters are adjusted based on the loss value, so that the adjustment of the model parameters is more accurate, thereby improving the model training efficiency.

[0105] Optionally, the computer device determines the loss value based on the difference between at least two predicted video frames and at least two sample video frames, including the following steps: the computer device determines the average of at least two differences to obtain the loss value; or, the computer device determines the average of the squares of at least two differences to obtain the loss value, that is, the obtained loss value is an L2 (least square error) loss value.

[0106] For example, the loss value is an L2 loss value, and the process can be expressed by the following formula (5): L = || YY′ || (5).

[0107] In formula (5), L represents the loss value, Y represents the first sample video, Y′ represents the predicted video, and ||YY′|| represents the calculation of the minimum square error between Y and Y′.

[0108] In an embodiment of the present application, the trained video generation model is used to generate a video based on speech, and the facial expressions of the virtual objects in the video match the speech content of the speech.

[0109] The video generation model provided in the embodiments of this application can be applied to scenarios such as virtual human live broadcasts and game character creation. For example, see Figure 7, which is a schematic diagram of an application scenario of a video generation model provided in the embodiments of this application. The front-end module is used to generate speech, which is then used to drive the video generation model in the facial service to produce a video. The facial expressions of the virtual object in the video match the speech content of the speech.

[0110] In some embodiments, the method provided by the embodiments of the present application is applied to a virtual human live broadcast scene. Among them, the live broadcast client captures the audience barrage on the live broadcast interface, generates a response dialogue through the dialogue generation service, and the dialogue content of the dialogue is generated by the speech synthesis service to generate a reply voice, and then the video generation module in the voice-driven facial service (implemented by the video generation model provided by this application) generates a video synchronized with the voice and displays the video. The virtual anchor in the video makes facial expressions that match the voice content to achieve dialogue interaction with the audience barrage. For example, see Figure 8, which is a flow chart of a virtual human live broadcast scene provided by an embodiment of the present application.

[0111] In other embodiments, the method provided by the embodiments of the present application is used for game NPC (Non-Player Character) scenes. Among them, the game object can have a conversation with the virtual object in the game, and the corresponding conversation content is sent to the conversation generation service through the game client to generate a response conversation. The conversation generates speech through the speech synthesis module, and the speech is synthesized by the speech synthesis module, and then sent to the video generation module in the voice-driven facial service (implemented by the same video generation model of this application), and then the video generation module is used to obtain a video synchronized with the speech. Finally, the video is rendered by the engine to display the video. The virtual object in the video makes a facial expression that matches the speech content, and the virtual object and the game object have a conversation. For example, see Figure 9, which is a schematic diagram of a game interface provided by an embodiment of the present application. Among them, a virtual object is displayed on the game interface, and the virtual object has a conversation with the game object, and the conversation content is displayed on the game interface. The facial expression of the virtual object matches the speech content of the virtual object.

[0112] The present embodiment utilizes a sample video processing model to expand the parallel data of sample speech and sample video, including objects. This expands the sample pairs used to train the video generation model, reducing the data cost required for training. Furthermore, combined with speech enhancement technology, this improves the model's generalizability when the amount of training data is insufficient.

[0113] In an embodiment of the present application, a sample video in a sample video including a sample speech and an object is processed by a sample video processing model to obtain a processed sample video. Since the sample speech is synchronized with the sample video including the object, a sample video including a virtual object that is synchronized with the sample speech is obtained, thereby obtaining parallel data of the sample speech and the sample video including the virtual object. By obtaining the parallel data of the sample speech and the sample video including the virtual object through the sample video processing model, it is no longer necessary for an animator to construct the facial expressions of the virtual objects in the sample video frame by frame, thereby improving the efficiency of obtaining sample videos including the facial expressions of the virtual objects. These data are used to train a video generation model for generating videos based on speech, thereby improving the efficiency of obtaining training data for the video generation model, thereby improving the training efficiency of the video generation model, and since the cost of obtaining training data for the video generation model is reduced, the training cost of the video generation model is reduced.

[0114] Refer to Figure 4, which is a flowchart of a training method for a sample video processing model provided in an embodiment of the present application. The method is executed by a computer device and includes the following steps.

[0115] 401. A computer device obtains at least two third sample pairs, each third sample pair including a second sample video frame and a third sample video frame, an object in the second sample video frame corresponds to a virtual object in the third sample video frame, and the object in the second sample video frame and the virtual object in the third sample video frame have the same facial expression.

[0116] In this embodiment of the present application, the third sample pair includes two sample video frames. The sample video frame including the object in the third sample pair is referred to as the second sample video frame, and the sample video frame including the virtual object in the third sample pair is referred to as the third sample video frame. The second sample video frame in the third sample pair can be the first sample video frame in the second sample video. That is, at least two third sample pairs can be derived from the at least two second sample pairs.

[0117] In some embodiments, the second sample video frame in the third sample pair can be an enhanced sample video frame or an original sample video frame. Illustratively, the computer device obtains at least two third sample videos, which are, for example, sample videos from the at least two second sample videos. Accordingly, the process of the computer device obtaining at least two third sample pairs includes the following steps:

[0118] The computer device enhances at least two fourth sample video frames included in each of the at least two third sample videos to obtain at least two enhanced fourth sample video frames, wherein the facial expression of the subject in each enhanced fourth sample video frame is the same as that in the fourth sample video frame before enhancement; and determines at least two third sample pairs based on the at least two fourth sample video frames included in each of the at least two third sample videos, the sample video frames including the virtual object corresponding to the at least two fourth sample video frames, the at least two enhanced fourth sample video frames, and the sample video frames including the virtual object corresponding to the at least two enhanced fourth sample video frames. The second sample video frame included in each third sample pair is the fourth sample video frame before enhancement or the fourth sample video frame after enhancement, and each enhanced fourth sample video frame has the same sample video frame data as that corresponding to the fourth sample video frame before enhancement.

[0119] This process means that after the computer device obtains at least two third sample videos, it can enhance the sample video frames in the third sample videos, and use the enhanced sample video frames as the second sample video frames in the third sample pair to participate in the training process of the sample video processing model, or it can use the sample video frames before enhancement (that is, the original sample video frames) as the second sample video frames in the third sample pair to participate in the training process of the sample video processing model.

[0120] In an embodiment of the present application, by enhancing the sample video frames included in each of at least two third sample videos, the diversity of samples is increased, and the sample video processing model is trained based on these multiple types of samples, which can improve the generalization of the sample video processing model.

[0121] In some embodiments, the computer device enhances the at least two fourth sample video frames included in each of the at least two third sample videos to obtain at least two enhanced fourth sample video frames, including at least one of the following: the computer device rotates the at least two fourth sample video frames included in each of the at least two third sample videos to obtain at least two enhanced fourth sample video frames; the computer device grayscales the at least two fourth sample video frames included in each of the at least two third sample videos to obtain at least two enhanced fourth sample video frames.

[0122] In some embodiments, rotating the sample video frame refers to transforming the coordinates of at least two pixels in the sample video frame through a transformation matrix. The coordinates of the pixel points are multiplied by the transformation matrix to obtain the transformed coordinates of the pixel points, and the sample video frame corresponding to the transformed coordinates of the at least two pixel points is the enhanced sample video frame. Alternatively, with a certain point in the sample video frame as the rotation center, at least two pixels in the sample video frame are rotated around the point by a preset angle. The rotation center and the preset angle can be set and changed as needed and are not specifically limited here. The computer device can rotate the sample video frame based on at least two rotation centers and at least two preset angles to obtain at least two enhanced sample video frames to further enhance the diversity of the sample.

[0123] It should be noted that the aforementioned methods for enhancing the sample video frames can be freely combined, such as grayscale conversion of the rotated sample video frames to obtain enhanced sample video frames. The aforementioned methods are merely optional implementations of enhancing the sample video frames. The computer device can also enhance the sample video frames using other optional implementations, which will not be detailed here.

[0124] In an embodiment of the present application, by rotating or graying the sample video frame, the sample is expanded without changing the facial expression of the object in the sample video frame, and the above-mentioned enhancement methods are highly convenient, thereby improving the efficiency of enhancing the sample video frame.

[0125] For example, see Figure 10, which is a schematic diagram of enhancing a sample video frame according to an embodiment of the present application. The original sample video frame is rotated and grayscaled to obtain an enhanced sample video frame. The original sample video frame is a sample video frame that has not been grayscaled.

[0126] 402. The computer device determines, through a sample video processing model, the processed second sample video frames corresponding to the at least two third sample pairs based on the respective second sample video frames of the at least two third sample pairs.

[0127] In some embodiments, the computer device processes the second sample video frame in each third sample pair using the sample video processing model to obtain a processed second sample video frame, including the following steps: the computer device processes the second sample video frame in each third sample pair using the feature extraction module in the sample video processing model to obtain video frame features of the second sample video frame, wherein the feature extraction module is used to extract video frame features. The computer device processes the video frame features using the regression module in the sample video processing model to obtain a processed second sample video frame, wherein the regression module is used to generate a processed video frame based on the video frame features.

[0128] The implementation of step 402 is described below with reference to formulas (6) to (8) and taking the second sample video frame in the third sample pair as the enhanced fourth sample video frame as an example.

[0129] Schematically, the third sample video is represented as X={x 1 ,…,x M}, M represents the number of fourth sample video frames included in the third sample video. M Represents the fourth sample video frame of the Mth frame.

[0130] The computer device enhances the fourth sample video frame in the third sample video to obtain an enhanced fourth sample video frame, which is also the second sample video frame in the third sample pair. This process can be expressed by the following formula (6). img (x) (6).

[0131] Wherein, x′ represents the fourth sample video frame after enhancement, that is, the second sample video frame in the third sample pair, x represents the fourth sample video frame before enhancement, aug img (x) indicates that the fourth sample video frame x is enhanced.

[0132] Optionally, the feature extraction module is a ResNet network. The second sample video frame is input into the feature extraction module to obtain the video frame features. This process can be expressed by the following formula (7). img =ResNet(x′) (7).

[0133] Among them, h img Represents the video frame features, and ResNet(x′) represents the features of the second sample video frame x′ extracted through the ResNet network.

[0134] In some embodiments, the input video frame of the ResNet network has a preset resolution, such as a preset resolution of 256×256. Therefore, before the second sample video frame is input into the ResNet network, the second sample video frame is cropped into a sample video frame with a resolution of 256×256.

[0135] Optionally, the regression module is a fully connected layer, which is used to perform nonlinear transformation on the video frame features to obtain the processed sample video frame. This process can be implemented by the following formula (8). img ) (8).

[0136] Among them, y′ represents the processed sample video frame, Linear(h img ) represents the video frame feature himg Perform nonlinear transformation.

[0137] 403. The computer device trains a sample video processing model based on the processed second sample video frames corresponding to the at least two third sample pairs and the third sample video frames corresponding to the at least two third sample pairs.

[0138] In some embodiments, the computer device determines, for each third sample pair, a loss value between the processed second sample video frame and the third sample video frame of the third sample pair, and adjusts model parameters of the sample video processing model based on the loss value.

[0139] The computer device iteratively trains the sample video processing model based on the loss value between the third sample video frame and the processed second sample video frame of each of at least two third sample pairs.

[0140] In one implementation, during each iteration, the computer device adjusts the model parameters of the sample video processing model based on the loss value of the third sample video frame and the processed second sample video frame of a third sample pair.

[0141] For example, the loss value is an L2 loss value, and the process can be expressed by the following formula (9): L = ||yy′|| (9).

[0142] In formula (9), L represents the loss value, y represents the third sample video frame, and y′ represents the processed second sample video frame. The third sample video frame and the processed second sample video frame are two vectors, each having the same dimension. ||yy′|| represents the calculation of the minimum square error between y and y′. The computer device determines the difference between each element in the processed second sample video frame and the corresponding element in the third sample video frame, determines the average of the squares of at least two of the differences, and obtains the loss value.

[0143] In another implementation, during each iterative training process, the computer device adjusts the model parameters of the sample video processing model based on the loss values ​​of some third sample pairs among the at least two third sample pairs. For example, the model parameters of the sample video processing model are adjusted based on the loss values ​​of at least two third sample pairs corresponding to the same sample video. The computer device adjusts the model parameters of the sample video processing model based on the average loss value between the loss values ​​of the at least two third sample pairs. Alternatively, the computer device determines the difference between the predicted sample video frame data and the sample video frame data of each of the at least two third sample pairs, determines the average of the squares of the at least two differences, obtains the loss value, that is, the obtained loss value is the L2 loss value, and adjusts the model parameters of the sample video processing model based on the loss value.

[0144] In the next iteration, the processed second sample video frame of the third sample pair used in this iteration is predicted based on the adjusted sample video processing model, and the model parameters of the sample video processing model are adjusted based on the loss value between the processed second sample video frame and the third sample video frame.

[0145] For example, see Figure 11, which is a training flow chart of a sample video processing model provided in an embodiment of the present application. The original sample video frame containing the object is first enhanced to obtain an enhanced sample video frame. The enhanced sample video frame is then processed by the ResNet network in the sample video processing model to obtain video frame features. The video frame features are then processed by the regression module in the sample video processing model to obtain a sample video frame containing the virtual object.

[0146] For example, refer to Figure 12, which is a flow chart for obtaining training data provided by an embodiment of the present application. In it, an object performs a performance, and the object makes facial expressions while uttering speech. Then, a sample speech is obtained by recording, and an original sample video is obtained by video recording. Then, the sample speech is aligned with the original sample video frame by frame to obtain at least two original sample video frames in the original sample video. Finally, an animator produces a sample video frame containing a virtual object corresponding to each original sample video frame, and obtains a processed sample video corresponding to the original sample video, and the processed sample video contains a virtual object. Since the sample speech is aligned with the original sample video frame by frame, and the original sample video is aligned with the processed sample video frame by frame, the sample speech is aligned with the processed sample video frame by frame. Accordingly, the computer device can also use the sample speech and the processed sample video as a group of first sample pairs to improve data utilization.

[0147] In some embodiments, a computer device obtains at least two initial sample pairs, which include sample speech and sample videos containing objects, such as at least two initial sample pairs including 100 sample speech with an average duration of 4 seconds and corresponding sample videos. Then, sample videos from some of the initial sample pairs are randomly selected to produce a small number of sample video frames containing virtual objects, thereby obtaining at least two third sample pairs. For example, 10 sample videos are selected to produce sample video frames containing virtual objects. Since each sample video containing an object includes at least two sample video frames, a large number of third sample pairs can be obtained. After training a sample video processing model based on at least two third sample pairs, the remaining initial sample pairs are processed based on the sample video processing model to obtain some first sample pairs. The previously produced sample video frames containing virtual objects are aligned with the sample speech frame by frame to obtain some first sample pairs. Both of these first sample pairs are used as training data for the video generation model, thereby achieving data reuse.

[0148] It should be noted that the execution entity of the training video generation model and the execution entity of the training sample video processing model can be the same or different, and there is no limitation here.

[0149] In an embodiment of the present application, the training process of the sample video processing model is implemented through the above steps 401-403. In this embodiment, the sample video processing model is trained based on the sample video frames containing objects and the sample video frames containing virtual objects, so that the sample video processing model can automatically generate the facial expressions of the virtual objects in the processed sample video frames based on the facial expressions of the real objects in the sample video frames. Therefore, when obtaining a training sample pair of sample voice-virtual sample video (that is, a sample video including a virtual object), based on a real video synchronized with the voice content, the real video is processed by the sample video processing model, and a virtual sample video corresponding to the real video can be obtained. Since the voice is synchronized with the real video, the virtual sample video obtained is synchronized with the voice. Obviously, the efficiency of obtaining the training sample pair of voice-virtual sample video is improved by the sample video processing model.

[0150] FIG13 is a block diagram of a training device for a video generation model according to an embodiment of the present application. Referring to FIG13 , the device includes:

[0151] An acquisition module 1301 is configured to acquire at least two first sample pairs, each first sample pair comprising a first sample voice and a first sample video corresponding to the first sample voice, wherein the facial expression of the virtual object in the first sample video matches the voice content of the first sample voice;

[0152] A processing module 1302 is configured to determine, using a video generation model, predicted videos corresponding to the first sample speech of each of the at least two first sample pairs based on the first sample speech of each of the at least two first sample pairs, where the video generation model is configured to generate video based on speech;

[0153] The training module 1303 is configured to train the video generation model based on the predicted video and the first sample video corresponding to the respective first sample voices of at least two first sample pairs.

[0154] In some embodiments, the acquisition module 1301 is configured to:

[0155] Obtain at least two second sample pairs, each second sample pair including a second sample voice and a second sample video corresponding to the second sample voice, wherein the facial expression of the subject in the second sample video matches the voice content of the second sample voice;

[0156] Determining, by a sample video processing model, processed second sample videos corresponding to the at least two second sample pairs based on respective second sample videos of the at least two second sample pairs, wherein a facial expression of the virtual object in the processed second sample videos matches the facial expression of the object in the second sample videos before processing, and the sample video processing model is used to generate a sample video including the virtual object based on the sample video including the object;

[0157] At least two first sample pairs are determined based on the respective second sample voices of the at least two second sample pairs and the processed second sample videos corresponding to the at least two second sample pairs.

[0158] In some embodiments, the acquisition module 1301 is configured to:

[0159] For each second sample pair, determining, by the sample video processing model, based on at least two first sample video frames in the second sample video of the second sample pair, processed first sample video frames corresponding to each of the at least two first sample video frames;

[0160] Based on the processed first sample video frames respectively corresponding to the at least two second sample pairs, the processed second sample videos respectively corresponding to the at least two second sample pairs are determined.

[0161] In some embodiments, the acquisition module 1301 is configured to:

[0162] enhancing the respective second sample speech of the at least two second samples to obtain at least two enhanced second sample speech pieces, wherein the speech content of each enhanced second sample speech piece is the same as that of the second sample speech piece before the enhancement;

[0163] Determining at least two first sample pairs based on the second sample voices of the at least two second sample pairs, the at least two enhanced second sample voices, and the processed second sample videos corresponding to the at least two second sample pairs;

[0164] The first sample speech in each first sample pair is the second sample speech before enhancement or the second sample speech after enhancement, and each second sample speech after enhancement is the same as the second sample video corresponding to the second sample speech before enhancement.

[0165] In some embodiments, the acquisition module 1301 is configured to perform at least one of the following:

[0166] Adjusting the pitch of the second sample speech of each of the at least two second sample pairs respectively to obtain at least two enhanced second sample speech pairs;

[0167] adding reverberation to the second sample speech of each of the at least two second sample pairs to obtain at least two enhanced second sample speech;

[0168] Noise is added to the second sample speech of each of the at least two second sample pairs to obtain at least two enhanced second sample speech.

[0169] In some embodiments, the training module 1303 is further configured to:

[0170] Acquire at least two third sample pairs, each third sample pair comprising a second sample video frame and a third sample video frame, wherein an object in the second sample video frame corresponds to a virtual object in the third sample video frame, and the object in the second sample video frame and the virtual object in the third sample video frame have the same facial expression;

[0171] Determining, by the sample video processing model, the processed second sample video frames corresponding to the at least two third sample pairs based on the respective second sample video frames of the at least two third sample pairs;

[0172] The sample video processing model is trained based on the processed second sample video frames corresponding to the at least two third sample pairs and the third sample video frames corresponding to the at least two third sample pairs.

[0173] In some embodiments, the acquisition module 1301 is further configured to:

[0174] enhancing at least two fourth sample video frames included in each of the at least two third sample videos to obtain at least two enhanced fourth sample video frames, wherein the facial expression of the subject in each enhanced fourth sample video frame is the same as that in the fourth sample video frame before enhancement;

[0175] Determining at least two third sample pairs based on at least two fourth sample video frames and at least two enhanced fourth sample video frames respectively included in the at least two third sample videos;

[0176] The second sample video frame in each third sample pair is the fourth sample video frame before enhancement or the fourth sample video frame after enhancement, and each enhanced fourth sample video frame has the same sample video frame data as the fourth sample video frame before enhancement.

[0177] In some embodiments, the acquisition module 1301 is configured to perform at least one of the following:

[0178] Rotating the at least two fourth sample video frames included in each of the at least two third sample videos to obtain at least two enhanced fourth sample video frames;

[0179] Grayscale is performed on the at least two fourth sample video frames included in each of the at least two third sample videos to obtain at least two enhanced fourth sample video frames.

[0180] In some embodiments, the predicted video includes at least two predicted video frames, the first sample video includes at least two sample video frames, and the at least two predicted video frames correspond one-to-one to the at least two sample video frames. The training module 1303 is configured to:

[0181] For each first sample pair, determining a loss value based on differences between the at least two predicted video frames and the at least two sample video frames respectively;

[0182] Based on the loss values ​​corresponding to the at least two first sample pairs, model parameters of the video generation model are adjusted.

[0183] The embodiment of the present application provides a training device for a video generation model, which trains the video generation model based on sample speech and sample videos corresponding to the sample speech, wherein the facial expressions of the virtual object in the sample video match the speech content of the sample speech, and then continuously trains the video generation model with the sample video as the target, so that the video generation model can gradually learn how to more accurately generate facial expressions that accurately match the speech content based on the speech. The video generation model trained in this way can generate a video synchronized with the speech, and in the video, the facial expressions of the virtual object accurately match the speech content, that is, using the trained video generation model, it is possible to generate facial expressions that accurately match the speech content for the virtual object based on the speech. For example, when the trained video generation model is applied to a virtual human live broadcast scene, it can generate more natural and accurate facial expressions for the virtual human, improving the user experience. In addition, generating facial expressions based on the video generation model can improve the convenience and efficiency of generating facial expressions.

[0184] In the embodiments of the present application, the computer device may be a terminal or a server. When the computer device is a terminal, the terminal serves as the execution subject to implement the technical solution provided in the embodiments of the present application; when the computer device is a server, the server serves as the execution subject to implement the technical solution provided in the embodiments of the present application; or, the technical solution provided in the present application may be implemented through interaction between the terminal and the server, which is not limited in the embodiments of the present application.

[0185] FIG14 shows a structural block diagram of a terminal 1400 provided by an exemplary embodiment of the present application.

[0186] Typically, the terminal 1400 includes a processor 1401 and a memory 1402 .

[0187] The processor 1401 may include one or at least two processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1401 may be implemented in at least one hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). The processor 1401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1401 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1401 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0188] The memory 1402 may include one or at least two computer-readable storage media, which may be non-transitory. The memory 1402 may also include a high-speed random access memory and a non-volatile memory, such as one or at least two disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1402 is used to store at least one program, which is executed by the processor 1401 to implement the training method of the video generation model provided in the method embodiment of the present application.

[0189] In some embodiments, terminal 1400 may optionally include a peripheral device interface 1403 and at least one peripheral device. Processor 1401, memory 1402, and peripheral device interface 1403 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 1403 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1404, a display screen 1405, a camera assembly 1406, an audio circuit 1407, and a power supply 1408.

[0190] The peripheral device interface 1403 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1401 and the memory 1402. In some embodiments, the processor 1401, the memory 1402, and the peripheral device interface 1403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1401, the memory 1402, and the peripheral device interface 1403 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0191] RF circuit 1404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. RF circuit 1404 communicates with communication networks and other communication devices via electromagnetic signals. RF circuit 1404 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. RF circuit 1404 optionally includes an antenna system, an RF transceiver, one or at least two amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. RF circuit 1404 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, RF circuit 1404 may also include circuitry related to Near Field Communication (NFC), although this application does not limit this.

[0192] The display screen 1405 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1405 is a touch screen display, the display screen 1405 also has the ability to collect touch signals on the surface or above the surface of the display screen 1405. The touch signal can be input as a control signal to the processor 1401 for processing. At this time, the display screen 1405 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there can be one display screen 1405, which is set on the front panel of the terminal 1400; in other embodiments, there can be at least two display screens 1405, which are respectively set on different surfaces of the terminal 1400 or in a folding design; in other embodiments, the display screen 1405 can be a flexible display screen, which is set on the curved surface or folding surface of the terminal 1400. Even more, the display screen 1405 can be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 1405 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0193] The camera assembly 1406 is used to capture images or videos. Optionally, the camera assembly 1406 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1406 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0194] The audio circuit 1407 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 1401 for processing, or input into the radio frequency circuit 1404 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be at least two microphones, each located in different parts of the terminal 1400. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 1401 or the radio frequency circuit 1404 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 1407 may also include a headphone jack.

[0195] Power supply 1408 is used to power various components in terminal 1400. Power supply 1408 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1408 includes a rechargeable battery, the rechargeable battery can be wired or wirelessly rechargeable. A wired rechargeable battery is charged via a wired line, while a wireless rechargeable battery is charged via a wireless coil. The rechargeable battery can also support fast charging technology.

[0196] In some embodiments, the terminal 1400 further includes one or at least two sensors 1409 . The one or at least two sensors 1409 include, but are not limited to, an acceleration sensor 1410 , a gyroscope sensor 1411 , a pressure sensor 1412 , an optical sensor 1413 , and a proximity sensor 1414 .

[0197] Accelerometer 1410 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by terminal 1400. For example, accelerometer 1410 can be used to detect the components of gravity acceleration along the three coordinate axes. Processor 1401 can control display screen 1405 to display the user interface in either a landscape or portrait view based on the gravity acceleration signal collected by accelerometer 1410. Accelerometer 1410 can also be used to collect game or user motion data.

[0198] The gyroscope sensor 1411 can detect the orientation and rotation angle of the terminal 1400. It can also work with the accelerometer 1410 to collect the user's 3D movements on the terminal 1400. Based on the data collected by the gyroscope sensor 1411, the processor 1401 can implement the following functions: motion sensing (for example, changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0199] The pressure sensor 1412 can be provided on the side frame of the terminal 1400 and / or below the display screen 1405. When the pressure sensor 1412 is provided on the side frame of the terminal 1400, it can detect the user's gripping signal of the terminal 1400, and the processor 1401 can perform left-hand or right-hand recognition or shortcut operations based on the gripping signal collected by the pressure sensor 1412. When the pressure sensor 1412 is provided below the display screen 1405, the processor 1401 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 1405. Operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0200] Optical sensor 1413 is used to detect ambient light intensity. In one embodiment, processor 1401 can control the display brightness of display screen 1405 based on the ambient light intensity detected by optical sensor 1413. Specifically, when the ambient light intensity is high, the display brightness of display screen 1405 is increased; when the ambient light intensity is low, the display brightness of display screen 1405 is decreased. In another embodiment, processor 1401 can also dynamically adjust the shooting parameters of camera assembly 1406 based on the ambient light intensity detected by optical sensor 1413.

[0201] Proximity sensor 1414, also known as a distance sensor, is typically located on the front panel of terminal 1400. Proximity sensor 1414 is used to detect the distance between the user and the front of terminal 1400. In one embodiment, when proximity sensor 1414 detects that the distance between the user and the front of terminal 1400 is gradually decreasing, processor 1401 controls display screen 1405 to switch from the screen-on state to the screen-off state. When proximity sensor 1414 detects that the distance between the user and the front of terminal 1400 is gradually increasing, processor 1401 controls display screen 1405 to switch from the screen-off state to the screen-on state.

[0202] Those skilled in the art will understand that the structure shown in FIG14 does not constitute a limitation on the terminal 1400 , and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0203] Figure 15 is a schematic diagram of the structure of a server provided according to an embodiment of the present application. The server 1500 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1501 and one or more memories 1502, wherein the memory 1502 is used to store at least one program, and the processor 1501 is configured to execute the at least one program to implement the training method of the video generation model provided by the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server may also include other components for implementing device functions, which will not be described in detail here.

[0204] An embodiment of the present application also provides a computer-readable storage medium, in which at least one program is stored. The at least one program is loaded and executed by a processor to implement a training method for a video generation model of any of the above-mentioned implementation methods.

[0205] An embodiment of the present application also provides a computer program product, which includes at least one program segment, and the at least one program segment is stored in a computer-readable storage medium. The processor of the computer device reads the at least one program segment from the computer-readable storage medium, and the processor executes the at least one program segment, so that the computer device executes the video generation model training method of any of the above-mentioned implementation methods.

[0206] In some embodiments, the computer program product involved in the embodiments of the present application can be deployed and executed on a computer device, or on at least two computer devices located at one location, or on at least two computer devices distributed at at least two locations and interconnected through a communication network. At least two computer devices distributed at at least two locations and interconnected through a communication network can constitute a blockchain system.

[0207] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here. The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for training a video generation model, performed by a computer device, comprising: Obtain at least two first sample pairs, each first sample pair including a first sample voice and a first sample video corresponding to the first sample voice, wherein the facial expression of the virtual object in the first sample video matches the voice content of the first sample voice; Determining, by a video generation model, based on the first sample speech of each of the at least two first sample pairs, a predicted video corresponding to the first sample speech of each of the at least two first sample pairs, wherein the video generation model is used to generate a video based on speech; The video generation model is trained based on the predicted video and the first sample video corresponding to the first sample speech of each of the at least two first sample pairs.

2. The method according to claim 1, wherein The obtaining of at least two first sample pairs comprises: Obtain at least two second sample pairs, each second sample pair including a second sample voice and a second sample video corresponding to the second sample voice, wherein a facial expression of an object in the second sample video matches the voice content of the second sample voice; Determining, by a sample video processing model, based on respective second sample videos of the at least two second sample pairs, processed second sample videos corresponding to the at least two second sample pairs, wherein the facial expression of the virtual object in the processed second sample videos matches the facial expression of the object in the second sample videos before processing, the sample video processing model being used to generate a sample video including the virtual object based on the sample video including the object; The at least two first sample pairs are determined based on the respective second sample voices of the at least two second sample pairs and the processed second sample videos corresponding to the at least two second sample pairs.

3. The method according to claim 2, wherein: The determining, using the sample video processing model, based on the respective second sample videos of the at least two second sample pairs, the processed second sample videos corresponding to the at least two second sample pairs, includes: For each second sample pair, determining, by the sample video processing model, based on at least two first sample video frames in the second sample video of the second sample pair, processed first sample video frames corresponding to each of the at least two first sample video frames; Based on the processed first sample video frames respectively corresponding to the at least two second sample pairs, the processed second sample videos respectively corresponding to the at least two second sample pairs are determined.

4. The method according to claim 2 or 3, wherein: The determining of the at least two first sample pairs based on the respective second sample voices of the at least two second sample pairs and the processed second sample videos corresponding to the at least two second sample pairs includes: Enhance the second sample speech of each of the at least two second sample pairs respectively to obtain at least two enhanced second sample speech pairs, wherein the speech content of each enhanced second sample speech pair is the same as that of the second sample speech before enhancement; Determining the at least two first sample pairs based on the respective second sample voices of the at least two second sample pairs, the at least two enhanced second sample voices, and the processed second sample videos corresponding to the at least two second sample pairs; The first sample speech in each first sample pair is the second sample speech before enhancement or the second sample speech after enhancement, and each second sample speech after enhancement is the same as the second sample video corresponding to the second sample speech before enhancement.

5. The method according to claim 4, wherein The step of enhancing the respective second sample speech of the at least two second sample pairs to obtain at least two enhanced second sample speech pairs includes at least one of the following: respectively adjusting the pitch of the second sample speech of each of the at least two second sample pairs to obtain the at least two enhanced second sample speech; adding reverberation to the second sample speech of each of the at least two second sample pairs to obtain the at least two enhanced second sample speech; Noise is added to the second sample speech of each of the at least two second sample pairs to obtain the at least two enhanced second sample speech.

6. The method according to any one of claims 2 to 5, wherein The training process of the sample video processing model includes: Acquire at least two third sample pairs, each third sample pair comprising a second sample video frame and a third sample video frame, wherein an object in the second sample video frame corresponds to a virtual object in the third sample video frame, and the object in the second sample video frame and the virtual object in the third sample video frame have the same facial expression; Determining, by the sample video processing model, the processed second sample video frames corresponding to the at least two third sample pairs based on the respective second sample video frames of the at least two third sample pairs; The sample video processing model is trained based on the processed second sample video frames corresponding to the at least two third sample pairs and the third sample video frames of the at least two third sample pairs.

7. The method according to claim 6, wherein: The obtaining of at least two third sample pairs comprises: enhancing at least two fourth sample video frames included in each of the at least two third sample videos to obtain at least two enhanced fourth sample video frames, wherein the facial expression of the subject in each enhanced fourth sample video frame is the same as that in the fourth sample video frame before enhancement; Determining the at least two third sample pairs based on the at least two fourth sample video frames and the at least two enhanced fourth sample video frames respectively included in the at least two third sample videos; The second sample video frame in each third sample pair is the fourth sample video frame before enhancement or the fourth sample video frame after enhancement, and each enhanced fourth sample video frame has the same sample video frame data as the fourth sample video frame before enhancement.

8. The method according to claim 7, wherein: The step of enhancing the at least two fourth video frames included in each of the at least two third sample videos to obtain at least two enhanced fourth sample video frames includes at least one of the following: Rotating the at least two fourth sample video frames included in each of the at least two third sample videos to obtain the at least two enhanced fourth sample video frames; Grayscale is performed on the at least two fourth sample video frames included in each of the at least two third sample videos to obtain the at least two enhanced fourth sample video frames.

9. The method according to any one of claims 1 to 8, wherein The predicted video includes at least two predicted video frames, the first sample video includes at least two sample video frames, and the at least two predicted video frames correspond one-to-one to the at least two sample video frames; The step of training the video generation model based on the predicted video and the first sample video corresponding to the first sample speech of each of the at least two first sample pairs includes: For each first sample pair, determining a loss value based on differences between the at least two predicted video frames and the at least two sample video frames respectively; Based on the loss values ​​corresponding to the at least two first sample pairs, the model parameters of the video generation model are adjusted.

10. A training device for a video generation model, comprising: An acquisition module is configured to acquire at least two first sample pairs, each first sample pair comprising a first sample voice and a first sample video corresponding to the first sample voice, wherein the facial expression of the virtual object in the first sample video matches the voice content of the first sample voice; a processing module, configured to determine, by means of a video generation model, a predicted video corresponding to the first sample speech of each of the at least two first sample pairs based on the first sample speech of each of the at least two first sample pairs, wherein the video generation model is configured to generate a video based on speech; A training module is used to train the video generation model based on the predicted video and the first sample video corresponding to the respective first sample voices of the at least two first sample pairs.

11. A computer device comprising a processor and a memory, wherein the memory is used to store at least one program, and the at least one program is loaded by the processor and executes the training method of the video generation model described in any one of claims 1 to 9.

12. A computer-readable storage medium, wherein the computer-readable storage medium is used to store at least one program, and the at least one program is used to execute the training method of the video generation model described in any one of claims 1 to 9.

13. A computer program product, comprising at least one program segment, wherein the at least one program segment is stored in a computer-readable storage medium, a processor of a computer device reads the at least one program segment from the computer-readable storage medium, and the processor executes the at least one program segment, so that the computer device executes the video generation model training method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Animation generation method, device and system and storage medium

    CN111968207A

  • Video generation method and device

    CN111988658A

  • Virtual character expression control method and device, electronic equipment and medium

    CN113633983A

  • Video generation method and device

    CN115209180A

  • Video generation model training method and device, equipment, storage medium and product

    CN117998166A