Video generation method and parameter generation model training method

Through the video generation method based on the diffusion model, the problem of pattern collapse in traditional video generation is solved, and high accuracy and vivid video generation is achieved, ensuring the synchronization of voice and expression.

WO2025123922A1PCT designated stage expired Publication Date: 2025-06-19ALIBABA (CHINA) CO LTD

Patent Information

Application Number
PCT/CN2024/125642
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-10-18
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Pattern collapse is prone to occur during the generation of traditional speaker videos, which makes it difficult to ensure the vividness and accuracy of the video.

Method used

Using a video generation method based on the diffusion model, we generate expression parameters by obtaining the emotional characteristics of the pending voice and target object, and input the object image and expression parameters into the video generation model to generate the target video.

Benefits of technology

Improve the accuracy and vividness of the target video, ensure the synchronization of voice and expressions, and incorporate diverse emotional information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024125642_19062025_PF_FP_ABST
    Figure CN2024125642_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a video generation method and a parameter generation model training method. The video generation method comprises: acquiring speech to be processed; inputting an emotion feature of a target object and said speech into a parameter generation model, so as to obtain an expression parameter, wherein the expression parameter is used for describing facial motion information of the target object under the influence of the emotion feature, the parameter generation model is obtained by means of performing training on the basis of a sample emotion feature, sample speech and an expression parameter label that corresponds to the sample speech, and the sample emotion feature and the sample speech are obtained on the basis of a sample video; and inputting an object image of the target object and the expression parameter into a video generation model, so as to obtain a target video of the target object. By means of generating an expression parameter on the basis of an emotion feature and speech to be processed, and further generating a target video on the basis of the expression parameter, diversified emotion information is fused into the target video while ensuring the synchronization between speech and expression in the target video, thereby improving the accuracy and vividness of the target video.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation method and parameter generation model training method

[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on December 14, 2023, with application number 2023117291605 and application name “Video Generation Method and Parameter Generation Model Training Method”, the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0002] The embodiments of the present disclosure relate to the field of computer technology, and in particular to a video generation method and a parameter generation model training method. Background Art

[0003] With the development of computer technology, speaker video generation has gradually become a research focus. Speaker video generation can analyze and process speech signals to help users create speaking videos, meeting their creative and entertainment needs. It is widely used in animation production, virtual agents, video conferencing, and other multimedia applications.

[0004] However, the traditional speaker video generation process will encounter the problem of pattern collapse, which makes it difficult to ensure the vividness and accuracy of the speaker video. Therefore, there is an urgent need for a vivid and highly accurate video generation solution.

[0005] Summary of the Invention

[0006] In view of this, embodiments of the present disclosure provide a video generation method. One or more embodiments of the present disclosure also involve a parameter generation model training method, a video generation apparatus, a parameter generation model training apparatus, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.

[0007] According to a first aspect of an embodiment of the present disclosure, a video generation method is provided, including:

[0008] Get the voice to be processed;

[0009] Generate a model based on the target object's emotional features and the speech input parameters to be processed to obtain expression parameters, where the expression parameters are used to describe the target object's facial movement information under the influence of the emotional features. The parameter generation model is trained based on sample emotional features, sample speech, and expression parameter labels corresponding to the sample speech. The sample emotional features and sample speech are obtained based on sample videos.

[0010] The object image and expression parameters of the target object are input into the video generation model to obtain the target video of the target object.

[0011] According to a second aspect of an embodiment of the present disclosure, a video generation method is provided, including:

[0012] Receive a video generation request sent by a user, wherein the video generation request carries the voice to be processed;

[0013] Generate a model based on the target object's emotional features and the speech input parameters to be processed to obtain expression parameters, where the expression parameters are used to describe the target object's facial movement information under the influence of the emotional features. The parameter generation model is trained based on sample emotional features, sample speech, and expression parameter labels corresponding to the sample speech. The sample emotional features and sample speech are obtained based on sample videos.

[0014] Inputting the object image and expression parameters of the target object into the video generation model to obtain the target video corresponding to the video generation request;

[0015] Send the target video corresponding to the video generation request to the user.

[0016] According to a third aspect of an embodiment of the present disclosure, a parameter generation model training method is provided, which is applied to a cloud-side device, including:

[0017] Acquire a plurality of sample videos including sample objects;

[0018] Extract sample speech and expression parameter labels corresponding to the sample speech from the sample video;

[0019] Inputting the sample emotion features and sample speech of the sample object into the initial parameter generation model to obtain the predicted expression parameters;

[0020] According to the predicted expression parameters and expression parameter labels, the model parameters of the initial parameter generation model are adjusted to obtain a trained parameter generation model.

[0021] According to a fourth aspect of the embodiments of the present disclosure, there is provided a video generating apparatus, including:

[0022] A first acquisition module is configured to acquire the speech to be processed;

[0023] A first input module is configured to generate a model based on the target object's emotional characteristics and the speech to be processed as input parameters to obtain expression parameters, wherein the expression parameters are used to describe the target object's facial movement information under the influence of the emotional characteristics, and the parameter generation model is trained based on sample emotional characteristics, sample speech, and expression parameter labels corresponding to the sample speech, and the sample emotional characteristics and sample speech are obtained based on sample videos;

[0024] The second input module is configured to input the object image and expression parameters of the target object into the video generation model to obtain the target video of the target object.

[0025] According to a fifth aspect of the embodiments of the present disclosure, there is provided a video generating apparatus, including:

[0026] A first receiving module is configured to receive a video generation request sent by a user, wherein the video generation request carries a voice to be processed;

[0027] A third input module is configured to generate a model based on the target object's emotional characteristics and the speech to be processed input parameters to obtain expression parameters, wherein the expression parameters are used to describe the target object's facial movement information under the influence of the emotional characteristics, and the parameter generation model is trained based on sample emotional characteristics, sample speech, and expression parameter labels corresponding to the sample speech, and the sample emotional characteristics and sample speech are obtained based on the sample video;

[0028] A fourth input module is configured to input the object image and expression parameters of the target object into the video generation model to obtain a target video corresponding to the video generation request;

[0029] The sending module is configured to send the target video corresponding to the video generation request to the user.

[0030] According to a sixth aspect of an embodiment of the present disclosure, a parameter generation model training apparatus is provided, which is applied to a cloud-side device and includes:

[0031] A second acquisition module is configured to acquire a plurality of sample videos including sample objects;

[0032] An extraction module is configured to extract sample speech and expression parameter labels corresponding to the sample speech from the sample video;

[0033] A fifth input module is configured to input the sample emotion features and sample speech of the sample object into the initial parameter generation model to obtain predicted expression parameters;

[0034] The adjustment module is configured to adjust the model parameters of the initial parameter generation model according to the predicted expression parameters and the expression parameter labels to obtain a trained parameter generation model.

[0035] According to a seventh aspect of an embodiment of the present disclosure, there is provided a computing device, including:

[0036] memory and processor;

[0037] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method provided in the first aspect, the second aspect, or the third aspect are implemented.

[0038] According to an eighth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the method provided in the first aspect, the second aspect or the third aspect are implemented.

[0039] According to a ninth aspect of an embodiment of the present disclosure, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the method provided in the first aspect, the second aspect, or the third aspect above.

[0040] One embodiment of the present disclosure provides a video generation method that obtains a speech to be processed; inputs the emotional characteristics of a target subject and the speech to be processed into a parameter generation model to obtain expression parameters, wherein the expression parameters are used to describe the facial movement information of the target subject under the influence of the emotional characteristics. The parameter generation model is trained based on sample emotional characteristics, sample speech, and expression parameter labels corresponding to the sample speech, wherein the sample emotional characteristics and sample speech are obtained based on sample videos; and inputs the object image and expression parameters of the target subject into the video generation model to obtain a target video of the target subject. By generating expression parameters based on the emotional characteristics and the speech to be processed, and further generating a target video based on the expression parameters, the target video is integrated with diverse emotional information while ensuring the synchronization of speech and expression in the target video, thereby improving the accuracy and vividness of the target video. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] FIG1 is an architecture diagram of a video generation system provided by one embodiment of the present disclosure;

[0042] FIG2 is an architecture diagram of another video generation system provided by an embodiment of the present disclosure;

[0043] FIG3 is a flow chart of a video generation method provided by one embodiment of the present disclosure;

[0044] FIG4 is a flowchart of another video generation method provided by an embodiment of the present disclosure;

[0045] FIG5 is a flow chart of a parameter generation model training method provided by one embodiment of the present disclosure;

[0046] FIG6 is a flowchart of a processing process of a video generation method provided by an embodiment of the present disclosure;

[0047] FIG7 is a schematic diagram of a video generation interface provided by an embodiment of the present disclosure;

[0048] FIG8 is a schematic structural diagram of a video generating device provided by an embodiment of the present disclosure;

[0049] FIG9 is a schematic structural diagram of another video generating device provided by an embodiment of the present disclosure;

[0050] FIG10 is a schematic structural diagram of a parameter generation model training device provided by one embodiment of the present disclosure;

[0051] FIG11 is a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0052] The following description sets forth many specific details to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.

[0053] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.

[0054] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0055] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0056] First, the terms involved in one or more embodiments of the present disclosure are explained.

[0057] Diffusion Model (DM): The diffusion model is an advanced generative model that focuses on generating high-quality samples and shows great potential in complex generation tasks. The diffusion model is essentially based on probability distribution and uses a series of iterative steps to gradually construct samples. The core mechanism of the diffusion model draws on Markov Chain Monte Carlo (MCMC) sampling, predicting the next state based on the current state in each iterative step. Through continuous local transformations, the diffusion model can gradually "diffuse" a simple initial state to a complex sample that conforms to the target probability distribution. More specifically, the diffusion model defines a transition probability function that guides the evolution from the current state to the next state. In the process of continuous iteration, the diffusion model proposes new candidate samples based on the current state and continuously adjusts these samples to make them closer to the desired target probability distribution.

[0058] Generative Adversarial Networks (GANs): Generative Adversarial Networks (GANs) are powerful deep learning models that consist of two parts: a generator and a discriminator, which compete with each other during training. The goal of the generator is to create realistic data samples, while the discriminator attempts to distinguish these generated samples from real data samples. This setting forms a dynamic adversarial process that encourages the generator to continuously improve the quality of the generated data. However, GANs suffer from a problem called mode collapse, which means that during training, the generator begins to generate highly similar or repetitive samples, failing to capture the diversity of the data distribution, resulting in limited diversity and quality of the generated data.

[0059] Emotional Speaker Video Generation: Emotional speaker video generation is a deep learning technology that generates synchronized facial animations by analyzing and processing input speech signals. It also supports controlling facial emotions using emotional reference data such as video and text. This technology uses a complex neural network model to extract audio features and translate them into accurate facial movements, including lip synchronization and other relevant facial expressions. This creates virtual speaker videos with accurate speech patterns and rich expressions, providing a more natural and realistic audiovisual experience. It is widely used in animation production, virtual agents, video conferencing, and other multimedia applications.

[0060] Expression parameters: In the fields of computer vision and graphics, 3D Morphable Modeling (3DMM) is a 3D model used to accurately capture and reconstruct human faces. Expression parameters are a set of parameters within this model that describe and control the expressive changes of a specific face. By adjusting these expression parameters, a variety of facial movements, such as smiling, frowning, and blinking, can be simulated, resulting in a 3D face model with rich expressions. The advantage of these parameters is that they can encode the complex variations of facial expressions in a relatively low-dimensional and efficient manner, making them a key technology for achieving facial animation and emotional expression simulation.

[0061] Self-attention pooling layer: The self-attention pooling layer can also be called a self-attention pooling unit. A self-attention pooling layer is a deep learning network layer that combines an attention mechanism with a pooling operation to extract and enhance important feature information in a neural network. This layer uses the self-attention mechanism to allow the network to focus more on key parts of the input features when performing dimensionality reduction pooling operations. Specifically, the self-attention pooling layer assigns a weight to each input feature, reflecting its importance to the task. It then performs weighted pooling based on these weights, ensuring that useful information is retained while reducing the data dimension.

[0062] Transformer: The Transformer architecture is a deep learning model primarily used for natural language processing (NLP) tasks. It is widely used in various tasks involving text and sequence data, such as text classification, question-answering systems, and summary generation. The core idea of ​​the Transformer is to use a self-attention mechanism to process input sequence data.

[0063] Emotional speaker video generation can help users create vivid speaking videos, satisfying their creative and entertainment needs. Currently, emotional speaker video generation is typically based on generative adversarial networks (GANs). However, because GANs are susceptible to mode collapse, these methods often struggle to ensure natural and vivid expressions and accurate mouth shapes when representing diverse emotions.

[0064] To address the aforementioned issues, the present disclosure leverages the diffusion model's superior data distribution learning capabilities and its superiority in learning speech-to-facial motion mapping across a wide range of emotional expressions. This approach attempts to utilize the diffusion model to generate emotional speaking videos. Specifically, a diffusion model-based emotional speaking video generation scheme is proposed. First, the diffusion model is used to generate expression parameters describing facial motion. A video generation model is then used to render these expression parameters into a speaking video. Specifically, a speech to be processed is obtained. The target subject's emotional features and the speech to be processed are input into a parameter generation model to obtain expression parameters, which describe the target subject's facial motion under the influence of the emotional features. The parameter generation model is trained based on sample emotional features, sample speech, and corresponding expression parameter labels. The sample emotional features and sample speech are obtained based on sample videos. Finally, an object image of the target subject and the expression parameters are input into the video generation model to obtain a target video of the target subject. By leveraging the diffusion model's powerful distributed learning capabilities, this approach generates more vivid emotions and mouth shapes that are more synchronized with the input speech, improving the accuracy and vividness of the target video.

[0065] In the present disclosure, a video generation method is provided. The present disclosure also relates to a parameter generation model training method, a video generation device, a parameter generation model training device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0066] Referring to FIG1 , FIG1 shows an architecture diagram of a video generation system provided by an embodiment of the present disclosure. The video generation system may include a client 100 and a server 200;

[0067] The client 100 is used to send the voice to be processed to the server 200;

[0068] The server 200 is configured to generate a model using the target object's emotional features and the speech to be processed as input parameters to obtain expression parameters, wherein the expression parameters are used to describe the target object's facial movement information under the influence of the emotional features. The parameter generation model is trained based on sample emotional features, sample speech, and expression parameter labels corresponding to the sample speech. The sample emotional features and sample speech are obtained based on sample videos. The target object's object image and expression parameters are input into the video generation model to obtain a target video of the target object. The target video is then sent to the client 100.

[0069] The client 100 is further configured to receive the target video sent by the server 200 .

[0070] By applying the solution of the embodiment of the present disclosure, expression parameters are generated based on emotional features and the speech to be processed, and a target video is further generated based on the expression parameters. While ensuring the synchronization of speech and expression in the target video, diverse emotional information is incorporated into the target video, thereby improving the accuracy and vividness of the target video.

[0071] Referring to FIG. 2 , FIG. 2 shows an architecture diagram of another video generation system provided by one embodiment of the present disclosure. The video generation system may include a server 200 and multiple clients 100. The clients 100 may include end-side devices, and the server 200 may include cloud-side devices. Multiple clients 100 may establish communication connections through the server 200. In a video generation scenario, the server 200 is used to provide video generation services between the multiple clients 100. The multiple clients 100 may act as senders or receivers, respectively, and communicate through the server 200.

[0072] Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the video generation scenario, the user can publish a data stream to the server 200 through the client 100, and the server 200 can generate a target video based on the data stream and push the target video to other clients with which communication has been established.

[0073] The client 100 and the server 200 are connected via a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, or other processing before being released to the server 200.

[0074] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language 5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application. The client 100 can be based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device to run or certain APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0075] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that provide background training to support models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server that is integrated with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0076] It is worth noting that the video generation method provided in the embodiments of the present disclosure is generally executed by the server. However, in other embodiments of the present disclosure, the client may also have similar functions to the server to execute the video generation method provided in the embodiments of the present disclosure. In other embodiments, the video generation method provided in the embodiments of the present disclosure may also be executed jointly by the client and the server.

[0077] 3 , which shows a flow chart of a video generation method provided by an embodiment of the present disclosure, specifically comprising the following steps:

[0078] Step 302: Acquire the speech to be processed.

[0079] In one or more embodiments of the present disclosure, a speech to be processed may be acquired, and the speech to be processed may be processed to generate a target video including the speech to be processed and a target object.

[0080] Specifically, the speech to be processed refers to audio data containing natural language. The speech to be processed can be speech in various scenarios, such as conference speech or music audio from a concert. Natural languages ​​in the speech to be processed include, but are not limited to, Chinese and English. For example, the speech to be processed can be audio data containing the sentence "I got first place on the exam."

[0081] In practical applications, there are multiple ways to obtain the speech to be processed, and the method to be selected depends on the actual situation. The embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, the speech to be processed can be received from the user through the client. In another possible implementation of the present disclosure, the speech to be processed can be read from other data acquisition devices or databases.

[0082] Step 304: Generate a model based on the emotional features of the target object and the input parameters of the speech to be processed to obtain expression parameters, wherein the expression parameters are used to describe the facial movement information of the target object under the influence of the emotional features. The parameter generation model is trained based on sample emotional features, sample speech, and expression parameter labels corresponding to the sample speech. The sample emotional features and sample speech are obtained based on the sample video.

[0083] In one or more embodiments of the present disclosure, after obtaining the speech to be processed, further, the emotional characteristics of the target object and the input parameters of the speech to be processed can be used to generate a model to obtain expression parameters.

[0084] It should be noted that since the processed speech usually includes the speaker's emotional information, for example, when happy, the speaker may raise their eyebrows and smile, while when sad, they may frown and pout, etc. Therefore, in order to ensure that the expression parameters can more vividly and naturally reflect the speaker's facial movement information, the target subject's emotional characteristics can be incorporated into the generation of expression parameters.

[0085] Specifically, the target object is the speaker that appears in the target video. The target object can be a person in real life or a character in a virtual scene. The target object can be the speaker of the speech to be processed, or it can be another object other than the speaker of the speech to be processed. For example, if the speaker of the speech to be processed is A, we can replace the speaker with target object B and generate a target video consisting of target object B and the speech to be processed. The emotional features of the target object are used to express the emotional information of the target object. The emotional features are obtained by extracting features from emotional reference data. Emotional reference data includes but is not limited to emotional reference speech, emotional reference images, emotional reference text, and emotional reference videos.

[0086] In practical applications, the parameter generation model can be a neural network model trained based on sample emotion features, sample speech, and the corresponding expression parameter labels. Alternatively, it can be a diffusion model trained based on sample emotion features, sample speech, and the corresponding expression parameter labels. This can be a speech-driven facial motion generator based on the diffusion model, which generates expression parameters using the diffusion model's sampling process.

[0087] In an optional embodiment of the present disclosure, the parameter generation model includes an encoding unit and a decoding unit; the above-mentioned generation model of the target object's emotional characteristics and the speech input parameters to be processed to obtain the expression parameters may include the following steps:

[0088] The speech to be processed is encoded by the encoding unit to obtain speech features;

[0089] After the decoding unit, the preset noise, emotional features and voice features are diffused to obtain expression parameters.

[0090] It should be noted that in the parameter generation model, the Transformer encoding unit can be used to encode the processed speech to obtain speech features. The speech features, preset noise, and emotion features are then input into the Transformer decoding unit, which outputs the denoised expression parameters.

[0091] Using the solution of the disclosed embodiment, the encoding unit encodes the processed speech to obtain speech features. The decoding unit then diffuses the preset noise, emotion features, and speech features to obtain expression parameters. Because the emotion features are incorporated into the expression parameter generation process, more realistic and vivid facial movement information is obtained, further improving the accuracy and vividness of the target video.

[0092] In practical applications, the emotional characteristics of the target object and the speech input parameters to be processed are used to generate a model. Before obtaining the expression parameters, the emotional characteristics of the target object can be obtained. There are multiple ways to obtain the emotional characteristics of the target object, and the specific method to be used depends on the actual situation. The embodiments of this disclosure do not impose any restrictions on this. In one possible implementation of this disclosure, the emotional characteristics of the target object can be extracted from multiple pre-generated emotional characteristics. For example, if the user specifies sadness as the emotion, the emotional characteristics corresponding to sadness are extracted as the emotional characteristics of the target object.

[0093] In another possible implementation of the present disclosure, the emotional characteristics of the target object can be generated using the emotional reference data. That is, before the above-mentioned generation of the model using the emotional characteristics of the target object and the speech input parameters to be processed and obtaining the expression parameters, the following steps may be further included:

[0094] Acquiring emotion reference data, wherein the emotion reference data includes emotion information;

[0095] Feature extraction is performed on the emotional reference data to obtain the emotional characteristics of the target object.

[0096] Specifically, emotion reference data refers to data that includes emotion information. There are multiple ways to obtain emotion reference data, and the specific method to be selected depends on the actual situation. The embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, emotion reference data sent by a user through a client can be received. In another possible implementation of the present disclosure, emotion reference data can be read from other data acquisition devices or databases.

[0097] In practical applications, when extracting features from emotional reference data, a corresponding feature extraction method can be selected based on the modality of the emotional reference data to obtain the target emotional features. For example, if the emotional reference data is an emotional reference video, the second feature extraction model can be used to extract features from the emotional reference video to obtain the emotional features of the target object.

[0098] In an optional embodiment of the present disclosure, the emotion reference data includes an emotion reference video; and the feature extraction of the emotion reference data to obtain the emotion features of the target object may include the following steps:

[0099] The emotional reference video is input into the second feature extraction model to obtain the emotional features of the target object.

[0100] Specifically, the second feature extraction model can be understood as a video emotion encoding model. The second feature extraction model takes the emotion reference video as input and extracts the emotion features contained in the emotion reference video.

[0101] By applying the solution of the embodiment of the present disclosure, the emotion reference video is input into the second feature extraction model to obtain the emotion features of the target object, thereby achieving the goal of obtaining the emotion features through the additional emotion reference video.

[0102] In another optional embodiment of the present disclosure, the emotion reference data includes an emotion reference image; and the feature extraction of the emotion reference data to obtain the emotion features of the target object may include the following steps:

[0103] The emotional reference image and the speech to be processed are input into the first feature extraction model to obtain the emotional features of the target object.

[0104] Specifically, the first feature extraction model can be understood as a speech emotion prediction model. It is a diffusion model based on a Transformer encoder, used to predict emotional features through the diffusion model's sampling process. The first feature extraction model takes as input an emotional reference image, the speech to be processed, and preset noise, and outputs the denoised emotional features.

[0105] It should be noted that the emotion reference image can be an object image of the target object. Since using the emotion reference video to generate the target object's emotional features during video generation relies heavily on the need for additional emotion reference videos, to reduce this dependence, the disclosed embodiments propose a first feature extraction model that directly uses the processed speech to predict the target object's emotional features, thereby avoiding the need for additional emotion reference videos.

[0106] Furthermore, during the video generation process, the data input into the video generation model includes the object image of the target object, and the portrait information in the object image can help the first feature extraction model predict the emotional characteristics that match the target object. Therefore, the emotional reference image, the speech to be processed and the preset noise can be input into the first feature extraction model to obtain the emotional characteristics of the target object.

[0107] By applying the solution of the embodiment of the present disclosure, the emotion reference image and the speech to be processed are input into the first feature extraction model to obtain the emotion characteristics of the target object, thereby reducing the dependence on additional emotion reference videos.

[0108] In an optional embodiment of the present disclosure, the second feature extraction model includes a sequence extraction unit, a sequence encoding unit, and a self-attention pooling unit; the above-mentioned inputting the emotion reference video into the second feature extraction model to obtain the emotion features of the target object may include the following steps:

[0109] Extracting an expression parameter reference sequence from the emotion reference video via a sequence extraction unit, wherein the expression parameter reference sequence includes reference expression parameters of each video frame in the emotion reference video;

[0110] The expression parameter reference sequence is encoded by the sequence encoding unit to obtain an expression parameter encoding sequence;

[0111] The expression parameter encoding sequence is pooled through the self-attention pooling unit to obtain the emotional characteristics of the target object.

[0112] It should be noted that the second feature extraction model includes a sequence extraction unit, a sequence encoding unit and a self-attention pooling unit. When the second feature extraction model is used to extract emotional features, the sequence extraction unit can extract the reference expression parameters of each video frame in the emotional reference video to form an expression parameter reference sequence; then, the Transformer-based sequence encoding unit encodes the expression parameter reference sequence to obtain an expression parameter encoding sequence; finally, the self-attention pooling unit performs pooling processing on the expression parameter encoding sequence to obtain emotional features.

[0113] By applying the solution of the embodiment of the present disclosure, the emotion reference video is input into the second feature extraction model, and the emotion characteristics of the target object are obtained after being processed by the sequence extraction unit, the sequence encoding unit and the self-attention pooling unit, thereby realizing the acquisition of emotion characteristics through additional emotion reference videos.

[0114] Step 306: Input the object image and expression parameters of the target object into the video generation model to obtain the target video of the target object.

[0115] In one or more embodiments of the present disclosure, a speech to be processed is obtained; after the emotional characteristics of the target object and the input parameters of the speech to be processed are used to generate a model and expression parameters are obtained, the object image and expression parameters of the target object can be further input into a video generation model to obtain a target video of the target object.

[0116] Specifically, the target image includes facial feature information of the target subject; therefore, the target image can be referred to as a portrait of the target subject. The target video can be understood as a video of an emotional speaker. The speech to be processed in the target video is synchronized with the facial movements of the target subject. The video generation model can be understood as a facial renderer, which generates a target video synchronized with the speech to be processed based on expression parameters and the target image. The facial expressions in the video frames of the target video are consistent with the label parameters, and the subject identity is consistent with the target image.

[0117] By applying the solution of the embodiment of the present disclosure, expression parameters are generated based on emotional features and the speech to be processed, and a target video is further generated based on the expression parameters. While ensuring the synchronization of speech and expression in the target video, diverse emotional information is incorporated into the target video, thereby improving the accuracy and vividness of the target video.

[0118] In an optional embodiment of the present disclosure, the object image and expression parameters of the target object are input into the video generation model. After the target video of the target object is obtained, model adjustment information sent by the user through the client can be received, and the model parameters of the parameter generation model and / or the video generation model can be adjusted based on the model adjustment information.

[0119] It should be noted that the model adjustment information includes, but is not limited to, instruction information for adjusting model parameters and an updated target video used to adjust the model parameters. For example, if the model adjustment information is an updated target video, the updated target video is a video that meets user requirements. The updated target video can be used as a real sample video to adjust the model parameters of the parameter generation model and / or the video generation model so that the model parameters of the parameter generation model and / or the video generation model can generate a video that is close to the real sample video.

[0120] By applying the solution of the embodiment of the present disclosure, model adjustment information sent by the user through the client is received, and model parameters of the parameter generation model and / or video generation model are adjusted based on the model adjustment information, thereby improving model accuracy and user satisfaction.

[0121] In an optional embodiment of the present disclosure, a first feature extraction model may be trained using sample speech, sample object images, and emotion feature labels. That is, before obtaining the expression parameters, the above-mentioned generation of a model using the emotion features of the target object and the speech input parameters to be processed may further include the following steps:

[0122] Acquire a plurality of sample videos including sample objects;

[0123] extracting a sample object image and a sample voice of a sample object from a sample video;

[0124] Input the sample video into the second feature extraction model to obtain the emotion feature label;

[0125] Inputting the sample object image and the sample speech into a first initial feature extraction model to obtain predicted emotion features;

[0126] According to the emotion feature labels and the predicted emotion features, the model parameters of the first initial feature extraction model are adjusted to obtain a trained first feature extraction model.

[0127] Specifically, the first initial feature extraction model is trained using supervised training, meaning that the training process includes real emotion feature labels. The emotion feature labels serve as generation targets for the first initial feature extraction model and guide its training. While adjusting the model parameters of the first initial feature extraction model, the model parameters of the second feature extraction model remain fixed.

[0128] The method of "obtaining multiple sample videos including sample objects" can refer to the implementation method of "obtaining emotional reference data" mentioned above. Since the second feature extraction model is pre-trained, the emotional features obtained by inputting the sample video into the second feature extraction model are accurate and can be used as emotional feature labels in the training process of the first feature extraction model. The method of "inputting the sample video into the second feature extraction model to obtain the emotional feature label" can refer to the implementation method of "inputting the emotional reference video into the second feature extraction model to obtain the emotional features of the target object" mentioned above, and the method of "inputting the sample object image and the sample voice into the first initial feature extraction model to obtain the predicted emotional features" can refer to the implementation method of "inputting the emotional reference image and the voice to be processed into the first feature extraction model to obtain the emotional features of the target object" mentioned above. The embodiments of the present disclosure do not impose any restrictions on this.

[0129] It should be noted that when extracting the sample object image and sample speech of the sample object from the sample video, an audio analysis tool can be used to extract the sample speech from the sample video, and any sample video frame of the sample video that includes the sample object can be randomly selected as the sample object image. Furthermore, if the sample object image contains emotional features, to prevent the first initial feature extraction model from directly extracting emotional features from the sample object image while ignoring the sample speech, in the disclosed embodiment, when selecting the sample object image from multiple sample video frames of the sample video, a sample video frame without obvious emotion can be selected, thereby improving the predictive ability of the first initial feature extraction model.

[0130] In actual applications, when adjusting the model parameters of the first initial feature extraction model according to the emotion feature labels and the predicted emotion features, the first loss value can be calculated according to the emotion feature labels and the predicted emotion features, and the model parameters of the first initial feature extraction model can be adjusted according to the first loss value until the first preset stop condition is reached, thereby obtaining the first feature extraction model that has completed training. Among them, there are many functions for calculating the first loss value, such as the cross entropy loss function, the L1 norm loss function, the maximum loss function, the mean square error loss function, the logarithmic loss function, etc. The specific selection is based on the actual situation, and the embodiments of the present disclosure do not impose any restrictions on this.

[0131] In one possible implementation of the present disclosure, the first preset stop condition includes the first loss value being less than or equal to a first preset threshold. After the first loss value is calculated based on the emotion feature label and the predicted emotion feature, the first loss value is compared with the first preset threshold.

[0132] Specifically, if the first loss value is greater than the first preset threshold, it means that the difference between the emotion feature label and the predicted emotion feature is large, and the first initial feature extraction model has poor prediction ability for the emotion feature. At this time, the model parameters of the first initial feature extraction model can be adjusted, and the first initial feature extraction model can continue to be trained until the first loss value is less than or equal to the first preset threshold, indicating that the difference between the emotion feature label and the predicted emotion feature is small, and the first preset stopping condition is reached, and the first feature extraction model that has completed training is obtained.

[0133] In another possible implementation of the present disclosure, in addition to comparing the first loss value and the first preset threshold, it is also possible to determine whether the current first initial feature extraction model is trained in combination with the first number of iterations.

[0134] Specifically, if the first loss value is greater than the first preset threshold, the model parameters of the first initial feature extraction model are adjusted, and the first initial feature extraction model is continued to be trained until the first preset number of iterations is reached, and the iteration is stopped to obtain the first feature extraction model that has completed the training. The first preset threshold and the first preset number of iterations are selected according to the actual situation, and the embodiments of the present disclosure do not impose any restrictions on this.

[0135] By applying the solution of the embodiment of the present disclosure, the model parameters of the first initial feature extraction model are adjusted according to the emotion feature labels and the predicted emotion features to obtain a trained first feature extraction model. By continuously adjusting the model parameters of the first initial feature extraction model, the final first feature extraction model can be made more accurate.

[0136] In an optional embodiment of the present disclosure, before generating a model based on the target object's emotional features and the speech input parameters to be processed and obtaining the expression parameters, the following steps may be further included:

[0137] Acquire a plurality of sample videos including sample objects;

[0138] Extract sample speech and expression parameter labels corresponding to the sample speech from the sample video;

[0139] Inputting the sample video into the second initial feature extraction model to obtain the sample emotion feature;

[0140] Input the sample emotion features and sample speech into the initial parameter generation model to obtain the predicted expression parameters;

[0141] According to the predicted expression parameters and expression parameter labels, the model parameters of the second initial feature extraction model and the initial parameter generation model are adjusted to obtain the trained second feature extraction model and parameter generation model.

[0142] Specifically, the training method of the second initial feature extraction model and the initial parameter generation model is supervised training, that is, the training process includes real expression parameter labels. The expression parameter labels are the generation targets of the initial parameter generation model, which are used to guide the training process of the second initial feature extraction model and the initial parameter generation model.

[0143] When extracting sample speech and expression parameter labels corresponding to the sample speech from a sample video, an audio parsing tool can be used to extract the sample speech from the sample video, and extract the facial motion information of the sample object from the sample video frame synchronized with the sample speech time, and generate expression parameter labels based on the facial motion information of the sample object.

[0144] It should be noted that before extracting the sample speech and the expression parameter labels corresponding to the sample speech from the sample video, a first sample sub-video and a second sample sub-video can be extracted from the sample video, where the time of the last sample video frame of the first sample sub-video is later than the time of the last sample video frame of the second sample sub-video. The first sample sub-video and the second sample sub-video may or may not have overlapping sample video frames. Furthermore, the sample speech and the expression parameter labels corresponding to the sample speech can be extracted from the first sample sub-video, and the second sample sub-video can be input into the second initial feature extraction model to obtain the sample emotion features; the sample emotion features and the sample speech can be input into the initial parameter generation model to obtain the predicted expression parameters.

[0145] For example, assuming that the sample video is a 10s video, the sample voice and the expression parameter label corresponding to the sample voice can be extracted from the first sample sub-video of the 6s-8s of the sample video, and the second sample sub-video of the 2s-5s of the sample video is input into the second initial feature extraction model to obtain the sample emotional features; the sample emotional features and the sample voice are input into the initial parameter generation model to obtain the predicted expression parameters.

[0146] The method of "inputting the sample video into the second initial feature extraction model to obtain the sample emotional features" can refer to the above-mentioned implementation method of "inputting the emotional reference video into the second feature extraction model to obtain the emotional features of the target object", and the method of "inputting the sample emotional features and the sample voice into the initial parameter generation model to obtain the predicted expression parameters" can refer to the above-mentioned implementation method of "inputting the emotional features of the target object and the voice to be processed into the parameter generation model to obtain the expression parameters", and the method of "adjusting the model parameters of the second initial feature extraction model and the initial parameter generation model according to the predicted expression parameters and expression parameter labels" can refer to the above-mentioned implementation method of "adjusting the model parameters of the first initial feature extraction model according to the emotional feature labels and predicted emotional features", and the embodiments of the present disclosure will not be repeated here.

[0147] By applying the solution of the embodiment of the present disclosure, the model parameters of the initial parameter generation model are adjusted according to the predicted expression parameters and expression parameter labels to obtain a trained parameter generation model. By incorporating sample emotion features into the training process of the parameter generation model, the parameter generation model can generate vivid and accurate expression parameters under diverse emotions, thereby improving the accuracy and flexibility of the parameter generation model.

[0148] In an optional embodiment of the present disclosure, before inputting the object image and expression parameters of the target object into the video generation model to obtain the target video of the target object, the following steps may be further included:

[0149] Acquire a plurality of sample videos including sample objects;

[0150] Extracting a first sample video frame and a second sample video frame from the sample video, and determining a sample expression parameter based on the first sample video frame;

[0151] Inputting the sample expression parameters and the second sample video frame into the initial video generation model to obtain a predicted video frame;

[0152] According to the predicted video frame and the first sample video frame, the model parameters of the initial video generation model are adjusted to obtain a trained video generation model.

[0153] Specifically, the video generation model is trained using supervised training, meaning the training process includes real training labels. These real training labels serve as the generation targets for the initial video generation model and guide its training. The initial video generation model is composed of a series of convolutional neural networks. The first sample video frame is temporally later than the second sample video frame. During the training of the video generation model, the second sample video frame is used as the real training label.

[0154] It should be noted that when extracting the first and second sample video frames from the sample video, two sample video frames can be randomly selected from the multiple sample video frames of the sample video, with the sample video frame that comes earlier in time being used as the second sample video frame, and the sample video frame that comes later in time being used as the first sample video frame. When determining the sample expression parameters based on the first sample video frame, facial motion information of the sample subject can be extracted from the first sample video frame, and the sample expression parameters can be generated based on the facial motion information of the sample subject.

[0155] The method of “inputting the sample expression parameters and the second sample video frame into the initial video generation model to obtain a predicted video frame” can refer to the above-mentioned implementation method of “inputting the object image and expression parameters of the target object into the video generation model to obtain the target video of the target object”, and the method of “adjusting the model parameters of the initial video generation model according to the predicted video frame and the first sample video frame” can refer to the above-mentioned implementation method of “adjusting the model parameters of the first initial feature extraction model according to the emotion feature label and the predicted emotion feature”, and the embodiments of the present disclosure will not be repeated here.

[0156] By applying the solution of the embodiment of this description, the model parameters of the initial video generation model are adjusted according to the predicted video frame and the first sample video frame to obtain a trained video generation model. By continuously adjusting the model parameters of the initial video generation model, the final video generation model can be made more accurate.

[0157] 4 , which shows a flow chart of another video generation method provided by an embodiment of the present disclosure, specifically comprising the following steps:

[0158] Step 402: Receive a video generation request sent by a user, wherein the video generation request carries the voice to be processed.

[0159] Step 404: Generate a model based on the emotional features of the target object and the input parameters of the speech to be processed to obtain expression parameters, wherein the expression parameters are used to describe the facial movement information of the target object under the influence of the emotional features. The parameter generation model is trained based on sample emotional features, sample speech and expression parameter labels corresponding to the sample speech. The sample emotional features and sample speech are obtained based on the sample video.

[0160] Step 406: Input the object image and expression parameters of the target object into the video generation model to obtain the target video corresponding to the video generation request.

[0161] Step 408: Send the target video corresponding to the video generation request to the user.

[0162] It should be noted that the implementation of steps 402 to 406 may refer to the implementation of steps 302 to 306 described above, and the embodiments of the present disclosure do not impose any limitation on this.

[0163] In actual applications, there are many ways to send the target video corresponding to the video generation request to the user, and the specific selection is based on the actual situation. The embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, the target video can be sent directly to the user. In another possible implementation of the present disclosure, the target video can be sent to the user based on the user's display demand information. The display demand information represents the user's demand for viewing the target video. The display demand information includes but is not limited to displaying only the target video, displaying the voice to be processed and the target video. The display demand information is specifically set according to the actual needs of the user, and the embodiments of the present disclosure do not impose any restrictions on this.

[0164] By applying the solution of the embodiment of the present disclosure, expression parameters are generated based on emotional features and the speech to be processed, and a target video is further generated based on the expression parameters. While ensuring the synchronization of speech and expression in the target video, diverse emotional information is incorporated into the target video, thereby improving the accuracy and vividness of the target video.

[0165] In an optional embodiment of the present disclosure, after sending the target video corresponding to the video generation request to the user, the following steps may also be included:

[0166] Receive video adjustment information sent by the user based on the target video, adjust the target video based on the video adjustment information, and obtain the adjusted target video.

[0167] It should be noted that after sending the target video corresponding to the video generation request to the user, video adjustment information sent by the user based on the target video can be received. Video adjustment information includes but is not limited to video filter adjustment information, video clarity adjustment information, video title generation information, etc. The specific selection is based on actual circumstances and is not limited in this embodiment.

[0168] Furthermore, after receiving the video adjustment information sent by the user based on the target video, when adjusting the target video based on the video adjustment information, the video adjustment information can be used as prompt information, and the video adjustment information and the target video can be input into the video adjustment model to obtain the adjusted target video.

[0169] By applying the solution of the embodiment of the present disclosure, video adjustment information sent by the user based on the target video is received, and the target video is adjusted based on the video adjustment information to obtain the adjusted target video, thereby realizing data interaction with the user and improving the user experience.

[0170] 5 , which shows a flow chart of a parameter generation model training method provided by one embodiment of the present disclosure. The parameter generation model training method is applied to a cloud-side device and specifically includes the following steps:

[0171] Step 502: Acquire multiple sample videos including sample objects.

[0172] Step 504: extracting sample speech and expression parameter labels corresponding to the sample speech from the sample video.

[0173] Step 506: Input the sample emotion features and sample speech of the sample object into the initial parameter generation model to obtain predicted expression parameters.

[0174] Step 508: Adjust the model parameters of the initial parameter generation model according to the predicted expression parameters and the expression parameter labels to obtain a trained parameter generation model.

[0175] It should be noted that the implementation method of steps 502 to 508 can refer to the training method of the parameter generation model in the above-mentioned video generation method, and the embodiment of the present disclosure does not impose any limitation on this.

[0176] In actual applications, after obtaining the trained parameter generation model, the model parameters of the trained parameter generation model can be sent to the terminal device, so that the user can build the parameter generation model locally based on the model parameters, and use the parameter generation model to generate expression parameters to realize video generation.

[0177] By applying the solution of the embodiment of the present disclosure, the model parameters of the initial parameter generation model are adjusted according to the predicted expression parameters and expression parameter labels to obtain a trained parameter generation model. By incorporating sample emotion features into the training process of the parameter generation model, the parameter generation model can generate vivid and accurate expression parameters under diverse emotions, thereby improving the accuracy and flexibility of the parameter generation model.

[0178] Referring to FIG6 , FIG6 shows a flowchart of a processing process of a video generation method provided by one embodiment of the present disclosure. During the video generation process, a parameter generation model and a video generation model are used to generate video frames, and the generated multiple video frames are used to form a target video. The processing flow of each model is described below:

[0179] Parameter generation model: Generates a model based on the preset noise, the emotional characteristics of the target object, and the input parameters of the speech to be processed to obtain the expression parameters;

[0180] Video generation model: The object image and expression parameters of the target object are input into the video generation model to obtain the target video of the target object.

[0181] It should be noted that the source of the emotional features can be obtained by extracting features from the emotional reference video using the second feature extraction model, or by extracting features from the processed speech and the emotional reference image using the first feature extraction model.

[0182] By applying the solution of the embodiment of the present disclosure, expression parameters are generated based on emotional features and the speech to be processed, and a target video is further generated based on the expression parameters. Under the premise of ensuring the synchronization of speech and expression in the target video, diverse emotional information is incorporated into the target video, thereby achieving the generation of vivid expressions and accurate mouth shapes under diverse emotions. At the same time, a first feature extraction model is proposed, which reduces the dependence on additional emotional reference videos.

[0183] Referring to Figure 7, a schematic diagram of a video generation interface provided by one embodiment of the present disclosure is shown. The video generation interface is divided into a request input interface and a result display interface. The request input interface includes a request input box, an "OK" control, and a "Cancel" control. The result display interface includes a result display box.

[0184] The user enters a video generation request in the request input box displayed on the client. The video generation request includes the audio to be processed. The user clicks the "OK" control. The server receives the audio to be processed from the client and inputs the target subject's emotional characteristics and the audio to be processed into a parameter generation model to obtain expression parameters. Expression parameters are used to describe the target subject's facial movements under the influence of emotional characteristics. The parameter generation model is trained based on sample emotional characteristics, sample audio, and expression parameter labels corresponding to the sample audio. The sample emotional characteristics and sample audio are obtained based on sample videos. The target subject's image and expression parameters are input into the video generation model to obtain the target video corresponding to the video generation request. The target video is then sent to the client. The client displays the target video in the result display box.

[0185] In actual applications, users can operate controls by clicking, double-clicking, touching, hovering the mouse, sliding, long pressing, voice control, or shaking, etc. The specific selection is based on the actual situation, and the embodiments of the present disclosure do not impose any restrictions on this.

[0186] Corresponding to the above-mentioned video generation method embodiment, the present disclosure also provides a video generation device embodiment. FIG8 shows a schematic structural diagram of a video generation device provided by one embodiment of the present disclosure. As shown in FIG8 , the device includes:

[0187] A first acquisition module 802 is configured to acquire speech to be processed;

[0188] A first input module 804 is configured to generate a model based on the target subject's emotional features and the speech to be processed as input parameters to obtain expression parameters, wherein the expression parameters are used to describe the target subject's facial movement information under the influence of the emotional features. The parameter generation model is trained based on sample emotional features, sample speech, and expression parameter labels corresponding to the sample speech. The sample emotional features and sample speech are obtained based on sample videos.

[0189] The second input module 806 is configured to input the object image and expression parameters of the target object into the video generation model to obtain a target video of the target object.

[0190] Optionally, the device further includes: a third acquisition module configured to acquire emotion reference data, wherein the emotion reference data includes emotion information; perform feature extraction on the emotion reference data to obtain emotion features of the target object.

[0191] Optionally, the emotion reference data includes an emotion reference video; and the third acquisition module is further configured to input the emotion reference video into a second feature extraction model to obtain the emotion features of the target object.

[0192] Optionally, the second feature extraction model includes a sequence extraction unit, a sequence encoding unit and a self-attention pooling unit; the third acquisition module is further configured to extract an expression parameter reference sequence from the emotion reference video through the sequence extraction unit, wherein the expression parameter reference sequence includes reference expression parameters of each video frame in the emotion reference video; encode the expression parameter reference sequence through the sequence encoding unit to obtain an expression parameter encoding sequence; and pool the expression parameter encoding sequence through the self-attention pooling unit to obtain the emotional characteristics of the target object.

[0193] Optionally, the parameter generation model includes an encoding unit and a decoding unit; the first input module 804 is further configured to encode the processed speech through the encoding unit to obtain speech features; and to diffuse the preset noise, emotional features and speech features through the decoding unit to obtain expression parameters.

[0194] Optionally, the device also includes: a first training module, configured to obtain multiple sample videos including sample objects; extract sample object images and sample voices of sample objects from the sample videos; input the sample videos into the second feature extraction model to obtain emotion feature labels; input the sample object images and sample voices into the first initial feature extraction model to obtain predicted emotion features; adjust the model parameters of the first initial feature extraction model according to the emotion feature labels and the predicted emotion features to obtain the trained first feature extraction model.

[0195] Optionally, the device also includes: a second training module, configured to obtain multiple sample videos including sample objects; extract sample speech and expression parameter labels corresponding to the sample speech from the sample videos; input the sample videos into a second initial feature extraction model to obtain sample emotional features; input the sample emotional features and the sample speech into an initial parameter generation model to obtain predicted emotional parameters; adjust the model parameters of the second initial feature extraction model and the initial parameter generation model according to the predicted emotional parameters and the expression parameter labels to obtain a trained second feature extraction model and parameter generation model.

[0196] Optionally, the device also includes: a third training module, configured to obtain multiple sample videos including sample objects; extract a first sample video frame and a second sample video frame from the sample video, and determine sample expression parameters based on the first sample video frame; input the sample expression parameters and the second sample video frame into the initial video generation model to obtain a predicted video frame; adjust the model parameters of the initial video generation model based on the predicted video frame and the first sample video frame to obtain a trained video generation model.

[0197] By applying the solution of the embodiment of the present disclosure, expression parameters are generated based on emotional features and the speech to be processed, and a target video is further generated based on the expression parameters. While ensuring the synchronization of speech and expression in the target video, diverse emotional information is incorporated into the target video, thereby improving the accuracy and vividness of the target video.

[0198] The above is a schematic diagram of a video generation device according to this embodiment. It should be noted that the technical solution of the video generation device and the technical solution of the above-mentioned video generation method are based on the same concept. For details not described in detail in the technical solution of the video generation device, please refer to the description of the technical solution of the above-mentioned video generation method.

[0199] Corresponding to the above-mentioned video generation method embodiment, the present disclosure also provides a video generation device embodiment. FIG9 shows a schematic structural diagram of another video generation device provided by one embodiment of the present disclosure. As shown in FIG9 , the device includes:

[0200] The first receiving module 902 is configured to receive a video generation request sent by a user, wherein the video generation request carries a voice to be processed;

[0201] The third input module 904 is configured to generate a model based on the target object's emotional characteristics and the speech input parameters to be processed to obtain expression parameters, wherein the expression parameters are used to describe the facial movement information of the target object under the influence of the emotional characteristics, and the parameter generation model is trained based on sample emotional characteristics, sample speech, and expression parameter labels corresponding to the sample speech, and the sample emotional characteristics and sample speech are obtained based on the sample video;

[0202] The fourth input module 906 is configured to input the object image and expression parameters of the target object into the video generation model to obtain a target video corresponding to the video generation request;

[0203] The sending module 908 is configured to send the target video corresponding to the video generation request to the user.

[0204] Optionally, the device further includes: a second receiving module configured to receive video adjustment information sent by the user based on the target video, and adjust the target video based on the video adjustment information to obtain the adjusted target video.

[0205] By applying the solution of the embodiment of the present disclosure, expression parameters are generated based on emotional features and the speech to be processed, and a target video is further generated based on the expression parameters. While ensuring the synchronization of speech and expression in the target video, diverse emotional information is incorporated into the target video, thereby improving the accuracy and vividness of the target video.

[0206] The above is a schematic diagram of a video generation device according to this embodiment. It should be noted that the technical solution of the video generation device and the technical solution of the above-mentioned video generation method are based on the same concept. For details not described in detail in the technical solution of the video generation device, please refer to the description of the technical solution of the above-mentioned video generation method.

[0207] Corresponding to the above-mentioned parameter generation model training method embodiment, the present disclosure also provides a parameter generation model training device embodiment. Figure 10 shows a schematic diagram of the structure of a parameter generation model training device provided by one embodiment of the present disclosure. As shown in Figure 10, the device is applied to a cloud-side device and includes:

[0208] The second acquisition module 1002 is configured to acquire a plurality of sample videos including sample objects;

[0209] Extraction module 1004, configured to extract sample speech and expression parameter labels corresponding to the sample speech from the sample video;

[0210] The fifth input module 1006 is configured to input the sample emotion feature and sample speech of the sample object into the initial parameter generation model to obtain the predicted expression parameter;

[0211] The adjustment module 1008 is configured to adjust the model parameters of the initial parameter generation model according to the predicted expression parameters and the expression parameter labels to obtain a trained parameter generation model.

[0212] By applying the solution of the embodiment of the present disclosure, the model parameters of the initial parameter generation model are adjusted according to the predicted expression parameters and expression parameter labels to obtain a trained parameter generation model. By incorporating sample emotion features into the training process of the parameter generation model, the parameter generation model can generate vivid and accurate expression parameters under diverse emotions, thereby improving the accuracy and flexibility of the parameter generation model.

[0213] The above is a schematic diagram of a parameter generation model training device according to this embodiment. It should be noted that the technical solution of the parameter generation model training device and the technical solution of the parameter generation model training method described above are based on the same concept. For details not described in detail in the technical solution of the parameter generation model training device, please refer to the description of the technical solution of the parameter generation model training method described above.

[0214] Figure 11 shows a block diagram of a computing device according to an embodiment of the present disclosure. Components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.

[0215] The computing device 1100 also includes an access device 1140 that enables the computing device 1100 to communicate via one or more networks 1160. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1140 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a World Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0216] In one embodiment of the present disclosure, the aforementioned components of the computing device 1100 and other components not shown in FIG11 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG11 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.

[0217] Computing device 1100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1100 may also be a mobile or stationary server.

[0218] Among them, the processor 1120 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned video generation method or parameter generation model training method.

[0219] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solutions of the aforementioned video generation method and parameter generation model training method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the aforementioned video generation method or parameter generation model training method.

[0220] An embodiment of the present disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned video generation method or parameter generation model training method.

[0221] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solutions of the aforementioned video generation method and parameter generation model training method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the aforementioned video generation method or parameter generation model training method.

[0222] An embodiment of the present disclosure further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned video generation method or parameter generation model training method.

[0223] The above is an illustrative embodiment of a computer program. It should be noted that the technical solution of this computer program is based on the same concept as the technical solutions of the video generation method and the parameter generation model training method described above. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solutions of the video generation method or the parameter generation model training method described above.

[0224] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0225] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0226] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present disclosure.

[0227] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0228] The preferred embodiments of the present disclosure disclosed above are only used to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

1. A video generation method, comprising: Get the voice to be processed; Generate a model based on the emotional features of the target object and the input parameters of the speech to be processed to obtain expression parameters, wherein the expression parameters are used to describe the facial movement information of the target object under the influence of the emotional features, and the parameter generation model is trained based on sample emotional features, sample speech, and expression parameter labels corresponding to the sample speech, and the sample emotional features and the sample speech are obtained based on sample videos; The object image of the target object and the expression parameters are input into a video generation model to obtain a target video of the target object.

2. The method according to claim 1, before generating a model using the emotional features of the target object and the speech input parameters to be processed to obtain the expression parameters, further comprises: Acquiring emotion reference data, wherein the emotion reference data includes emotion information; Feature extraction is performed on the emotion reference data to obtain the emotion features of the target object.

3. The method according to claim 2, wherein the emotion reference data comprises an emotion reference video; The step of extracting features from the emotion reference data to obtain the emotion features of the target object includes: The emotion reference video is input into a second feature extraction model to obtain the emotion features of the target object.

4. According to the method of claim 3, the second feature extraction model comprises a sequence extraction unit, a sequence encoding unit and a self-attention pooling unit; The step of inputting the emotion reference video into a second feature extraction model to obtain the emotion feature of the target object includes: Extracting an expression parameter reference sequence from the emotion reference video via the sequence extraction unit, wherein the expression parameter reference sequence includes reference expression parameters of each video frame in the emotion reference video; The expression parameter reference sequence is encoded by the sequence encoding unit to obtain an expression parameter encoding sequence; The expression parameter encoding sequence is pooled through the self-attention pooling unit to obtain the emotional characteristics of the target object.

5. The method according to claim 1, wherein the parameter generation model comprises an encoding unit and a decoding unit; The step of generating a model from the emotional features of the target object and the speech input parameters to be processed to obtain expression parameters includes: The encoding unit encodes the speech to be processed to obtain speech features; The decoding unit performs diffusion processing on the preset noise, the emotion feature and the voice feature to obtain expression parameters.

6. The method according to claim 1, before generating a model using the emotional features of the target object and the input parameters of the speech to be processed to obtain the expression parameters, further comprises: Acquire a plurality of sample videos including sample objects; extracting a sample object image and a sample voice of a sample object from the sample video; Inputting the sample video into a second feature extraction model to obtain an emotion feature label; Inputting the sample object image and the sample speech into a first initial feature extraction model to obtain predicted emotion features; According to the emotion feature label and the predicted emotion feature, the model parameters of the first initial feature extraction model are adjusted to obtain a trained first feature extraction model.

7. The method according to claim 1, before generating a model using the emotional features of the target object and the speech input parameters to be processed to obtain the expression parameters, further comprises: Acquire a plurality of sample videos including sample objects; Extracting sample speech and expression parameter labels corresponding to the sample speech from the sample video; Inputting the sample video into a second initial feature extraction model to obtain sample emotion features; Inputting the sample emotion feature and the sample speech into an initial parameter generation model to obtain predicted expression parameters; According to the predicted expression parameters and the expression parameter labels, the model parameters of the second initial feature extraction model and the initial parameter generation model are adjusted to obtain the trained second feature extraction model and parameter generation model.

8. The method according to claim 1, before inputting the object image of the target object and the expression parameters into a video generation model to obtain a target video of the target object, further comprising: Acquire a plurality of sample videos including sample objects; Extracting a first sample video frame and a second sample video frame from the sample video, and determining a sample expression parameter according to the first sample video frame; Inputting the sample expression parameter and the second sample video frame into an initial video generation model to obtain a predicted video frame; According to the predicted video frame and the first sample video frame, the model parameters of the initial video generation model are adjusted to obtain a trained video generation model.

9. A video generation method, comprising: Receiving a video generation request sent by a user, wherein the video generation request carries the voice to be processed; Generate a model based on the emotional features of the target object and the input parameters of the speech to be processed to obtain expression parameters, wherein the expression parameters are used to describe the facial movement information of the target object under the influence of the emotional features, and the parameter generation model is trained based on sample emotional features, sample speech, and expression parameter labels corresponding to the sample speech, and the sample emotional features and the sample speech are obtained based on sample videos; Inputting the object image of the target object and the expression parameter into a video generation model to obtain a target video corresponding to the video generation request; The target video corresponding to the video generation request is sent to the user.

10. The method according to claim 9, after sending the target video corresponding to the video generation request to the user, further comprising: The video adjustment information sent by the user based on the target video is received, and the target video is adjusted based on the video adjustment information to obtain an adjusted target video.

11. A parameter generation model training method, applied to a cloud-side device, comprising: Acquire a plurality of sample videos including sample objects; Extracting sample speech and expression parameter labels corresponding to the sample speech from the sample video; Inputting the sample emotion characteristics of the sample object and the sample speech into an initial parameter generation model to obtain predicted expression parameters; According to the predicted expression parameters and the expression parameter labels, the model parameters of the initial parameter generation model are adjusted to obtain a trained parameter generation model.

12. A computing device comprising: Memory and processor; The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by a processor, the steps of the method described in any one of claims 1 to 8 or any one of claims 9 to 10 or claim 11 are implemented.

13. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the method described in any one of claims 1 to 8 or any one of claims 9 to 10 or claim 11.

14. A computer program, wherein when the computer program is executed by a computer, the method according to any one of claims 1 to 8 or any one of claims 9 to 10 is implemented.

Citation Information

Patent Citations

  • Virtual character expression generation method and device, virtual character expression control method and device and terminal equipment

    CN111489424A

  • Speaking video generation method and device, electronic equipment, medium and product

    CN114245215A

  • Video generation method and parameter generation model training method

    CN117893652A

  • Speech and text driven HMM-based body animation synthesis

    US20100082345A1

Cited By

  • Risk detection method

    CN120673317A

  • Adaptive evolution video data processing method, device and system

    CN120726543A

  • Adaptive evolutionary video data processing method, device and system

    CN120726543B

  • Video generation method and device, electronic equipment, storage medium and program product

    CN121000952A

  • Video generation method, video live broadcast method and training method of video generation model

    CN121644922A