Video generation method and interaction method based on digital human, and device, storage medium and program product
Through text-to-speech processing and voice-to-expression processing based on user's voice characteristics and emotional labels, the problem of insufficient realism of traditional digital people is solved, personalized digital people driving is realized, and user experience is improved.
Patent Information
- Application Number
- PCT/CN2024/130676
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-04
- Filing Date
- 2024-11-08
- Publication Date
- 2025-08-07
AI Technical Summary
Traditional digital people mainly use single and standardized tones to drive sound and dynamic images, which is poor in fidelity, reducing the user's immersion and authenticity of interactive experience.
Text-to-speech processing is performed based on the user's voice characteristics and emotional tags, and voice-to-speech processing is performed based on the mapping relationship between the user's voice characteristics and expression coefficients, and the digital human model is rendered to generate video data.
It realizes personalized driving of digital people, improves the realism of sound and dynamic images, improves the user experience, and enhances the interactivity and sense of reality.
Smart Images

Figure CN2024130676_07082025_PF_FP_ABST
Abstract
Description
Digital human-based video generation and interaction method, device, storage medium and program product
[0001] Cross-references
[0002] This application refers to Chinese Patent Application No. 2024101578690 filed on February 4, 2024, entitled “Video Generation and Interaction Method, Device, Storage Medium and Program Product Based on Digital Human”, which is incorporated into this application in its entirety by reference. Technical Field
[0003] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, storage medium and program product for video generation and interaction based on digital humans. Background Art
[0004] With the rapid advancement of artificial intelligence and large-scale modeling technologies, digital human technology has emerged. Digital humans are virtual characters with digital appearances, possessing the ability to visualize, perceive, express, and interact. They are widely used in various fields, such as live streaming, short videos, and online customer service, to enhance service quality and user experience.
[0005] However, traditional digital humans rely primarily on a single, standardized timbre to drive their voices and dynamic images, resulting in poor fidelity and diminishing user immersion and the authenticity of the interactive experience. Therefore, a new digital human driving solution is urgently needed.
[0006] Summary of the Invention
[0007] Multiple aspects of the present application provide a method, device, storage medium, and program product for video generation and interaction based on digital humans, which are used to achieve personalized driving of digital humans, improve the realism of digital humans in terms of sound and dynamic image, and thus improve user experience.
[0008] An embodiment of the present application provides a video generation method based on a digital human, comprising: receiving target text information for driving a digital human model in a target scene, the target scene corresponding to a target emotion label, and the digital human model corresponding to a target user; performing voice conversion on the target text information according to the voice characteristics and target emotion label of the target user to obtain a target voice signal adapted to the target user; performing expression mapping on the target voice signal according to the mapping relationship between the voice characteristics and expression coefficients of the target user to obtain a target expression coefficient sequence adapted to the target voice signal; and rendering the digital human model according to the target voice signal and the target expression coefficient sequence to obtain video data of the digital human model.
[0009] An embodiment of the present application also provides an interaction method based on a digital human, including: receiving question information initiated to a digital human model in a target scene, generating reply text information according to the question information, the target scene corresponds to a target emotion label, and the digital human model corresponds to a target user; performing voice conversion on the reply text information according to the voice characteristics and target emotion label of the target user to obtain a target voice signal adapted to the target user; performing expression mapping on the target voice signal according to the mapping relationship between the voice characteristics and expression coefficients of the target user to obtain a target expression coefficient sequence adapted to the target voice signal; rendering the digital human model according to the target voice signal and the target expression coefficient sequence to obtain video data of the digital human model.
[0010] An embodiment of the present application also provides an electronic device, comprising: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory, and is used to execute the computer program to execute the steps in the digital human-based video generation method or the digital human-based interaction method.
[0011] An embodiment of the present application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is enabled to implement the steps in the digital human-based video generation method or the digital human-based interaction method.
[0012] An embodiment of the present application also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is enabled to implement the steps in the digital human-based video generation method or the digital human-based interaction method.
[0013] In this embodiment of the application, text-to-speech processing is performed based on the user's voice characteristics and emotion tags, and speech-to-expression processing is performed based on the mapping relationship between the user's voice characteristics and expression coefficients. The digital human model is then rendered based on the voice signal and expression coefficients to obtain video data of the digital human model. This accurately simulates the user's voice characteristics, ensuring that the digital human's voice output is not only natural but also highly personalized, enabling personalized driving of the digital human, improving the fidelity of the digital human's voice and dynamic image, and thus enhancing the user experience and increasing the interactivity, realism, and immersion of the digital human. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0015] FIG1 is a flow chart of a method for generating a video based on a digital human provided in an embodiment of the present application;
[0016] FIG2 is a schematic diagram of the structure of an exemplary target text-to-speech model provided in an embodiment of the present application;
[0017] FIG3 is a schematic diagram of the structure of an exemplary target speech-to-expression model provided in an embodiment of the present application;
[0018] FIG4 is an exemplary rendering process diagram provided in an embodiment of the present application;
[0019] FIG5 is a diagram illustrating an exemplary training principle for an initial text-to-speech model according to an embodiment of the present application;
[0020] FIG6 is a diagram illustrating a training principle of an exemplary general speech-to-expression model according to an embodiment of the present application;
[0021] FIG7 is a diagram illustrating an exemplary training principle of a 3D Gaussian model according to an embodiment of the present application;
[0022] FIG8 is a diagram illustrating an exemplary training principle of end-cloud interaction according to an embodiment of the present application;
[0023] FIG9 is a diagram illustrating an exemplary reasoning principle of end-cloud interaction provided in an embodiment of the present application;
[0024] FIG10 is a flow chart of an interactive method based on a digital human provided in an embodiment of the present application;
[0025] FIG11 is a diagram illustrating an exemplary application scenario of end-cloud interaction provided by an embodiment of the present application;
[0026] FIG12 is a schematic structural diagram of a video generation device based on a digital human provided in an embodiment of the present application;
[0027] FIG13 is a schematic structural diagram of an interactive device based on a digital human provided in an embodiment of the present application;
[0028] FIG14 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0030] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation portals for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) are in compliance with relevant laws and standards.
[0031] With the rapid advancement of artificial intelligence and large-scale modeling technologies, digital human technology has emerged. Digital humans are virtual characters with digital appearances, possessing the ability to visualize, perceive, express, and interact. They are widely used in various fields, such as live streaming, short videos, and online customer service, to enhance service quality and user experience.
[0032] However, traditional digital humans rely primarily on a single, standardized timbre to drive their voices and dynamic images, resulting in poor fidelity and diminishing user immersion and the authenticity of the interactive experience. Therefore, a new digital human driving solution is urgently needed.
[0033] To this end, embodiments of the present application provide a method, device, storage medium, and program product for video generation and interaction based on a digital human. These methods perform text-to-speech processing based on the user's voice characteristics and emotional tags, as well as voice-to-expression processing based on the mapping relationship between the user's voice characteristics and expression coefficients. Furthermore, a digital human model is rendered based on the voice signal and expression coefficients to obtain video data of the digital human model. This accurately simulates the user's voice characteristics, ensuring that the digital human's voice output is not only natural but also highly personalized, enabling personalized driving of the digital human and improving the fidelity of the digital human's voice and dynamic image. This, in turn, enhances the user experience and increases the interactivity, realism, and immersion of the digital human.
[0034] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0035] FIG1 is a flow chart of a method for generating a video based on a digital human provided by an embodiment of the present application. Referring to FIG1 , the method may include the following steps:
[0036] 101. Receive target text information for driving a digital human model in a target scene, where the target scene corresponds to a target emotion tag, and the digital human model corresponds to a target user.
[0037] In this embodiment, the target scene can be any video scene, such as but not limited to live broadcast scenes, online customer service scenes, short video scenes, etc. These scenes have different emotional expressions. Different emotional expressions are classified in advance to obtain different emotional labels. Emotional labels include but are not limited to positive emotions and negative emotions. Positive emotions are the emotions generated by people's increase in positive value or decrease in negative value; negative emotions are the emotions generated by people's decrease in positive value or increase in negative value. Of course, emotional labels can also express more fine-grained emotions. Emotional labels include but are not limited to: positive energy, happiness, happiness, trust, gratitude, happiness, depression, sadness, loss, pain, contempt, hatred, jealousy, etc.
[0038] In this embodiment, the digital human model is rendered to produce a virtualized image of the target user (also referred to as an avatar). The digital human model can be any 3D model obtained by three-dimensionally reconstructing the target user. It is understood that the digital human model is a 3D model of the target user. The 3D model is composed of many vertices, which form triangles or quadrilaterals. Multiple triangles or quadrilaterals form a three-dimensional 3D model. Preferably, the digital human model is a 3D Gaussian model.
[0039] 102. Perform speech conversion on the target text information according to the voice characteristics and target emotion label of the target user to obtain a target speech signal adapted to the target user.
[0040] In this embodiment, TTS (Text To Speech) technology is used to convert target text information into a target voice signal tailored to the target user. The TTS process integrates the target user's voice characteristics and target emotion tags to simulate the target user's actual speaking intonation, rhythm, and emotion, resulting in a target voice signal that closely resembles a real person's voice. Voice characteristics include, but are not limited to, pitch, loudness, timbre, rhythm, and pronunciation habits.
[0041] In practical applications, a TTS model can be trained that reflects the target user's voice characteristics and produces a speech signal that closely resembles the target user's actual voice. This trained TTS model is referred to herein as the target text-to-speech model. Based on the above, step 102 is implemented by inputting the target text information and target emotion tag into the target text-to-speech model, performing speech conversion on the target text information, and obtaining the target speech signal.
[0042] This embodiment does not restrict the model structure of the TTS model. Optionally, as shown in Figure 2 , the target text-to-speech model includes a text feature encoding network and a speech feature decoding network. Based on this, step 102 is implemented by inputting the target text information and target emotion tags into the text feature encoding network to obtain target text features; and then inputting the target text features into the speech feature decoding network to obtain the target speech signal.
[0043] In this embodiment, the text feature encoding network has the function of encoding text information to obtain text features. The text feature encoding network can be an encoding network of any structure, and there is no limitation on this. For example, referring to Figure 2, the text feature encoding network may include: a text feature learning module, a vectorization module corresponding to the target emotion label, and a text encoder. For another example, referring to Figure 2, the text feature encoding network may include: a text feature learning module, a vectorization module corresponding to the emotion label, other vectorization modules, and a text encoder. Other vectorization modules include one or more of the following: a vectorization module corresponding to the language type, a vectorization module corresponding to the phoneme information, and a vectorization module corresponding to the user's identification information. It can be understood that the more vectorization modules, the higher the accuracy of the text feature encoding network and the better the network performance.
[0044] In this embodiment, the text feature learning module may be a feature extraction module of any structure, which is used to learn the features of text information (ie, text features) from text information.
[0045] In this embodiment, the vectorization module may be a network module with any structure and a vectorization processing function.
[0046] In this embodiment, the text encoder may be an encoder of any structure.
[0047] As an example, the text feature encoding network inputs the target text information and the target emotion label into the text feature encoding network. When obtaining the target text feature, it is specifically used for: inputting the target text information into the text feature learning module to learn the linguistic feature information to obtain multi-level initial text features; inputting the target emotion label into the vectorization module corresponding to the emotion label for vectorization processing to obtain the emotion feature; inputting the initial text feature and emotion feature into the text encoder for feature encoding to obtain the target text feature.
[0048] As another example, the text feature encoding network inputs the target text information and the target emotion label into the text feature encoding network. When obtaining the target text feature, it is specifically used to: input the target text information, the target emotion label and other model input parameters into the text feature encoding network for text feature fusion encoding to obtain the target text feature.
[0049] Optionally, the text feature encoding network includes: a text feature learning module, a vectorization module and a text encoder; the target text information, target emotion label and other model input parameters are input into the text feature encoding network for text feature fusion encoding to obtain target text features, including: inputting the target text information into the text feature learning module for learning linguistic feature information to obtain multi-level initial text features; inputting the target emotion label and other model input parameters into the corresponding vectorization module for vectorization processing to obtain emotion features and features of other model input parameters; inputting the initial text features, emotion features and features of other model input parameters into the text encoder for feature encoding to obtain target text features.
[0050] Optionally, other model input parameters may be determined based on the target text information, where the other model input parameters include at least language type information of the target text information, target phoneme information contained in the target text information, and / or identification information of the target user.
[0051] The language type information includes, but is not limited to, Chinese, English, Japanese, and the like.
[0052] Among them, phoneme information is the smallest speech unit divided according to the natural properties of speech. It is analyzed based on the pronunciation actions in the syllable, and one action constitutes a phoneme.
[0053] The user's identification information (ID) corresponds to the timbre category to which they belong, and is used to distinguish different users. The timbre categories include, but are not limited to, the timbre of a teenage boy, the timbre of a middle-aged person, the timbre of a girl, and the like.
[0054] In this embodiment, the speech feature decoding network has the function of decoding text features into speech signals. The speech feature decoding network can be a decoding network of any structure. Further optionally, referring to Figure 2, to improve TTS performance, the speech feature decoding network includes: a text feature projection module, a speech length predictor, and a speech decoder.
[0055] In the present embodiment, the text feature projection module has the function of projecting text features into speech features, and the text feature projection module can be a feature projection module of any structure. Specifically, the feature projection module can learn the mapping relationship between input features and output features through the defined projection function, and can project the input features into different spaces, and perform learning in the space to map the input features to the output features. Optionally, the text feature projection module can be a neural projection network (Neural Projection Networks). In the present embodiment, the text feature projection module first projects the input target text features into a speech feature space with the same semantics, and performs learning in the speech feature space to map the target text features to the target speech features.
[0056] In this embodiment, the speech length predictor has a speech duration prediction function and can be a network module of any structure. The speech length predictor can predict the speech duration of a speech signal and generate a random noise signal with the speech duration distribution.
[0057] In this embodiment, the speech decoder can be a decoding module of any structure, and there is no limitation on this. In each embodiment of the present application, the decoding module is also called a decoder, and the encoding module is also called an encoder.
[0058] Based on the above, the speech feature decoding network is used to perform speech decoding operations on the target text features to obtain the target speech signal. The implementation method is as follows: the target text features are input into the text feature projection module, and the speech features are learned in the speech feature space with the same semantics for the target text features to obtain the target speech features; the target text features are input into the speech length predictor, and the speech length is predicted for the target text features to obtain the target noise signal; the target speech features and the target noise signal are input into the speech decoder, and the target speech features are decoded using the length of the target noise signal to obtain the target speech signal.
[0059] 103. Perform expression mapping on the target speech signal according to the mapping relationship between the target user's voice features and expression coefficients to obtain a target expression coefficient sequence adapted to the target speech signal.
[0060] In this embodiment, the expression coefficients are coefficients that control the digital human model's facial expressions, such as smiling, frowning, blinking, and speaking. A mapping relationship can be established between the target user's voice features and the expression coefficients. This mapping relationship can be used to map the voice features corresponding to the target voice signal into a target expression coefficient sequence adapted to the target voice signal. The target expression coefficient sequence includes multiple expression coefficients.
[0061] Alternatively, a target speech-to-expression model can be trained that maps voice features to expression coefficients. The target speech-to-expression model reflects the mapping relationship between the target user's voice features and expression coefficients. Based on this, step 103 is implemented by inputting the target speech signal into the target speech-to-expression model and mapping the target speech signal to the expression coefficients to obtain a target expression coefficient sequence.
[0062] In this embodiment, there is no restriction on the model structure of the target speech-to-expression model. Further optionally, referring to FIG3 , the target speech-to-expression model may include: a speech feature encoding network, an expression feature decoding network, and a linearization network.
[0063] Based on the above, the implementation method of step 103 is as follows: input the target speech signal into the speech feature encoding network, extract the speech features of the target speech signal to obtain the target potential speech features; input the target potential speech features into the expression feature decoding network for expression decoding to obtain the target vertex offset sequence; generate a target vertex sequence adapted to the target speech signal based on the reference vertex sequence and the target vertex offset sequence, the reference vertex sequence is the vertex sequence corresponding to the digital human model in the expressionless state; use the linearization network to linearize the target vertex sequence to obtain the target expression coefficient sequence.
[0064] In this embodiment, the speech feature encoding network can be any encoding network that extracts speech features, such as a wav2vec (audio vectorization) model, a wav2vec2 (audio vectorization) model, etc. Referring to FIG3 , when a target speech signal is input into the speech feature encoding network, the speech feature encoding network outputs target latent speech features that reflect the state of the facial expression.
[0065] In this embodiment, the expression feature decoding network can be any decoding network with expression decoding function. Referring to FIG3 , when the target potential speech feature is input into the expression feature decoding network, a target vertex offset sequence is output, which includes offset information of multiple target vertices.
[0066] It is understood that the digital human model is composed of many vertices, and the target vertex is part or all of these vertices. A corresponding reference vertex sequence is provided for the digital human model. The reference vertex sequence includes the position coordinates of multiple reference vertices. The reference vertex sequence is the vertex sequence corresponding to the digital human model in its neutral state. The neutral state can be flexibly set as needed.
[0067] In this embodiment, the offset information of the target vertex refers to the offset of the position coordinates of the target vertex relative to the position coordinates of the corresponding reference vertex. The offset information of the target vertex is the offset caused by the facial expression. Based on this, the target vertex offset sequence is the vertex offset sequence related to the facial expression.
[0068] In this embodiment, referring to FIG3 , a target vertex sequence adapted to the target speech signal is generated based on the reference vertex sequence and the target vertex offset sequence. For example, the offset information of the target vertex and the position coordinates of the corresponding reference vertex are added to obtain the position coordinates of the target vertex, and the position coordinates of multiple target vertices constitute the target vertex sequence. For another example, the offset information of the target vertex and the position coordinates of the corresponding reference vertex are subtracted to obtain the position coordinates of the target vertex, and the position coordinates of multiple target vertices constitute the target vertex sequence. For another example, the offset information of the target vertex and the position coordinates of the corresponding reference vertex are added, and the addition result is multiplied by a constant coefficient to obtain the position coordinates of the target vertex, and the position coordinates of multiple target vertices constitute the target vertex sequence, and there is no limitation on this.
[0069] In this embodiment, the linearization network is a network module with a linearization processing function. Referring to FIG3 , the target vertex sequence is input into the linearization network, and the linearization network outputs a target expression coefficient sequence obtained by linearizing the target vertex sequence. The target expression coefficient sequence includes multiple expression coefficients.
[0070] 104. Render the digital human model according to the target speech signal and the target expression coefficient sequence to obtain video data of the digital human model.
[0071] In this embodiment, since the expression coefficients in the target expression coefficient sequence are essentially related to the position coordinates of the vertices in the digital human model, when rendering the digital human model, the expression coefficients in the target expression coefficient sequence are used to adjust the position coordinates of the vertices in the digital human model, that is, the digital human model is deformed so that the digital human model presents corresponding expression movements. The deformed digital human model is rendered based on rendering parameters such as vertex color and texture information to obtain a rendered digital human model. Through perspective transformation, the rendered digital human model is converted into an action image sequence, which includes multiple action images, and the action images carry expression information. The action image sequence is merged with the target voice signal to obtain video data of the digital human model.
[0072] Further optionally, the digital human model is a 3D Gaussian model, and step 104 is implemented as follows: using the target expression coefficient sequence to deform the 3D Gaussian model in the initial space to obtain a 3D Gaussian model in the standard space; using the bone transformation coefficient to linearly skin the 3D Gaussian model in the standard space to obtain a 3D Gaussian model in the deformed space; using Gaussian splatting to convert the 3D Gaussian model in the deformed space into a sequence of images to be rendered; rendering the sequence of images to be rendered to obtain a sequence of action images; merging the action image sequence with the target voice signal to obtain video data of the digital human model.
[0073] Specifically, the 3D Gaussian model includes point cloud data, which includes vertex data for multiple vertices. Vertex data includes, but is not limited to, vertex position, size, rotation, transparency, and color characteristics. Deformation coefficients that cause the 3D Gaussian model to deform include, but are not limited to, expression coefficients, shape coefficients, head motion position, mouth motion position, and skeletal transformation coefficients (also known as skeletal transformation matrices).
[0074] In this embodiment, the initial space, the standard space, and the deformed space are spaces in different coordinate systems and can be converted between different spaces through coordinate transformation.
[0075] In this embodiment, referring to FIG4 , the target expression coefficient sequence is used to adjust the position coordinates of the corresponding vertices in the 3D Gaussian model in the initial space, that is, the 3D Gaussian model in the initial space is deformed, and the 3D Gaussian model in the initial space can be converted to a 3D Gaussian model in the standard space. The 3D Gaussian model in the standard space is linearly skinned using the bone transformation coefficient, and the 3D Gaussian model in the standard space can be converted to a 3D Gaussian model in the deformed space; the 3D Gaussian model in the three-dimensional deformed space can be converted into a two-dimensional sequence of images to be rendered by Gaussian sputtering, and the two-dimensional sequence of images to be rendered is rendered based on rendering parameters such as color and texture information to obtain a sequence of action images. It is worth noting that the faces in the action images in FIG4 are cartoon faces, which are only for illustration purposes and are actually real faces.
[0076] The technical solution provided in the embodiments of this application performs text-to-speech processing based on the user's voice characteristics and emotion tags, and performs speech-to-expression processing based on the mapping relationship between the user's voice characteristics and expression coefficients. Furthermore, the digital human model is rendered based on the voice signal and expression coefficients to obtain video data of the digital human model. This accurately simulates the user's voice characteristics, ensuring that the digital human's voice output is not only natural but also highly personalized, enabling personalized driving of the digital human, improving the fidelity of the digital human's voice and dynamic image, and thus enhancing the user experience and enhancing the digital human's interactivity, realism, and immersion.
[0077] The following describes the training process of the target text-to-speech model.
[0078] In an embodiment, when training a target text-to-speech model, the initial text-to-speech model can be trained as a general model using the first general sample data to obtain a general text-to-speech model; and the general text-to-speech model can be trained as a personalized model using the first personalized sample data to obtain a target text-to-speech model. The first personalized sample data includes the target user's personalized sample voice signal, personalized sample text information, and personalized sample emotion label.
[0079] Specifically, model training based on the first general sample data can improve the generalization performance of the target text-to-speech model. Model training based on the first personalized sample data can improve the model accuracy of the target text-to-speech model.
[0080] In this embodiment, the initial text-to-speech model may be a TTS model of any structure, and there is no limitation on this.
[0081] In this embodiment, the first general sample data includes general sample speech signals, general sample text information, and general sample emotion tags corresponding to multiple first sample users, but is not limited thereto. For example, the first general sample data may also include the language type corresponding to the general sample text information of the first sample user, phoneme information included in the general sample text information of the first sample user, and identification information of the first sample user.
[0082] In this embodiment, the first sample user is any user participating in the model training, the general sample voice signal of the first sample user is also the voice signal of the first sample user when participating in the model training, the general sample text information of the first sample user is also the text information of the first sample user when participating in the model training, and the general sample emotion label of the first sample user is also the emotion label of the scene corresponding to the model training.
[0083] In this embodiment, the initial text-to-speech model is trained using the first general sample data to obtain a general text-to-speech model, including: filtering the general sample speech signal to obtain general sample spectral features, and performing speech feature extraction on the general sample spectral features to obtain general sample speech features; inputting the general sample text information and the general sample emotion label into the initial text feature encoding network in the initial text-to-speech model to obtain general sample text features; inputting the general sample text features and the general sample speech features into the initial speech feature decoding network in the initial text-to-speech model to perform speech decoding operations to obtain converted speech signals; and in the above-mentioned model training process, calculating at least one loss function, and ending the general model training process when at least one loss function meets the set end condition.
[0084] Specifically, referring to FIG5 , the general sample speech signal of the first sample user is filtered and subjected to speech feature extraction to obtain the general sample speech feature.
[0085] Referring to FIG5 , as an example, the general sample text information and the general sample emotion label of the first sample user are input into the initial text feature encoding network in the initial text-to-speech model to obtain the general sample text feature.
[0086] Optionally, if the initial text feature encoding network includes a text feature learning module, a vectorization module corresponding to the emotion label, and a text encoder, the general sample text information is input into the text feature learning module to learn the linguistic feature information to obtain multi-level initial text features; the general sample emotion label is input into the vectorization module corresponding to the emotion label for vectorization processing to obtain the emotion feature; the initial text feature and the emotion feature are input into the text encoder for feature encoding to obtain the general sample text feature.
[0087] Referring to Figure 5 , as another example, the first sample user's general sample text information, general sample emotion labels, and other model input data are input into the initial text feature encoding network of the initial text-to-speech model to obtain general sample text features. The other model input data may include, for example, at least one of the following: the language type corresponding to the first sample user's general sample text information, phoneme information included in the first sample user's general sample text information, and identification information of the first sample user.
[0088] Optionally, the text feature encoding network includes: a text feature learning module, a vectorization module and a text encoder; the general sample text information, general sample emotion labels and other model input data of the first sample user are input into the initial text feature encoding network in the initial text-to-speech model to obtain general sample text features, including: inputting the general sample text information into the text feature learning module to learn linguistic feature information to obtain multi-level initial text features; inputting the general sample emotion labels and other model input parameters into the corresponding vectorization module for vectorization processing to obtain emotion features and features of other model input parameters; inputting the initial text features, emotion features and features of other model input parameters into the text encoder for feature encoding to obtain target text features.
[0089] Referring to Figure 5 , the generic sample text features and generic sample speech features are input into the initial speech feature decoding network in the initial text-to-speech model for speech decoding to obtain a converted speech signal. Further, optionally, referring to Figure 5 , to improve TTS performance, the initial speech feature decoding network includes a text feature projection module, a speech length predictor, and a speech decoder. Based on the above, the general sample text features and the general sample speech features are input into the initial speech feature decoding network in the initial text-to-speech model for speech decoding operation to obtain the converted speech signal. The implementation method is as follows: the general sample text features are input into the text feature projection module, and the speech features are learned in the speech feature space with the same semantics for the general sample text features to obtain the intermediate sample speech features; the general sample speech features and the intermediate sample speech features are processed using the monotonic alignment search (MAS) algorithm to obtain a monotonic alignment relationship, which reflects the feature point pairs that match the general sample speech features and the intermediate sample speech features; the general sample text features and the monotonic alignment relationship are input into the speech length predictor for prediction processing to obtain the predicted speech duration and the sample noise signal that satisfies the speech duration distribution; the general sample speech features and the sample noise signal are input into the speech decoder to obtain the converted speech signal.
[0090] In this embodiment, during the model training process, at least one loss function is calculated, and the general model training process ends when at least one loss function meets the set end condition. The at least one loss function includes, but is not limited to:
[0091] (1) KL divergence loss function between the general sample speech features and the intermediate sample speech features obtained by feature projection of the general sample text features in the initial speech feature decoding network.
[0092] The Kullback-Leibler Divergence (KL Divergence) is a metric that measures the difference between two probability distributions. It is also known as relative entropy. KL Divergence is widely used in fields such as information theory, statistics, machine learning, and data science.
[0093] (2) Length loss function between the general sample speech signal and the converted speech signal.
[0094] Specifically, the length loss function reflects the difference information between the speech duration of the general sample speech signal and the speech duration of the converted speech signal. The length loss function includes but is not limited to: Negative Log Likelihood Loss (NLL Loss), Cross Entropy Loss, Reconstruction Error Loss, etc.
[0095] (3) Reconstruction loss function between the general sample speech signal and the converted speech signal.
[0096] Specifically, the reconstruction loss function reflects the difference information between the speech features of the general sample speech signal and the speech features of the converted speech signal. The reconstruction loss function includes but is not limited to: negative log-likelihood loss function, cross-entropy loss function, reconstruction error loss function, etc.
[0097] (4) A loss function between the general sample spectral features and the spectral features obtained by filtering the converted speech signal. The loss function includes, but is not limited to, a negative log-likelihood loss function, a cross-entropy loss function, and the like.
[0098] (5) The loss function between the general sample speech signal and the real speech signal selected by the discriminator from the general sample speech signal and the converted speech signal.
[0099] Specifically, the discriminator can be a discriminator in a generative adversarial network (GAN). The loss function represents the difference between the discrimination result output by the discriminant network and the universal sample speech signal. The discrimination result is whether the universal sample speech signal is a true speech signal or whether the converted speech signal is a true speech signal. Examples of loss functions include, but are not limited to, negative log-likelihood loss, cross-entropy loss, and least squares loss.
[0100] In this embodiment, there is no restriction on the set end condition. For example, various operations such as weighted summation, averaging, accumulation, etc. are performed on the at least one loss function mentioned above to obtain a total loss function. If the total loss function is less than or equal to the preset loss value, it is considered that the set end condition is met. If the total loss function is greater than the preset loss value, it is considered that the set end condition is not met. For another example, if each loss function in the at least one loss function mentioned above is less than or equal to its corresponding preset loss value, it is considered that the set end condition is met; if there is a loss function greater than the corresponding preset loss value, it is considered that the set end condition is not met. For another example, if each loss function in the at least one loss function mentioned above is less than or equal to its corresponding preset loss value, and the total loss function is less than or equal to the preset loss value, it is considered that the set end condition is met; if there is a loss function greater than the corresponding preset loss value, or the total loss function is greater than the preset loss value, it is considered that the set end condition is not met.
[0101] In this embodiment, the first personalized sample data is used to perform personalized model training on the general text-to-speech model to obtain a target text-to-speech model. The first personalized sample data includes the personalized sample voice signal, personalized sample text information and personalized sample emotion label of the target user.
[0102] In this embodiment, the process of performing general model training on the initial text-to-speech model is the same as the process of performing personalized model training on the general text-to-speech model. The only difference is the model input data, which will not be repeated here.
[0103] In this embodiment, the first personalized sample data includes the target user's personalized sample voice signal, personalized sample text information, and personalized sample emotion tag, but is not limited thereto. For example, the first personalized sample data may also include the language type corresponding to the target user's personalized sample text information, the phoneme information corresponding to the target user's personalized sample text information, and the target user's identification information.
[0104] In this embodiment, the personalized sample voice signal of the target user is also the voice signal of the target user when participating in the model training, the personalized sample text information of the target user is also the text information of the target user when participating in the model training, and the personalized sample emotion label of the target user is also the emotion label of the scene corresponding to the model training.
[0105] The following describes the training process of the target speech-to-expression model.
[0106] In this embodiment, when training the target speech-to-expression model, the initial speech-to-expression model is trained as a general model using the second general sample data to obtain a general speech-to-expression model, and the second general sample data includes general sample speech signals and general sample expression coefficient sequences corresponding to multiple second sample users; the general speech-to-expression model is trained as a personalized model using the second personalized sample data to obtain the target speech-to-expression model, and the second personalized sample data includes the personalized sample speech signals and personalized sample expression coefficient sequences of the target users.
[0107] Specifically, model training based on the second general sample data can improve the generalization performance of the target speech-to-expression model. Model training based on the second personalized sample data can improve the model accuracy of the target speech-to-expression model.
[0108] In this embodiment, the second sample user is any user participating in model training, the second sample user's universal sample speech signal is also the second sample user's speech signal when participating in model training, and the second sample user's universal sample expression coefficient sequence is also the second sample user's expression coefficient sequence when participating in model training. The target user's personalized sample speech signal is also the target user's speech signal when participating in model training, and the personalized sample expression coefficient sequence is also the target user's expression coefficient sequence when participating in model training.
[0109] In this embodiment, the network structure of the initial speech-to-expression model is not restricted. The process of general model training for the initial speech-to-expression model is the same as the process of personalized model training for the general speech-to-expression model, with the only difference being the model input data. Here, personalized model training for the general speech-to-expression model is described. The general model training process for the initial speech-to-expression model can refer to the training process for personalized model training for the general speech-to-expression model.
[0110] Based on the above, referring to Figure 6, the second personalized sample data is used to perform personalized model training on the universal speech-to-expression model to achieve the target speech-to-expression model as follows: the personalized sample speech signal is input into the speech feature encoding network in the universal speech-to-expression model, and the speech feature of the personalized sample speech signal is extracted to obtain a first potential speech feature; the first potential speech feature is input into the expression feature decoding network in the universal speech-to-expression model for expression decoding to obtain a first vertex offset sequence; based on the reference vertex sequence and the first vertex offset sequence, a first vertex sequence adapted to the personalized sample speech signal is generated; the first vertex sequence is linearized using the linearization network in the universal speech-to-expression model to obtain a first expression coefficient sequence; in the above model training process, at least one loss function is calculated, and the personalized model training process is ended when at least one loss function meets the set end condition.
[0111] In this embodiment, the at least one loss function includes, but is not limited to:
[0112] (1) A KL divergence loss function between the first latent speech feature and the second latent speech feature; wherein the second latent speech feature is obtained by speech encoding the vertex sequence of the mouth region in the first vertex sequence using a mouth region encoding network.
[0113] As shown in Figure 6, the personalized sample speech signal is input into the speech feature encoding network of the universal speech-to-expression model, which outputs a first latent speech feature. The universal speech-to-expression model also includes a mouth region encoding network, a decoding network of arbitrary structure, which outputs a second latent speech feature and calculates the KL divergence loss function between the first and second latent speech features.
[0114] (2) The L1 loss function between the first vertex sequence and the second vertex sequence, where the second vertex sequence is the vertex sequence corresponding to the digital human model when it emits the personalized sample speech signal. The L1 loss function is also the mean absolute error function.
[0115] (3) An L1 loss function between the second vertex sequence and the third vertex sequence, wherein the expression coefficients in the first expression coefficient sequence are decoded using an expression coefficient decoder to obtain the third vertex sequence.
[0116] 6 , the general speech-to-expression model further includes an expression coefficient decoder, which is a decoding network of arbitrary structure. The first expression coefficient sequence is input into the expression coefficient decoder, which outputs a third vertex sequence.
[0117] (4) L1 loss function between the first expression coefficient sequence and the corresponding expression coefficients in the personalized sample expression coefficient sequence.
[0118] The following describes the training method of the 3D Gaussian model.
[0119] Specifically, the initial 3D Gaussian model is trained using the third personalized sample data to obtain a 3D Gaussian model; the third personalized sample data includes the personalized sample expression coefficient of the target user and the corresponding personalized sample image.
[0120] Specifically, the personalized sample expression coefficient of the target user is the expression coefficient of the target user during model training, and the personalized sample image is an image that includes the target user's face during model training. Referring to Figure 4, during the training phase, the personalized sample expression coefficient is used to adjust the position coordinates of corresponding vertices in the initial 3D Gaussian model in the initial space, that is, the initial 3D Gaussian model in the initial space is deformed, thereby converting the initial 3D Gaussian model in the initial space into a 3D Gaussian model in the standard space. The 3D Gaussian model in the standard space is linearly skinned using the bone transformation coefficient to convert the 3D Gaussian model in the standard space into a 3D Gaussian model in the deformed space. The 3D Gaussian model in the three-dimensional deformed space is converted into a two-dimensional image sequence to be rendered using Gaussian sputtering. The two-dimensional image sequence to be rendered is rendered based on rendering parameters such as color and texture information to obtain an action image sequence. A loss function is calculated between the action image sequence and the personalized sample image sequence, and the vertex data of the vertices in the initial 3D Gaussian model in the initial space is adjusted based on the loss function. The above steps are repeated until the model training termination condition is met. After the model training is completed, the 3D Gaussian model in the initial space is used as the final digital human model, which can be used in the inference stage.
[0121] In this embodiment, the personalized sample image sequence includes multiple personalized sample images. When calculating the loss function between the action image sequence and the personalized sample image sequence, the loss function is calculated between each personalized sample image and its corresponding action image. The loss function between each personalized sample image and its corresponding action image is subjected to various operations such as weighted summation, averaging, and accumulation to obtain a final loss function. If the final loss function is less than or equal to a preset loss value, the model training termination condition is considered to have been met. If the final loss function is greater than the preset loss value, the model training termination condition is considered not to have been met. For another example, when the number of model training times reaches a specified number, the model training termination condition is considered to have been met.
[0122] The technical solution provided in the embodiments of the present application can be executed by the end side, or by the cloud side, or by the end-cloud collaboration, without limitation. The cloud side can be understood as a system that provides cloud services, and the end side can be understood as a system that provides local services or edge services. Compared with the cloud side, the end side is closer to the user. Exemplarily, when executed by the end-cloud collaboration, steps 101 to 103 can be executed by the cloud side, and step 104 can be executed by the end side. In this way, the expression coefficient is decoupled from the process of digital human video rendering, the generation of expression coefficients is placed in the cloud, and the digital human video rendering is placed on the client side of the end side, which can effectively reduce the computing load on the cloud side. In addition, the cloud side only needs to transmit a small amount of expression coefficients, voice signals, etc. to the end side, and the client side renders the digital human video based on the expression coefficients. This can effectively reduce the response delay caused by transmitting a large amount of data, that is, reduce the data transmission delay.
[0123] In order to better understand the end-cloud interaction, a specific scenario embodiment is introduced below with reference to Figures 8 and 9.
[0124] As shown in Figure 8, model training is performed by a cloud server in the cloud. The client on the end can upload a target user's speech video to the cloud. The cloud then prepares the target user's first, second, and third personalized sample data based on the target user's speech video. The cloud uses the first general sample data and the first personalized sample data to train the target text-to-speech model; the cloud uses the second general sample data and the second personalized sample data to train the target speech-to-expression model; and the cloud uses the third personalized sample data to train the 3D Gaussian model (also known as the digital human model). The cloud server then sends the digital human model to the client.
[0125] As shown in Figure 9, during the inference phase, end-to-end cloud interaction completes the rendering of the digital human video. First, the client sends the target text information to the cloud server on the cloud side. The cloud server then uses the target speech-to-expression model to convert the target text information into a target speech signal. The cloud server also uses the target speech-to-expression model to convert the target speech signal into a target expression coefficient sequence. The cloud server then returns the target expression coefficient sequence and target speech signal to the client on the end side, allowing the client to render the digital human model based on the target expression coefficient sequence and target speech signal, resulting in the digital human video.
[0126] FIG10 is a flow chart of a digital human-based interaction method provided in an embodiment of the present application. Referring to FIG10 , the method may include:
[0127] 201. Receive question information initiated to the digital human model in the target scene, generate reply text information according to the question information, the target scene corresponds to the target emotion label, and the digital human model corresponds to the target user.
[0128] In this embodiment, the question information is used to describe the target user's question. The question information can be in text or voice format, without limitation. The question information is understood for its intent, and based on the intent understanding result, a response is processed to obtain a reply text message.
[0129] Optionally, if the question information is voice-type question information, reply text information is generated based on the question information, including: inputting the voice-type question information into a speech-to-text model for text information conversion to obtain text-type question information; at least inputting the text-type question information into a question-answering model for intent understanding and answer generation to obtain reply text information.
[0130] The speech-to-text model may be any model with an automatic speech recognition (ASR) function, and there is no restriction on this.
[0131] Among them, question-answering models include, but are not limited to: DBQA (Document-Based Question Answering) model, artificial intelligence generated content (AIGC) model, and large language model (LLM).
[0132] 202. Perform voice conversion on the target text information according to the voice characteristics and target emotion label of the target user to obtain a target voice signal adapted to the target user.
[0133] 203. Perform expression mapping on the target speech signal according to the mapping relationship between the target user's voice features and expression coefficients to obtain a target expression coefficient sequence adapted to the target speech signal.
[0134] 204. Render the digital human model according to the target speech signal and the target expression coefficient sequence to obtain video data of the digital human model.
[0135] In this embodiment, the reply text message can be understood as the target text message in the above embodiment. Regarding steps 202 to 204, reference can be made to steps 102 to 104 in the above embodiment, which will not be repeated here.
[0136] For more details about the embodiments of this application, please refer to the aforementioned embodiments.
[0137] The technical solution provided by the embodiments of this application converts questions addressed to a digital human model in a target scenario into text responses, performs text-to-speech processing based on the user's voice characteristics and emotion tags, and performs speech-to-expression processing based on the mapping relationship between the user's voice characteristics and expression coefficients. Furthermore, the digital human model is rendered based on the voice signal and expression coefficients to obtain video data of the digital human model. This accurately simulates the user's voice characteristics, ensuring that the digital human's voice output is not only natural but also highly personalized, enabling personalized driving of the digital human, improving the fidelity of the digital human's voice and dynamic image, and thus enhancing the user experience and increasing the interactivity, realism, and immersion of the digital human.
[0138] The following describes a specific scenario embodiment with reference to Figure 11. First, the client inputs voice into the cloud, that is, inputs the user's voice-type question information. Next, the cloud server on the cloud uses the ASR model to perform speech recognition on the voice-type question information to obtain text-type question information; then, the text-type question information is input into the large language model, and the large language model outputs reply text information; then, the cloud server on the cloud uses the target voice-to-expression model to convert the reply text information into a target voice signal. The cloud also uses the target voice-to-expression model to convert the target voice signal into a target expression coefficient sequence. Next, the cloud server on the cloud returns the target expression coefficient sequence, target voice signal, and reply text information to the client on the end side. The client on the end side renders the action image sequence based on the target expression coefficient sequence, merges the action image sequence, target voice signal, and reply text information, and obtains the digital human video.
[0139] FIG12 is a schematic diagram of the structure of a digital human-based video generation device provided in an embodiment of the present application. Referring to FIG12 , the device may include:
[0140] A receiving module 11 is configured to receive target text information for driving a digital human model in a target scene, wherein the target scene corresponds to a target emotion tag and the digital human model corresponds to a target user;
[0141] The speech conversion module 12 is used to perform speech conversion on the target text information according to the voice characteristics and target emotion label of the target user to obtain a target speech signal adapted to the target user;
[0142] An expression mapping module 13 is used to perform expression mapping on the target speech signal according to the mapping relationship between the target user's voice characteristics and expression coefficients to obtain a target expression coefficient sequence adapted to the target speech signal;
[0143] The rendering module 14 is used to render the digital human model according to the target speech signal and the target expression coefficient sequence to obtain video data of the digital human model.
[0144] Optionally, the speech conversion module 12 is specifically used to: input the target text information and the target emotion label into the target text-to-speech model, and perform speech conversion on the target text information to obtain a target speech signal; wherein, the target text-to-speech model is obtained by personalized training of the general text-to-speech model based on the first personalized sample data of the target user, and is used to reflect the voice characteristics of the target user.
[0145] Optionally, the speech conversion module 12 is specifically used to: determine other model input parameters based on the target text information, the other model input parameters at least including the language type information of the target text information, the target phoneme information contained in the target text information and / or the identification information of the target user; input the target text information, the target emotion label and other model input parameters into the target text-to-speech model, and perform speech conversion on the target text information to obtain the target speech signal.
[0146] Optionally, the target text-to-speech model includes: a text feature encoding network and a speech feature decoding network; the speech conversion module 12 is specifically used to: input the target text information, target emotion label and other model input parameters into the text feature encoding network for text feature fusion encoding to obtain the target text features; use the speech feature decoding network to perform speech decoding operations on the target text features to obtain the target speech signal.
[0147] Optionally, the text feature encoding network includes: a text feature learning module, a vectorization module and a text encoder; when the speech conversion module 12 performs text feature fusion encoding, it is specifically used to: input the target text information into the text feature learning module to learn the linguistic feature information to obtain multi-level initial text features; input the target emotion label and other model input parameters into the corresponding vectorization module for vectorization processing to obtain emotion features and features of other model input parameters; input the initial text features, emotion features and features of other model input parameters into the text encoder for feature encoding to obtain target text features.
[0148] Optionally, the speech feature decoding network includes: a text feature projection module, a speech length predictor and a speech decoder; when the speech conversion module 12 performs a speech decoding operation, it is specifically used to: input the target text feature into the text feature projection module, and learn the speech feature in the speech feature space with the same semantics for the target text feature to obtain the target speech feature; input the target text feature into the speech length predictor, and predict the speech length for the target text feature to obtain the target noise signal; input the target speech feature and the target noise signal into the speech decoder, and use the length of the target noise signal to decode the target speech feature to obtain the target speech signal.
[0149] Optionally, the above-mentioned device also includes: a training module, which is used to use the first general sample data to perform general model training on the initial text-to-speech model to obtain a general text-to-speech model, and the first general sample data includes general sample voice signals, general sample text information and general sample emotion labels corresponding to multiple first sample users; and use the first personalized sample data to perform personalized model training on the general text-to-speech model to obtain a target text-to-speech model, and the first personalized sample data includes the personalized sample voice signals, personalized sample text information and personalized sample emotion labels of the target user.
[0150] Optionally, the process of general model training for the initial text-to-speech model is the same as the process of personalized model training for the general text-to-speech model;
[0151] Among them, when the training module performs universal model training, it is specifically used to: filter the universal sample speech signal to obtain universal sample spectrum features, and extract speech features from the universal sample spectrum features to obtain universal sample speech features; input the universal sample text information and the universal sample emotion label into the initial text feature encoding network in the initial text-to-speech model to obtain universal sample text features; input the universal sample text features and the universal sample speech features into the initial speech feature decoding network in the initial text-to-speech model to perform speech decoding operations to obtain converted speech signals; and in the above-mentioned model training process, calculate at least one loss function, and end the universal model training process when at least one loss function meets the set end condition:
[0152] (1) KL divergence loss function between the general sample speech features and the intermediate sample speech features obtained by feature projection of the general sample text features in the initial speech feature decoding network;
[0153] (2) Length loss function between the general sample speech signal and the converted speech signal;
[0154] (3) reconstruction loss function between the general sample speech signal and the converted speech signal;
[0155] (4) the loss function between the spectral features of the general sample and the spectral features obtained by filtering the converted speech signal;
[0156] (5) The loss function between the general sample speech signal and the real speech signal selected by the discriminator from the general sample speech signal and the converted speech signal.
[0157] Optionally, when the expression mapping module 13 performs expression mapping, it is specifically used to: input the target speech signal into the target speech-to-expression model, map the expression coefficient of the target speech signal to obtain a target expression coefficient sequence; wherein, the target speech-to-expression model is obtained by personalized training of the general speech-to-expression model based on the second personalized sample data of the target user, and is used to reflect the mapping relationship between the sound characteristics and expression coefficients of the target user.
[0158] Optionally, the target speech to expression model includes: a speech feature encoding network, an expression feature decoding network and a linearization network; when the expression mapping module 13 performs expression mapping, it is specifically used to: input the target speech signal into the speech feature encoding network, extract the speech features of the target speech signal to obtain the target potential speech features; input the target potential speech features into the expression feature decoding network for expression decoding to obtain the target vertex offset sequence; generate a target vertex sequence adapted to the target speech signal based on the reference vertex sequence and the target vertex offset sequence, the reference vertex sequence is the vertex sequence corresponding to the digital human model in the expressionless state; use the linearization network to linearize the target vertex sequence to obtain the target expression coefficient sequence.
[0159] Optionally, the training module is also used to: use the second general sample data to perform general model training on the initial speech-to-expression model to obtain a general speech-to-expression model, the second general sample data includes general sample speech signals and general sample expression coefficient sequences corresponding to multiple second sample users; use the second personalized sample data to perform personalized model training on the general speech-to-expression model to obtain a target speech-to-expression model, the second personalized sample data includes the personalized sample speech signals and personalized sample expression coefficient sequences of the target user.
[0160] Optionally, the process of performing universal model training on the initial speech-to-expression model is the same as the process of performing personalized model training on the universal speech-to-expression model;
[0161] Among them, when the training module performs personalized model training, it is specifically used to: input the personalized sample voice signal into the voice feature encoding network in the universal voice-to-expression model, extract the voice features of the personalized sample voice signal to obtain a first potential voice feature; input the first potential voice feature into the expression feature decoding network in the universal voice-to-expression model for expression decoding to obtain a first vertex offset sequence; generate a first vertex sequence adapted to the personalized sample voice signal based on the reference vertex sequence and the first vertex offset sequence; use the linearization network in the universal voice-to-expression model to linearize the first vertex sequence to obtain a first expression coefficient sequence; in the above-mentioned model training process, calculate at least one loss function, and end the personalized model training process when at least one loss function meets the set end condition:
[0162] (1) a KL divergence loss function between the first latent speech feature and the second latent speech feature; wherein the second latent speech feature is obtained by speech encoding the vertex sequence of the mouth region in the first vertex sequence using a mouth region encoding network;
[0163] (2) an L1 loss function between the first vertex sequence and the second vertex sequence, where the second vertex sequence is the vertex sequence corresponding to the digital human model when emitting the personalized sample speech signal;
[0164] (3) an L1 loss function between the second vertex sequence and the third vertex sequence, wherein the expression coefficients in the first expression coefficient sequence are decoded using an expression coefficient decoder to obtain the third vertex sequence;
[0165] (4) L1 loss function between the first expression coefficient sequence and the corresponding expression coefficients in the personalized sample expression coefficient sequence.
[0166] Optionally, the digital human model is a 3D Gaussian model, and the rendering module 14 is specifically used to: use the target expression coefficient sequence to deform the 3D Gaussian model in the initial space to obtain a 3D Gaussian model in the standard space; use the bone transformation coefficient to perform linear skinning on the 3D Gaussian model in the standard space to obtain a 3D Gaussian model in the deformed space; use Gaussian sputtering to convert the 3D Gaussian model in the deformed space into a sequence of images to be rendered; render the sequence of images to be rendered to obtain a sequence of action images; and merge the action image sequence with the target voice signal to obtain video data of the digital human model.
[0167] Optionally, the training module is also used to: use third personalized sample data to perform model training on the initial 3D Gaussian model to obtain a 3D Gaussian model; the third personalized sample data includes the personalized sample expression coefficient of the target user and the corresponding personalized sample image.
[0168] The device shown in FIG12 can execute the method shown in the embodiment shown in FIG1 , and its implementation principle and technical effects are not described in detail here. The specific manner in which each module and unit performs the operation of the device shown in FIG12 in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.
[0169] FIG13 is a schematic diagram of the structure of an interactive device based on a digital human provided in an embodiment of the present application. Referring to FIG13 , the device may include:
[0170] The acquisition module 21 is used to receive question information sent to the digital human model in the target scene, and generate a reply text message according to the question information, where the target scene corresponds to the target emotion label and the digital human model corresponds to the target user;
[0171] The voice conversion module 22 is used to perform voice conversion on the reply text message according to the voice characteristics and target emotion label of the target user to obtain a target voice signal adapted to the target user;
[0172] An expression mapping module 23 is used to perform expression mapping on the target speech signal according to the mapping relationship between the target user's voice characteristics and expression coefficients to obtain a target expression coefficient sequence adapted to the target speech signal;
[0173] The rendering module 24 is used to render the digital human model according to the target speech signal and the target expression coefficient sequence to obtain video data of the digital human model.
[0174] Optionally, when the question information is voice-type question information, the acquisition module 21 generates reply text information based on the question information, and is specifically used to: input the voice-type question information into the speech-to-text model for text information conversion to obtain text-type question information; at least input the text-type question information into the question-answering model for intent understanding and answer generation to obtain reply text information.
[0175] The device shown in FIG13 can execute the method shown in the embodiment shown in FIG10, and its implementation principle and technical effects are not described in detail here. The specific manner in which each module and unit performs the operation of the device shown in FIG13 in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.
[0176] It should be noted that the execution entity of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution entity of steps 101 to 104 can be device A; for another example, the execution entity of steps 101 and 102 can be device A, and the execution entity of steps 102 and 103 can be device B; and so on.
[0177] In addition, some of the processes described in the above embodiments and the accompanying drawings include multiple operations that appear in a specific order, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0178] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0179] FIG14 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. As shown in FIG14 , the electronic device includes: a memory 11 and a processor 12;
[0180] The memory 11 is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc.
[0181] The memory 11 can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0182] The processor 12 is coupled to the memory 11 and is configured to execute the computer program in the memory 11 for performing the steps in the digital human-based video generation method or the digital human-based interaction method.
[0183] Optionally, as shown in Figure 14, the electronic device also includes: a communication component 13, a display 14, a power supply component 15, an audio component 16 and other components. Figure 14 only schematically shows some components, which does not mean that the electronic device only includes the components shown in Figure 14. In addition, the components in the dotted box in Figure 14 are optional components, not mandatory components, and the specific product form of the electronic device may depend on the product form. The electronic device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone or an IOT (Internet of things) device, or it can be a server-side device such as a conventional server, a cloud server or a server array. If the electronic device of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, etc., it may include the components in the dotted box in Figure 14; if the electronic device of this embodiment is implemented as a server-side device such as a conventional server, a cloud server or a server array, it may not include the components in the dotted box in Figure 14.
[0184] The detailed implementation process of the processor executing each action can be found in the relevant description in the aforementioned method embodiment or device embodiment, and will not be repeated here.
[0185] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be performed by the electronic device in the above method embodiment.
[0186] Accordingly, an embodiment of the present application also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is enabled to implement the steps in the above method embodiment that can be performed by an electronic device.
[0187] The above-mentioned communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi (Wireless Fidelity), 2G (2Generation, 2nd generation), 3G (3Generation, 3rd generation), 4G (4Generation, 4th generation) / LTE (long Term Evolution, long term evolution), 5G (5Generation, 5th generation) and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wide band (UWB) technology, Bluetooth (BT) technology and other technologies.
[0188] The above-mentioned display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundary of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.
[0189] The power supply assembly provides power to various components of the device in which the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly is located.
[0190] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as call mode, recording mode, and voice recognition mode, the microphone is configured to receive external audio signals. The received audio signal can be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0191] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0192] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.
[0193] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0194] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0195] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interfaces, network interfaces, and memory.
[0196] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0197] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change RAM (PRAM), static random-access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0198] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0199] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for generating a video based on a digital human, characterized in that: include: Receiving target text information for driving a digital human model in a target scene, wherein the target scene corresponds to a target emotion tag, and the digital human model corresponds to a target user; Performing voice conversion on the target text information according to the voice characteristics of the target user and the target emotion label to obtain a target voice signal adapted to the target user; Performing expression mapping on the target speech signal according to a mapping relationship between the target user's voice features and expression coefficients to obtain a target expression coefficient sequence adapted to the target speech signal; The digital human model is rendered according to the target speech signal and the target expression coefficient sequence to obtain video data of the digital human model.
2. The method according to claim 1, characterized in that According to the voice characteristics of the target user and the target emotion label, the target text information is converted into speech to obtain a target speech signal adapted to the target user, including: Inputting the target text information and the target emotion tag into a target text-to-speech model, performing speech conversion on the target text information to obtain the target speech signal; The target text-to-speech model is obtained by performing personalized training on a general text-to-speech model based on the first personalized sample data of the target user, and is used to reflect the voice characteristics of the target user.
3. The method according to claim 2, characterized in that Inputting the target text information and the target emotion tag into a target text-to-speech model, and performing speech conversion on the target text information to obtain the target speech signal, including: Determining other model input parameters according to the target text information, the other model input parameters including at least language type information of the target text information, target phoneme information contained in the target text information, and / or identification information of the target user; The target text information, the target emotion label and the other model input parameters are input into the target text-to-speech model, and the target text information is converted into speech to obtain the target speech signal.
4. The method according to claim 3, characterized in that The target text-to-speech model includes: a text feature encoding network and a speech feature decoding network; Inputting the target text information, the target emotion label, and the other model input parameters into the target text-to-speech model, and performing speech conversion on the target text information to obtain the target speech signal, including: Inputting the target text information, the target emotion label and the other model input parameters into the text feature encoding network for text feature fusion encoding to obtain target text features; The speech feature decoding network is used to perform a speech decoding operation on the target text feature to obtain the target speech signal.
5. The method according to claim 4, characterized in that The text feature encoding network includes: a text feature learning module, a vectorization module and a text encoder; The target text information, the target emotion label and the other model input parameters are input into the text feature encoding The coding network performs text feature fusion encoding to obtain the target text features, including: Inputting the target text information into the text feature learning module to learn linguistic feature information to obtain multi-level initial text features; Input the target emotion label and the other model input parameters into the corresponding vectorization module for vectorization processing to obtain emotion features and features of the other model input parameters; The initial text features, sentiment features and features of other model input parameters are input into the text encoder for feature encoding to obtain the target text features.
6. The method according to claim 4, characterized in that The speech feature decoding network includes: a text feature projection module, a speech length predictor and a speech decoder; Performing a speech decoding operation on the target text feature using the speech feature decoding network to obtain the target speech signal includes: Inputting the target text feature into the text feature projection module, and performing speech feature learning on the target text feature in a speech feature space with the same semantics to obtain a target speech feature; Inputting the target text feature into the speech length predictor, and predicting the speech length based on the target text feature to obtain a target noise signal; The target speech feature and the target noise signal are input into the speech decoder, and the target speech feature is decoded using the length of the target noise signal to obtain the target speech signal.
7. The method according to any one of claims 2 to 6, characterized in that: Also includes: Performing universal model training on the initial text-to-speech model using first universal sample data to obtain a universal text-to-speech model, wherein the first universal sample data includes universal sample speech signals, universal sample text information, and universal sample emotion labels corresponding to a plurality of first sample users; The first personalized sample data is used to perform personalized model training on the general text-to-speech model to obtain the target text-to-speech model, wherein the first personalized sample data includes the personalized sample voice signal, personalized sample text information and personalized sample emotion label of the target user.
8. The method according to any one of claims 1 to 6, characterized in that Performing expression mapping on the target speech signal according to a mapping relationship between the target user's voice features and expression coefficients to obtain a target expression coefficient sequence adapted to the target speech signal includes: Inputting the target speech signal into a target speech-to-expression model, mapping the target speech signal to expression coefficients to obtain the target expression coefficient sequence; The target speech-to-expression model is obtained by personalized training of a general speech-to-expression model based on the second personalized sample data of the target user, and is used to reflect the mapping relationship between the voice characteristics and expression coefficients of the target user.
9. The method according to claim 8, characterized in that The target speech-to-expression model includes: a speech feature encoding network, an expression feature decoding network and a linearization network; Inputting the target speech signal into a target speech-to-expression model, and mapping the target speech signal to expression coefficients to obtain the target expression coefficient sequence, comprising: Inputting the target speech signal into the speech feature coding network, performing speech feature extraction on the target speech signal to obtain target latent speech features; Inputting the target potential speech feature into the expression feature decoding network for expression decoding to obtain a target vertex offset sequence; Generating a target vertex sequence adapted to the target speech signal according to a reference vertex sequence and the target vertex offset sequence, wherein the reference vertex sequence is a vertex sequence corresponding to the digital human model in a neutral state; The target vertex sequence is linearized using the linearization network to obtain the target expression coefficient sequence.
10. The method according to claim 8, characterized in that Also includes: Performing universal model training on the initial speech-to-expression model using second universal sample data to obtain a universal speech-to-expression model, wherein the second universal sample data includes universal sample speech signals and universal sample expression coefficient sequences corresponding to a plurality of second sample users; The second personalized sample data is used to perform personalized model training on the general speech-to-expression model to obtain the target speech-to-expression model, wherein the second personalized sample data includes the personalized sample speech signal and personalized sample expression coefficient sequence of the target user.
11. The method according to any one of claims 1 to 6, characterized in that: The digital human model is a 3D Gaussian model, and the rendering of the digital human model according to the target speech signal and the target expression coefficient sequence to obtain video data of the digital human model includes: Using the target expression coefficient sequence to deform the 3D Gaussian model in the initial space to obtain a 3D Gaussian model in the standard space; The 3D Gaussian model in the standard space is linearly skinned using the bone transformation coefficient to obtain the 3D Gaussian model in the deformed space; The 3D Gaussian model in the deformation space is converted into a sequence of images to be rendered using Gaussian sputtering; Rendering the image sequence to be rendered to obtain an action image sequence; The action image sequence is combined with the target voice signal to obtain video data of the digital human model.
12. An interactive method based on digital human, characterized in that: include: Receive a question message sent to a digital human model in a target scene, and generate a reply text message according to the question message, wherein the target scene corresponds to a target emotion tag, and the digital human model corresponds to a target user; Performing voice conversion on the reply text message according to the voice characteristics of the target user and the target emotion label to obtain a target voice signal adapted to the target user; Performing expression mapping on the target speech signal according to a mapping relationship between the target user's voice features and expression coefficients to obtain a target expression coefficient sequence adapted to the target speech signal; The digital human model is rendered according to the target speech signal and the target expression coefficient sequence to obtain video data of the digital human model.
13. An electronic device, characterized in that: include: memory and processor; The memory is used to store computer programs; The processor is coupled to the memory and configured to execute the computer program for performing the steps of the method according to any one of claims 1 to 12.
14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is enabled to implement the steps of the method according to any one of claims 1 to 12.
15. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, causes the processor to implement the steps of the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Voice generation method and device, equipment, medium and product
CN114882861A
Face image generation method and device, equipment, medium and product
CN115761075A
Human-computer interaction method and device based on digital human, electronic equipment and storage medium
CN116775824A
Method for generating digital human voice and facial animation through text
CN116863038A
Video generation and interaction method and device based on digital human, storage medium and program product
CN117710543A
Cited By
Three-dimensional Gaussian digital human generation system and method and electronic equipment
CN120823342A
Three-dimensional digital human generation method and system capable of voice interaction
CN120931773A
AI digital human interactive response method based on large language model
CN121144484A
Large model-based audio and video data generation method, training method and intelligent agent
CN121306089A
Digital human video generation method, training method and device of digital human generation model, and computer equipment
CN121330444A