Digital human-driven model training and application method, device, equipment, storage medium and product

By integrating frame-level and sentence-level emotional classification results in the digital human-driven model, and optimizing the generation of digital human-driven parameters using long-term and short-term memory networks and self-attention networks, the problems of insufficient digital verbal coherence and expression richness in the prior art are solved, and higher anthropomorphic fidelity and user experience are achieved.

CN119940459APending Publication Date: 2025-05-06CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411998858.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the driving parameters of digital people generated based on speech have problems such as insufficient lip coherence and expression richness, resulting in insufficient fidelity of digital people and affecting user experience.

Method used

By obtaining a training sample set constructed based on speech data, including frame-level emotion classification results, sentence-level emotion classification results and digital human-driven parameters, a digital human-driven model with a long and short-term memory network structure is used for training. This model combines self-attention networks and integrates frame-level and sentence-level emotional classification results to optimize the generation of digital human-driven parameters.

Benefits of technology

The emotional expression ability of the digital human-driven model is improved, making the emotional expression of the output digital human-driven parameters more delicate, enhancing the anthropomorphic fidelity of the digital human and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940459A_ABST
    Figure CN119940459A_ABST
Patent Text Reader

Abstract

The invention discloses a digital human-driven model training and application method and device, equipment, a storage medium and a product. The method comprises the following steps: acquiring a training sample set, wherein each voice sample in the training sample set comprises a first sample label, a second sample label and a third sample label; training the digital human-driven model based on the training sample set until a trained digital human-driven model is obtained; a total loss function trained by the digital human-driven model is determined based on a first loss function, a second loss function and a third loss function, the first loss function represents a loss value corresponding to a frame-level sentiment classification result, and the second loss function represents a loss value corresponding to a sentence-level sentiment classification result; the third loss function represents a loss value corresponding to the digital human driving parameter. The personification fidelity of the digital human can be enhanced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a method, device, equipment, storage medium and product for training and applying a digital human driven model. Background Art

[0002] Digital Human or Meta Human is a digital human image close to human image created by digital technology. As a virtual character with digital appearance, Digital Human breaks the physical boundaries to provide anthropomorphic services and experiences, which is its core value. Its development trend is hyper-realism, tool-based, and strong interaction.

[0003] The technology of generating driving parameters of digital humans (for example, lip shape and expression driving parameters) based on speech is an important part of digital human technology. That is, through speech processing, the speech is mapped and the corresponding body driving parameters of the digital human are obtained. Based on the driving parameters, the effects of digital humans "speaking" can be achieved through post-drive rendering.

[0004] In the related art, the driving parameters of digital humans generated based on speech often have defects such as insufficient lip shape coherence and insufficient expression richness, that is, the digital humans are not realistic enough, which affects the user experience. Summary of the invention

[0005] In view of this, the embodiments of the present application provide a digital human driving model training and application method, device, equipment, storage medium and product, aiming to enhance the anthropomorphic realism of digital humans and improve user experience.

[0006] The technical solution of the embodiment of the present application is implemented as follows:

[0007] In a first aspect, an embodiment of the present application provides a digital human driving model training method, comprising:

[0008] Acquire a training sample set constructed based on speech data, wherein each speech sample in the training sample set includes: a first sample label, a second sample label, and a third sample label, wherein the first sample label represents a frame-level sentiment classification result of the speech sample after framing processing, the second sample label represents a sentence-level sentiment classification result of the speech sample, and the third sample label represents a digital human driving parameter of the speech sample;

[0009] The digital human driving model is trained based on the training sample set until a trained digital human driving model is obtained; wherein the total loss function of the digital human driving model training is determined based on a first loss function, a second loss function and a third loss function, wherein the first loss function represents the loss value corresponding to the frame-level sentiment classification result, the second loss function represents the loss value corresponding to the sentence-level sentiment classification result, and the third loss function represents the loss value corresponding to the digital human driving parameter.

[0010] In the above scheme, the digital human driving model adopts a long short-term memory network structure, the first loss function is used to obtain the loss value corresponding to the frame-level sentiment classification result for each network layer before the output layer of the long short-term memory network, the second loss function is used to obtain the loss value corresponding to the sentence-level sentiment classification result based on the output layer of the long short-term memory network, and the third loss function is used to obtain the loss value corresponding to the digital human driving parameter based on the output layer of the long short-term memory network.

[0011] In the above scheme, the digital human driving model also includes a self-attention network, which is used to determine the weight value corresponding to each frame after frame processing, wherein the predicted sentence-level sentiment classification result is obtained based on the data of the output layer of the long short-term memory network and the weight value corresponding to each frame.

[0012] In the above solution, the digital human driving model is trained based on the training sample set until a trained digital human driving model is obtained, including:

[0013] During the training process, the loss value of the total loss function determined based on the first loss function, the second loss function and the third loss function is obtained, and the model parameters of the digital human driving model are updated based on the loss value of the total loss function until a trained digital human driving model is obtained.

[0014] In the above solution, the step of obtaining a training sample set constructed based on speech data includes:

[0015] Acquire voice data and digital human driving parameters corresponding to the voice data, wherein the voice data has character emotion annotations and whole sentence emotion annotations;

[0016] Performing frame processing on the speech data, and determining the first sample label of each frame after the frame processing based on the character emotion annotation;

[0017] The second sample label is determined based on the whole sentence emotion annotation, and the third sample label is determined based on the digital human driving parameters corresponding to the voice data.

[0018] In the above solution, the character emotion annotation includes an indicator for distinguishing the intensity under the emotion category, so that the first sample label can indicate the intensity corresponding to the emotion category of each frame.

[0019] In a second aspect, an embodiment of the present application provides a digital human driving method, including:

[0020] Perform frame processing on the voice data to be driven;

[0021] The feature data after the frame processing is input into the digital human driving model trained by the method described in the first aspect of the embodiment of the present application to obtain the digital human driving parameters.

[0022] In a third aspect, the present application embodiment provides a digital human driving model training device, including:

[0023] An acquisition module is used to acquire a training sample set constructed based on speech data, wherein each speech sample in the training sample set includes: a first sample label, a second sample label and a third sample label, wherein the first sample label represents a frame-level sentiment classification result of the speech sample after framing processing, the second sample label represents a sentence-level sentiment classification result of the speech sample, and the third sample label represents a digital human driving parameter of the speech sample;

[0024] A training module is used to train the digital human driving model based on the training sample set until a trained digital human driving model is obtained; wherein the total loss function of the digital human driving model training is determined based on a first loss function, a second loss function and a third loss function, wherein the first loss function represents the loss value corresponding to the frame-level sentiment classification result, the second loss function represents the loss value corresponding to the sentence-level sentiment classification result, and the third loss function represents the loss value corresponding to the digital human driving parameter.

[0025] In a fourth aspect, an embodiment of the present application provides a digital human driving device, comprising:

[0026] A pre-processing module, used for performing frame processing on the voice data to be driven;

[0027] The driving module is used to input the feature data after frame processing into the digital human driving model trained by the digital human driving model training device described in the third aspect of the embodiment of the present application to obtain digital human driving parameters.

[0028] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory for storing a computer program that can be run on the processor, wherein when the processor is used to run the computer program, it executes the steps of the method described in the first aspect or the second aspect of the embodiment of the present application.

[0029] In a sixth aspect, an embodiment of the present application provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first aspect or the second aspect of the embodiment of the present application are implemented.

[0030] In a seventh aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the method described in the first aspect or the second aspect of the embodiment of the present application.

[0031] The technical solution provided by the embodiment of the present application is to obtain a training sample set constructed based on voice data, wherein each voice sample in the training sample set includes: a first sample label, a second sample label and a third sample label, wherein the first sample label represents the frame-level emotion classification result of the voice sample after frame processing, the second sample label represents the sentence-level emotion classification result of the voice sample, and the third sample label represents the digital human driving parameter of the voice sample; the digital human driving model is trained based on the training sample set until a trained digital human driving model is obtained; wherein the total loss function of the digital human driving model training is determined based on the first loss function, the second loss function and the third loss function, wherein the first loss function represents the loss value corresponding to the frame-level emotion classification result, the second loss function represents the loss value corresponding to the sentence-level emotion classification result, and the third loss function represents the loss value corresponding to the digital human driving parameter. In this way, the digital human driving model can be jointly trained based on the loss value corresponding to the frame-level emotion classification result, the loss value corresponding to the sentence-level emotion classification result and the loss value corresponding to the digital human driving parameter, which can effectively improve the training effect of the model, so that the emotional expression of the digital human driving parameter output by the trained digital human driving model is more delicate, the anthropomorphic fidelity of the digital human is enhanced, and the user experience is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a flow chart of a digital human driving model training method according to an embodiment of the present application;

[0033] Figure 2 This is a flow chart of the digital human driving method according to an embodiment of the present application;

[0034] Figure 3 This is a schematic diagram of the principle of emotion annotation in speech samples in the application embodiment of this application;

[0035] Figure 4 A schematic diagram of the principles of model construction and training for the application embodiment of this application;

[0036] Figure 5 This is a structural schematic diagram of a digital human driving model training device according to an embodiment of the present application;

[0037] Figure 6This is a schematic diagram of the structure of the digital human driving device according to the embodiment of the present application;

[0038] Figure 7 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0039] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0041] In the related technology, the method of generating digital human mouth shape and expression driving parameters based on speech is mainly divided into two technical routes, namely the rule mapping method and the deep learning method. Among them, the rule mapping method is to decompose the speech into basic pronunciation units such as vowels and vowels, and retrieve the corresponding mouth shape and expression parameters based on the established pronunciation basic units and mouth shape expression mapping library, and then solve problems such as stiff mouth shape switching through smoothing and other technologies. The deep learning method is based on speech data and the corresponding expression categories and mouth shape expression parameters, and obtains the emotion recognition model and the mouth shape expression parameter generation model through deep learning technology training. In the use stage, the emotion recognition model is first used to obtain the emotion category of the input speech, and then the speech and emotion category are simultaneously input into the mouth shape expression parameter generation model, and finally the mouth shape and expression driving parameters corresponding to the digital human speech are obtained.

[0042] The above-mentioned method for generating digital mouth shape and expression-driven parameters based on rule mapping has a simple process, but the generated mouth shape is not coherent enough in complex scenes, and the expression is not rich enough. Compared with the rule mapping method, the above-mentioned method based on deep learning improves the naturalness of mouth shape and expression, but the process is relatively complicated, and in the step of generating mouth shape and expression parameters, a single emotion category label is usually used as input and input into the network at the same time as the voice data, without target modeling of emotion features, and a single emotion category is used as a global parameter, which weakens the local strong emotion and the facial expression movement is not delicate enough.

[0043] Based on this, in various embodiments of the present application, in order to better train the digital human driving model, sentence-level emotion classification results and frame-level emotion classification results of speech samples are introduced, and the sentence-level emotion classification results and frame-level emotion classification results are integrated into the training of the digital human driving model, which can effectively improve the emotional expression ability of the digital human driving model, and make the emotional expression of the driving parameters output by the trained digital human driving model more delicate, enhance the anthropomorphic realism of the digital human, and improve the user experience.

[0044] The embodiment of the present application provides a digital human driving model training method, which can be independently applied to electronic devices with data processing capabilities, such as terminal devices, servers, etc., and can also be implemented by the cooperation between terminal devices and servers; wherein the terminal device can specifically be a computer, a smart phone, a personal digital assistant (PDA), etc.; the server can specifically be an application server or a Web server. In actual deployment, the server can be an independent server or a cluster server.

[0045] like Figure 1 As shown, the training method of the embodiment of the present application includes:

[0046] Step 101, obtaining a training sample set constructed based on speech data, wherein each speech sample in the training sample set includes: a first sample label, a second sample label and a third sample label, wherein the first sample label represents a frame-level sentiment classification result of the speech sample based on frame processing, the second sample label represents a sentence-level sentiment classification result of the speech sample, and the third sample label represents a digital human driving parameter of the speech sample.

[0047] Here, the speech data used for training can be selected from a specific database, and the speech data has corresponding digital human driving parameters, so that it is easy to construct the input and output data for training the digital driving model, that is, to obtain the speech sample for training the digital driving model, and the digital human driving parameters in the database can be used as the third sample label of the speech sample. In order to improve the emotional expression ability of the digital human driving model, the speech sample in the embodiment of the present application introduces a local emotion label (i.e., the first sample label) and a global emotion label (i.e., the second sample label), wherein the first sample label can be understood as the emotion classification result corresponding to the feature data after frame processing, and the second sample label can be understood as the emotion classification result corresponding to the entire speech data, so that the emotion parameters can be integrated into the subsequent model training as a training target based on the first sample label and the second sample label of the speech sample.

[0048] Step 102, training the digital human driving model based on the training sample set until a trained digital human driving model is obtained; wherein the total loss function of the digital human driving model training is determined based on a first loss function, a second loss function and a third loss function, wherein the first loss function represents the loss value corresponding to the frame-level sentiment classification result, the second loss function represents the loss value corresponding to the sentence-level sentiment classification result, and the third loss function represents the loss value corresponding to the digital human driving parameter.

[0049] It can be understood that the training method of the embodiment of the present application can jointly train the digital human driving model based on the loss value corresponding to the frame-level sentiment classification result, the loss value corresponding to the sentence-level sentiment classification result and the loss value corresponding to the digital human driving parameter, which can effectively improve the training effect of the model, so that the emotional expression of the driving parameters output by the trained digital human driving model is more delicate, thereby enhancing the anthropomorphic realism of the digital human and improving the user experience.

[0050] Exemplarily, the digital human driving model adopts a long short-term memory (LSTM) network structure, the first loss function is used to obtain the loss value corresponding to the frame-level sentiment classification result for each network layer before the output layer of the long short-term memory network, the second loss function is used to obtain the loss value corresponding to the sentence-level sentiment classification result based on the output layer of the long short-term memory network, and the third loss function is used to obtain the loss value corresponding to the digital human driving parameter based on the output layer of the long short-term memory network.

[0051] Here, the long short-term memory network is a special recurrent neural network (RNN) that aims to solve the long-term dependency problem faced by traditional RNN when processing long sequence data. It controls the inflow, retention and output of information by introducing a gating mechanism, thereby capturing the dependency of a longer sequence without gradient vanishing or exploding. Specifically, the long short-term memory network controls the flow of information by introducing cell states and gates, where the gates include forget gates, input gates and output gates. The forget gate is used to determine whether the cell state of the previous time step needs to be retained or forgotten, the input gate is used to determine whether the new information of the current time step needs to be updated to the cell state, and the output gate is used to control the output of LSTM (i.e., the new hidden state). The cell state is updated by combining the output of the forget gate and the input gate. The long short-term memory network can better handle important events with long intervals and delays in the time series, which is conducive to refining context-related emotional parameters, thereby improving the emotional expression ability of the digital human-driven model.

[0052] It should be noted that for the long short-term memory network adopted by the digital human driving model, the embodiment of the present application obtains the loss value corresponding to the frame-level sentiment classification result of each network layer based on the first loss function, and obtains the loss value of the sentence-level sentiment classification result corresponding to the output layer based on the second loss function. In this way, the loss values ​​of the global and local sentiment classification can be fused with the loss values ​​corresponding to the driving parameters, so that the digital human driving model can learn the emotional parameters that match the driving parameters, and thus have stronger emotional expression capabilities.

[0053] Exemplarily, the digital human driving model also includes a self-attention network, which is used to determine the weight value corresponding to each frame after frame processing, wherein the predicted sentence-level sentiment classification result is obtained based on the data of the output layer of the long short-term memory network and the weight value corresponding to each frame.

[0054] Here, when mapping frame-level data to sentence-level sentiment classification output, by introducing the self-attention mechanism corresponding to the self-attention network, different weight values ​​can be assigned to the data features after framing based on the self-attention network, thereby highlighting the local emotions corresponding to the framing, making the expression of the digital human richer, more delicate and realistic.

[0055] Exemplarily, the step of training the digital human driving model based on the training sample set until a trained digital human driving model is obtained includes:

[0056] During the training process, the loss value of the total loss function determined based on the first loss function, the second loss function and the third loss function is obtained, and the model parameters of the digital human driving model are updated based on the loss value of the total loss function until a trained digital human driving model is obtained.

[0057] Exemplarily, the parameters of the digital human driving model can be back propagated and learned based on the loss value of the total loss function. The back propagation learning adopts the back propagation algorithm (BP) to train the model. The back propagation algorithm mainly consists of two links (excitation propagation and weight update) that are repeatedly iterated until the network's response to the input reaches a predetermined target range. The learning process of the BP algorithm consists of a forward propagation process and a back propagation process. In the forward propagation process, the input information passes through the input layer and the hidden layer, and is processed layer by layer and transmitted to the output layer. If the expected output value is not obtained in the output layer, the square sum of the error between the output and the expected value is taken as the objective function, and the back propagation is turned to, and the partial derivative of the objective function to each neuron weight is obtained layer by layer, forming the gradient of the objective function to the weight vector, which is used as the basis for modifying the weight. The learning of the network is completed in the weight modification process. When the error reaches the expected value, the network learning ends. In this way, a trained digital human driving model can be obtained.

[0058] Exemplarily, the obtaining of a training sample set constructed based on speech data includes:

[0059] Acquire voice data and digital human driving parameters corresponding to the voice data, wherein the voice data has character emotion annotations and whole sentence emotion annotations;

[0060] Performing frame processing on the speech data, and determining the first sample label of each frame after the frame processing based on the character emotion annotation;

[0061] The second sample label is determined based on the whole sentence emotion annotation, and the third sample label is determined based on the digital human driving parameters corresponding to the voice data.

[0062] Here, in the model training stage, the speech data obtained by the embodiment of the present application needs to have character emotion annotations and whole sentence emotion annotations, and the speech data is framed based on preprocessing. For example, the speech data can be framed with a frame length of 25ms (milliseconds) and a frame shift of 10ms, that is, the length of each frame is 25ms, and there is an overlap of 25-10=15ms between two frames for frame processing. The feature data of each frame can use the corresponding character emotion annotation to determine the frame-level emotion classification result, that is, determine the first sample label of each frame.

[0063] Exemplarily, the character emotion annotation includes an indicator for distinguishing the intensity under the emotion category, so that the first sample label can indicate the intensity corresponding to the emotion category of each frame. In an application example, assuming that the sentence-level emotion classification results are divided into positive (H), neutral (N) and negative (S), the character emotion annotation can introduce the distinction of intensity based on the above positive and negative emotion categories, for example, divided into five levels from high to low according to the intensity, so as to highlight the strong emotion and better fit the emotion change. In an application example, the character emotion annotation can be divided into H5, H4, H3, H2, H1, N, S1, S2, S3, S4, S5 based on the emotion category and intensity. Taking H5 as an example, it means that the emotion category is "positive" and the intensity is "5". It can be understood that in other examples, the character emotion annotation can also distinguish the emotion category and intensity based on the continuous numerical coding method, and the embodiment of the present application is not limited to this. Accordingly, the first sample label (i.e., the frame-level emotion classification result) can also realize the distinction between the emotion category and the intensity under the corresponding category. In this way, the intensity of frame-level emotions can be learned during model training, making the digital human driving parameters more realistic.

[0064] Exemplarily, the present application also provides a digital human driving method, such as Figure 2 As shown, the method includes:

[0065] Step 201, performing frame processing on the voice data to be driven;

[0066] Step 202: Input the feature data after frame processing into the digital human driving model trained by the digital human driving model training method to obtain digital human driving parameters.

[0067] It is understandable that after the digital human driving model is obtained through training based on the aforementioned method, the embodiment of the present application can perform frame processing on the voice data to be driven, and input the feature data after the frame processing into the digital human driving model. The digital human driving model can obtain the frame-level emotion classification result and the sentence-level emotion classification result based on the feature data after the frame processing, and fuse the frame-level emotion classification result and the sentence-level emotion classification result to generate the digital human driving parameter. Exemplarily, the digital human driving parameter can include lip-shaped driving parameters and / or expression driving parameters, so that the digital human can be rendered based on the digital human driving parameters to achieve the effect of the digital human "speaking".

[0068] It should be pointed out that the digital human driving model trained by the training method of the embodiment of the present application has a more delicate emotional expression of the digital human driving parameters output by it, which enhances the anthropomorphic realism of the digital human and improves the user experience.

[0069] The present application is further described in detail below in conjunction with an application example.

[0070] This application embodiment relates to a method for generating digital population and expression driving parameters based on speech data. This application embodiment divides emotions into sentence-level emotions and character-level emotions, and integrates the emotion parameters into the digital population and expression generation network as one of the training targets. A frame-level emotion loss function is introduced in the middle layer of the training network to improve the model's emotion representation ability. A self-attention mechanism is introduced in the output layer to highlight local strong emotions, making the performance more realistic.

[0071] In the model training phase, the speech data is first annotated with whole sentence and character emotions to obtain global and local emotion type labels respectively. Based on the digital human mouth shape and expression driving parameter generation network (such as LSTM), the frame-level emotion classification loss function (corresponding to the first loss function mentioned above) is introduced in the middle layer, and the sentence-level emotion classification loss function (corresponding to the second loss function mentioned above) is introduced in the final output layer to increase the emotion modeling ability of the model. Finally, the frame-level emotion classification loss function, sentence-level emotion classification loss function and lip shape expression parameter prediction loss function (corresponding to the third loss function mentioned above) are combined to make the generated lip shape and expression emotion more realistic. When the frame-level data is mapped to the sentence-level emotion classification output, the self-attention mechanism is introduced to highlight the local expression, making the digital human expression more varied. In the testing phase, the speech data is directly used without emotion annotation, and directly input into the neural network to obtain the corresponding speaker mouth shape and expression driving parameters.

[0072] The method of this application embodiment includes the following steps: training data preprocessing, model construction and training, and model application. The following is a detailed description of each step:

[0073] 1) Training data preprocessing

[0074] In this application embodiment, the corresponding model training input and output data can be obtained based on the "person in the middle" voice data and the digital population type and expression parameters (blendshape) driven by it, and the voice data can be annotated with whole sentence emotions and character emotions, and then the voice data can be processed by frame division, and the speech frame-level emotion category can be obtained based on the character emotion annotation results. It should be pointed out that this application embodiment introduces character-level emotion annotation on the basis of whole sentence emotion annotation, and in the character-level emotion annotation, each emotion (except "neutral") is divided into five levels ("5", "4", "3", "2", "1") from high to low according to the intensity, highlighting the strong emotions and better fitting the emotional changes. Among them, the sentence-level emotion types are divided into "positive" (H), "neutral" (N), and "negative" (S), and the frame-level emotions are divided into H5, H4, H3, H2, H1, N, S1, S2, S3, S4, S5 according to the type and intensity. Taking H5 as an example, it means that the emotion type is "positive" and the degree is "5". In an application example, Figure 3 As shown in the example, the sentence-level sentiment category (i.e., the sentence-level sentiment classification result) is H. Then, each character is sentimentally labeled to obtain the character-level sentiment category. Finally, the speech is framed to obtain the frame-level sentiment category (i.e., the frame-level sentiment classification result).

[0075] 2) Model construction and training

[0076] This application embodiment introduces emotion classification as the target loss function, and divides it into the inter-layer frame-level loss function (i.e., the first loss function mentioned above) and the final sentence-level loss function (i.e., the second loss function mentioned above), and finally integrates it with the lip shape and expression loss function (i.e., the third loss function mentioned above), so that the model has a stronger emotion representation ability, and introduces a self-attention mechanism when the final frame-level structure is mapped to the sentence-level emotion classification, which can better learn the hierarchical emotions in the training data, making the model emotion fitting richer. The detailed structure is as follows Figure 4 shown.

[0077] Figure 4 The frame-level emotion Loss1 is the frame-level emotion classification loss function of each layer. The specific calculation formula is as follows:

[0078]

[0079] Among them, t is the number of frames, that is, the number of time steps of a single-layer network in the LSTN network, N is the maximum value of the time step, k is the frame-level emotion category number, K is the total number of emotion categories, and p tk C is the feature data after frame processing tkThe output after softmax transformation is the predicted frame-level sentiment classification result, y tk is the training target category label vector, that is, the first sample label mentioned above.

[0080] Figure 4 The sentence-level sentiment Loss2 maps the frame-level output to the sentence-level output and calculates the sentence-level sentiment loss function. The specific formula is as follows:

[0081]

[0082] Among them, m is the sentence-level sentiment category number, M is the total number of corresponding sentiment categories, and x m is the sentence-level output vector U after softmax transformation, that is, the predicted sentence-level sentiment classification result, q m is the sentence-level training target category label vector, i.e., the second sample label mentioned above.

[0083] For the calculation of the sentence-level output vector U, firstly, the weighted summation of each frame output is performed, and the formula is as follows:

[0084]

[0085] Where t=1...N represents the number of frames, β is the weighting coefficient, and O is the frame-level output vector, that is, the frame-level output is mapped to the sentence-level output through weighted summation, which is used to calculate the sentence-level sentiment loss function. Since in a speech, the emotions in different time periods are strong or weak, and the contribution to the final sentiment output is different, the self-attention mechanism is introduced here. Through parameter learning, a greater weight is given to the emotionally strong frame, and vice versa. The formula is as follows:

[0086] α t =Vf(WO t +b)+k

[0087]

[0088] Among them, α t is the frame-level output vector O t The corresponding initial weight values, V, W, b, k are all matrices or vectors obtained through training, and f is the activation function.

[0089] Figure 4 The lip shape and expression parameter Loss3 is the standard lip shape and expression parameter loss function, and the formula is as follows:

[0090]

[0091] Among them, y t is the predicted digital human driving parameter, y t` is the digital human driving parameter corresponding to the speech sample, that is, the third sample label, which integrates the inter-layer frame-level sentiment classification loss function, sentence-level sentiment classification loss function and lip expression parameter loss function. The specific formula is as follows

[0092]

[0093] Among them, a 1L , a2, a3 are Loss fusion parameters obtained by model training, Loss Ll represents the frame-level sentiment classification loss corresponding to each network layer, and L=1,...,l is the number of model layers.

[0094] 3) Model application

[0095] The constructed model is trained based on the training sample set obtained by the aforementioned training preprocessing to obtain a lip shape and expression parameter generation model (i.e., a digital human driven model) with high emotional representation. During the use phase, the speech is framed and input into the model to obtain the digital human shape and expression driven parameters. Due to the use of sentence-level and frame-level emotional labels, and the generation of digital human shape and expression driven parameters assisted by emotion recognition training, and the introduction of global and hierarchical loss functions to improve the model training effect, and the addition of a self-attention mechanism between layers to highlight emotional expression, the lip shape and expression driven parameters generated by the model of this application embodiment have delicate emotional expression and high anthropomorphic fidelity, thereby improving the user experience.

[0096] In order to implement the method of the embodiment of the present application, the embodiment of the present application also provides a digital human driving model training device, which corresponds to the above-mentioned digital human driving model training method, and each step in the above-mentioned digital human driving model training method embodiment is also fully applicable to the embodiment of the digital human driving model training device.

[0097] like Figure 5As shown, the digital human driving model training device includes: an acquisition module 501 and a training module 502. The acquisition module 501 is used to acquire a training sample set constructed based on speech data, and each speech sample in the training sample set includes: a first sample label, a second sample label and a third sample label, wherein the first sample label represents the frame-level emotion classification result of the speech sample after frame processing, the second sample label represents the sentence-level emotion classification result of the speech sample, and the third sample label represents the digital human driving parameter of the speech sample; the training module 502 is used to train the digital human driving model based on the training sample set until a trained digital human driving model is obtained; wherein the total loss function of the digital human driving model training is determined based on the first loss function, the second loss function and the third loss function, wherein the first loss function represents the loss value corresponding to the frame-level emotion classification result, the second loss function represents the loss value corresponding to the sentence-level emotion classification result, and the third loss function represents the loss value corresponding to the digital human driving parameter.

[0098] Exemplarily, the digital human driving model adopts a long short-term memory network structure, the first loss function is used to obtain the loss value corresponding to the frame-level sentiment classification result for each network layer before the output layer of the long short-term memory network, the second loss function is used to obtain the loss value corresponding to the sentence-level sentiment classification result based on the output layer of the long short-term memory network, and the third loss function is used to obtain the loss value corresponding to the digital human driving parameter based on the output layer of the long short-term memory network.

[0099] Exemplarily, the digital human driving model also includes a self-attention network, which is used to determine the weight value corresponding to each frame after frame processing, wherein the predicted sentence-level sentiment classification result is obtained based on the data of the output layer of the long short-term memory network and the weight value corresponding to each frame.

[0100] Exemplarily, the training module 502 is specifically used for:

[0101] During the training process, the loss value of the total loss function determined based on the first loss function, the second loss function and the third loss function is obtained, and the model parameters of the digital human driving model are updated based on the loss value of the total loss function until a trained digital human driving model is obtained.

[0102] Exemplarily, the acquisition module 501 is specifically used for:

[0103] Acquire voice data and digital human driving parameters corresponding to the voice data, wherein the voice data has character emotion annotations and whole sentence emotion annotations;

[0104] Performing frame processing on the speech data, and determining the first sample label of each frame after the frame processing based on the character emotion annotation;

[0105] The second sample label is determined based on the whole sentence emotion annotation, and the third sample label is determined based on the digital human driving parameters corresponding to the voice data.

[0106] Exemplarily, the character emotion annotation includes an indicator for distinguishing the intensity under the emotion category, so that the first sample label can indicate the intensity corresponding to the emotion category of each frame.

[0107] In actual application, the acquisition module 501 and the training module 502 can be implemented by a processor in the digital human driving model training device. Of course, the processor needs to run the computer program in the memory to implement its function.

[0108] It should be noted that: the digital human driving model training device provided in the above embodiment only uses the division of the above program modules as an example when performing digital human driving model training. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device is divided into different program modules to complete all or part of the above-described processing. In addition, the digital human driving model training device provided in the above embodiment and the digital human driving model training method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0109] In order to implement the method of the embodiment of the present application, the embodiment of the present application also provides a digital human driving device, which corresponds to the above-mentioned digital human driving method, and each step in the above-mentioned digital human driving method embodiment is also fully applicable to the embodiment of the digital human driving device.

[0110] like Figure 6 As shown, the digital human driving device includes: a preprocessing module 601 and a driving module 602. The preprocessing module 601 is used to perform frame processing on the voice data to be driven; the driving module 602 is used to input the feature data after the frame processing into the digital human driving model trained by the digital human driving model training device described in the third aspect of the embodiment of the present application to obtain the digital human driving parameters.

[0111] In actual application, the pre-processing module 601 and the driving module 602 can be implemented by a processor in the digital human driving device. Of course, the processor needs to run the computer program in the memory to realize its function.

[0112] Based on the hardware implementation of the above program modules and in order to implement the method of the embodiment of the present application, the embodiment of the present application also provides an electronic device. Figure 7Only an exemplary structure of the electronic device is shown, not all structures, and it can be implemented as needed. Figure 7 Partial or complete structure shown.

[0113] like Figure 7 As shown, the electronic device 700 provided in the embodiment of the present application includes: at least one processor 701, a memory 702, a user interface 703 and at least one network interface 704. The various components in the electronic device 700 are coupled together through a bus system 705. It can be understood that the bus system 705 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 705 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, in Figure 7 Various buses are labeled as bus system 705.

[0114] The user interface 703 may include a display, a keyboard, a mouse, a trackball, a click wheel, keys, buttons, a touch pad or a touch screen.

[0115] The memory 702 in the embodiment of the present application is used to store various types of data to support the operation of the electronic device. Examples of such data include: any computer program used to operate on the electronic device.

[0116] The digital human driving model training method and / or digital human driving method disclosed in the embodiment of the present application can be applied to the processor 701, or implemented by the processor 701. The processor 701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the digital human driving model training method and / or the digital human driving method can be completed by the hardware integrated logic circuit or software instructions in the processor 701. The above-mentioned processor 701 can be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 701 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiment of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in the memory 702. The processor 701 reads the information in the memory 702 and, in combination with its hardware, completes the steps of the digital human driving model training method and / or digital human driving method provided in the embodiment of the present application.

[0117] In an exemplary embodiment, the electronic device may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field programmable gate array (FPGA), general processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned method.

[0118] It can be understood that the memory 702 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.

[0119] In an exemplary embodiment, the present application also provides a computer storage medium, which can be a computer-readable storage medium, for example, a memory 702 storing a computer program, and the computer program can be executed by a processor 701 of an electronic device to complete the steps described in the method of the present application embodiment. The computer-readable storage medium can be a memory such as a ROM, a PROM, an EPROM, an EEPROM, a Flash Memory, a magnetic surface memory, an optical disk, or a CD-ROM.

[0120] In an exemplary embodiment, the embodiment of the present application further provides a computer program product, including a computer program, which can be executed by the processor 701 of the electronic device 700 to complete the steps described in the method of the embodiment of the present application.

[0121] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0122] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.

[0123] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A digital human driving model training method, characterized in that: include: Acquire a training sample set constructed based on speech data, wherein each speech sample in the training sample set includes: a first sample label, a second sample label, and a third sample label, wherein the first sample label represents a frame-level sentiment classification result of the speech sample after framing processing, the second sample label represents a sentence-level sentiment classification result of the speech sample, and the third sample label represents a digital human driving parameter of the speech sample; The digital human driving model is trained based on the training sample set until a trained digital human driving model is obtained; wherein the total loss function of the digital human driving model training is determined based on a first loss function, a second loss function and a third loss function, wherein the first loss function represents the loss value corresponding to the frame-level sentiment classification result, the second loss function represents the loss value corresponding to the sentence-level sentiment classification result, and the third loss function represents the loss value corresponding to the digital human driving parameter.

2. The method according to claim 1, characterized in that The digital human driving model adopts a long short-term memory network structure. The first loss function is used to obtain the loss value corresponding to the frame-level sentiment classification result for each network layer before the output layer of the long short-term memory network. The second loss function is used to obtain the loss value corresponding to the sentence-level sentiment classification result based on the output layer of the long short-term memory network. The third loss function is used to obtain the loss value corresponding to the digital human driving parameter based on the output layer of the long short-term memory network.

3. The method according to claim 2, characterized in that The digital human driving model also includes a self-attention network, which is used to determine the weight value corresponding to each frame after frame processing, wherein the predicted sentence-level sentiment classification result is obtained based on the data of the output layer of the long short-term memory network and the weight value corresponding to each frame.

4. The method according to claim 1, characterized in that The step of training the digital human driving model based on the training sample set until a trained digital human driving model is obtained includes: During the training process, the loss value of the total loss function determined based on the first loss function, the second loss function and the third loss function is obtained, and the model parameters of the digital human driving model are updated based on the loss value of the total loss function until a trained digital human driving model is obtained.

5. The method according to claim 1, characterized in that The obtaining of a training sample set constructed based on speech data comprises: Acquire voice data and digital human driving parameters corresponding to the voice data, wherein the voice data has character emotion annotations and whole sentence emotion annotations; Performing frame processing on the speech data, and determining the first sample label of each frame after the frame processing based on the character emotion annotation; The second sample label is determined based on the whole sentence emotion annotation, and the third sample label is determined based on the digital human driving parameters corresponding to the voice data.

6. The method according to claim 5, characterized in that The character emotion annotation includes an indicator for distinguishing the intensity under the emotion category, so that the first sample label can indicate the intensity corresponding to the emotion category of each frame.

7. A digital human driving method, characterized in that: include: Perform frame processing on the voice data to be driven; The feature data after frame processing is input into the digital human driving model trained by the method according to any one of claims 1 to 6 to obtain the digital human driving parameters.

8. A digital human driving model training device, characterized in that: include: An acquisition module is used to acquire a training sample set constructed based on speech data, wherein each speech sample in the training sample set includes: a first sample label, a second sample label and a third sample label, wherein the first sample label represents a frame-level sentiment classification result of the speech sample after framing processing, the second sample label represents a sentence-level sentiment classification result of the speech sample, and the third sample label represents a digital human driving parameter of the speech sample; A training module is used to train the digital human driving model based on the training sample set until a trained digital human driving model is obtained; wherein the total loss function of the digital human driving model training is determined based on a first loss function, a second loss function and a third loss function, wherein the first loss function represents the loss value corresponding to the frame-level sentiment classification result, the second loss function represents the loss value corresponding to the sentence-level sentiment classification result, and the third loss function represents the loss value corresponding to the digital human driving parameter.

9. A digital human driving device, characterized in that: include: A pre-processing module, used for performing frame processing on the voice data to be driven; The driving module is used to input the feature data after frame processing into the digital human driving model trained by the digital human driving model training device as claimed in claim 8 to obtain digital human driving parameters.

10. An electronic device, characterized in that: include: A processor and a memory for storing a computer program that can be executed on the processor, wherein: The processor is used to execute the steps of the method according to any one of claims 1 to 7 when running a computer program.

11. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.