Digital human emotion modeling method based on deep learning

By obtaining multimodal stimulation data for emotional judgment and initialization, and combining emotional dynamic values ​​for emotional updates, the optimal emotional sequence is finally output, which solves the problem of time changes in digital human emotions modeling, and improves the accuracy of emotional modeling and the sense of realism of human-computer interaction.

CN120068628AActive Publication Date: 2025-05-30HANGZHOU DIGITAL SPACE TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510146630.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively model and simulate the time changes of digital human emotions, resulting in low accuracy in emotional modeling.

Method used

By obtaining multimodal stimulation data, including stimulation intensity, stimulation frequency, stimulation duration and emotion type, emotional judgment and initialization are performed, emotional updates are performed in combination with emotional dynamic values, and finally a pre-trained emotion optimization model is input to output the optimal emotion sequence.

Benefits of technology

It improves the accuracy of digital human emotion modeling, and can more accurately simulate and understand user emotional changes, thereby improving the authenticity and user experience of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068628A_ABST
    Figure CN120068628A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital human modeling, and discloses a digital human emotion modeling method based on deep learning, and the method comprises the steps: obtaining multi-modal stimulation data and a current emotion state, and the multi-modal stimulation data comprises the stimulation intensity, the stimulation frequency, the stimulation duration and the emotion type; performing emotion judgment according to the multi-modal stimulation data to obtain positive and negative emotion types; performing emotion initialization according to the emotion positive and negative type and the current emotion state to obtain an emotion initial value; performing emotion accumulation simulation according to the stimulation frequency to obtain an emotion dynamic value; performing emotion updating according to the emotion dynamic value and the emotion initial value to obtain an initial emotion sequence; and inputting the initial emotion sequence into a pre-trained emotion optimization model, and outputting a final optimal emotion sequence. According to the method, the digital human emotion modeling accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital human modeling, and in particular, to a digital human emotion modeling method based on deep learning. Background Art

[0002] With the development of artificial intelligence technology, the digital human, as a new type of human-computer interaction interface, is gradually becoming a research hotspot. The digital human can not only imitate human appearance and behavior, but also simulate human emotional responses through emotion modeling, so as to enhance the interaction experience with users. Especially in the fields of virtual assistants, customer service, and education and entertainment, digital humans with emotion recognition and expression capabilities can provide a more natural and smooth communication experience. However, to enable digital humans to accurately understand and respond to the emotional states of users, an efficient and accurate emotion modeling method is required.

[0003] In the prior art, the digital human emotion modeling method based on deep learning includes two main steps: First, a deep neural network model is trained using a large amount of labeled data to achieve the understanding of various modal inputs such as speech, facial expressions, and text. In the emotion recognition process, structures such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs) are used to analyze the speech features or facial expression changes of users, so as to judge their current emotional states. Second, based on the recognized emotional states, generative adversarial networks (GANs) or other types of generative models are used to synthesize corresponding expressions or language responses, enabling the digital human to adjust its response mode according to the emotions of users. This method can effectively improve the interaction realism and user experience of digital humans.

[0004] Although the above methods have made significant progress, there are still some challenges and limitations. Over time, the emotional intensity of digital humans will change. The existing methods cannot model the emotional changes caused by time changes, resulting in low accuracy of digital human emotion modeling. Summary of the Invention

[0005] The present invention provides a digital human emotion modeling method, system, electronic device, and storage medium based on deep learning to achieve the goal of improving the accuracy of digital human emotion modeling.

[0006] In a first aspect, to solve the above technical problems, the present invention provides a digital human emotion modeling method based on deep learning, including:

[0007] Obtain multimodal stimulus data and the current emotional state, where the multimodal stimulus data includes: stimulus intensity, stimulus frequency, stimulus duration, and emotion type;

[0008] Perform sentiment judgment based on the multi-modal stimulus data to obtain the positive / negative sentiment type;

[0009] Perform sentiment initialization based on the positive / negative sentiment type and the current sentiment state to obtain the initial sentiment value;

[0010] Perform sentiment accumulation simulation based on the stimulus frequency to obtain the dynamic sentiment value;

[0011] Perform sentiment update based on the dynamic sentiment value and the initial sentiment value to obtain the initial sentiment sequence;

[0012] Input the initial sentiment sequence into a pre-trained sentiment optimization model and output the final optimal sentiment sequence.

[0013] In an alternative embodiment, the performing sentiment judgment based on the multi-modal stimulus data to obtain the positive / negative sentiment type includes:

[0014] Perform numerical encoding on the multi-modal stimulus data to obtain a stimulus feature vector;

[0015] Input the stimulus feature vector into a pre-trained sentiment classification model and output the positive probability and the negative probability;

[0016] When the positive probability is greater than the negative probability, determine that the positive / negative sentiment type of the current stimulus is positive sentiment;

[0017] When the positive probability is less than the negative probability, determine that the positive / negative sentiment type of the current stimulus is negative sentiment;

[0018] Among them, the training process of the sentiment classification model includes:

[0019] Train based on a large-scale labeled sentiment corpus, and obtain the trained model after monitoring that the loss function of the model meets the conditions.

[0020] In an alternative embodiment, the performing sentiment initialization based on the positive / negative sentiment type and the current sentiment state to obtain the initial sentiment value includes:

[0021] Calculate the weight matrix through the following formula:

[0022] W T =α·W base +(1-α)·S current

[0023] Among them, W T represents the weight matrix, W base represents the pre-trained matrix, S current represents the current sentiment state, and α represents the decay coefficient;

[0024] Construct an input vector based on the weight matrix and the positive / negative sentiment type;

[0025] Input the input vector into a pre-trained sentiment impact value model to output an initial sentiment value;

[0026] Among them, the training process of the sentiment impact value model includes:

[0027] Construct a sentiment impact value model based on historical sentiment vectors, train the model, and determine that the training is completed after detecting that the loss function of the model meets the conditions, and obtain the trained sentiment impact value model.

[0028] In an alternative embodiment, the simulating sentiment accumulation according to the stimulation frequency to obtain a sentiment dynamic value includes:

[0029] The sentiment dynamic value is calculated by the following formula:

[0030]

[0031] Among them, I dynamic represents the sentiment dynamic value, β represents the accumulation coefficient, γ represents the accumulation effect index, N represents the total number of frequency samplings, k represents the frequency number, and f k represents the k-th stimulation frequency.

[0032] In an alternative embodiment, the updating the sentiment according to the sentiment dynamic value and the initial sentiment value to obtain an initial sentiment sequence includes:

[0033] Calculate the sentiment intensity by the following formula:

[0034] I t = I init ·e -λt + I dynamic

[0035] Among them, I t represents the sentiment intensity at the t-th time point, I init represents the initial sentiment value, e represents the natural constant, λ represents the natural decay coefficient, t is the time point number, and I dynamic represents the sentiment dynamic value;

[0036] Sort the sentiment intensities in ascending order of time points to obtain an initial sentiment sequence.

[0037] In an alternative embodiment, the inputting the initial sentiment sequence into a pre-trained sentiment optimization model to output a final optimal sentiment sequence includes:

[0038] When the stimulation duration is greater than a preset time threshold, perform hysteresis accumulation according to the emotion sequence function to obtain an initial hysteresis accumulation value;

[0039] Perform weighted averaging according to the emotion sequence function and the initial hysteresis accumulation value to obtain an updated emotion sequence;

[0040] Perform double-threshold detection according to the updated emotion sequence to obtain a discrete emotion sequence;

[0041] Perform gradient optimization according to the discrete emotion sequence to obtain an optimal emotion sequence.

[0042] In an alternative embodiment, the performing gradient optimization according to the discrete emotion sequence to obtain an optimal emotion sequence includes:

[0043] Calculate the gradient of the loss function through the following formula:

[0044]

[0045] Wherein, represents the gradient of the loss function, θ represents the hyperparameters of the emotion optimization model, m represents the number of samples, i represents the sample number, and h θ (x (i) ) represents the predicted value of the emotion optimization model for the i-th sample, y (i) represents the value in the discrete emotion sequence of the i-th sample, and x (i) represents the feature vector of the i-th sample;

[0046] Update the hyperparameters of the emotion optimization model according to the gradient of the loss function and the learning rate;

[0047] When it is detected that the loss function of the emotion optimization model is less than a preset detection threshold, output the optimal emotion sequence.

[0048] In a second aspect, the present invention provides a digital human emotion modeling system based on deep learning, including:

[0049] A data acquisition module, configured to acquire multi-modal stimulation data and the current emotion state, where the multi-modal stimulation data includes: stimulation intensity, stimulation frequency, stimulation duration, and emotion type;

[0050] An emotion positive and negative module, configured to perform emotion judgment according to the multi-modal stimulation data to obtain an emotion positive and negative type;

[0051] An initial emotion module, configured to perform emotion initialization according to the emotion positive and negative type and the current emotion state to obtain an initial emotion value;

[0052] A dynamic emotion module, configured to perform emotion accumulation simulation according to the stimulation frequency to obtain an emotion dynamic value;

[0053] An emotion sequence module, configured to perform emotion update according to the emotion dynamic value and the initial emotion value to obtain an initial emotion sequence;

[0054] A sequence optimization module, configured to input the initial emotion sequence into a pre-trained emotion optimization model and output a final optimal emotion sequence.

[0055] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for digital human emotion modeling based on deep learning described in any one of the above is implemented.

[0056] In a fourth aspect, the present invention further provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method for digital human emotion modeling based on deep learning described in any one of the above.

[0057] Compared with the prior art, the present invention has the following beneficial effects: The present invention discloses a method for digital human emotion modeling based on deep learning. The method includes obtaining multi-modal stimulation data and the current emotion state, where the multi-modal stimulation data includes: stimulation intensity, stimulation frequency, stimulation duration, and emotion type; performing emotion judgment according to the multi-modal stimulation data to obtain the positive and negative emotion types; performing emotion initialization according to the positive and negative emotion types and the current emotion state to obtain an initial emotion value; performing emotion accumulation simulation according to the stimulation frequency to obtain an emotion dynamic value; performing emotion update according to the emotion dynamic value and the initial emotion value to obtain an initial emotion sequence; inputting the initial emotion sequence into a pre-trained emotion optimization model and outputting a final optimal emotion sequence. This method can improve the accuracy of digital human emotion modeling.

[0058] Specifically, this method proposes an emotion judgment step for multi-modal stimulation data. First, perform numerical encoding on the collected multi-modal stimulation data to convert it into a stimulation feature vector that can be processed by a machine. This conversion helps to capture and represent the complex information in the original data and lays a foundation for subsequent emotion analysis.

[0059] Next, the obtained stimulus feature vector is input into a pre-trained sentiment classification model. This model aims to output a positive probability and a negative probability based on the input feature vector, so as to help determine the positive and negative types of sentiment. If the positive probability is higher than the negative probability, it is determined that the current stimulus has a positive sentiment; conversely, if the positive probability is lower than the negative probability, it is determined that the stimulus has a negative sentiment.

[0060] The improvement of this method in technology is mainly reflected in the accuracy of sentiment recognition. By combining multi-modal data and deep learning models, it can understand the user's emotional state more comprehensively and accurately. Compared with traditional methods, it not only considers a single type of input, but integrates multi-dimensional data, improving the accuracy of sentiment type judgment.

[0061] Furthermore, the present invention proposes a method for determining an emotional dynamic value and updating the emotion based on this value and the initial emotional state. First, for the calculation of the emotional dynamic value, a formula is proposed to quantify the emotional changes caused by external stimuli of different frequencies. Here, the "emotional dynamic value" represents the cumulative effect of emotions changing over time; the "cumulative coefficient" adjusts the speed or intensity of this emotional accumulation; the "cumulative effect index" affects the degree of contribution of different stimulus frequencies to emotional accumulation; the "total number of frequency samples" is the number of external stimulus frequencies considered; the "stimulus frequency" refers to each individual external stimulus event. The design of this calculation method aims to capture and quantify the emotional changes caused by external stimuli, providing a mathematical way to simulate how emotions change dynamically with external stimuli in real situations.

[0062] Secondly, for the calculation of emotional intensity, another formula is used to achieve the transition from the basic emotional state to the current emotional state. Here, the "emotional intensity" represents the emotional level at a certain point in time; the "initial emotional value" reflects the basic emotional state without external stimuli; a natural decay function describes the process of emotions naturally declining over time; the "natural decay coefficient" determines the speed at which the emotional intensity decreases over time; the "time point number" is used to identify different time points; the "emotional dynamic value" is the result obtained from the previous calculation.

[0063] This step realizes the transition from the basic emotional state to the current emotional state, and combines the natural decay of emotions over time and the dynamic changes in emotions brought about by recent stimuli. Finally, the calculated emotional intensity values at each time point are sorted in chronological order to generate an initial emotional sequence. This method allows the system to dynamically adjust its emotional expression to adapt to the changing environment and interaction needs, thus achieving a more natural and harmonious human-computer interaction experience. In addition, this strategy improves the scientificity and accuracy of emotional modeling, enabling the machine to more delicately and realistically imitate the emotional response patterns of humans. Throughout the process, through specific numerical calculations, the abstract emotional changes are transformed into quantifiable indicators, improving the accuracy of emotional modeling. Brief Description of the Drawings

[0064] Figure 1 is a schematic flowchart of a digital human emotional modeling method based on deep learning provided by the first embodiment of the present invention;

[0065] Figure 2 is a schematic structural diagram of a digital human emotional modeling system based on deep learning provided by the second embodiment of the present invention. Detailed Embodiments

[0066] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0067] With the development of artificial intelligence technology, the digital human, as a new type of human-computer interaction interface, is gradually becoming a research hotspot. The digital human can not only imitate the appearance and behavior of humans, but also simulate the emotional responses of humans through emotional modeling, thereby enhancing the interaction experience with users. Especially in the fields of virtual assistants, customer service, and education and entertainment, digital humans with the ability to recognize and express emotions can provide a more natural and smooth communication experience. However, to enable the digital human to accurately understand and respond to the emotional state of the user, an efficient and accurate emotional modeling method is required.

[0068] In the prior art, the digital human emotion modeling method based on deep learning includes two main steps: First, a deep neural network model is trained using a large amount of labeled data to achieve the understanding of various modal inputs such as speech, facial expressions, and text. In the emotion recognition process, structures such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs) are used to analyze the speech features or facial expression changes of users, so as to judge their current emotional state. Second, based on the recognized emotional state, generative adversarial networks (GANs) or other types of generative models are used to synthesize corresponding expressions or language responses, enabling the digital human to adjust its response mode according to the user's emotion. This method can effectively improve the interaction realism and user experience of the digital human.

[0069] Although the above methods have made remarkable progress, there are still some challenges and limitations. Over time, the emotional intensity of the digital human will change. The existing methods cannot model the emotional changes caused by time changes, resulting in low accuracy of digital human emotion modeling.

[0070] To solve the above problems, referring to Figure 1 , the first embodiment of the present invention provides a digital human emotion modeling method based on deep learning, including the following steps:

[0071] S11, obtaining multi-modal stimulus data and the current emotional state, where the multi-modal stimulus data includes: stimulus intensity, stimulus frequency, stimulus duration, and emotion type;

[0072] S12, making an emotion judgment according to the multi-modal stimulus data to obtain the positive and negative emotion types;

[0073] S13, performing emotion initialization according to the positive and negative emotion types and the current emotional state to obtain an initial emotion value;

[0074] S14, performing emotion accumulation simulation according to the stimulus frequency to obtain an emotion dynamic value;

[0075] S15, performing emotion update according to the emotion dynamic value and the initial emotion value to obtain an initial emotion sequence;

[0076] S16, inputting the initial emotion sequence into a pre-trained emotion optimization model and outputting the final optimal emotion sequence.

[0077] In step S11, multi-modal stimulus data and the current emotional state are obtained, where the multi-modal stimulus data includes: stimulus intensity, stimulus frequency, stimulus duration, and emotion type.

[0078] In one implementation, the acquisition of multi-modal stimulus data can be achieved in various ways: Visual data can capture the user's facial expressions and body movements through a camera. In a virtual reality or augmented reality environment, the user's behaviors and reactions can be recorded in real time; Auditory data is collected through a microphone to obtain the user's speech information, including features such as volume, intonation, and rhythm, which helps analyze the emotional tendency when the user is speaking; Text data extracts emotional words and semantic information from the user's input (such as chat records, comments) to understand their emotional state. The current emotional state is obtained by directly asking the user to self-report their emotional state. Through these methods, the system can comprehensively understand the user's emotional state and make corresponding responses accordingly.

[0079] It should be noted that the parameter explanations are as follows: The stimulus intensity reflects the strength or intensity of the external stimulus. For example, the decibel level of sound, image contrast, etc. can all be used as measurement criteria; The stimulus frequency describes the periodicity of the occurrence of a specific stimulus. For example, the frequency of a flashing light source, repeated sound patterns, etc. are all typical examples; The stimulus duration represents the length of time a certain stimulus acts. For example, the playing time of a piece of music or the display duration of a scene all belong to this category; The emotion type reflects the emotional reaction triggered by the external stimulus, such as pleasure, fear, anger, etc.

[0080] In step S12, emotional judgment is performed based on the multi-modal stimulus data to obtain the positive and negative emotion types.

[0081] In one implementation, numerical encoding is performed on the multi-modal stimulus data to obtain a stimulus feature vector; The stimulus feature vector is input into a pre-trained emotion classification model, and the positive probability and negative probability are output; When the positive probability is greater than the negative probability, it is determined that the positive and negative emotion type of the current stimulus is a positive emotion; When the positive probability is less than the negative probability, it is determined that the positive and negative emotion type of the current stimulus is a negative emotion; Among them, the training process of the emotion classification model includes: training based on a large-scale labeled emotion corpus, and obtaining the trained model after monitoring that the loss function of the model meets the conditions.

[0082] In one implementation, the emotion classification model uses a pre-trained multi-layer perceptron model. This multi-layer perceptron network is trained based on a large-scale labeled emotion corpus. During the training process, the model learns how to identify emotion-related patterns from the input features. During training, the network weights are adjusted by monitoring the performance of the loss function to ensure that the model can achieve a satisfactory performance level on the given data set. Once the loss function meets the predetermined conditions, the training is completed, and a pre-trained model that can be used for actual applications is obtained.

[0083] It should be noted that, for better illustration, an example is given here: Analyze a user comment from a social media platform with the aim of determining whether the sentiment expressed in this comment is positive or negative. The specific process is as follows: First, preprocess the comment text, such as removing stop words, tokenizing, etc., and then use the word embedding method to convert each word into its corresponding vector representation to form the feature vector of the entire comment. Pass this feature vector as input to a pre-trained multi-layer perceptron network. After a series of non-linear transformations by the multi-layer perceptron network, the probabilities of positive sentiment and negative sentiment are given at the output layer. For example, for a comment "The food in this restaurant is very delicious", the positive sentiment probability is output as 0.85 and the negative sentiment probability is 0.15. Since the positive sentiment probability is greater than the negative sentiment probability, it is determined to be positive sentiment.

[0084] In step S13, perform sentiment initialization according to the positive / negative sentiment type and the current sentiment state to obtain the initial sentiment value.

[0085] In one implementation, the weight matrix is calculated through the following formula:

[0086] W T = α·W base +(1 - α)·S current

[0087] where, W T represents the weight matrix, W base represents the pre-trained matrix, S current represents the current sentiment state, and α represents the decay coefficient;

[0088] Construct an input vector according to the weight matrix and the positive / negative sentiment type;

[0089] Input the input vector into a pre-trained sentiment impact value model to output the initial sentiment value;

[0090] where, the training process of the sentiment impact value model includes:

[0091] Construct a sentiment impact value model based on historical sentiment vectors, train the model, and determine that the training is completed after detecting that the loss function of the model meets the conditions to obtain the trained sentiment impact value model.

[0092] It should be noted that the weight matrix is used to combine the pre-trained matrix and the current emotional state to generate a new matrix that comprehensively reflects the influence of both. The pre-trained matrix is pre-trained based on a large amount of historical emotional data and contains the basic response patterns or weight distributions for different types of emotional stimuli. The current emotional state represents the emotional state currently recognized by the system and is a vector representation inferred from multi-modal data. The attenuation coefficient, a coefficient between 0 and 1, determines the relative contribution of the pre-trained matrix and the current emotional state in the newly generated weight matrix. When the attenuation coefficient is close to 1, the pre-trained matrix is more preferred; when the attenuation coefficient is close to 0, it is more dependent on the current emotional state.

[0093] In one implementation, the emotional influence value model is implemented using an LSTM network. This model is constructed and trained based on historical emotional vectors (i.e., emotional states and their changes at different past time points). During the training process, the model learns how to extract patterns related to emotional changes from the input features and minimize the difference between the predicted value and the actual value (loss function). Once it is detected that the loss function of the model reaches a certain predetermined standard, the model is considered to be trained and ready for predicting new initial emotional values.

[0094] In step S14, emotional accumulation simulation is performed according to the stimulation frequency to obtain an emotional dynamic value.

[0095] In one implementation, the emotional dynamic value is calculated by the following formula:

[0096]

[0097] where, I dynamic represents the emotional dynamic value, β represents the accumulation coefficient, γ represents the accumulation effect index, N represents the total number of frequency samplings, k represents the frequency number, and f k represents the k-th stimulation frequency.

[0098] It should be noted that the emotional dynamic value represents the cumulative emotional effect caused by the change in the external stimulus frequency. This value can be used to describe the trend of emotions gradually increasing or decreasing over time. Cumulative coefficient. Adjusts the speed or intensity of this emotional accumulation. Different application scenarios require adjusting this coefficient to better match the actual observed emotional responses. Cumulative effect index. Affects the degree to which different stimulus frequencies contribute to emotional accumulation. By adjusting this index, the different effects of high-frequency and low-frequency stimuli on the emotional state can be controlled. For example, if the value is large, high-frequency stimuli will have a greater impact on emotional accumulation. Total number of frequency samplings. Refers to the number of external stimulus frequencies considered. By summing up the contributions of all frequencies, this formula can reflect the changing trend of the emotional state under the combined action of multiple frequency stimuli over a long period of time. This is very important for understanding the emotional evolution under long-term exposure to certain environments.

[0099] In step S15, emotional update is performed based on the emotional dynamic value and the initial emotional value to obtain an initial emotional sequence.

[0100] In one implementation, the emotional intensity is calculated by the following formula:

[0101] I t =I init ·e -λt +I dynamic

[0102] Where, I t represents the emotional intensity at the t-th time point, i init represents the initial emotional value, e represents the natural constant, λ represents the natural decay coefficient, t is the time point number, and I dynamic represents the emotional dynamic value;

[0103] Sort the initial emotional sequence according to the emotional intensity from smallest to largest time point.

[0104] It should be noted that for the natural constant (Euler's number), the value taken in this method is 2.7. Natural decay coefficient. This coefficient determines the speed at which the emotional intensity naturally decreases over time. A larger value indicates that the emotional intensity will weaken faster. The purpose of this formula is to describe how the emotional intensity evolves over time. The first part shows the process of the emotional intensity naturally decaying over time. This decay reflects the phenomenon that human emotions do not remain at a high level indefinitely but will gradually return to the baseline level over time. The second part is the cumulative emotional effect calculated based on the external stimulus frequency. It does not directly depend on time but exists as an additional influencing factor on the emotional intensity. Even without new stimulus inputs, previously experienced high-frequency stimuli will still have a persistent impact on the current emotional state.

[0105] In step S16, the initial emotion sequence is input into a pre-trained emotion optimization model to output the final optimal emotion sequence.

[0106] In one implementation, when the stimulation duration is greater than a preset time threshold, hysteresis accumulation is performed according to the emotion sequence function to obtain an initial hysteresis accumulation value;

[0107] Weighted averaging is performed according to the emotion sequence function and the initial hysteresis accumulation value to obtain an updated emotion sequence;

[0108] Double-threshold detection is performed according to the updated emotion sequence to obtain a discrete emotion sequence;

[0109] Gradient optimization is performed according to the discrete emotion sequence, including:

[0110] The gradient of the loss function is calculated by the following formula:

[0111]

[0112] where, represents the gradient of the loss function, θ represents the hyperparameter of the emotion optimization model, m represents the number of samples, i represents the sample number, and h θ (x (i) ) represents the predicted value of the emotion optimization model for the i-th sample, y (i) represents the value in the discrete emotion sequence of the i-th sample, and x (i) represents the feature vector of the i-th sample;

[0113] The hyperparameters of the emotion optimization model are updated according to the gradient of the loss function and the learning rate;

[0114] When it is detected that the loss function of the emotion optimization model is less than a preset detection threshold, the optimal emotion sequence is output.

[0115] It should be noted that upper and lower threshold values are set to detect whether a significant change has occurred in the emotional state. For example, the upper threshold is set to twice the normal fluctuation, and the lower threshold is - twice. If the emotional intensity fluctuation within a certain period exceeds these thresholds, it is considered that a sudden change has occurred in the emotional state. The discrete emotion sequence is obtained after double-threshold detection, where continuous emotional intensity values are converted into a series of discrete emotion labels (such as happy, sad, etc.).

[0116] It should be noted that the hysteresis effect refers to the situation where when the duration of an external stimulus (such as vision, hearing, etc.) acting on the system exceeds a certain preset threshold, this long-term stimulus will cause a delay or cumulative effect in the change of the emotional state. This means that even when the external stimulus stops, the individual's emotional response will not immediately return to normal, but will continue for a period of time or take a longer time to return to the baseline level. Therefore, by setting a time threshold, it is used to judge whether the external stimulus is long enough to cause the hysteresis effect. If the duration of the stimulus exceeds this threshold, it is considered that the stimulus will trigger the hysteresis effect and subsequent steps will be carried out. When the duration of the stimulus exceeds the preset threshold, the system performs additional calculations according to the previously calculated emotional sequence function to determine the cumulative emotional intensity caused by the long-term stimulus. This reflects the accumulation of emotional changes caused by long-term exposure to certain stimuli.

[0117] It should be noted that the gradient descent method is used for optimization. The partial derivative is used to measure the direction and magnitude of the gap between the current model prediction value and the actual value. By minimizing this gradient, the model can better fit the training data. Hyperparameters represent the set of all adjustable parameters of the emotion optimization model, including weights, biases, etc. Further, mapping the residuals back to the input space can show how to adjust the model parameters corresponding to the input features to reduce the error. Different features have different impacts on the final prediction result, so their roles need to be considered separately.

[0118] In summary, the present invention discloses a digital human emotion modeling method based on deep learning, aiming to improve the accuracy of digital human emotion simulation through a series of steps. First, the method starts from obtaining multi-modal stimulus data, which includes stimulus intensity, stimulus frequency, stimulus duration, and emotion type, and simultaneously records the current emotional state. This process lays the foundation for subsequent emotion analysis. Then, based on the collected multi-modal stimulus data, emotion judgment is carried out to determine the positive and negative types of emotions. Here, the numerical coding method is used to convert the original data into a machine-processable stimulus feature vector, and then the pre-trained emotion classification model is used to analyze these feature vectors, and the probability values of positive or negative emotions are output. According to these probability values, the emotional tendency of the current stimulus can be judged as positive or negative.

[0119] Furthermore, the method proposed by the present invention completes emotion initialization according to the identified positive and negative emotion types and the current emotional state to obtain an initial emotion value. In this step, by calculating the weight matrix and combining the positive and negative emotion types to construct an input vector, and then inputting this input vector into the pre-trained emotion influence value model, the initial emotion value can be obtained. This process takes into account the influence of historical emotion data on the current emotional state, making emotion initialization more accurate.

[0120] Subsequently, the method performs emotional accumulation simulation based on the stimulation frequency and calculates the emotional dynamic value. This part introduces a formula to quantify the emotional changes caused by external stimuli, which covers parameters such as the accumulation coefficient and the accumulation effect index to describe the contribution degree of different frequency stimuli to emotional accumulation. Through this method, the changing trend of emotions over time can be effectively captured.

[0121] After obtaining the emotional dynamic value, it combines with the initial emotional value to update the emotion and generate the initial emotional sequence. Specifically, a natural decay function is used to describe the decay process of emotions over time, and the emotional intensity at each time point is calculated accordingly. Finally, it is arranged in chronological order to form the initial emotional sequence. This step realizes the transition from the basic emotional state to the current emotional state, and also takes into account the dynamic changes brought by the recent stimuli.

[0122] Finally, the initial emotional sequence is input into a pre-trained emotion optimization model. After a series of processes such as lag accumulation, weighted average, double-threshold detection, and gradient optimization, the final optimal emotional sequence is output. During this process, by calculating the gradient of the loss function and adjusting the hyperparameters, it is ensured that the result output by the model is close to the actual emotional state. It should be noted that for the long-term external stimulus effect, this method also considers the lag effect. Even when the external stimulus stops, the individual's emotional response will not immediately return to normal, but there will be a certain delay or cumulative effect.

[0123] In summary, the present invention provides a digital human emotion modeling method, which not only considers a single type of input, but also integrates multi-dimensional data such as vision, audition, and text, improving the accuracy of emotion type judgment. In addition, through the in-depth understanding and precise simulation of the emotional change law, the digital human improves the accuracy in understanding and responding to user emotions.

[0124] Referring to Figure 2 , the second embodiment of the present invention provides a digital human emotion modeling system based on deep learning, including:

[0125] A data acquisition module for acquiring multi-modal stimulus data and the current emotional state, where the multi-modal stimulus data includes: stimulus intensity, stimulus frequency, stimulus duration, and emotion type;

[0126] An emotion positive and negative module for making an emotion judgment based on the multi-modal stimulus data to obtain the emotion positive and negative type;

[0127] An initial emotion module for initializing the emotion based on the emotion positive and negative type and the current emotional state to obtain the initial emotional value;

[0128] A dynamic emotion module for performing emotional accumulation simulation based on the stimulus frequency to obtain the emotional dynamic value;

[0129] An emotional sequence module, configured to perform emotional update based on the emotional dynamic value and the emotional initial value to obtain an initial emotional sequence;

[0130] A sequence optimization module, configured to input the initial emotional sequence into a pre-trained emotional optimization model and output a final optimal emotional sequence.

[0131] Preferably, the data acquisition module is configured to:

[0132] Obtain multimodal stimulus data and the current emotional state, where the multimodal stimulus data includes: stimulus intensity, stimulus frequency, stimulus duration, and emotional type.

[0133] Preferably, the positive and negative emotion module is configured to:

[0134] Perform emotional judgment based on the multimodal stimulus data to obtain the positive and negative emotion types, including:

[0135] Perform numerical encoding based on the multimodal stimulus data to obtain a stimulus feature vector;

[0136] Input the stimulus feature vector into a pre-trained emotional classification model and output a positive probability and a negative probability;

[0137] When the positive probability is greater than the negative probability, determine that the positive and negative emotion type of the current stimulus is a positive emotion;

[0138] When the positive probability is less than the negative probability, determine that the positive and negative emotion type of the current stimulus is a negative emotion;

[0139] Wherein, the training process of the emotional classification model includes:

[0140] Training based on a large-scale labeled emotional corpus, and obtaining a trained model after monitoring that the loss function of the model meets the conditions.

[0141] Preferably, the initial emotion module is configured to:

[0142] Perform emotional initialization based on the positive and negative emotion type and the current emotional state to obtain an emotional initial value, including:

[0143] Calculate a weight matrix through the following formula:

[0144] W T = α·W base +(1 - α)·S current

[0145] Wherein, W T represents the weight matrix, W basedenotes the pre-trained matrix, S current denotes the current emotional state, and α denotes the decay coefficient;

[0146] Construct an input vector according to the weight matrix and the positive / negative type of emotion;

[0147] Input the input vector into a pre-trained emotion impact value model to output the initial emotion value;

[0148] Among them, the training process of the emotion impact value model includes:

[0149] Construct an emotion impact value model based on historical emotion vectors, train the model, and determine that the training is completed after detecting that the loss function of the model meets the conditions, and obtain the trained emotion impact value model.

[0150] Preferably, the dynamic emotion module is used for:

[0151] Perform emotion accumulation simulation according to the stimulation frequency to obtain an emotion dynamic value, including:

[0152] The emotion dynamic value is calculated by the following formula:

[0153]

[0154] Among them, I dynamic denotes the emotion dynamic value, β denotes the accumulation coefficient, γ denotes the accumulation effect index, N denotes the total number of frequency samplings, k denotes the frequency number, and f k denotes the k-th stimulation frequency.

[0155] Preferably, the emotion sequence module is used for:

[0156] Perform emotion update according to the emotion dynamic value and the initial emotion value to obtain an initial emotion sequence, including:

[0157] Calculate the emotion intensity by the following formula:

[0158] I t =I init ·e -λt +I dynamic

[0159] Among them, I t denotes the emotion intensity at the t-th time point, I init denotes the initial emotion value, e denotes the natural constant, λ denotes the natural decay coefficient, t is the time point number, and I dynamic denotes the emotion dynamic value;

[0160] Sort the emotion intensities in ascending order of time points to obtain an initial emotion sequence.

[0161] Preferably, the sequence optimization module is configured to:

[0162] Input the initial emotion sequence into a pre-trained emotion optimization model, and output the final optimal emotion sequence, including:

[0163] When the stimulation duration is greater than a preset time threshold, perform hysteresis accumulation according to the emotion sequence function to obtain an initial hysteresis accumulation value;

[0164] Perform weighted average according to the emotion sequence function and the initial hysteresis accumulation value to obtain an updated emotion sequence;

[0165] Perform double-threshold detection according to the updated emotion sequence to obtain a discrete emotion sequence;

[0166] Perform gradient optimization according to the discrete emotion sequence to obtain an optimal emotion sequence, including:

[0167] Calculate the loss function gradient through the following formula:

[0168]

[0169] Wherein, represents the loss function gradient, θ represents the hyperparameter of the emotion optimization model, m represents the number of samples, i represents the sample number, and h θ (x (i) ) represents the predicted value of the emotion optimization model for the i-th sample, y (i) represents the value in the discrete emotion sequence of the i-th sample, and x (i) represents the feature vector of the i-th sample;

[0170] Update the hyperparameters of the emotion optimization model according to the loss function gradient and the learning rate;

[0171] When it is detected that the loss function of the emotion optimization model is less than a preset detection threshold, output the optimal emotion sequence.

[0172] It should be noted that the digital human emotion modeling system based on deep learning provided in the embodiments of the present invention is used to execute all the process steps of the digital human emotion modeling method based on deep learning in the above embodiments, and the working principles and beneficial effects of the two correspond one by one, so they will not be elaborated here.

[0173] The embodiments of the present invention also provide an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a data acquisition program. When the processor executes the computer program, it implements the steps in the above embodiments of the digital human emotion modeling method based on deep learning, such as Figure 1The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above-described device embodiments, such as the data acquisition module.

[0174] Exemplarily, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.

[0175] The electronic device can be a computing device such as a desktop computer, a notebook, a palm computer, and a smart tablet. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above components are only examples of the electronic device and do not constitute a limitation on the electronic device. It may include more or fewer components than the above, or combine some components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.

[0176] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device and connects various parts of the entire electronic device using various interfaces and lines.

[0177] The memory can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory and invoking the data stored in the memory, the processor realizes various functions of the electronic device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0178] Among them, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0179] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0180] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A digital human emotion modeling method based on deep learning, characterized in that: include: Acquiring multimodal stimulation data and current emotional state, wherein the multimodal stimulation data includes: stimulation intensity, stimulation frequency, stimulation duration and emotional type; Performing emotional judgment according to the multimodal stimulation data to obtain positive and negative emotional types; Performing emotion initialization according to the positive and negative emotion types and the current emotion state to obtain an initial emotion value; Performing emotion accumulation simulation according to the stimulation frequency to obtain an emotion dynamic value; Performing emotion updating according to the emotion dynamic value and the emotion initial value to obtain an initial emotion sequence; The initial emotion sequence is input into a pre-trained emotion optimization model, and a final optimal emotion sequence is output.

2. The digital human emotion modeling method based on deep learning according to claim 1 is characterized in that: The emotional judgment is performed according to the multimodal stimulation data to obtain the positive and negative types of emotions, including: Performing numerical encoding according to the multimodal stimulation data to obtain a stimulation feature vector; Inputting the stimulus feature vector into a pre-trained sentiment classification model and outputting positive probability and negative probability; When the positive probability is greater than the negative probability, the positive or negative type of the emotion of the current stimulus is determined to be a positive emotion; When the positive probability is less than the negative probability, the positive or negative type of the emotion of the current stimulus is determined to be a negative emotion; The training process of the sentiment classification model includes: Based on large-scale labeled sentiment corpus training, after monitoring that the model's loss function meets the conditions, the trained model is obtained.

3. The digital human emotion modeling method based on deep learning according to claim 1 is characterized in that: The performing of emotion initialization according to the emotion positive and negative type and the current emotion state to obtain an emotion initial value includes: The weight matrix is ​​calculated by the following formula: W T =α·W base +(1-α)·S current Among them, W T represents the weight matrix, W base represents the pre-training matrix, S current represents the current emotional state, α represents the attenuation coefficient; Constructing an input vector according to the weight matrix and the positive and negative types of emotions; Inputting the input vector into a pre-trained sentiment impact value model and outputting a sentiment initial value; The training process of the emotion impact value model includes: A sentiment influence value model is constructed based on historical sentiment vectors, and the model is trained. After detecting that the loss function of the model meets the conditions, the training is determined to be completed, and the trained sentiment influence value model is obtained.

4. The digital human emotion modeling method based on deep learning according to claim 1 is characterized in that: The performing emotion accumulation simulation according to the stimulation frequency to obtain the emotion dynamic value includes: The emotional dynamic value is calculated by the following formula: Among them, I dynamic represents the emotional dynamic value, β represents the cumulative coefficient, γ represents the cumulative effect index, N represents the total number of frequency samples, k represents the frequency number, and f k represents the stimulation frequency k.

5. The digital human emotion modeling method based on deep learning according to claim 1 is characterized in that: The updating of emotions according to the emotion dynamic value and the emotion initial value to obtain an initial emotion sequence comprises: The sentiment intensity is calculated by the following formula: I t =I init ·e -λt +I dynamic Among them, I t Indicates the emotional intensity at time t, I init represents the initial value of emotion, e represents the natural constant, λ represents the natural attenuation coefficient, t is the time point number, I dynamic Indicates the dynamic value of emotion; According to the emotion intensity, the time points are sorted from small to large to obtain an initial emotion sequence.

6. The method for digital human emotion modeling based on deep learning according to claim 1, characterized in that: The step of inputting the initial emotion sequence into a pre-trained emotion optimization model and outputting a final optimal emotion sequence comprises: When the stimulus duration is greater than a preset time threshold, performing hysteresis accumulation according to the emotion sequence function to obtain an initial hysteresis accumulation value; Performing weighted averaging based on the emotion sequence function and the initial lag cumulative value to obtain an updated emotion sequence; Performing double threshold detection according to the updated emotion sequence to obtain a discrete emotion sequence; Gradient optimization is performed according to the discrete emotion sequence to obtain an optimal emotion sequence.

7. The method for digital human emotion modeling based on deep learning according to claim 6 is characterized in that: The step of performing gradient optimization according to the discrete emotion sequence to obtain an optimal emotion sequence comprises: The loss function gradient is calculated using the following formula: in, represents the gradient of the loss function, θ represents the hyperparameter of the sentiment optimization model, m represents the number of samples, i represents the sample number, and h θ (x(i)) represents the predicted value of the sentiment optimization model for the i-th sample, y(i) represents the value in the discrete sentiment sequence of the i-th sample, and x(i) represents the feature vector of the i-th sample; Update the hyperparameters of the sentiment optimization model according to the loss function gradient and the learning rate; When it is detected that the loss function of the emotion optimization model is less than the preset detection threshold, the optimal emotion sequence is output.

8. A digital human emotion modeling system based on deep learning, characterized in that: include: A data acquisition module, used to acquire multimodal stimulation data and current emotional state, wherein the multimodal stimulation data includes: stimulation intensity, stimulation frequency, stimulation duration and emotion type; A positive and negative emotion module, used to make an emotion judgment based on the multimodal stimulation data to obtain a positive and negative emotion type; An initial emotion module, used to perform emotion initialization according to the positive and negative emotion types and the current emotion state to obtain an initial emotion value; A dynamic emotion module, used to perform emotion accumulation simulation according to the stimulation frequency to obtain an emotion dynamic value; An emotion sequence module, used for updating the emotion according to the emotion dynamic value and the emotion initial value to obtain an initial emotion sequence; The sequence optimization module is used to input the initial emotion sequence into a pre-trained emotion optimization model and output a final optimal emotion sequence.

9. An electronic device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the deep learning-based digital human emotion modeling method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the deep learning-based digital human emotion modeling method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Micro-blog emotion prediction method based on weak supervised type multi-modal deep learning

    CN108108849A

  • Human-computer interaction method and system based on artificial intelligence

    CN118171662A

  • Human emotion detection

    US10592609B1

  • Generating synthetic sentiment using multiple transactions and bias criteria

    US20130297546A1

  • Method for estimating human emotions using deep psychological affect network and system therefor

    US20190347476A1