A Voice Sentiment Analysis Method Driven by Facial Expressions of Virtual Characters

By constructing inter-frame difference coefficients and emotion transfer coefficients, and adjusting the probability distribution of speech signal frames, the problem of ignoring the emotional correlation of the time dimension in speech emotion analysis is solved, thereby improving the coherence and naturalness of the facial expressions of virtual characters.

CN122090885AActive Publication Date: 2026-05-26BEIJING MIAOYIN ANIMATION CULTURE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING MIAOYIN ANIMATION CULTURE CO LTD
Filing Date
2026-04-24
Publication Date
2026-05-26

Smart Images

  • Figure CN122090885A_ABST
    Figure CN122090885A_ABST
Patent Text Reader

Abstract

This application relates to the field of speech analysis technology and proposes a speech emotion analysis method driven by facial expressions of virtual characters. The method includes: acquiring speech signals in different emotional states to obtain different speech signal frames; extracting different speech emotion features from the speech signal frames, pre-setting different time scales, and constructing inter-frame difference coefficients for the speech signal frames; associating the speech signals with emotional states to obtain the probability distribution of all emotional states and the dominant emotion of the speech signal frames, calculating the prior probability of the dominant emotion of the speech signal frames, and calculating the emotion transfer coefficient of the speech signal frames; and obtaining an adjustment probability distribution based on the inter-frame difference coefficients and emotion transfer coefficients of the speech signal frames, wherein the adjustment probability distribution is the result of speech emotion analysis. This application aims to improve the coherence and naturalness of facial expressions of virtual characters driven by speech emotion recognition results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech analysis technology, specifically to a speech emotion analysis method driven by the facial expressions of virtual characters. Background Technology

[0002] Voice emotion analysis driven by virtual character facial expressions is a cross-modal emotion computing technology. It aims to drive the real-time facial expressions of virtual characters by extracting, recognizing and quantifying the speaker's emotional features such as happiness, sadness, anger and calmness from the speech signal. These emotional features are then mapped to the corresponding facial expression parameters of the virtual character, ultimately achieving precise synchronization between voice emotion and virtual facial expression. Facial expression parameters include the degree of mouth corner raising, eyebrow arch raising angle, and eyelid opening and closing.

[0003] In the process of sentiment analysis of speech, speech segments are generally treated as isolated units for sentiment recognition, ignoring the contextual sentiment association and lacking modeling of the evolution of emotions over time. This can cause frequent emotional "jitter" in the sentiment recognition results corresponding to different consecutive speech segments, resulting in abrupt and mechanical switching of facial expressions of the driven virtual characters, which seriously weakens the naturalness and credibility of the virtual characters' facial expressions. Summary of the Invention

[0004] This application provides a speech emotion analysis method driven by facial expressions of virtual characters to solve the problem that neglecting the emotional correlation of speech in the time dimension leads to insufficient naturalness of facial expressions of virtual characters driven by speech emotion recognition results. The specific technical solution adopted is as follows:

[0005] One embodiment of this application provides a voice emotion analysis method driven by facial expressions of virtual characters, the method comprising the following steps:

[0006] Collect speech signals in different emotional states and obtain different speech signal frames;

[0007] Different speech emotion features of speech signal frames are extracted. Different time scales are preset. Under the same preset time scale, the differences in the changing trends of the same type of speech emotion features of speech signal frames are analyzed. The same-direction change difference of speech signal frames is calculated. Based on the same-direction change difference of speech signal frames determined under all preset time scales, the inter-frame difference coefficient of speech signal frames is constructed. The inter-frame difference coefficient is used to characterize the same-direction monotonicity and the intensity of fluctuation of speech emotion features of speech signal frames under all time scales.

[0008] Based on all speech samples in the time-annotated speech database, the speech signals are associated with emotional states to obtain the probability distribution of all emotional states and the dominant emotion of the speech signal frames. The prior probability of the dominant emotion from the first speech signal frame to the last speech signal frame in two speech signal frames with a preset time interval is calculated. Based on the prior probability of the dominant emotion of the speech signal frame and other speech signal frames within all preset time intervals, as well as the difference between the probability distribution of all emotional states of the speech signal frame, the emotion transfer coefficient of the speech signal frame is calculated. The emotion transfer coefficient is used to characterize the credibility of the dominant emotion of the speech signal frame.

[0009] Based on the inter-frame difference coefficient and emotion transfer coefficient of the speech signal frame, the probability distribution of all emotional states of the speech signal frame is adjusted to obtain the adjusted probability distribution, which is the result of speech emotion analysis.

[0010] Furthermore, the specific method for determining the differences in the same direction of change of the speech signal frames is as follows:

[0011] Establish speech emotion feature sequences at different preset time scales, calculate the slope of each sequential position in the speech emotion feature sequence, and calculate the feature slope ratio of the speech emotion feature sequence based on the proportion of positive and negative slopes.

[0012] Any position in the speech emotion feature sequence is denoted as the target position. The first absolute value of the target position is determined based on the difference in slope between the target position and all previous positions in the speech emotion feature sequence. The first absolute value is used to characterize the monotonicity of the target position in the same direction.

[0013] Determine the speech signal frame corresponding to the target sequence position. The product of the first absolute value of the sequence position corresponding to the determined speech signal frame and the feature slope ratio of the speech emotion feature sequence in which the sequence position is located is denoted as the first product of the speech emotion features of the corresponding category with respect to the speech signal frame. The mean of the first products of all categories of speech emotion features with respect to the speech signal frame is denoted as the same-direction change difference of the speech signal frame.

[0014] Furthermore, the method for constructing the feature slope ratio is as follows:

[0015] The maximum ratio of positive to negative slopes of the fitted curve at all positions in the speech emotion feature sequence is used as the denominator, and the difference between the number 1 and the denominator is used as the numerator. The result of the fraction calculation is recorded as the feature slope ratio of the speech emotion feature sequence.

[0016] Furthermore, the method for determining the first absolute value of the target order position is as follows:

[0017] The average slope of the target position is the average slope of the target position, which is the mean of the slopes of the target position and all positions preceding the target position in the speech emotion feature sequence. The absolute value of the difference between the slope of the target position and the average slope is the first absolute value of the target position.

[0018] Furthermore, the inter-frame difference coefficient of the speech signal frame is: the positive correlation processing result of the same-direction change difference of the speech signal frames determined under all preset time scales.

[0019] Furthermore, the specific method for associating speech signals with emotional states and obtaining the probability distribution of all emotional states of speech signal frames based on all speech samples from a time-annotated speech database includes:

[0020] Obtain sentiment labels from speech samples in a time-annotated speech database, where each sentiment label corresponds one-to-one with the sentiment state of the speech signal; train a sentiment probability distribution model using the speech samples, and obtain the trained sentiment probability distribution model.

[0021] Input the speech signal frame into the emotion probability distribution model to obtain the probability distribution of all emotional states of the speech signal frame.

[0022] Furthermore, the method for determining the dominant emotion is as follows:

[0023] The emotional state with the highest probability in the probability distribution is taken as the dominant emotion of the speech signal frame.

[0024] Furthermore, the emotion transfer coefficient of the speech signal frame is positively correlated with the prior probability of the dominant emotion of the speech signal frame and other speech signal frames within all preset time scales, and the emotion transfer coefficient of the speech signal frame is negatively correlated with the difference in probability distribution of all emotional states of the speech signal frame and other speech signal frames within all preset time scales.

[0025] Furthermore, the specific method for adjusting the probability distribution of all emotional states in the speech signal frame based on the inter-frame difference coefficient and emotion transfer coefficient to obtain the adjusted probability distribution includes:

[0026] The transfer weights of the speech signal frames are calculated. The transfer weights of the speech signal frames are positively correlated with the emotion transfer coefficients of the speech signal frames, and negatively correlated with the inter-frame difference coefficients of the speech signal frames.

[0027] The probability distribution of all emotional states in the first frame of the speech signal is directly used as the adjusted probability distribution of all emotional states in the first frame of the speech signal.

[0028] Starting from the second speech signal frame, the adjustment probability distribution of all emotional states in the speech signal frame is determined based on the probability distribution and transition weight of all emotional states in the speech signal frame, as well as the adjustment probability distribution of all emotional states in the previous speech signal frame.

[0029] Furthermore, the adjustment probability distribution of all emotional states in the speech signal frames starting from the second speech signal frame is: the result of weighted summation of the adjustment probability distribution of all emotional states in the previous speech signal frame and the probability distribution of all emotional states in the speech signal frame, using the transition weights of the speech signal frames.

[0030] The beneficial effects of this application are:

[0031] Although inter-frame emotional differences in different speech signals can lead to overall signal instability, this application considers the close correlation between the emotions carried by the speech in different speech signal frames. It pre-sets different time scales for analysis, evaluating the monotonicity and fluctuation intensity of the emotional features of speech signal frames at different time scales, obtaining the inter-frame difference coefficient, and realizing the quantitative analysis of inter-frame speech data differences and emotional changes. The larger the inter-frame difference coefficient, the more significant the monotonicity trend of the emotional changes in the speech signal frame. Furthermore, it learns the emotional transfer rules through a speech database and correlates the learned emotional transfer rules with inter-frame speech data fluctuations, improving the accuracy of speech emotion recognition. The prior probability of the dominant emotion of a frame is the emotion transfer law. Therefore, based on the difference between the prior probability of the dominant emotion of a speech signal frame and other speech signal frames within all preset time scales, as well as the probability distribution of all emotion states, the credibility of the dominant emotion of the speech signal frame is evaluated, and the emotion transfer coefficient of the speech signal frame is obtained. Finally, based on the inter-frame difference coefficient and the emotion transfer coefficient of the speech signal frame, the probability distribution of all emotion states of the speech signal frame is adjusted to obtain the results of speech emotion analysis. This solves the problem of neglecting the emotional correlation of speech in the time dimension, which leads to insufficient naturalness of the facial expressions of virtual characters driven by the speech emotion recognition results, and improves the coherence and naturalness of the facial expressions of virtual characters driven by the speech emotion recognition results. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1This is a schematic flowchart of a voice emotion analysis method driven by facial expressions of virtual characters, provided as an embodiment of this application. Detailed Implementation

[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0035] Please see Figure 1 The diagram illustrates a flowchart of a voice emotion analysis method driven by facial expressions of virtual characters, provided in one embodiment of this application. The method includes the following steps:

[0036] Step S001: Collect speech signals of different emotional states and obtain different speech signal frames.

[0037] Voice signals of different emotional states are extracted from the voice emotion database that drives the facial expressions of virtual characters. All voice signals are then segmented and windowed to obtain different voice signal frames.

[0038] In this embodiment, five emotional states are selected: joy, sadness, anger, neutrality, and disgust. Framing and windowing of the speech signal are well-known techniques and will not be described further. This embodiment uses a first-order high-pass filter to filter the speech signal and then frames the filtered speech signal, with each frame having a length of 30ms and a frame shift of 10ms. A Hamming window is used as the window function to perform windowing operations on the framed speech signal frames.

[0039] At this point, different audio signal frames have been obtained.

[0040] Step S002: Extract different speech emotion features from the speech signal frames, preset different time scales, and analyze the differences in the changing trends of the same type of speech emotion features in the speech signal frames under the same preset time scale. Calculate the same-direction change difference of the speech signal frames. Based on the same-direction change difference of the speech signal frames determined under all preset time scales, construct the inter-frame difference coefficient of the speech signal frames. The inter-frame difference coefficient is used to characterize the same-direction monotonicity and the intensity of fluctuation of the speech emotion features of the speech signal frames under the corresponding time scale.

[0041] With the development of human-computer interaction technology, a large number of digital human virtual avatars have emerged, effectively improving the human-computer interaction experience. The core step in the generation of digital human virtual avatars is facial expression generation. To improve the quality of facial expression generation, it is necessary to accurately identify emotional changes in speech signals. Current speech emotion analysis methods mostly adopt frame-by-frame processing, performing independent emotion recognition on short frames obtained after frame-by-frame processing. However, this process severs the emotional connections between frames, easily leading to significant temporal discontinuities and jumps in the emotion recognition results, ultimately resulting in inconsistent facial expression transitions in the generated digital human. Therefore, there is an urgent need for a coherent and dynamic speech emotion analysis method to improve the accuracy of recognizing emotional states and continuous changes in speech signals.

[0042] The aforementioned method of performing independent sentiment analysis on frame-by-frame processed speech signals stems from the characteristics of speech signals: macroscopically, speech signals exhibit unstable characteristics, but microscopically, they can be approximated as stable over a short period, typically 10ms-30ms. However, in reality, although the emotional differences between frames lead to overall signal instability, the emotions carried by different frames of speech are not completely isolated but closely related, specifically manifested in the continuous nature of emotional changes.

[0043] First, emotional changes exhibit a gradual and continuous nature. The emotional differences between different speech signal frames transition gradually, with very few instantaneous abrupt changes. This ensures a continuous correlation between the emotions carried by adjacent speech signal frames. Second, the continuity of emotional changes determines the monotonous and unidirectional nature of emotional evolution between frames. Specifically, speech features or combinations of speech features will show an evolutionary trend in the same direction, and the overall feature intensity changes monotonically. Finally, the analysis of emotional continuity needs to consider multiple time scales. Specifically, using a smaller time scale can improve the accuracy of measuring inter-frame differences and facilitate the identification of the influence of instantaneous noise; using a larger time scale can better measure the overall trend of emotional feature changes. Therefore, introducing multi-time-scale comparative analysis helps to strengthen the extraction of inter-frame emotional change features while reducing the interference of uncontrollable instantaneous noise.

[0044] First, different speech emotion features are extracted from the speech signal frames. Based on the extracted speech emotion features, speech emotion feature sequences are established at different preset time scales. The order of the speech emotion features in the speech emotion feature sequence is used as the independent variable, and the speech emotion features are used as the dependent variable. Curve fitting is performed to obtain the fitted curve, and the slope of the fitted curve at each position in the speech emotion feature sequence is calculated. The maximum ratio of positive to negative slopes at all positions in the speech emotion feature sequence is used as the denominator, and the difference between the digit 1 and the denominator is used as the numerator. The result of the fraction calculation is recorded as the feature slope ratio of the speech emotion feature sequence.

[0045] The extraction of speech emotion features and curve fitting are well-known techniques and will not be elaborated further. The speech emotion features extracted in this embodiment include the short-time average zero-crossing rate, short-time average energy, short-time energy change rate, short-time average amplitude, maximum fundamental frequency, average fundamental frequency, average fundamental frequency, average fundamental frequency change rate, maximum, average, and average change rate of the first, second, and third formant frequencies, and 12th-order MFCC coefficients. As other implementation methods, while achieving the goal of speech emotion feature extraction, the implementer may select other speech emotion features for extraction; this application does not impose any special limitations. This embodiment uses polynomial fitting technology to achieve curve fitting.

[0046] In this embodiment, the preset time scale includes 10 frames of audio signal, 20 frames of audio signal, and 30 frames of audio signal.

[0047] The process of determining the speech emotion feature sequence under different preset time scales is as follows: for each preset time scale, the sequence of all speech emotion features of the same type in all speech signal frames under the time scale is obtained by arranging them in sequence.

[0048] Under a preset time scale, any position in the speech emotion feature sequence is denoted as the target position. The average slope of the target position and all previous positions in the speech emotion feature sequence is denoted as the average slope of the target position. The absolute value of the difference between the slope of the target position and the average slope is denoted as the first absolute value of the target position. The speech signal frame corresponding to the target position is determined. The product of the first absolute value of the determined speech signal frame and the ratio of the feature slope of the speech emotion feature sequence in which the position is located is denoted as the first product of the speech emotion features of the corresponding category with respect to the speech signal frame. The average of the first products of all categories of speech emotion features with respect to the speech signal frame is denoted as the same-direction change difference of the speech signal frame. The positive correlation processing result of the same-direction change differences of all speech signal frames determined under the preset time scale is denoted as the inter-frame difference coefficient of the speech signal frame.

[0049] It is understood that positive correlation processing is applied to the differences in the same direction of speech signal frames determined under all preset time scales, that is, to ensure that the differences in the same direction of speech signal frames determined under all preset time scales are positively correlated with the inter-frame difference coefficient of the speech signal frames. It is understood that the positive correlation in this application refers to the relationship between the independent variable and the dependent variable, where the independent variable is the differences in the same direction of speech signal frames determined under all preset time scales, and the dependent variable is the inter-frame difference coefficient of the speech signal frames. The positive correlation means that the dependent variable increases (decreases) as the independent variable increases (decreases), and can be an additive relationship, a multiplicative relationship, etc.

[0050] Preferably, as an embodiment of this application, the normalization function is set as follows: ,in, Normalization function The independent variable, The sum of the normalized values ​​of the differences in the same direction of speech signal frames determined under all preset time scales, which is a natural constant, is denoted as the inter-frame difference coefficient of the speech signal frame. The calculation of the normalized value is achieved through the normalization function.

[0051] The first absolute value of the target sequence position is used to evaluate the monotonicity of the target sequence position in the same direction; the difference between the slopes of adjacent sequence positions in the speech emotion feature sequence is used to evaluate the monotonicity trend of adjacent speech emotion features in the same direction and the intensity of emotion feature fluctuations. The larger the inter-frame difference coefficient of the speech signal frame, the more significant the monotonicity trend of the emotion in the speech signal frame is.

[0052] At this point, the inter-frame difference coefficient for each audio signal frame is obtained.

[0053] Step S003: Based on all speech samples in the time-annotated speech database, associate speech signals with emotional states, obtain the probability distribution of all emotional states and dominant emotions of speech signal frames, calculate the prior probability of the dominant emotion from the first speech signal frame to the last speech signal frame in two speech signal frames at a preset time interval, and calculate the emotion transfer coefficient of the speech signal frame based on the prior probability of the dominant emotions of the speech signal frame and other speech signal frames within all preset time scales, as well as the difference in the probability distribution of all emotional states of the speech signal frame. The emotion transfer coefficient is used to characterize the credibility of the dominant emotion of the speech signal frame.

[0054] The inter-frame difference coefficient of speech signal frames enables quantitative analysis of differences in speech data and emotional changes between frames, but it has significant limitations. Specifically, the calculation of the inter-frame difference coefficient needs to take into account the differences in multiple speech features, while the feature change patterns corresponding to different emotional states differ. Using an averaging difference measurement method may result in inter-frame difference coefficients obtained from different emotional state changes being similar. Therefore, relying solely on the inter-frame difference coefficient is insufficient to accurately measure the emotional state and dynamic evolution of speech signal frames.

[0055] In a continuous emotional space, the speed and shift of emotions follow a regular pattern, with almost no large jumps. Therefore, emotional shift patterns can be learned from a speech database, and the learned patterns can be correlated with fluctuations in inter-frame speech data to improve the accuracy of speech emotion recognition.

[0056] Speech samples and corresponding emotion tags are obtained from a time-annotated speech database by those skilled in the art. Each emotion tag corresponds one-to-one with the emotion state of the speech signal. In this embodiment, the emotion states are joy, sadness, anger, neutrality, and disgust. The time-annotated speech database refers to a database where each speech sample is equipped with a frame-level emotion category tag sequence aligned with the time axis, meaning that any speech frame can uniquely determine its emotion category. Typical databases include IEMOCAP and RECOLA. All speech samples are divided into training and testing sets in a 7:3 ratio. A one-dimensional CNN model is used to train the emotion probability distribution model, obtaining the trained emotion probability distribution model. The optimizer is Adam, and the activation function is ReLU. Speech signal frames are input into the emotion probability distribution model to obtain the probability distribution of all emotion states of the speech signal frames. The emotion state with the highest probability in the probability distribution is taken as the dominant emotion of the speech signal frame.

[0057] It is important to note that the selected speech samples must ensure that the calculation of the prior probabilities and emotion transfer coefficients is meaningful.

[0058] Based on the dominant sentiment of the speech samples, the prior probability of the dominant sentiment from the first speech signal frame to the last speech signal frame is calculated between two speech signal frames at a preset time interval. The calculation formula is as follows: ;

[0059] in, Indicates the first Frame of audio signal to the first frame Prior probability of dominant sentiment in frame-by-frame speech signal The preset time scale is indicated. In this embodiment, the preset time scale includes 10 frames of voice signal, 20 frames of voice signal, and 30 frames of voice signal. Indicates the first The dominant emotion of a frame of speech signal; Indicates the first The dominant emotion of a frame of speech signal; Indicates the first The dominant emotion of a frame of speech signal; Indicates the first The dominant emotion of a frame of speech signal; Indicates the speech sample from Transfer to with The number of different emotional states; Indicates the speech sample from arrive The number of times that situation occurs.

[0060] Emotional changes are continuous and usually do not involve large-scale abrupt changes. They generally follow a certain trajectory of emotional transfer. Therefore, the prior probability of different dominant emotional transitions can be determined based on the probability distribution of emotional states.

[0061] The emotion transfer coefficient of a speech signal frame is calculated based on the prior probability of the dominant emotion between the speech signal frame and other speech signal frames within all preset time scales, and the difference in probability distribution of all emotion states of the speech signal frame. The emotion transfer coefficient of the speech signal frame is positively correlated with the prior probability of the dominant emotion between the speech signal frame and other speech signal frames within all preset time scales, and negatively correlated with the difference in probability distribution of all emotion states between the speech signal frame and other speech signal frames within all preset time scales.

[0062] Preferably, as an embodiment of this application, the negative correlation processing result of the KL divergence of the probability distribution of all emotional states of a speech signal frame and other speech signal frames within a preset time scale is recorded as the first difference of the speech signal frame at the preset time scale. The product of the prior probability of the dominant emotion of the speech signal frame and other speech signal frames within the preset time scale with the first difference of the speech signal frame at the same preset time scale is recorded as the emotional change continuity of the speech signal frame at the preset time scale. The positive correlation processing result of the emotional change continuity of all speech signal frames at the preset time scale is recorded as the emotional transfer coefficient of the speech signal frame.

[0063] The calculation of KL divergence is a well-known technique and will not be elaborated further.

[0064] It is understood that negative correlation processing is applied to the KL divergence of the probability distribution of all emotional states of the speech signal frame and other speech signal frames within a preset time scale, that is, to ensure that the KL divergence of the probability distribution of the emotional states is negatively correlated with the first difference. It is understood that the negative correlation in this application refers to the relationship between the independent variable and the dependent variable, where the independent variable is the KL divergence of the probability distribution of the emotional states, and the dependent variable is the first difference. The negative correlation means that the dependent variable decreases (increases) as the independent variable increases (decreases), and can be an inverse relationship, a subtraction relationship, etc.

[0065] Preferably, as an embodiment of this application, the negative of the KL divergence of the probability distribution of all emotional states of a speech signal frame and other speech signal frames within a preset time scale is used as the exponent of an exponential function with the natural constant as the base, and the calculation result of the exponential function is recorded as the first difference of the speech signal frames under the preset time scale.

[0066] Preferably, as an embodiment of this application, the summation of the continuity of emotional changes in all preset time scale speech signal frames is recorded as the emotional transfer coefficient of the speech signal frame.

[0067] Based on the KL divergence of the probability distributions of different emotional states, the risk of abrupt changes between different emotional states can be assessed. The more the emotional transfer from the dominant emotion of a speech signal frame to other dominant emotions conforms to the prior evolutionary laws and maintains a smooth distribution transition, the higher the credibility of the emotional change in the dominant emotion of the speech signal frame. In determining the facial expression of the virtual character corresponding to the speech signal frame, the higher the credibility of the dominant emotion of the speech signal frame, and the larger the emotional transfer coefficient of the speech signal frame.

[0068] At this point, the emotion transfer coefficients of the speech signal frames are obtained.

[0069] Step S004: Based on the inter-frame difference coefficient and emotion transfer coefficient of the speech signal frame, adjust the probability distribution of all emotional states of the speech signal frame to obtain the adjusted probability distribution, which is the result of speech emotion analysis.

[0070] The transfer weight of a speech signal frame is calculated based on the inter-frame difference coefficient and the emotion transfer coefficient of the speech signal frame. The transfer weight of the speech signal frame is positively correlated with the emotion transfer coefficient of the speech signal frame, and negatively correlated with the inter-frame difference coefficient of the speech signal frame.

[0071] Preferably, as an embodiment of this application, the negative of the inter-frame difference coefficient of the speech signal frame is used as the exponent of an exponential function with the natural constant as the base. The calculation result of the exponential function is recorded as the first coefficient of the speech signal frame. The normalized value of the product of the first coefficient of the speech signal frame and the emotion transfer coefficient of the speech signal frame is recorded as the second coefficient of the speech signal frame. The maximum value of the second coefficient of the speech signal frame and the preset constant is recorded as the transfer weight of the speech signal frame.

[0072] It should be noted that this embodiment uses the maximum-minimum normalization method to calculate the normalized value. In actual application, implementers may use other methods of existing technology, such as the sigmoid function, to calculate the normalized value, which is not limited here; in this embodiment, the value of the preset constant is 0.3.

[0073] The probability distribution of all emotional states in the first audio signal frame is directly used as the adjusted probability distribution of all emotional states in the first audio signal frame. Starting from the second audio signal frame, the adjusted probability distribution of all emotional states in the previous audio signal frame and the probability distribution of all emotional states in the audio signal frame are weighted and summed according to the transition weight of the audio signal frame to obtain the adjusted probability distribution of all emotional states in the audio signal frame.

[0074] Specifically, in the weighted summation process, the weights of the probability distributions of all emotional states in the speech signal frame are the transition weights of the speech signal frame, and the weights of the adjusted probability distributions of all emotional states in the previous speech signal frame are the difference between the digit 1 and the transition weights of the previous speech signal frame. The formula for calculating the adjusted probability distributions of all emotional states in the speech signal frame is: ;

[0075] in, Indicates the first The adjusted probability distribution of all emotional states in a frame of audio signal, where, It is an integer greater than or equal to 2; Indicates the first The transition weights of the audio signal frames; Indicates the first The probability distribution of all emotional states in a frame of audio signal; Indicates the first The adjusted probability distribution of all emotional states in a frame of speech signal.

[0076] In speech sentiment analysis, abrupt changes in emotion and interference from noise can lead to significant shifts in the emotional state identified from the speech signal. This can cause the facial expressions of virtual characters driven by the speech signal to flicker, resulting in abrupt and mechanical facial expression transitions. To improve the accuracy of speech sentiment analysis and suppress abrupt changes that do not conform to prior rules, a weighted fusion method is used, based on the probability distribution of all emotional states in different speech signal frames preceding each previous frame and the reliability evaluation results of the emotional states in each speech signal frame. This method ensures the continuity of emotional state changes while avoiding temporal fragmentation and jumps in the emotion recognition results, thereby improving the quality of emotional state recognition corresponding to the speech signal.

[0077] At this point, the adjusted probability distribution of all emotional states in all speech signal frames has been obtained, and the adjusted probability distribution is the result of speech sentiment analysis.

[0078] At this point, based on the speech signal, the speech emotion analysis results used to drive the facial expressions of the virtual character are obtained.

[0079] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A speech emotion analysis method for virtual character facial expression driving, characterized in that, The method includes the following steps: Collect speech signals in different emotional states and obtain different speech signal frames; Different speech emotion features are extracted from speech signal frames. Different time scales are preset. Under the same preset time scale, the differences in the changing trends of the same type of speech emotion features in the speech signal frames are analyzed. The same-direction change difference of the speech signal frames is calculated. Based on the same-direction change difference of the speech signal frames determined under all preset time scales, the inter-frame difference coefficient of the speech signal frames is constructed. The inter-frame difference coefficient is used to characterize the same-direction monotonicity and the intensity of fluctuation of the speech emotion features of the speech signal frames under all time scales. Based on all speech samples in a time-annotated speech database, the speech signals are associated with emotional states to obtain the probability distribution of all emotional states and the dominant emotion of the speech signal frames. The prior probability of the dominant emotion from the first speech signal frame to the last speech signal frame in two speech signal frames separated by a preset time scale is calculated. Based on the prior probability of the dominant emotion of the speech signal frame and other speech signal frames within all preset time scales, as well as the difference in the probability distribution of all emotional states, the emotion transfer coefficient of the speech signal frame is calculated. The emotion transfer coefficient is used to characterize the credibility of the dominant emotion of the speech signal frame. Based on the inter-frame difference coefficient and emotion transfer coefficient of the speech signal frame, the probability distribution of all emotional states of the speech signal frame is adjusted to obtain the adjusted probability distribution, which is the result of speech emotion analysis.

2. The voice emotion analysis method driven by facial expressions of virtual characters according to claim 1, characterized in that, The specific method for determining the difference in the same direction of change of the speech signal frames is as follows: Establish speech emotion feature sequences at different preset time scales, calculate the slope of each sequential position in the speech emotion feature sequence, and calculate the feature slope ratio of the speech emotion feature sequence based on the proportion of positive and negative slopes. Any position in the speech emotion feature sequence is denoted as the target position. The first absolute value of the target position is determined based on the difference in slope between the target position and all previous positions in the speech emotion feature sequence. The first absolute value is used to characterize the monotonicity of the target position in the same direction. Determine the speech signal frame corresponding to the target sequence position. The product of the first absolute value of the sequence position corresponding to the determined speech signal frame and the feature slope ratio of the speech emotion feature sequence in which the sequence position is located is denoted as the first product of the speech emotion features of the corresponding category with respect to the speech signal frame. The mean of the first products of all categories of speech emotion features with respect to the speech signal frame is denoted as the same-direction change difference of the speech signal frame.

3. The voice emotion analysis method driven by facial expressions of virtual characters according to claim 2, characterized in that, The method for constructing the characteristic slope ratio is as follows: The maximum ratio of positive to negative slopes of the fitted curve at all positions in the speech emotion feature sequence is used as the denominator, and the difference between the number 1 and the denominator is used as the numerator. The result of the fraction calculation is recorded as the feature slope ratio of the speech emotion feature sequence.

4. The voice emotion analysis method driven by facial expressions of virtual characters according to claim 2, characterized in that, The method for determining the first absolute value of the target order position is as follows: The average slope of the target position is the average slope of the target position, which is the mean of the slopes of the target position and all positions preceding the target position in the speech emotion feature sequence. The absolute value of the difference between the slope of the target position and the average slope is the first absolute value of the target position.

5. The voice emotion analysis method driven by facial expressions of virtual characters according to claim 1, characterized in that, The inter-frame difference coefficient of the speech signal frame is the positive correlation processing result of the same-direction change difference of the speech signal frames determined under all preset time scales.

6. The voice emotion analysis method driven by facial expressions of virtual characters according to claim 1, characterized in that, The specific method for associating speech signals with emotional states and obtaining the probability distribution of all emotional states of speech signal frames based on all speech samples from a time-annotated speech database includes: Obtain sentiment labels from speech samples in a time-annotated speech database, where each sentiment label corresponds one-to-one with the sentiment state of the speech signal; train a sentiment probability distribution model using the speech samples, and obtain the trained sentiment probability distribution model. Input the speech signal frame into the emotion probability distribution model to obtain the probability distribution of all emotional states of the speech signal frame.

7. The voice emotion analysis method driven by facial expressions of virtual characters according to claim 1, characterized in that, The method for determining the dominant emotion is as follows: The emotional state with the highest probability in the probability distribution is taken as the dominant emotion of the speech signal frame.

8. The voice emotion analysis method driven by facial expressions of virtual characters according to claim 1, characterized in that, The emotion transfer coefficient of the speech signal frame is positively correlated with the prior probability of the dominant emotion of the speech signal frame and other speech signal frames within all preset time scales, and the emotion transfer coefficient of the speech signal frame is negatively correlated with the difference in probability distribution of all emotional states of the speech signal frame and other speech signal frames within all preset time scales.

9. A voice emotion analysis method driven by facial expressions of virtual characters according to claim 1, characterized in that, The method for adjusting the probability distribution of all emotional states in a speech signal frame based on the inter-frame difference coefficient and emotion transfer coefficient to obtain the adjusted probability distribution includes the following specific methods: The transfer weights of the speech signal frames are calculated. The transfer weights of the speech signal frames are positively correlated with the emotion transfer coefficients of the speech signal frames, and negatively correlated with the inter-frame difference coefficients of the speech signal frames. The probability distribution of all emotional states in the first frame of the speech signal is directly used as the adjusted probability distribution of all emotional states in the first frame of the speech signal. Starting from the second speech signal frame, the adjustment probability distribution of all emotional states in the speech signal frame is determined based on the probability distribution and transition weight of all emotional states in the speech signal frame, as well as the adjustment probability distribution of all emotional states in the previous speech signal frame.

10. A voice emotion analysis method driven by facial expressions of virtual characters according to claim 9, characterized in that, The adjusted probability distribution of all emotional states in the speech signal frames starting from the second speech signal frame is: the result of weighted summation of the adjusted probability distribution of all emotional states in the previous speech signal frame and the probability distribution of all emotional states in the speech signal frame, using the transition weights of the speech signal frames.