Emotion category determination method and device, equipment, storage medium and product
By extracting and fusion of various features of audio data, the problem of low accuracy of customer service voice emotion recognition in the prior art is solved, and a higher accuracy of emotional category determination is achieved.
Patent Information
- Application Number
- CN202510248752.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The existing multimodal model has low accuracy when determining the emotional categories of audio data corresponding to customer service voice, mainly due to the poor correlation between text and audio data.
By extracting the Mel cepspectral coefficient MFCC features, pinyin syllable features, tone embedding features, word embedding features, position embedding features and segment embedding features of the audio data, and fuse them and input them into the emotion recognition model to determine the target emotion category.
It improves the accuracy of determining the emotional categories of audio data, integrates various features such as pronunciation characteristics, text characteristics, tone characteristics and tone characteristics, and enhances the recognition ability of the model.
Smart Images

Figure CN119993215A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition technology, and in particular to a method, device, equipment, storage medium and product for determining an emotion category. Background Art
[0002] With the development of society, people's demand for high-quality services is also increasing. As the first window directly facing users, the service quality of customer service outbound call service can directly affect user satisfaction. In order to improve user satisfaction, customer service voice can be quality inspected. Emotion is one of the factors that represent service quality. Therefore, quality inspection can be performed by emotion recognition of the audio data corresponding to the customer service voice.
[0003] At present, emotion recognition of audio data corresponding to customer service speech is mainly carried out through multimodal models to determine the emotion category of audio data. However, the existing multimodal models only use the audio features of audio data and the text features of the text converted from audio data to determine the emotion category of audio data. The correlation between text and audio data is not strong, resulting in low accuracy in determining the emotion category of audio data. Summary of the invention
[0004] The embodiments of the present application provide a method, apparatus, device, storage medium and product for determining an emotion category, which can improve the accuracy of determining the emotion category of audio data.
[0005] In a first aspect, an embodiment of the present application provides a method for determining an emotion category, comprising:
[0006] Get audio data;
[0007] Extract the Mel-frequency cepstral coefficient MFCC feature of the audio data and the first feature of the pinyin syllable of the audio data;
[0008] Convert the audio data into text and determine the intonation embedding features of each word in the text, as well as the word embedding features, position embedding features, and segment embedding features of the text;
[0009] The MFCC feature, the first feature, the intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature are fused to obtain the second feature;
[0010] The second feature is input into the emotion recognition model, and the target emotion category corresponding to the second feature is determined according to the relationship information between the preset features and the preset emotion categories in the emotion recognition model, and the target emotion category is determined to be the emotion category of the audio data.
[0011] In a possible implementation, obtaining audio data includes:
[0012] Get the original voice data;
[0013] Perform channel segmentation on the original speech data and extract the audio data of the preset channel.
[0014] In a possible implementation, extracting a first feature of a pinyin syllable of audio data includes:
[0015] Extracting pinyin syllables of audio data, where the pinyin syllables include pinyin and tones;
[0016] The pinyin and the tone are encoded according to a first preset encoding method to obtain a first feature.
[0017] In one possible implementation, determining the intonation embedding feature of each word in the text includes:
[0018] Get the mean volume, pitch and duration of each word in the text;
[0019] Cluster the words in the text according to the mean volume, mean pitch and sound length to obtain the intonation category of each word in the text;
[0020] The intonation category is encoded according to the second preset encoding method to obtain the intonation embedding feature of each word in the text.
[0021] In a possible implementation, the MFCC feature, the first feature, the intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature are fused to obtain the second feature, including:
[0022] The intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature are integrated to obtain the third feature;
[0023] The MFCC feature, the first feature and the third feature are fused to obtain the second feature.
[0024] In one possible embodiment, before inputting the second feature into the emotion recognition model and determining the target emotion category corresponding to the second feature according to the relationship information between the preset feature and the preset emotion category in the emotion recognition model, the method further includes:
[0025] Obtain audio data samples and actual emotion categories of the audio data samples;
[0026] Extracting MFCC feature samples of the audio data samples and fourth features of the pinyin syllables of the audio data samples;
[0027] Converting the audio data sample into a text sample, and determining the intonation embedding feature sample of each word in the text sample, as well as the word embedding feature sample, the position embedding feature sample and the segment embedding feature sample of the text sample;
[0028] The MFCC feature sample, the fourth feature, the intonation embedding feature sample, the word embedding feature sample, the position embedding feature sample and the segment embedding feature sample are fused to obtain the fifth feature;
[0029] Inputting the fifth feature into the initial emotion recognition model, and determining a predicted emotion category corresponding to the fifth feature according to initial relationship information between the preset features and the preset emotion categories in the initial emotion recognition model;
[0030] Determine the loss value of the initial emotion recognition model based on the actual emotion category and the predicted emotion category;
[0031] When the loss value does not meet the training stop condition, the parameters of the initial emotion recognition model are adjusted, the initial relationship information of the preset features and the preset emotion categories is updated, the predicted emotion categories are updated using the updated initial relationship information, and the loss value is updated according to the actual emotion categories and the updated predicted emotion categories until the updated loss value meets the training stop condition to obtain the emotion recognition model.
[0032] In a second aspect, an embodiment of the present application provides a device for determining an emotion category, comprising:
[0033] An acquisition module, used to acquire audio data;
[0034] An extraction module, used for extracting Mel-frequency cepstral coefficient (MFCC) features of audio data and first features of pinyin syllables of audio data;
[0035] A determination module, for converting audio data into text, and determining the intonation embedding feature of each word in the text, as well as the word embedding feature, position embedding feature and segment embedding feature of the text;
[0036] A fusion module, used for fusing the MFCC feature, the first feature, the intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature to obtain a second feature;
[0037] The determination module is also used to input the second feature into the emotion recognition model, determine the target emotion category corresponding to the second feature according to the relationship information between the preset features and the preset emotion categories in the emotion recognition model, and determine the target emotion category as the emotion category of the audio data.
[0038] In a third aspect, an embodiment of the present application provides an electronic device, the device comprising:
[0039] A processor and a memory storing computer program instructions; when the processor executes the computer program instructions, any one of the above-mentioned methods for determining the emotion category is implemented.
[0040] In a fourth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, any of the above-mentioned methods for determining an emotion category is implemented.
[0041] In a fifth aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute any of the above-mentioned methods for determining an emotion category.
[0042] The method, device, equipment, storage medium and product for determining the emotion category of the embodiment of the present application obtain audio data; extract the Mel-frequency cepstral coefficient MFCC feature of the audio data and the first feature of the pinyin syllable; convert the audio data into text, and determine the intonation embedding feature of each word in the text, as well as the word embedding feature, position embedding feature and segment embedding feature of the text; fuse the MFCC feature, the first feature, the intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature to obtain the second feature; input the second feature into the emotion recognition model, and determine the target emotion category corresponding to the second feature according to the relationship information between the preset feature and the preset emotion category in the emotion recognition model, and determine the target emotion category as the emotion category of the audio data. The emotion category is determined by the second feature after the fusion of the MFCC feature, the first feature, the intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature, which fuses multiple features such as speech features, text features, intonation features and tone features based on the auditory characteristics of the human ear, thereby improving the accuracy of determining the emotion category of the audio data. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] Figure 1 is a schematic diagram of the structure of a system for determining an emotion category provided by an embodiment of the present application;
[0045] Figure 2 is a flowchart of a method for determining an emotion category provided by another embodiment of the present application;
[0046] Figure 3 is a schematic diagram of converting audio to pinyin and tone provided by another embodiment of the present application;
[0047] Figure 4 is a flowchart of a method for determining an emotion category provided in yet another embodiment of the present application;
[0048] Figure 5 is a schematic diagram of a feature example provided by yet another embodiment of the present application;
[0049] Figure 6 This is an overall architecture diagram of a bidirectional encoder characterization model based on a converter provided in yet another embodiment of the present application;
[0050] Figure 7 is a flowchart of a method for determining an emotion category provided in yet another embodiment of the present application;
[0051] Figure 8 is a schematic diagram of the structure of a device for determining an emotion category provided in yet another embodiment of the present application;
[0052] Fig. 9 It is a structural diagram of an electronic device provided in yet another embodiment of the present application. DETAILED DESCRIPTION
[0053] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.
[0054] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0055] With the development of society, people's demand for high-quality services is also increasing. As the first window directly facing users, the service quality of customer service outbound call service can directly affect user satisfaction. In order to improve user satisfaction, customer service voice can be quality inspected. Emotion is one of the factors that represent service quality. Therefore, by identifying the emotions of the audio data corresponding to the customer service voice, the quality inspection effect can be quantified, and irregular service content and personnel can be discovered in a timely and effective manner, thus realizing quality inspection of service defects and service quality.
[0056] At present, emotion recognition of audio data corresponding to customer service speech is mainly carried out through multimodal models to determine the emotion category of audio data. However, the existing multimodal models only use the audio features of audio data and the text features of the text converted from audio data to determine the emotion category of audio data. The correlation between text and audio data is not strong, resulting in low accuracy in determining the emotion category of audio data.
[0057] In order to solve the problems of the prior art, the embodiments of the present application provide a method, apparatus, device, storage medium and product for determining an emotion category. The method for determining an emotion category provided in the embodiments of the present application can be applied to a system for determining an emotion category. Figure 1 As shown, the emotion category determination system 100 includes a data acquisition module 110, a speech transcription module 120, a speech quality inspection module 130 and a data storage module 140. Among them, the data acquisition module 110 is used to identify the customer service call voice data and collect the customer service voice data. The speech transcription module 120 is used to convert the collected customer service voice data into text, pinyin and tone, and identify the volume mean, pitch mean and sound length of the audio interval of the corresponding word. The speech quality inspection module 130 is used to clean the recognized text with special symbols, cluster each word of the text based on the volume mean, pitch mean and sound length of the corresponding audio interval, determine the audio tone category of the corresponding word, and input it into the emotion recognition model to realize the recognition of customer service voice emotions. The data storage module 140 is used to store the relevant data of the customer service voice in layers according to the data type requirements, and realize the association matching of the relevant data of the customer service voice. The data storage module 140 may include a computing memory.
[0058] The method for determining the emotion category provided in the embodiment of the present application determines the emotion category through the second feature obtained by fusing the MFCC feature, the first feature, the intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature. It integrates multiple features such as speech features, text features, intonation features and tone features based on the auditory characteristics of the human ear, thereby improving the accuracy of determining the emotion category of audio data.
[0059] The following is an introduction to the method for determining the emotion category provided in the embodiment of the present application. Figure 2FIG. 1 is a flow chart of a method for determining an emotion category provided by an embodiment of the present application. Figure 2 As shown, the method for determining the emotion category provided in the embodiment of the present application includes the following steps.
[0060] S210: Acquire audio data.
[0061] In some embodiments, the audio data includes customer service voice data.
[0062] S220, extracting Mel-frequency cepstral coefficient (MFCC) features of the audio data and the first features of the pinyin syllables of the audio data.
[0063] Among them, the Mel Frequency Cepstral Coefficents (MFCC) feature is a speech feature based on the auditory characteristics of the human ear. Pinyin syllables include pinyin and tone.
[0064] In some embodiments, the librosa toolkit is used to pre-emphasize the audio data to increase the volume of the audio data, and then pre-processing such as framing and windowing is performed, and then a fast Fourier transform is performed to convert the time domain into the frequency domain. After the change, it is input into a Mel filter to simulate the sound characteristics heard by the human ear, and then the logarithm is taken, and finally the cepstral coefficients are calculated through discrete cosine transform to obtain MFCC features. It should be noted that MFCC feature extraction belongs to the prior art and will not be described in detail here.
[0065] In some embodiments, the MFCC features may be vectorized features. For the extracted initial MFCC features, a convolutional neural network (CNN) model is used to extract high-dimensional features. By setting three groups of convolution kernels of different sizes, the kernel sizes are 2, 3, and 4, respectively, the temporal and spatial features in the audio data can be captured more effectively, and finally the output of the CNN model, i.e., the MFCC features, is obtained, denoted as C′=(c1, c2, ..., c n ).
[0066] In some embodiments, since speech cannot be accurately transcribed into text, in order to retain the original information of the audio data and improve the model's tolerance to this situation, the pinyin features and tone features of the audio data are obtained through the wave package, and the first feature is obtained after vectorization.
[0067] Among them, the high-level features of pinyin features and tone features are extracted through 1-dimensional CNN and MaxPooling, and three different sizes of convolution kernels of 2, 3, and 4 are set to obtain the first feature containing pinyin features and tone features.
[0068] As an example, Figure 3 As shown, the pinyin features and tone features of the audio data are extracted. For example, Figure 3 the pinyin features and tone features of "敬" in it are "jing4". The pinyin of "敬" is "jing", and the tone is the fourth tone. Therefore, the pinyin features and tone features extracted from "敬" are "jing4".
[0069] S230. Convert the audio data into text, and determine the intonation embedding features of each character in the text, as well as the character embedding features, position embedding features, and segment embedding features of the text.
[0070] In some embodiments, the audio data is converted into initial text, and the punctuation marks in the initial text are removed to obtain the text.
[0071] In some embodiments, the text and the audio data are strongly aligned to determine the intonation category of each character in the text. And the intonation category is encoded using a Bidirectional Encoder Representations from Transformers (BERT) model to obtain the intonation embedding features.
[0072] In some embodiments, the BERT model is used to encode the text. For the text d = {d1, d2,..., d n}, the input of the BERT model is [CLS], d1, d2,..., d n , [SEP], and the character embedding features, position embedding features, and segment embedding features of the text sequence are generated. Among them, the intonation categories of [CLS] and [SEP] are 0.
[0073] Among them, the basic architecture of BERT is an encoder based on Transformer. Transformer is an architecture based on the attention mechanism, which can perform parallel computing, efficiently process long sequence data, and can capture semantic information from both the front and back directions of the text simultaneously.
[0074] 240. Fuse the MFCC features, the first features, the intonation embedding features, the character embedding features, the position embedding features, and the segment embedding features to obtain the second features.
[0075] In some embodiments, the fusion method of the MFCC features, the first features, the intonation embedding features, the character embedding features, the position embedding features, and the segment embedding features includes adding the MFCC features, the first features, the intonation embedding features, the character embedding features, the position embedding features, and the segment embedding features to obtain the second features.
[0076] S250, inputting the second feature into the emotion recognition model, determining the target emotion category corresponding to the second feature according to the relationship information between the preset features and the preset emotion categories in the emotion recognition model, and determining the target emotion category as the emotion category of the audio data.
[0077] Here, the emotion recognition model is trained in advance.
[0078] In some embodiments, the emotion recognition model includes a multi-head self-attention mechanism, that is, includes multiple single-head attention models.
[0079] In some embodiments, emotion categories may include, but are not limited to, normal, impatient, negative, and indifferent.
[0080] The embodiment of the present application determines the emotion category through the second feature obtained by fusing the MFCC feature, the first feature, the intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature, and integrates multiple features such as speech features, text features, intonation features and tone features based on the auditory characteristics of the human ear, thereby improving the accuracy of determining the emotion category of the audio data.
[0081] Based on this, in some embodiments, the above S210 may specifically include:
[0082] Get the original voice data;
[0083] Perform channel segmentation on the original speech data and extract the audio data of the preset channel.
[0084] In some embodiments, in the customer service and customer voices, the customer service voice signal and the customer voice signal are located in different channels respectively, so as to record the different voices of the two roles of customer service and customer.
[0085] The original voice data is segmented into channels to extract the audio data of the preset channels. The embodiment of the present application only focuses on the customer service voice emotions, so the audio data of the customer service segment needs to be identified and extracted from the complete original voice data.
[0086] As an example, the FFmpeg audio processing tool and the audio processing toolkit in Python are used to perform channel segmentation on the original voice data and extract the audio data of the preset channel.
[0087] The embodiment of the present application can extract audio data required for emotion recognition by performing channel segmentation on speech data, and then determine the corresponding emotion category, which not only improves the accuracy of emotion category determination but also saves computing resources.
[0088] Based on this, in some embodiments, in the above S220, extracting the first feature of the pinyin syllable of the audio data may specifically include:
[0089] Extract the pinyin syllables of the audio data, where the pinyin syllables include pinyin and tones.
[0090] Encode the pinyin and tones according to the first preset encoding method to obtain the first feature.
[0091] Among them, the first preset encoding method can be set in advance. For example, the first preset encoding method can be a randomly generated encoding method. It can be understood that during use, the first preset encoding method always remains unchanged.
[0092] In some embodiments, as Figure 3 shown, extract the pinyin and tones of the audio data. For example, the pinyin of "敬" is "jing", and the tone is the fourth tone. Therefore, the pinyin and tone extracted from "敬" are "jing4". Encode the pinyin and tones extracted from the audio data through the first preset encoding method, and extract the high-level features of the encoded pinyin and tones through 1D CNN and MaxPooling. Set three different sizes of convolutional kernels of 2, 3, and 4 to obtain the first feature S'=(s1, s2,..., s n ).
[0093] In the embodiments of the present application, by encoding the pinyin and tones of the extracted audio data and performing feature fusion after encoding, the accuracy of the fused features is improved, and thus the accuracy of emotion category determination is improved.
[0094] Based on this, in some embodiments, as Figure 4 shown, in the above S230, determine the intonation embedding features of each word in the text, which may specifically include S231 to S233.
[0095] S231. Obtain the volume mean, pitch mean, and duration of each word in the text.
[0096] In some embodiments, strongly align the text and the audio data to obtain the volume mean, pitch mean, and duration of the audio interval corresponding to each word.
[0097] S232. Cluster the words in the text according to the volume mean, pitch mean, and duration to obtain the intonation category of each word in the text.
[0098] In some embodiments, perform clustering on the volume mean, pitch mean, and duration of each word through Kmeans clustering. For example, cluster into 4 categories, corresponding to intonation categories of rising tone, falling tone, flat tone, and zigzag tone, to obtain the intonation category of each word and achieve strong alignment of the text and audio intonation.
[0099] S233: Encode the intonation category according to a second preset encoding method to obtain an intonation embedding feature of each word in the text.
[0100] The second preset encoding method is set in advance.
[0101] In some embodiments, for the intonation category of each word in the text, the corresponding intonation code is determined as 0, 1, 4, ..., 3, 0 according to the text sequence.
[0102] It should be noted that intonation relies on volume, pitch, and sound length as its expression. In the embodiment of the present application, the volume mean, pitch mean, and sound length of the text and the corresponding audio interval are clustered through Kmeans to obtain the intonation category of each word, thereby achieving strong alignment of text and intonation.
[0103] The embodiment of the present application determines the intonation category through the mean volume, mean pitch and sound length of each word in the text, and integrates the intonation category into the text after encoding, thereby achieving strong alignment of text and intonation, thereby improving the accuracy of emotion category determination.
[0104] Based on this, in some embodiments, the above S240 may specifically include:
[0105] The intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature are integrated to obtain the third feature;
[0106] The MFCC feature, the first feature and the third feature are fused to obtain the second feature.
[0107] In some embodiments, intonation embedding features, word embedding features, position embedding features, and segment embedding features are as follows: Figure 5 As shown in the figure, by adding the intonation embedding features, word embedding features, position embedding features and segment embedding features, and then inputting them into the BERT model, the BERT model can integrate the intonation embedding features into the corresponding text. It should be noted that Figure 5 The first line is the word embedding feature, the second line is the position embedding feature, the third line is the segment embedding feature, and the fourth line is the intonation embedding feature.
[0108] It is understandable that if Figure 6As shown in the figure, the overall structure of the BERT model consists of a stack of multiple layers of Transformer encoder layers. Each encoder layer consists of a multi-head self-attention layer, a residual connection, a fully connected layer, and an activation function. These layers process the input text layer by layer and convert it into a feature vector representation. In the multi-head self-attention layer, the model assigns different importance to each word in the input sequence by learning attention weights, so that the model can consider the semantic information of the entire sentence without losing the context. The residual connection of the BERT model ensures that after multi-layer encoding, information will not be lost due to too many layers. At the same time, the fully connected layer and the activation function are responsible for dimensionality transformation and nonlinear transformation of the feature vector to obtain richer semantic information. Among them, trm represents the encoder end of the Transformer. Finally, the output of the BERT model is T'=(t1,t2,...t n ,).
[0109] In some embodiments, the model extracts the audio MFCC features, the text features integrated with intonation, and the features of pinyin and tone as matrices C', T', and S', respectively, and achieves the purpose of feature fusion by concatenating the three feature representations. The fusion is recorded as X = Concat (C', T', S').
[0110] The embodiment of the present application achieves strong alignment of text and intonation by fusing the intonation embedding feature into the embedding feature of the text, and then fuses the MFCC feature, the first feature and the third feature to avoid interference when fusing the intonation embedding feature and the embedding feature of the text, thereby further improving the accuracy of determining the emotion category.
[0111] Based on this, in some embodiments, such as Figure 7 As shown, before the above S250, the method may further include:
[0112] S310, obtaining an audio data sample and an actual emotion category of the audio data sample;
[0113] S320, extracting the MFCC feature sample of the audio data sample and the fourth feature of the pinyin syllable of the audio data sample;
[0114] S330, converting the audio data sample into a text sample, and determining the intonation embedding feature sample of each word in the text sample, as well as the word embedding feature sample, the position embedding feature sample and the segment embedding feature sample of the text sample;
[0115] S340, fusing the MFCC feature sample, the fourth feature, the intonation embedding feature sample, the word embedding feature sample, the position embedding feature sample, and the segment embedding feature sample to obtain a fifth feature;
[0116] S350, inputting the fifth feature into the initial emotion recognition model, and determining the predicted emotion category corresponding to the fifth feature according to the initial relationship information between the preset features and the preset emotion categories in the initial emotion recognition model;
[0117] S360, determining a loss value of the initial emotion recognition model according to the actual emotion category and the predicted emotion category;
[0118] S370. When the loss value does not meet the training stop condition, adjust the parameters of the initial emotion recognition model, update the initial relationship information between the preset features and the preset emotion categories, use the updated initial relationship information to update the predicted emotion categories, and update the loss value according to the actual emotion categories and the updated predicted emotion categories until the updated loss value meets the training stop condition, thereby obtaining the emotion recognition model.
[0119] In some embodiments, the actual emotion categories are annotated by the user. For example, the audio data samples are annotated into four categories of emotions: normal, impatient, negative, and indifferent.
[0120] In some embodiments, feature extraction is performed on the audio data sample, and the extracted audio MFCC features, text features integrated with intonation, and features of pinyin and intonation are respectively matrix C, matrix T, and matrix S. The MFCC feature samples, the fourth feature, the intonation embedding feature samples, the word embedding feature samples, the position embedding feature samples, and the segment embedding feature samples are fused to obtain the fifth feature, which is recorded as X=Concat(C, T, S). X is input into the initial emotion recognition model to train it.
[0121] For each x in input X i , map it to three different spaces, and get the query vector q i , key vector k i Sum value vector v i , for the entire input X=(x1,x2,...,x n ), the matrices Q, K, and V formed by the three vectors are shown in formula (1):
[0122]
[0123] Among them, W q ,W k ,W v Represent the parameter matrices of the linear mapping. The query vector q i , key vector k i Sum value vector v i Respectively represent the information about x in the database i Additional information, fields, and field values.
[0124] For the query matrix Q, the similarity matrix is obtained by taking the dot product of the query matrix Q and the key matrix K. The similarity matrix can be represented by H. The calculation of the matrix H is shown in formula (2):
[0125]
[0126] Among them, softmax represents the column-normalized function, d k It means that each x i Dimension.
[0127] The output results H calculated by multiple parallel self-attention mechanisms are concatenated and transformed linearly to obtain the output results of the multi-head self-attention mechanism, as shown in the following formulas (3) and (4):
[0128] head i =Attention(QW i Q ,KW i K ,VW i V ) (3)
[0129] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W o (4)
[0130] Among them, head i is a single-head attention model, and h is the number of splicing. i Q , W i K , W i V and W o It is understandable that by substituting formula (2) into formula (3), a model including a trainable matrix, an output result and an input X can be obtained.
[0131] Next, for the output a of the multi-head self-attention mechanism cls =MultiHead(Q,K,V), the emotion category is obtained through the fully connected layer and softmax activation function, and the calculation method is shown in formula (5):
[0132] P = softmax(Wa cls +b) (5)
[0133] It should be noted that the model uses the cross entropy loss function to calculate the loss value. The loss function is shown in formula (6):
[0134]
[0135] Where N is the number of emotion categories, n is the length of the sample, and p ic represents the probability that sample i belongs to category c, y ic Represents a sign function, that is, if the true category of sample i is c, it is 1, otherwise it is 0.
[0136] The embodiment of the present application trains the initial emotion recognition model through a large number of samples, and can accurately determine the relationship information between the preset features and the preset emotion categories, thereby improving the accuracy of emotion category determination.
[0137] The embodiment of the present application obtains the intonation category of the corresponding word by clustering the volume, pitch and length of the word corresponding to the audio data, designs the intonation embedding vector of the intonation category, and adds it to the word embedding vector, position embedding vector and segment embedding vector of the BERT model to obtain the features of the audio text incorporating the intonation features, thereby integrating the intonation features of each word into the text and improving the accuracy of the model's emotion recognition.
[0138] The embodiment of the present application takes into account that speech-to-text conversion cannot be completely accurate, while pinyin can well preserve the original information of the audio, and tone is also a factor reflecting emotions. Therefore, the pinyin and tone features corresponding to the audio data are combined, which can not only well preserve the original information of the audio, but also improve the model's tolerance for audio-to-text errors, and enrich the model's feature expression, thereby improving the accuracy of emotion recognition.
[0139] Due to the different intonations of the audio, even if the same text content is expressed through different intonations, its emotions will be very different. For example, high rising tones are used in sentences expressing anger, tension, and warnings; falling tones are used in exclamatory sentences, imperative sentences, or sentences expressing determination, confidence, praise, etc.; flat tones are mostly used in narration, explanation, or sentences expressing hesitation, thinking, indifference, etc.; tortuous tones are used to express special emotions, such as sarcasm, ridicule, exaggeration, emphasis, special surprise, etc. Therefore, aligning the audio intonation and text can integrate the intonation features of the audio into the text, which plays an important role in improving the accuracy of speech emotion recognition. Moreover, in the process of voice-to-text conversion, the audio cannot be completely transcribed into correct text, and the embodiments of the present application consider how to retain the audio information of the incorrectly transcribed text, which improves the model's tolerance for incorrect text.
[0140] The embodiment of the present application belongs to the field of artificial intelligence and speech emotion analysis. The method is based on customer service speech, and extracts multiple information from audio data, including 1. MFCC feature extraction of audio data, and high-level features corresponding to MFCC features based on CNN model; 2. In order to integrate audio into the text, the audio is transcribed into text through audio recognition, and starting from the three elements of intonation, volume, pitch, and length, the text is strongly aligned with the volume mean, pitch mean, and length of the corresponding audio interval. Kmeans clustering is performed for the volume, pitch, and length features of each word, and the clustering results of the word are used as the intonation category of the corresponding word. The intonation embedding vector of the intonation category is designed, and the word embedding vector, position embedding vector, and segment embedding vector of the BERT model are added to obtain text features that incorporate intonation features. 3. In order to alleviate the problem of text errors in audio transcription, enrich the feature expression of the model, extract pinyin and intonation from the audio, and extract high-level features of pinyin and intonation based on TextCNN. Finally, the above three features are spliced, and the fused features are extracted based on the attention mechanism to realize speech emotion recognition.
[0141] Based on the method for determining the emotion category provided in the above embodiment, the present application also provides a specific implementation of the device for determining the emotion category. Please refer to the following embodiment.
[0142] See first Figure 8 The emotion category determination device 400 provided in the embodiment of the present application includes:
[0143] An acquisition module 410 is used to acquire audio data;
[0144] An extraction module 420, used to extract Mel-frequency cepstral coefficient (MFCC) features of the audio data and a first feature of a pinyin syllable of the audio data;
[0145] A determination module 430, for converting the audio data into text, and determining the intonation embedding feature of each word in the text, as well as the word embedding feature, position embedding feature and segment embedding feature of the text;
[0146] A fusion module 440 is used to fuse the MFCC feature, the first feature, the intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature to obtain a second feature;
[0147] The determination module 430 is also used to input the second feature into the emotion recognition model, determine the target emotion category corresponding to the second feature according to the relationship information between the preset features and the preset emotion categories in the emotion recognition model, and determine the target emotion category as the emotion category of the audio data.
[0148] Based on this, in some embodiments, the acquisition module 410 may be specifically used to:
[0149] Get the original voice data;
[0150] Perform channel segmentation on the original speech data and extract the audio data of the preset channel.
[0151] Based on this, in some embodiments, the extraction module 420 may be specifically used to:
[0152] Extracting pinyin syllables, pinyin syllables and tones of audio data;
[0153] According to a first preset encoding method, the pinyin and the tone are encoded to obtain a first feature.
[0154] Based on this, in some embodiments, the determination module 430 may be specifically used to:
[0155] Get the mean volume, pitch and duration of each word in the text;
[0156] Cluster the words in the text according to the mean volume, mean pitch and sound length to obtain the intonation category of each word in the text;
[0157] The intonation category is encoded according to the second preset encoding method to obtain the intonation embedding feature of each word in the text.
[0158] Based on this, in some embodiments, the fusion module 440 may be specifically used to:
[0159] The intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature are integrated to obtain the third feature;
[0160] The MFCC feature, the first feature and the third feature are fused to obtain the second feature.
[0161] Based on this, in some embodiments, the apparatus 400 may further include:
[0162] The acquisition module 410 is further used to obtain the audio data sample and the actual emotion category of the audio data sample before inputting the second feature into the emotion recognition model and determining the target emotion category corresponding to the second feature according to the relationship information between the preset feature and the preset emotion category in the emotion recognition model;
[0163] The extraction module 420 is further used to extract the MFCC feature samples of the audio data samples, and the fourth features of the pinyin and tone of the audio data samples;
[0164] The determination module 430 is further used to convert the audio data sample into a text sample, and determine the intonation embedding feature sample of each word in the text sample, as well as the word embedding feature sample, the position embedding feature sample and the segment embedding feature sample of the text sample;
[0165] The fusion module 440 is further used to fuse the MFCC feature sample, the fourth feature, the intonation embedding feature sample, the word embedding feature sample, the position embedding feature sample and the segment embedding feature sample to obtain a fifth feature;
[0166] The determination module 430 is further configured to input the fifth feature into the initial emotion recognition model, and determine the predicted emotion category corresponding to the fifth feature according to the initial relationship information between the preset features and the preset emotion categories in the initial emotion recognition model;
[0167] The determination module 430 is further used to determine the loss value of the initial emotion recognition model according to the actual emotion category and the predicted emotion category;
[0168] Determination module 430 is also used to adjust the parameters of the initial emotion recognition model when the loss value does not meet the training stop condition, update the initial relationship information between the preset features and the preset emotion categories, update the predicted emotion categories using the updated initial relationship information, and update the loss value according to the actual emotion categories and the updated predicted emotion categories until the updated loss value meets the training stop condition to obtain the emotion recognition model.
[0169] The various modules of the device for determining the emotion category provided in the embodiment of the present application can realize the functions of the various steps of the method for determining the emotion category provided above and can achieve the corresponding technical effects, which will not be described in detail here for the sake of brevity.
[0170] Based on the same inventive concept, an embodiment of the present application also provides an electronic device.
[0171] Fig. 9 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.
[0172] The electronic device may include a processor 501 and a memory 502 storing computer program instructions.
[0173] Specifically, the processor 501 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0174] The memory 502 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 502 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 502 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 502 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 502 is a non-volatile solid-state memory.
[0175] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Therefore, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of the present disclosure.
[0176] The processor 501 reads and executes the computer program instructions stored in the memory 502 to implement any one of the emotion category determination methods in the above embodiments.
[0177] In one example, the electronic device may further include a communication interface 503 and a bus 510. Fig. 9 As shown, the processor 501, the memory 502, and the communication interface 503 are connected via a bus 510 and communicate with each other.
[0178] The communication interface 503 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0179] The bus 510 includes hardware, software or both, coupling the components of the electronic device to each other. For example, but not limitation, the bus may include an accelerated graphics port (Accelerated Graphics Port, AGP) or other graphics bus, an enhanced industry standard architecture (Extended Industry Standard Architecture, EISA) bus, a front side bus (Front Side Bus, FSB), a hypertransport (Hyper Transport, HT) interconnection, an industry standard architecture (Industry Standard Architecture, ISA) bus, an infinite bandwidth interconnection, a low pin count (Linear Predictive Coding, LPC) bus, a memory bus, a micro channel architecture (MicroChannel Architecture, MCA) bus, a peripheral component interconnect (Peripheral Component Interconnect, PCI) bus, a PCI-Express (Peripheral Component Interconnect-X, PCI-X) bus, a serial advanced technology attachment (Serial Advanced Technology Attachment, SATA) bus, a video electronics standard association local (VESA Local Bus, VLB) bus or other suitable bus or a combination of two or more of these. Where appropriate, the bus 510 may include one or more buses. Although the embodiments of the present application describe and illustrate a specific bus, the present application considers any suitable bus or interconnection. The electronic device can execute the method for determining the emotion category in the embodiments of the present invention, thereby realizing the above-mentioned method for determining the emotion category.
[0180] In addition, in combination with the method for determining the emotion category in the above embodiment, the embodiment of the present application can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by the processor, any one of the methods for determining the emotion category in the above embodiment is implemented.
[0181] The present application also provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes each process of implementing any one of the above-mentioned emotion category determination method embodiments.
[0182] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0183] The functional blocks shown in the structural block diagram described above can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), appropriate firmware, plug-in, function card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or communication link by a data signal carried in a carrier. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memory (Read-Only Memory, ROM), flash memory, erasable read-only memory (Erasable ReadOnly Memory, EROM), floppy disks, compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical discs, hard disks, optical fiber media, radio frequency (Radio Frequency, RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0184] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0185] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0186] The above are only specific implementation methods of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. A method for determining an emotion category, characterized in that: include: Get audio data; Extracting Mel-frequency cepstral coefficient (MFCC) features of the audio data and first features of pinyin syllables of the audio data; Converting the audio data into text, and determining the intonation embedding feature of each word in the text, as well as the word embedding feature, position embedding feature and segment embedding feature of the text; fusing the MFCC feature, the first feature, the intonation embedding feature, the character embedding feature, the position embedding feature and the segment embedding feature to obtain a second feature; The second feature is input into an emotion recognition model, and a target emotion category corresponding to the second feature is determined based on relationship information between preset features and preset emotion categories in the emotion recognition model, and the target emotion category is determined to be the emotion category of the audio data.
2. The method for determining the emotion category according to claim 1, characterized in that: The obtaining of audio data comprises: Get the original voice data; The original voice data is segmented into channels, and audio data of a preset channel is extracted.
3. The method for determining the emotion category according to claim 1, characterized in that: Extracting a first feature of the pinyin syllable of the audio data includes: Extracting the pinyin syllables of the audio data, wherein the pinyin syllables include pinyin and tones; The pinyin and the tone are encoded according to a first preset encoding method to obtain a first feature.
4. The method for determining an emotion category according to claim 1, characterized in that: Determining the intonation embedding feature of each word in the text includes: Obtaining the mean volume, mean pitch and length of each word in the text; Clustering the characters in the text according to the volume mean, the pitch mean and the tone length to obtain the intonation category of each character in the text; The intonation category is encoded according to a second preset encoding method to obtain an intonation embedding feature of each word in the text.
5. The method for determining the emotion category according to claim 1, characterized in that: The step of fusing the MFCC feature, the first feature, the intonation embedding feature, the character embedding feature, the position embedding feature and the segment embedding feature to obtain a second feature includes: fusing the intonation embedding feature, the word embedding feature, the position embedding feature and the segment embedding feature to obtain a third feature; The MFCC feature, the first feature and the third feature are fused to obtain a second feature.
6. The method for determining an emotion category according to claim 1, characterized in that: Before inputting the second feature into the emotion recognition model and determining the target emotion category corresponding to the second feature according to the relationship information between the preset feature and the preset emotion category in the emotion recognition model, the method further includes: Acquire an audio data sample and an actual emotion category of the audio data sample; Extracting the MFCC feature sample of the audio data sample and the fourth feature of the pinyin syllable of the audio data sample; Converting the audio data sample into a text sample, and determining the intonation embedding feature sample of each word in the text sample, as well as the word embedding feature sample, the position embedding feature sample and the segment embedding feature sample of the text sample; Fusing the MFCC feature sample, the fourth feature, the intonation embedding feature sample, the word embedding feature sample, the position embedding feature sample and the segment embedding feature sample to obtain a fifth feature; Inputting the fifth feature into an initial emotion recognition model, and determining a predicted emotion category corresponding to the fifth feature according to initial relationship information between preset features and preset emotion categories in the initial emotion recognition model; Determining a loss value of the initial emotion recognition model according to the actual emotion category and the predicted emotion category; When the loss value does not meet the training stop condition, the parameters of the initial emotion recognition model are adjusted, the initial relationship information between the preset features and the preset emotion categories is updated, the predicted emotion categories are updated using the updated initial relationship information, and the loss value is updated according to the actual emotion categories and the updated predicted emotion categories until the updated loss value meets the training stop condition, thereby obtaining the emotion recognition model.
7. A device for determining an emotion category, characterized in that: include: An acquisition module, used to acquire audio data; An extraction module, used to extract the Mel-frequency cepstral coefficient (MFCC) feature of the audio data and the first feature of the pinyin syllable of the audio data; A determination module, configured to convert the audio data into text, and determine the intonation embedding feature of each word in the text, as well as the word embedding feature, position embedding feature and segment embedding feature of the text; A fusion module, configured to fuse the MFCC feature, the first feature, the intonation embedding feature, the character embedding feature, the position embedding feature and the segment embedding feature to obtain a second feature; The determination module is also used to input the second feature into the emotion recognition model, determine the target emotion category corresponding to the second feature based on the relationship information between the preset features and the preset emotion categories in the emotion recognition model, and determine the target emotion category as the emotion category of the audio data.
8. An electronic device, characterized in that: The device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method for determining the emotion category according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the method for determining an emotion category according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the method for determining an emotion category as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Character sequence recognition method fusing dictionary and character features
CN114662476A
Chinese and English cross-language speech synthesis method and device, electronic equipment and storage medium
CN114664282A
Artificial intelligence customer service response system
CN115643341A
Multi-mode speech emotion recognition method and device, equipment and storage medium
CN116631450A
Digital human voice generation system based on large language model
CN119314465A
Cited By
Method and device for determining audio emotion and generating digital human video, and related product
CN121483310A