A Multimodal Emotion Recognition Method Based on Acoustic and Text Features
Voice features are extracted through OpenSMILE and Transformer networks, and combined with the multimodal emotion recognition method of DC-BERT and BiLSTM networks, the problem of insufficient combination of speech and text modality is solved, and more accurate emotion recognition is achieved.
Patent Information
- Application Number
- CN202210108118.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-01-28
AI Technical Summary
The prior art cannot effectively combine speech and text modes in speech emotion recognition, resulting in limited recognition performance, and the BERT model cannot make up for the insufficient potential emotional information in the transcription text when the amount of emotional data is insufficient.
OpenSMILE is used to extract the shallow features of speech, combine with the Transformer network to generate deep acoustic features, and obtain pause information through forced alignment, and extract text features using the improved DC-BERT model. Finally, emotional classification is performed through BiLSTM network with attention mechanism.
It improves the accuracy of emotion recognition, obtains rich semantic information through multimodal fusion, corrects the ambiguity of pure text recognition, and enhances the diversity and accuracy of emotion recognition.
Smart Images

Figure CN114446324B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-modal emotion recognition method based on acoustic and text features, which is applicable to the extraction of speech and text emotion features, and belongs to the technical fields of artificial intelligence and speech emotion recognition. Background Art
[0002] With the development of technology, great progress has been made in speech emotion recognition and natural language processing, but humans are still unable to communicate naturally with machines. Therefore, it is crucial to establish a system that can detect emotions in human-computer interaction. However, due to the variability and complexity of human emotions, this remains a challenging task.
[0003] Traditional emotion recognition mainly focuses on a single modality, such as text, speech, images, etc., and there are certain limitations in recognition performance. For example, in early speech emotion recognition tasks, researchers mainly used acoustic features in speech and some related prosodic features, often ignoring the specific semantic information (text information) contained in speech. However, in daily conversations and social media, the voice is often a repetition of a text content, and the two are closely related. Considering the identity, complementarity, and strong correlation between the speech and text modalities, many researchers have shifted from single-modal to multi-modal emotion recognition research. Among them, fusing the two different modality information of speech and text for emotion recognition has also become a hot research direction. Compared with a single modality, considering multiple modality information simultaneously can capture emotions more accurately.
[0004] Many research institutions are also constantly exploring new language models. In 2019, Google Research first proposed a new type of language representation model, BERT, which can generate deep bidirectional representations of language and greatly improve the results of various natural language processing tasks. Although BERT can be used to obtain context word embeddings to represent the information contained in the transcribed text, it does not consider the problem of mismatch due to the complex network structure of BERT and the insufficient amount of emotion corpus data. Although BERT can be used to generate representations of text information, it cannot make up for the deficiency that the transcribed text itself ignores some potential emotion information.
[0005] The pause information during the speaking process is not reflected in the transcribed text. After investigating the relationship between speaking pause information and emotions, it is found that compared with happiness and positivity, in the emotional states of sadness and fear, the proportion of the average duration of silent pauses in the whole speech increases, and it is also noted that when in different emotional states, the frequency, duration, and position of speaking pauses will also be different.
[0006] On the other hand, deep networks based on the attention mechanism have shown superior performance in the decoding stage and have been widely used in the fields of natural language processing and speech recognition. In speech emotion recognition, since emotional features are not evenly distributed in sentences, many researchers have added an attention mechanism to the emotion recognition task, enabling the network to have a guiding mechanism for parts containing more emotional information and highlighting the most emotional local information. For this reason, the present invention proposes a multimodal emotion recognition method that can effectively extract speech and text emotional features and add pause information at the same time, and designs a BiLSTM network model with an attention mechanism to classify emotions. Summary of the Invention
[0007] Aiming at the deficiencies of the prior art, a multimodal recognition method based on acoustic and text features is provided, which combines two modalities of data, namely speech and text. It can obtain rich semantic information in the transcribed text and perceive the fluctuations of speech through the speech audio perception task, so as to further obtain accurate emotions and correct the ambiguity of emotion recognition solely through text.
[0008] To achieve the above technical objectives, a multimodal emotion recognition method based on acoustic and text features of the present invention is characterized in that: the emotion shallow features of the input speech are extracted by using OpenSMILE and fused with the deep features obtained by the Transformer network learning the shallow features to generate multi-level acoustic features; the pause information is obtained by forced alignment of the speech with the same content and the transcribed text, and then the speaking pause information in the speech is encoded and added to the transcribed text, and then sent into the hierarchical dense connection DC-BERT model to obtain the text features, which are then fused with the acoustic features; the bidirectional long short-term memory neural network BiLSTM-ATT based on the attention mechanism is used as a classifier. The BiLSTM network uses prior knowledge to obtain effective context information, and extracts the part that highlights the emotional information in the features through the attention mechanism to avoid information redundancy. A global average pooling layer is added after the attention mechanism to replace the traditional fully connected layer, which can effectively prevent the overfitting problem, and finally sent into the softmax layer for emotion classification;
[0009] The specific steps are as follows:
[0010] S1: Input the original speech audio to be judged into OpenSMILE, and use the emobase feature set in the OpenSMILE toolbox to extract the shallow acoustic features in the original speech data;
[0011] S2: Input the extracted shallow acoustic features into the Transformer network, and utilize the encoder structure of the Transformer network to effectively learn the relationships between the input shallow acoustic features, thereby outputting a feature sequence related to emotion, that is, deep features with global information;
[0012] S3: Concatenate and fuse the sequence of shallow acoustic features and the sequence of deep features to obtain a deep - shallow fusion feature sequence, with the content of the shallow feature sequence in the front and the deep features in the back for concatenation;
[0013] S4: Pre - process the text of the original speech transcription: delete the punctuation marks in the text and unify the writing forms of the words and phrases formed by the transcription;
[0014] S5: Use the Penn Phonetics Lab ForcedAligner (P2FA) to perform forced alignment on the pre - processed transcription text and the original speech in step S4, thereby determining the positions and durations of pauses;
[0015] S6: Divide the different pause durations in the speech audio into six intervals: 0.05 - 0.1s, 0.1 - 0.3s, 0.3 - 0.6s, 0.6 - 1.0s, 1.0 - 2.0s, and greater than 2.0s. Use "..", "…", "…", "…", "……", "……" respectively to mark the pause durations in the six intervals in the transcription text, match the pause durations in the speech audio at the marked positions in the transcription text, and add the mark "." at the end of each speaker's sentence in the text as an end sign;
[0016] S7: Input the transcription text with pause encoding marked into the trained improved DC - BERT, and the improved DC - BERT outputs the emotion features of the utterance - level text according to the pause encoding marks in the transcription text;
[0017] S8: Concatenate and fuse the deep - shallow fusion feature sequence corresponding to the speech audio and the emotion features of the utterance - level text to obtain the acoustic - text fusion features of each sentence in this section of the audio;
[0018] S9: Finally, send the acoustic - text fusion features into a BiLSTM network with an attention mechanism for emotion classification, and output the corresponding emotion classification to achieve emotion recognition.
[0019] Furthermore, use the built - in file to extract shallow acoustic features from the original speech signal input into OpenSMILE, including intensity, loudness, Mel - Frequency Cepstral Coefficients, pitch, and their statistical values (such as maximum value, minimum value, average value, and standard deviation) for each short frame at the utterance level;
[0020] The shallow acoustic features consist of a sequence of low-level descriptors; only the audio and transcribed texts representing anger, happiness, neutrality, and sadness in the emotion dataset are selected for recognition, and happiness is composed of the merged emotions of gladness and excitement.
[0021] Furthermore, the transcribed text that has been forced alignment and encoded by the University of Pennsylvania Speech Label Forced Alignment Tool is fed into the improved DC-BERT, and the 768-dimensional output sequence of the penultimate layer of DC-BERT is selected as the utterance-level text feature;
[0022] The improved DC-BERT model retains the residual connections inside each multi-head self-attention layer of the Transformer in the traditional BERT model, and adds dense connections between layers, that is, the input of each multi-head self-attention layer is additionally added with the feature information of the first two layers to accelerate the convergence speed of the model, make the loss function of the network smoother, and the features extracted by each layer can also be reused between different attention layers, improving the utilization rate of features;
[0023] The internal form of the improved DC-BERT is: assuming a given input feature sequence X, then x i = H(x i-1 ) + αx i-1 + βx i-2 , where x i is the i-th element of the input feature sequence X, H is a non-linear function, and α and β are weight coefficients for retaining the information of the first two layers, so that each layer can obtain the results processed by the first two layers but does not dominate; the improved DC-BERT model consists of 12 layers of Transformer, and the output of each layer can theoretically be used as the utterance-level text feature.
[0024] Furthermore, after fusing the acoustic features and text features, they are fed into a BiLSTM network with an attention mechanism for emotion classification. There are three attention mechanisms in the BiLSTM network, namely local attention mechanism, self-attention mechanism, and multi-head attention mechanism;
[0025] Local attention mechanism: This mechanism only focuses on a part of the encoded hidden layer. The local attention first generates an alignment position p t at time t for the current node, and then selectively sets a context window with a fixed size of 2D + 1. The formula is as follows:
[0026]
[0027] where D is selected according to experience; p t is the window center, which is the h of the current hidden state tThe decision is a real number; the calculation process of the alignment weights is similar to that of traditional attention:
[0028]
[0029] where the standard deviation σ is set empirically, and h t is the hidden state at the t-th time step of the current decoder, is the hidden state at the i-th time step of the encoder, i represents the position of the input sequence, and T x represents the sequence length;
[0030] The self-attention mechanism utilizes the weighted correlation between the elements of the input feature sequence, that is, each element of the input sequence can be projected into three different representations through a linear function: query, key, and value. The calculation formula is as follows:
[0031]
[0032] where x i represents the i-th element in the input feature sequence, q i , v i , k i represent the query vector, value vector, and key vector of the i-th element in the input feature sequence. represents the transpose of the three weight matrices for obtaining the query vector, value vector, and key vector.
[0033] The final attention matrix is shown in the formula:
[0034]
[0035] where Q is the query matrix, K is the key matrix, V is the value matrix of the sentence, and d k is the scaling factor;
[0036] On the basis of the self-attention mechanism, the impact of the multi-head self-attention mechanism on the speech emotion recognition task is compared. Multi-head means that the number of projections for each variable of the input feature sequence: query, key, and value is more than one group. That is, on the premise of non-sharing parameters, after mapping Q, K, and V through the parameter matrix, a single-layer self-attention is performed, and then the self-attention is stacked layer by layer. The multi-head self-attention calculation formula is:
[0037] head i = attention(QW i Q , KW i K , VW iV )
[0038] Multihead(Q, K, V) = Concat(head1, ..., head n )
[0039] Beneficial effects:
[0040] Regarding the problem that shallow features only contain global information and are insufficient in expressing emotions, by using the deep features obtained from the secondary learning of the Transformer network, the two are fused to obtain shallow and deep features. After the fusion of shallow and deep features, they have multi-level acoustic features. At the same time, considering the correlation between pause information and emotions in speech, pause information is obtained by aligning audio with the transcribed text, and different pause information is encoded and added to the transcribed text, adding a new connection between semantics and pause information, making the transcribed text information more diversified, and effectively improving the accuracy of emotion recognition;
[0041] To make up for the mismatch between the complex network structure of BERT and the small amount of emotion data, the DC-BERT model is used to extract discourse-level text features, which speeds up the convergence rate of the model and improves the utilization rate of features. The best one is selected after comparing the impacts of three attention mechanisms in the emotion recognition task.
[0042] This method uses two-modal data of speech and text. During the emotion recognition process, it can obtain rich semantic information in the transcribed text and can also perceive the fluctuations of speech through the speech audio task, thereby further obtaining accurate emotions and correcting the ambiguity of emotion recognition solely through text.
[0043] Technical advantages of this application:
[0044] In terms of the speech modality, this method uses the Transformer Encoder to perform secondary learning on low-level descriptor features, mines deeper emotion information therein, and fuses it with low-level descriptor features to form multi-level and multi-faceted acoustic features. In terms of the text modality, the present invention adds pause information to the transcribed text, supplementing other subordinate information of the text modality in addition to semantic information, making the text information more diversified. The fusion of acoustic and text features can complement each other's missing information while mining the emotion information hidden in the features in multiple directions. The emotion of a sentence often appears in a certain paragraph or a certain word in the sentence. Therefore, using a BiLSTM network with an attention mechanism as a classifier can make the network pay more attention to the parts with strong emotions and ignore some unimportant information, resulting in better classification effects.
[0045] 1) The common emotion recognition feature set is extracted using the OpenSMILE toolbox. Here, the emobase feature set is used to extract 988-dimensional shallow acoustic features. OpenSMILE is fast and effective in feature extraction;
[0046] 2) Due to the multi-head self-attention mechanism, Transformer has a method for global speech emotion analysis;
[0047] 3) The calculation speed of Transformer overcomes the slow training characteristic of RNN and can perform parallel computing;
[0048] 4) DC-BERT retains the residual connections inside each multi-head self-attention layer in Transformer and adds dense connections between layers, that is, the input of each multi-head self-attention layer additionally adds the feature information of the previous two layers. The purpose is to accelerate the convergence speed of the model, make the loss function of the network smoother, and the features extracted by each layer can also be reused between different attention layers, improving the utilization rate of features;
[0049] 5) The BiLSTM model with an attention mechanism has good feature learning ability and good generalization ability of the model. Description of the Drawings
[0050] Figure 1 is the system framework diagram of the multi-modal emotion recognition method of the present invention;
[0051] Figure 2 is the internal structure diagram of the DC-BERT model used in the present invention;
[0052] Figure 3 is the flowchart of the pause encoding for the transcribed text in the present invention. Detailed Embodiments
[0053] To more fully explain the present invention, the present invention will be described in detail below in conjunction with the drawings and specific embodiments.
[0054] Such as Figure 1As shown in the figure, the multi-modal emotion recognition method based on acoustic and text features of the present invention uses OpenSMILE to extract the shallow emotion features of the input speech, and fuses them with the deep features obtained by the Transformer network learning the shallow features to generate multi-level acoustic features; uses forced alignment of the speech with the same content and the transcribed text to obtain pause information, then encodes the speaking pause information in the speech and adds it to the transcribed text, and sends it into the hierarchical densely connected DC-BERT model to obtain the text features, and then fuses them with the acoustic features; uses the bidirectional long short-term memory neural network BiLSTM-ATT based on the attention mechanism as the classifier, uses a priori knowledge through the BiLSTM network to obtain effective context information, and extracts the part that highlights the emotion information in the features through the attention mechanism to avoid information redundancy. A global average pooling layer is added behind the attention mechanism to replace the traditional fully connected layer, which can effectively prevent the overfitting problem, and finally sends it into the softmax layer for emotion classification; adopts the joint training method of Transformer and BiLSTM, and through artificial observation, it is found that the effect of 10 network iterations is the best. Therefore, the model after 10 iterations is selected as the final classifier model of the present invention.
[0055] The specific steps are as follows:
[0056] The first step: Send the original speech signal into OpenSMILE, and use its internal configuration file to extract the features of the speech, including intensity, loudness, Mel-frequency cepstral coefficients, pitch, and their statistical values for each short frame at the utterance level, such as maximum value, minimum value, average value, and standard deviation, etc.;
[0057] The second step: Send the shallow acoustic features extracted in the first step into the Transformer network to obtain deep features with global information;
[0058] The third step: Fuse the features obtained in the first and second steps to obtain shallow and deep features;
[0059] The fourth step: Use the Penn Phonetics Lab Forced Aligner (P2FA) to perform forced alignment on the preprocessed transcribed text and audio. After alignment, the timestamp of each word will be generated. According to the interval length between words, use "." to encode the pauses;
[0060] The fifth step: Send the pause-encoded text obtained in the fourth step into DC-BERT. The present invention selects the 768-dimensional output sequence of the penultimate layer of DC-BERT as the utterance-level text feature; specifically as Figure 3 shown
[0061] Step 6: After fusing the acoustic features and text features, send them into a BiLSTM network with an attention mechanism for sentiment classification;
[0062] Specifically, the internal form of DC-BERT in Step 5 is as follows: Assume a given input feature sequence X, then x i = H(x i-1 ) + αx i-1 + βx i-2 , where x i is the i-th element of the input feature sequence X, H is a non-linear function, and α and β are weight coefficients for retaining the information of the first two layers, so that each layer can obtain the results processed by the first two layers without dominating. The DC-BERT model consists of 12 layers of Transformer, and the output of each layer can theoretically be used as the text features at the utterance level, as shown in Figure 2 .
[0063] There are three attention mechanisms used in Step 6, namely the local attention mechanism, the self-attention mechanism, and the multi-head attention mechanism.
[0064] Local attention mechanism, which only focuses on a part of the encoded hidden layer. Local attention first generates an alignment position p t at time t for the current node, and then selectively sets a context window with a fixed size of 2D + 1. The formula is as follows:
[0065]
[0066] where D is selected according to experience; p t is the center of the window, determined by h t of the current hidden state, and is a real number; the calculation process of the alignment weights is similar to that of traditional attention:
[0067]
[0068] where the standard deviation σ is set according to experience.
[0069] The self-attention mechanism utilizes the weighted correlation between the elements of the input feature sequence. Specifically, each element of the input sequence can be projected into three different representation forms through a linear function: query, key, and value, and their calculation formulas are as follows:
[0070]
[0071] The final attention matrix is as shown in the formula:
[0072]
[0073] where Q is the query matrix, K is the key matrix, V is the value matrix of the sentence, and d k is the scaling factor.
[0074] Based on the self-attention mechanism, the present invention compares the influence of the multi-head self-attention mechanism on the speech emotion recognition task. Multi-head means that there is more than one set of projections for each variable (query, key, and value) of the input feature sequence. That is to say, on the premise of non-sharing of parameters, after mapping Q, K, and V through the parameter matrix, single-layer self-attention is performed, and then the self-attention is stacked layer by layer. The calculation formula of multi-head self-attention is:
[0075] head i = attention(QW i Q , KW i K , VW i V )
[0076] Multihead(Q, K, V) = Concat(head1,..., head n )
[0077] It is found through experiments that the BiLSTM network based on the local attention mechanism performs better than the BiLSTM network based on the self-attention mechanism or the multi-head self-attention mechanism. After analysis, in terms of the network structure, the local attention mechanism has fewer model parameters than the other two attention mechanisms, and for the emotion recognition task with a small amount of data, a relatively large network structure may not achieve the expected effect. Therefore, it is preferably to use the BiLSTM network based on the local attention mechanism as the classifier.
Claims
1. A multi-modal sentiment recognition method based on acoustic and text features, characterized in that: Extract the shallow emotional features of the input speech using OpenSMILE, and fuse them with the deep features obtained after the Transformer network learns the shallow features to generate multi-level acoustic features; use the speech with the same content and the transcription text to perform forced alignment to obtain pause information, then encode the speaking pause information in the speech and add it to the transcription text, and send it into the hierarchical dense connection DC-BERT model to obtain text features, and then fuse them with the acoustic features; Among them, input the transcription text with the marked pause encoding into the trained improved DC-BERT, and the improved DC-BERT outputs the emotional features of the utterance-level text according to the pause encoding annotation in the transcription text; splice and fuse the deep and shallow fusion feature sequences corresponding to the speech audio with the emotional features of the utterance-level text to obtain the acoustic-text fusion features of each sentence in this segment of audio; Use the bidirectional long short-term memory neural network BiLSTM-ATT based on the attention mechanism as a classifier. The BiLSTM network uses prior knowledge to obtain effective context information, and extracts the part of the feature that highlights the emotional information through the attention mechanism to avoid information redundancy. Add a global average pooling layer after the attention mechanism to replace the traditional fully connected layer, effectively preventing the overfitting problem, and finally send it into the softmax layer for emotion classification.
2. The multimodal emotion recognition method based on acoustic and text features according to claim 1, wherein The specific steps are as follows: S1: Input the original speech audio to be judged into OpenSMILE, and use the emobase feature set in the OpenSMILE toolbox to extract the shallow acoustic features in the original speech data; S2: Input the extracted shallow acoustic features into the Transformer network, and use the encoder structure of the Transformer network to effectively learn the relationship between the input shallow acoustic features, so as to output a feature sequence related to emotion, that is, a deep feature with global information; S3: Splice and fuse the sequence of shallow acoustic features and the sequence of deep features to obtain a deep and shallow fusion feature sequence, with the content of the shallow feature sequence in the front and the deep feature in the back for splicing; S4: Preprocess the text transcribed from the original speech: delete the punctuation marks in the text, and unify the writing form of the word format formed by the transcription; S5: Use the Penn Phonetic Aligner P2FA to perform forced alignment on the transcription text preprocessed in step S4 and the original speech, so as to determine the position and duration of the pause; S6: Divide the different pause durations in the speech audio into six intervals: 0.05 - 0.1s, 0.1 - 0.3s, 0.3 - 0.6s, 0.6 - 1.0s, 1.0 - 2.0s, and greater than 2.0s. Use: "..”, "…”, "…”, "…”, "……”, "……” to mark the pause durations of the six intervals in the transcription text, match the pause durations of the speech audio at the marked positions in the transcription text, and add the mark "." at the end of each speaker's sentence in the text as the end mark; S7: Input the transcribed text with marked pause codes into the trained improved DC-BERT. The improved DC-BERT outputs the emotional features of the utterance-level text according to the pause code annotations in the transcribed text; S8: Concatenate and fuse the deep and shallow fusion feature sequences corresponding to the speech audio with the emotional features of the utterance-level text to obtain the acoustic-text fusion features for each sentence in this segment of audio; S9: Finally, send the acoustic-text fusion features into a BiLSTM network with an attention mechanism for emotion classification, and output the corresponding emotion classification to achieve emotion recognition.
3. The multimodal emotion recognition method based on acoustic and text features according to claim 1, wherein: Use the built-in file to extract shallow acoustic features from the original speech signal sent into OpenSMILE, including intensity, loudness, mel-frequency cepstral coefficients, pitch, and their statistical values for each short frame at the utterance level, including maximum value, minimum value, average value, and standard deviation; The shallow acoustic features consist of a sequence of low-level descriptors; only select the audio and transcribed text representing anger, happiness, neutrality, and sadness in the emotion dataset for recognition, and happiness is combined by the emotions of gladness and excitement.
4. The multimodal emotion recognition method based on acoustic and text features according to claim 1, characterized in that: Send the transcribed text that has been forced-aligned and encoded by the University of Pennsylvania Speech Label Forced Alignment Tool into the improved DC-BERT, and select the 768-dimensional output sequence of the penultimate layer of DC-BERT as the utterance-level text features; The improved DC-BERT model retains the residual connections inside each multi-head self-attention layer of the Transformer in the traditional BERT model, and adds dense connections between layers, that is, the input of each multi-head self-attention layer additionally adds the feature information of the previous two layers to accelerate the convergence speed of the model, make the loss function of the network smoother, and the features extracted by each layer can also be reused between different attention layers, improving the utilization rate of features; The internal form of the improved DC-BERT is as follows: Given an input feature sequence X, then x i = H(x i-1 ) + αx i-1 + βx i-2 , where x i is the i-th element of the input feature sequence X, H is a non-linear function, and α and β are weight coefficients for retaining the information of the first two layers, such that each layer can obtain the results of the first two layers' processing without dominating; the improved DC-BERT model consists of 12 layers of Transformers, and the output of each layer can be used as the text feature at the discourse level.
5. The multimodal emotion recognition method based on acoustic and text features according to claim 1, characterized in that: After fusing the acoustic features and text features, send them into a BiLSTM network with an attention mechanism for emotion classification. There are three types of attention mechanisms in the BiLSTM network, namely local attention mechanism, self-attention mechanism, and multi-head attention mechanism; Local attention mechanism: This mechanism only focuses on a part of the encoded hidden layer. Local attention first generates an alignment position p for the current node at time t t , and then selectively sets a context window with a fixed size of 2D + 1, as shown in the following formula: where D is selected empirically; p t is the window center, determined by h of the current hidden state t and is a real number; the calculation process of the alignment weights is similar to that of traditional attention: where the standard deviation σ is set empirically, h t is the hidden state of the current decoder at the t-th time step, is the hidden state of the encoder at the i-th time step, where i represents the position of the input sequence, and T x represents the sequence length; The self-attention mechanism utilizes the weighted correlation between the elements of the input feature sequence, that is, each element of the input sequence can be projected into three different representation forms through a linear function: query, key, value, and its calculation formula is as follows: where x i represents the i-th element in the input feature sequence, q i , v i , k i represent the query vector, value vector, and key vector of the i-th element in the input feature sequence, represent the transposes of the three weight matrices for obtaining the query vector, value vector, and key vector, The final attention matrix is as shown in the formula: where Q is the query matrix, K is the key matrix, V is the value matrix of the sentence, and d k is the scaling factor; On the basis of the self-attention mechanism, the influence of the multi-head self-attention mechanism on the speech emotion recognition task is compared. Multi-head means that the number of projections of each variable of the input feature sequence: query, key, value is more than one group. That is, on the premise of non-shared parameters, after mapping Q, K, V through the parameter matrix, perform single-layer self-attention, and then stack the self-attention layers. The calculation formula of the multi-head self-attention is: Multihead(Q,k,V) = Concat(head1,...,head n )。
Citation Information
Patent Citations
Multi-modal emotion recognition method, device and equipment and storage medium
CN111898670A
Speech emotion feature extraction method based on transformer model encoder
CN112466326A