Parallel network speech emotion recognition method based on mixed attention mechanism

By adopting a parallel network structure with a hybrid attention mechanism in speech emotion recognition and using multi-scale dilated convolution and Bi-LSTM for feature extraction and modeling, the problems of insufficient feature extraction and lack of temporal modeling in the existing technology are solved, and higher recognition accuracy and model generalization ability are achieved.

CN120600054APending Publication Date: 2025-09-05CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510609900.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing speech emotion recognition methods find it difficult to effectively extract the emotional features of speech signals, and a single network structure is unable to handle local and global dependencies, resulting in the model being unable to learn deep features. In addition, the lack of a clear temporal modeling mechanism makes it difficult to fully capture temporal information.

Method used

A parallel network structure based on a hybrid attention mechanism is adopted to achieve more comprehensive feature capture and fusion through multi-scale void convolution feature extraction, multi-channel Bi-LSTM sequence modeling and hybrid attention mechanism, combined with self-attention and channel attention.

Benefits of technology

It improves the accuracy of speech emotion recognition and enhances the generalization ability of the model, which can better capture the emotional changes in speech signals and achieve more accurate classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600054A_ABST
    Figure CN120600054A_ABST
Patent Text Reader

Abstract

The invention provides a parallel network speech emotion recognition method based on a mixed attention mechanism, and belongs to the field of speech emotion recognition. Comprising the following steps: S1, extracting combined features of an audio; s2, mapping the low-dimensional vector into a higher-dimensional vector by using an embedding layer; s3, performing multi-scale cavity convolution feature extraction on the combined features; s4, carrying out sequence modeling by using a multi-channel LSTM; s5, capturing a deep relationship by using a mixed attention mechanism; s6, fusing features by using feature splicing and convolution; and S7, classifying and outputting emotion categories by using a SoftMax layer. The method has high accuracy in the speech emotion recognition task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of speech emotion recognition, and specifically provides a parallel network speech emotion recognition method based on a hybrid attention mechanism. Background Art

[0002] Everyday speech contains not only semantic information but also rich emotional information. Speech emotion recognition aims to identify the speaker's emotional state by analyzing the emotional features in speech signals. This technology can be applied in many fields, including human-computer interaction, intelligent customer service, and mental health analysis. For example, identifying the user's emotional state in intelligent voice assistants can improve the naturalness and intelligence of interactions. In the medical field, speech emotion analysis can help monitor mental illnesses such as depression. Therefore, researching effective speech emotion recognition methods is crucial.

[0003] In the task of speech emotion recognition, how to effectively extract emotional features from speech signals and improve the model's generalization ability is a major research challenge. Early speech emotion recognition methods primarily relied on hand-crafted features and traditional machine learning algorithms for classification, which have limitations when processing complex emotional information. In recent years, deep learning technology has been rapidly applied to speech emotion recognition. Techniques such as convolutional neural networks, recurrent neural networks, long short-term memory networks, and self-attention mechanisms have been widely used for emotion classification and feature extraction. However, a single network architecture struggles to adequately model and process local and global dependencies, resulting in the model's inability to learn the deep-level characteristics of speech emotion. Furthermore, the representation of different emotion categories in speech signals may exhibit cross-scale characteristics, and a single feature may not fully capture subtle emotional variations in speech.

[0004] After searching, I found application publication number CN113793627B, titled "A Method and Apparatus for Attention-Based Multi-Scale Convolutional Speech Emotion Recognition." The patent describes a method for speech emotion recognition based on attention, including constructing a speech emotion recognition model comprising a first convolutional neural network layer, two parallel channels consisting of an attention layer and a second convolutional neural network layer, a first fully connected layer, a spatial attention layer, a second fully connected layer, and a softmax classifier; and inputting the spectrogram corresponding to the speech to be recognized into the trained speech emotion recognition model to obtain an emotion classification result for the speech to be recognized. The patent describes a method for speech emotion recognition based on attention, embedding two parallel channel attention mechanisms and a spatial attention mechanism that fuses channels within a deep learning neural network to enhance useful information and suppress information that is not useful for the task, resulting in more accurate recognition results. However, this architecture lacks a clear temporal modeling mechanism, making it difficult to fully capture temporal information. The present invention employs a long short-term memory network for explicit temporal modeling in each channel and employs a hybrid attention mechanism to fuse information across different channels. This not only implements the functions of a self-attention mechanism but also a channel-attention mechanism, further improving the accuracy of speech emotion recognition. Summary of the Invention

[0005] The present invention aims to solve the above problems in the prior art. It proposes a parallel network speech emotion recognition method based on a hybrid attention mechanism. The technical solution of the present invention is as follows:

[0006] A parallel network speech emotion recognition method based on a hybrid attention mechanism comprises the following steps:

[0007] S1, preprocesses and frames the input speech signal and extracts combined features;

[0008] S2, uses the embedding layer to map the low-dimensional vector to a higher-dimensional vector;

[0009] S3, multi-scale hole convolution feature extraction is performed on the combined features;

[0010] S4, sequence modeling using multi-channel Bi-LSTM;

[0011] S5, uses hybrid attention mechanism to capture deep relationships;

[0012] S6, uses feature concatenation and convolution to fuse features;

[0013] S7, use the SoftMax layer to classify and output the emotion category.

[0014] Furthermore, the step S1 pre-processes and frames the input speech signal to extract combined features, specifically including:

[0015] First, the input speech signal is expanded by time stretching, pitch change, and time shift, and then the data set is divided. Each speech is then divided into frames.

[0016] Secondly, five features are extracted from each speech segment after framing: Mel-frequency cepstral coefficient feature, first-order Mel-frequency cepstral coefficient feature, fundamental frequency, energy, and zero-crossing rate;

[0017] Finally, different features of the same frame are concatenated into a feature vector

[0018] Furthermore, the step S2 uses an embedding layer to map the low-dimensional vector to a higher-dimensional vector, specifically including:

[0019] The combined features extracted from each frame are mapped through the fully connected layer, and the feature information is input into the subsequent network layer. The calculation formula is:

[0020] X out =X in w

[0021] Among them, w is the weight matrix, X in 、X out are the features of each frame in the input and output feature maps, respectively.

[0022] Furthermore, the step 3 performs multi-scale dilated convolution feature extraction on the combined features, specifically including:

[0023] The feature tensor obtained in step S2 is input into three convolutional networks of different scales but the same structure. Each parallel network consists of a convolutional layer, a batch normalization layer, and an activation layer. After passing through a convolutional layer, batch normalization is performed, and then the final activation is performed.

[0024] Furthermore, step S4 uses a multi-channel Bi-LSTM to perform sequence modeling, specifically including:

[0025] The outputs of the three channels obtained in step S3 are fed into three identical Bi-LSTM networks. Bi-LSTM is an extended form of LSTM. Unlike the traditional unidirectional LSTM, Bi-LSTM has two hidden states obtained by latent propagation in two directions at each time step. The other hidden state is obtained by backward propagation from the end of the sequence. n The forward hidden layer output and the reverse hidden layer output at the moment will be spliced ​​to obtain the hidden layer output y at the current moment n ,Bi-LSTM is used to capture the past and future context information of the sequence.

[0026] Furthermore, step S5 uses a hybrid attention mechanism to capture deep relationships, specifically including:

[0027] When calculating the attention of the X1 channel, unlike the self-attention mechanism, the concatenated output of the three channels is used to calculate the k vector. This allows us to focus on the temporal relationship of the X1 channel while also introducing attention between the X2 and X3 channels, enabling the model to learn more emotional information. The Resnet network structure is applied to the hybrid attention mechanism, directly adding the input and attention output. The hybrid attention mechanism is applied to the three channels after the Bi-LSTM layer to obtain the hybrid attention value for each channel.

[0028] Furthermore, step S6 utilizes feature concatenation and convolution to fuse features, specifically including:

[0029] The output features of the three channels are convolved through a convolutional layer with three parameters, using dilated convolution instead of average pooling; the expansion coefficient of the dilated convolution is equal to the output feature dimension of a single channel, so that the features at corresponding positions of different channels can be multiplied by a learnable parameter and then added together.

[0030] Furthermore, step S7 utilizes the SoftMax layer to classify and output the emotion category, specifically including:

[0031] The fused features are flattened and then mapped to the probability values ​​of the eight emotions through the SoftMax layer. The calculation formula is:

[0032]

[0033] Among them, z i is the i-th component in the input vector, and n is the length of the input vector.

[0034] An electronic device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for parallel network speech emotion recognition based on a hybrid attention mechanism as described in any one of the above is implemented.

[0035] A non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the parallel network speech emotion recognition method based on a hybrid attention mechanism as described in any one of the above is implemented.

[0036] The advantages and beneficial effects of the present invention are as follows:

[0037] 1. The present invention combines five features of speech extraction in the feature extraction part, and uses the combined features to make up for the defect that a single feature is insufficient in expressing speech emotion information, thereby improving the accuracy of speech emotion recognition.

[0038] 2. The present invention adopts a parallel network channel structure, which calculates the deep features of speech features separately and then fully integrates them, so that the emotional expression ability of the model is enhanced.

[0039] 3. The present invention uses dilated convolution and ordinary convolution in the multi-scale dilated convolution module, making full use of their different information capture capabilities to achieve better deep feature capture and further improve the accuracy of speech emotion recognition.

[0040] 4. The present invention improves the traditional attention mechanism so that the improved attention mechanism can not only realize the function of the self-attention mechanism, but also capture the correlation information of different channels, further enhance the ability to express speech emotion information, and achieve more accurate classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 The present invention provides a flow chart of a preferred embodiment of a parallel network speech emotion recognition method based on a hybrid attention mechanism.

[0042] Figure 2 It is a multi-scale dilated convolution module diagram.

[0043] Figure 3 This is the Bi-LSTM network structure diagram.

[0044] Figure 4 This is the structure diagram of the mixed attention machine.

[0045] Figure 5 It is a feature fusion structure diagram. DETAILED DESCRIPTION

[0046] The following will describe the technical solutions in the embodiments of the present invention in detail with reference to the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention.

[0047] The technical solution of the present invention to solve the above technical problems is:

[0048] like Figure 1 As shown, a parallel network speech emotion recognition method based on a hybrid attention mechanism includes:

[0049] S1, extract combined features from audio.

[0050] First, the input speech signal is expanded by time stretching, pitch change, and time shift, and then the data set is divided. Each speech is then divided into frames.

[0051] Secondly, five features are extracted from each speech segment after framing: Mel-frequency cepstral coefficient feature, first-order Mel-frequency cepstral coefficient feature, fundamental frequency, energy, and zero-crossing rate;

[0052] Finally, different features of the same frame are concatenated into a feature vector.

[0053] S2 uses the embedding layer to map the low-dimensional vector into a higher-dimensional vector.

[0054] The combined features extracted from each frame are mapped into higher-dimensional vectors through the embedding layer to enhance the feature representation capability and provide richer feature information to be input to the subsequent network layers. The calculation formula is:

[0055] X out =X in w

[0056] Among them, w is the weight matrix, X in 、X out are the features of each frame in the input and output feature maps, respectively.

[0057] S3, perform multi-scale dilated convolution feature extraction on the combined features.

[0058] The feature tensor obtained in step S2 is input into three convolutional networks with different scales but the same structure, such as Figure 2 As shown in the figure, each parallel network consists of a convolutional layer, a batch normalization layer, and an activation layer. After passing through a convolutional layer, batch normalization is performed and the final activation is performed.

[0059] S4, sequence modeling using multi-channel LSTM.

[0060] The outputs of the three channels obtained in step S3 are fed into three identical Bi-LSTM networks. Bi-LSTM has two hidden states at each time step: one is the hidden state obtained by forward propagation from the beginning of the sequence, and the other is the hidden state obtained by backward propagation from the end of the sequence. Figure 3 As shown, X n The forward hidden layer output and the reverse hidden layer output at the moment will be spliced ​​to obtain the hidden layer output y at the current moment n Due to this bidirectional processing, Bi-LSTM is able to more comprehensively capture the past and future contextual information of the sequence, thereby gaining a deeper understanding of the information in the sequence and helping to more accurately model long-term dependencies in sequence data.

[0061] S5, uses hybrid attention mechanism to capture deep relationships.

[0062] like Figure 4As shown, when calculating the attention of the X1 channel, unlike the self-attention mechanism, the present invention uses the spliced ​​output of the three channels to calculate the k vector. In this way, while paying attention to the temporal relationship of the X1 channel, it can also introduce the attention between the X2 and X3 channels, so that the model can learn more emotional information. In addition, in order to make the model converge quickly, the structure of the Resnet network is applied to the hybrid attention mechanism, and the input and attention output are directly added. The hybrid attention mechanism is applied to the three channels after the Bi-LSTM layer to obtain the hybrid attention value of each channel, so as to achieve better feature representation and thus improve the accuracy of emotion classification.

[0063] S6, uses feature concatenation and convolution to fuse features.

[0064] like Figure 5 As shown in the figure, the output features of the three channels are convolved through a convolutional layer with three parameters. Dilated convolution is used instead of average pooling, which allows for better feature fusion using learnable parameters. The dilation coefficient of the dilated convolution is equal to the output feature dimension of a single channel. This allows the features at corresponding positions in different channels to be multiplied by a learnable parameter and then added together.

[0065] S7, use the SoftMax layer to classify and output the emotion category.

[0066] The fused features are flattened and then mapped to the probability values ​​of the eight emotions through the SoftMax layer. The calculation formula is:

[0067]

[0068] Among them, z i is the i-th component in the input vector, and n is the length of the input vector.

[0069] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions.

[0070] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0071] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0072] The above embodiments should be understood as merely illustrating the present invention and not as limiting the scope of protection of the present invention. After reading the contents of the present invention, technicians may make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. A parallel network speech emotion recognition method based on a hybrid attention mechanism, characterized in that: The following steps are involved: S1, preprocesses and frames the input speech signal and extracts combined features; S2, uses the embedding layer to map the low-dimensional vector to a higher-dimensional vector; S3, multi-scale hole convolution feature extraction is performed on the combined features; S4, sequence modeling using multi-channel Bi-LSTM; S5, uses hybrid attention mechanism to capture deep relationships; S6, uses feature concatenation and convolution to fuse features; S7, use the SoftMax layer to classify and output the emotion category.

2. A parallel network speech emotion recognition method based on a hybrid attention mechanism according to claim 1, characterized in that: The step S1 pre-processes and frames the input speech signal to extract combined features, specifically including: First, the input speech signal is expanded by time stretching, pitch change, and time shift, and then the data set is divided. Each speech is then divided into frames. Secondly, five features are extracted from each speech segment after framing: Mel-frequency cepstral coefficient feature, first-order Mel-frequency cepstral coefficient feature, fundamental frequency, energy, and zero-crossing rate; Finally, different features of the same frame are concatenated into a feature vector.

3. A parallel network speech emotion recognition method based on a hybrid attention mechanism according to claim 1, characterized in that: The step S2 uses an embedding layer to map the low-dimensional vector to a higher-dimensional vector, specifically including: The combined features extracted from each frame are mapped through the fully connected layer, and the feature information is input into the subsequent network layer. The calculation formula is: X out =X in ·w Among them, w is the weight matrix, X in 、X out are the features of each frame in the input and output feature maps, respectively.

4. A parallel network speech emotion recognition method based on a hybrid attention mechanism according to claim 1, characterized in that: The step 3 performs multi-scale dilated convolution feature extraction on the combined features, specifically including: The feature tensor obtained in step S2 is input into three convolutional networks of different scales but the same structure. Each parallel network consists of a convolutional layer, a batch normalization layer, and an activation layer. After passing through a convolutional layer, batch normalization is performed, and then the final activation is performed.

5. A parallel network speech emotion recognition method based on a hybrid attention mechanism according to claim 4, characterized in that: Step S4 uses a multi-channel Bi-LSTM to perform sequence modeling, specifically including: The outputs of the three channels obtained in step S3 are fed into three identical Bi-LSTM networks. Bi-LSTM is an extended form of LSTM. Unlike the traditional unidirectional LSTM, Bi-LSTM has two hidden states in each time step: one is the hidden state obtained by forward propagation from the beginning of the sequence, and the other is the hidden state obtained by backward propagation from the end of the sequence. n The forward hidden layer output and the reverse hidden layer output at the moment will be spliced ​​to obtain the hidden layer output y at the current moment n ,Bi-LSTM is used to capture the past and future context information of the sequence.

6. A parallel network speech emotion recognition method based on a hybrid attention mechanism according to claim 1, characterized in that: Step S5 uses a hybrid attention mechanism to capture deep relationships, specifically including: When calculating the attention of the X1 channel, unlike the self-attention mechanism, the concatenated output of the three channels is used to calculate the k vector. This allows us to focus on the temporal relationship of the X1 channel while also introducing attention between the X2 and X3 channels, enabling the model to learn more emotional information. The Resnet network structure is applied to the hybrid attention mechanism, directly adding the input and attention output. The hybrid attention mechanism is applied to the three channels after the Bi-LSTM layer to obtain the hybrid attention value for each channel.

7. A parallel network speech emotion recognition method based on a hybrid attention mechanism according to claim 1, characterized in that: The step S6 utilizes feature concatenation and convolution to fuse features, specifically including: The output features of the three channels are convolved through a convolutional layer with three parameters, using dilated convolution instead of average pooling; the expansion coefficient of the dilated convolution is equal to the output feature dimension of a single channel, so that the features at corresponding positions of different channels can be multiplied by a learnable parameter and then added together.

8. A parallel network speech emotion recognition method based on a hybrid attention mechanism according to claim 1, characterized in that: The step S7 uses the SoftMax layer to classify and output the emotion category, specifically including: The fused features are flattened and then mapped to the probability values ​​of the eight emotions through the SoftMax layer. The calculation formula is: Among them, z i is the i-th component in the input vector, and n is the length of the input vector.

9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for parallel network speech emotion recognition based on a hybrid attention mechanism as described in any one of claims 1 to 8 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the parallel network speech emotion recognition method based on the hybrid attention mechanism is implemented as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • A Multi-Scale Convolutional Speech Emotion Recognition Method and Device Based on Attention

    CN113793627B

Cited By

  • Scraper disconnection identifying and monitoring method based on sound array and spatial-temporal characteristic network

    CN121708956A