A method and system for speech emotion recognition
By using the parallel channel attention mechanism and the double nested residual structure to extract weighted fusion emotional features in speech emotion recognition, the problems of insufficient feature extraction and improved network complexity in the prior art are solved, and the effect of efficiently extracting speech emotion features and reducing model complexity is achieved.
Patent Information
- Application Number
- CN202210816317.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-12
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-07-12
AI Technical Summary
When extracting voice emotional characteristics, the existing speech emotion recognition algorithms are not perfect enough, which can easily lead to the loss of useful information in the speech signal. At the same time, the network complexity is increased, which is not conducive to model optimization.
The parallel channel attention mechanism and the double nested residual structure were used to extract weighted fusion affective features, and the deep fusion and shallow emotional features were extracted through small-size and large-scale residual connections, and fused them to obtain a weighted fusion affective feature map.
Effectively extract speech emotional features, reduce global features loss, reduce model complexity, improve the efficiency of deep learning network parameter training, and avoid slow training and inefficiency caused by excessive network complexity.
Smart Images

Figure CN115273902B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech emotion recognition, and particularly relates to a deep learning-based speech emotion recognition method and system for extracting weighted fusion emotion features based on a parallel channel attention mechanism and a double nested residual structure. Background Art
[0002] Speech emotion recognition is one of the key technologies for human-computer communication, and the extraction of speech emotion features is an important basis for emotion discrimination, which will directly determine the effectiveness of emotion recognition. At present, speech emotion recognition algorithms are mainly divided into two categories: non-deep learning and deep learning.
[0003] For non-deep learning speech emotion recognition algorithms, speech is first input into a machine learning model to manually extract specific features, and then the result is obtained through a classifier. The features extracted by this method are not perfect enough, and it is easy to cause the loss of useful information in the speech signal. For deep learning-based speech emotion recognition algorithms, speech signals are first converted into spectrograms, and then input into a deep learning network model for feature extraction and classification, which can usually extract features more efficiently and obtain a higher recognition rate than traditional algorithms.
[0004] In deep learning-based speech emotion recognition algorithms, a convolutional neural network (CNN) is usually used as a feature extraction network. In order to better extract emotion features, CNN is also combined with other networks to form a joint feature extraction network, but at the same time, it also leads to an increase in network complexity, which is not conducive to the further optimization of the model. Therefore, the architecture of the emotion feature extraction network is particularly important for whether emotion features can be effectively extracted.
[0005] Inserting an attention mechanism into the network is an effective weighted operation to enhance the connection between features. In current related research, it is roughly divided into: channel attention mechanism, spatial attention mechanism, and hybrid attention mechanism.
[0006] The channel attention mechanism obtains a global receptive field through compression and excitation operations, and at the same time emphasizes the weight of each channel. However, the fully connected layer in its excitation operation consumes a large amount of network resources. The later proposed ECANet, whose ECA module changes the fully connected excitation operation to a non-fully connected method, greatly reduces the model complexity, but at the same time results in the inability to obtain a global receptive field and is prone to losing some global features.
[0007] The spatial attention mechanism expands the receptive field through convolutions of different sizes to capture context associations, which is beneficial to strengthening the associations between feature image pixels and usually needs to be used in combination with other structures such as pyramids.
[0008] The hybrid attention mechanism combines the channel attention mechanism and the spatial attention mechanism. Compared with using only one of these attention mechanisms alone, this method can achieve higher accuracy, but at the same time increases the spatial complexity. Summary of the Invention
[0009] Aiming at the above-mentioned defects existing in the prior art, with the purpose of effectively extracting speech emotion features and reducing the model complexity, the present invention designs a double-nested residual structure for extracting fused emotion features, and combines a parallel channel attention mechanism to weight the feature map, and proposes a speech emotion recognition method and system for extracting weighted fused emotion features based on the parallel channel attention mechanism and the double-nested residual structure.
[0010] To achieve the above object, the present invention adopts the following technical solutions:
[0011] A speech emotion recognition method, comprising the following steps:
[0012] S1. Perform parallel channel attention weighting on the input speech feature map to obtain a weighted feature map;
[0013] S2. Perform feature extraction on the weighted feature map through small-size residual connections and feature convolutions to obtain a deeply fused emotion feature map;
[0014] Perform shallow feature extraction on the weighted feature map through large-scale residual connections to obtain a shallow emotion feature map;
[0015] S3. Fuse the deeply fused emotion feature map and the shallow emotion feature map to obtain a weighted fused emotion feature map.
[0016] As a preferred solution, step S1 specifically includes the following steps:
[0017] S11. Adopt a channel attention mechanism to perform squeeze-and-excitation operations on the input speech feature map, extract the attention weights of each channel, and obtain the weighted parameters of each channel of the feature map;
[0018] S12. Reset the output channel dimension parameter of the excitation operation in the channel attention mechanism, and repeat the above step S11 until N groups of weight parameters of different scales X = {X1, X2,...., X N};
[0019] S13. Weight the input speech feature map with the weights obtained by taking the average of the weight parameters in X to obtain a weighted feature map.
[0020] As a preferred solution, in step S2, the extraction process of the deeply fused emotion feature map includes:
[0021] The weighted feature map is further subjected to feature extraction through six layers of 3×3 convolutions. Every other convolutional layer adjusts the length and width of the output feature map to 1 / 2 of the original, and the number of channels is adjusted to 2 times the original. Moreover, residual connections are inserted into the middle two convolutional layers for small-scale feature fusion. Meanwhile, 1×1 convolutions are used for dimensionality increase operations, and finally, a deeply fused emotion feature map is obtained in the last convolutional layer.
[0022] As a preferred solution, in step S2, the extraction process of the shallow emotion feature map includes:
[0023] Insert large-scale residual connections, and perform shallow feature extraction on the weighted feature map through three layers of convolutions: 1×1 convolution, 3×3 convolution, and 1×1 convolution to obtain a shallow emotion feature map; among them, each convolutional operation adjusts the length and width of the output feature map to 1 / 2 of the original, and the number of channels is adjusted to 2 times the original.
[0024] As a preferred solution, batch normalization and activation functions are also performed after each convolutional operation.
[0025] As a preferred solution, the activation function uses the Leaky ReLU function.
[0026] As a preferred solution, step S3 specifically includes:
[0027]
[0028] Among them, Z1 is the deeply fused emotion feature map, Z2 is the shallow emotion feature map, and Z is the weighted fused emotion feature map.
[0029] The present invention also provides a speech emotion recognition system, which applies the speech emotion recognition method described in any of the above solutions. The speech emotion recognition system includes:
[0030] A parallel channel attention weighting module, which is used to perform parallel channel attention weighting on the input speech feature map to obtain a weighted feature map;
[0031] A deep fusion feature extraction module, which is used to perform feature extraction on the weighted feature map through small-size residual connections and feature convolutions to obtain a deeply fused emotion feature map;
[0032] A shallow feature extraction module, which is used to perform shallow feature extraction on the weighted feature map through large-scale residual connections to obtain a shallow emotion feature map;
[0033] A feature fusion module, which is used to fuse the deeply fused emotion feature map and the shallow emotion feature map to obtain a weighted fused emotion feature map.
[0034] As a preferred solution, the process of extracting the deep fusion emotion feature map by the deep fusion feature extraction module includes: further extracting features from the weighted feature map through six layers of 3×3 convolution. Every other convolutional layer adjusts the length and width of the output feature map to 1 / 2 of the original, and the number of channels is adjusted to 2 times the original. And residual connections are inserted into the middle two convolutional layers for small-scale feature fusion. At the same time, 1×1 convolution is used for dimension elevation operation, and finally the deep fusion emotion feature map is obtained in the last layer of convolution;
[0035] The process of extracting the shallow emotion feature map by the shallow feature extraction module includes: inserting large-scale residual connections, and performing shallow feature extraction on the weighted feature map through three layers of convolution: 1×1 convolution, 3×3 convolution, and 1×1 convolution to obtain the shallow emotion feature map; wherein, each convolution operation adjusts the length and width of the output feature map to 1 / 2 of the original, and the number of channels is adjusted to 2 times the original.
[0036] As a preferred solution, the deep fusion feature extraction module and the shallow feature extraction module constitute a double-nested residual structure for extracting fused emotion features.
[0037] Compared with the prior art, the present invention has the following advantages:
[0038] (1) The present invention proposes a parallel channel attention mechanism that improves the ECA module, takes the mean of weight information at different scales and then weights the input speech feature map, minimizing the loss of global features to the greatest extent. At the same time, the model only performs a single dot product operation during weighting, ensuring that the addition of this module will not cause a large increase in the complexity of the entire system, thus slowing down the network process.
[0039] (2) The present invention designs a double-nested residual structure for effectively extracting fused emotion features, extracts more comprehensive emotion features by increasing the randomness of feature extraction; realizes fused emotion features through residual connections at different scales, generating new feature maps. This structure uses fewer network layers, effectively extracts fused emotion features, and is beneficial to improving the efficiency of deep learning network parameter training.
[0040] (3) The present invention effectively realizes the extraction of emotion features in the speech feature map, reduces the loss of global features, effectively realizes the extraction of fused emotion features, and at the same time has a low model complexity, avoiding problems such as slow training and low efficiency caused by an overly complex network. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a flowchart of the speech emotion recognition method according to an embodiment of the present invention;
[0042] Figure 2 is a model block diagram of the parallel channel attention machine weighted according to an embodiment of the present invention;
[0043] Figure 3 It is the double-nested residual structure diagram of the embodiment of the present invention;
[0044] Figure 4 It is the module architecture diagram of the speech emotion recognition system of the embodiment of the present invention. Specific implementation
[0045] The technical solution of the present invention will be further explained below through specific embodiments.
[0046] For convenience of description, an input speech feature map (abbreviated as input feature map) I ∈ R is set C×H×W .
[0047] As Figure 1 shown, the speech emotion recognition method for extracting weighted fusion emotion features based on the parallel channel attention mechanism and the double-nested residual structure in the embodiment of the present invention includes the following steps:
[0048] S1. Input the input feature map I into the channel attention mechanism of different scales to obtain N groups of attention weights of different scales:
[0049] X i = ECAweight(I), i ∈ [1, N]
[0050] Take the average of these weights to obtain the overall attention weight:
[0051]
[0052] Use this weight to weight the input speech feature map through dot product operation to obtain the weighted feature map Y.
[0053] Specifically, as Figure 2 shown, the above step S1 includes:
[0054] S11. First, perform a squeezing operation on the input speech feature map through global average pooling, and then perform a local excitation operation for non-full connection through two 1×1 convolutional kernels to obtain the weighted parameters of each channel of the input speech feature map;
[0055] S12. Set the output channel dimension parameter of the first layer of convolution of the excitation operation to different sizes, repeat the above step S11, and obtain N groups of weight parameters X of different scales after N operations:
[0056] X = {X1, X2,...., X N};
[0057] S13. Take the average of the weight parameters in X and weight the input speech feature map I to obtain the weighted feature map Y:
[0058]
[0059] S2. As shown in Figure 3 , the weighted feature map Y is further used to extract features through six 3×3 convolutional kernels. In order to improve the randomness of feature extraction and ensure that the complexity will not be increased accordingly, every other convolutional layer adjusts the length and width of the output feature map to 1 / 2 of the original, and the number of channels is adjusted to 2 times the original. Among them, the middle two convolutional layers adopt residual connections for small-scale feature fusion, which plays a role in feature fusion. At the same time, 1×1 convolution is used for dimensionality increase operation, and finally the deeply fused emotion feature map Z1 is obtained in the last convolutional layer.
[0060] The size of the output feature map is adjusted every other convolutional layer as follows:
[0061]
[0062] where l represents the current convolutional layer number, and l ∈ {2, 4, 6}, and the deeply fused feature map (i.e., the deeply fused emotion feature map) is obtained.
[0063] In addition, in the embodiment of the present invention, shallow features are also extracted from the weighted feature map Y by using three convolutional kernels of 1×1, 3×3, and 1×1 respectively through large-scale residuals. To ensure that the dimension of the extracted shallow feature map is consistent with that of the deeply fused feature map, the dimension of the output feature map is adjusted according to the formula described in step S2 for each convolutional layer, that is, the length and width of the output feature map are adjusted to 1 / 2 of the original for each convolutional operation, and the number of channels is adjusted to 2 times the original to ensure the unity of feature dimensions. Finally, the shallow feature map (i.e., the shallow emotion feature map) is obtained.
[0064] In order to extract the emotion features of the weighted feature map under convolutional operations of different scales, a relatively small number of network layers are used to extract the useful information contained in Y to the greatest extent. It should be noted that after each convolutional operation, batch normalization and an activation function are passed through, and the Leaky ReLU function is used as the activation function.
[0065] The extraction of the above shallow features and deeply fused emotion features constitutes a double-nested residual structure for fused emotion feature extraction.
[0066] S3. The deeply fused feature map and the shallow feature map in step S2 are fused to obtain a weighted fused emotion feature map:
[0067]
[0068] Based on the emotion recognition method of the embodiment of the present invention, as shown in Figure 4 , the emotion recognition system provided by the embodiment of the present invention includes:
[0069] Parallel channel attention weighting module: It performs channel attention weighting on the input speech feature map at different scales to obtain N groups of feature weights, takes the mean to obtain the overall weight, and uses this weight to weight the input speech feature map to obtain a weighted feature map.
[0070] Deep fusion feature extraction module: It further extracts the emotional features of the weighted feature map through six layers of 3×3 convolution operations. At the same time, it performs small-scale feature fusion on the features extracted by the middle two convolutional layers and ensures the unity of the feature dimensions through dimensionality increase operations to obtain a deep fusion emotional feature map.
[0071] Shallow feature extraction module: Through three convolutional kernels of 1×1, 3×3, and 1×1, it performs shallow feature extraction on the fused feature Y and keeps the feature dimensions consistent with the deep fusion features Figure 1 to obtain a shallow feature map.
[0072] Feature fusion module: It adds the deep fusion feature map and the shallow feature map to obtain a weighted fusion feature map output by the double nested residual structure.
[0073] Among them, the deep fusion feature extraction module and the shallow feature extraction module constitute a double nested residual structure for fusing emotional feature extraction.
[0074] In summary, the present invention is an emotional recognition method and system for effectively extracting emotional features based on parallel channel attention weighting and double nested residual structure. The input speech feature map is first weighted through the parallel channel attention mechanism to strengthen useful information; then it passes through the double nested residual structure, extracts the deep fusion feature map and the shallow feature map, and fuses them to obtain the required emotional feature map, thereby performing emotional recognition. The present invention realizes the effective extraction of emotional features, reduces the loss of global features, and avoids problems such as slow training and low efficiency caused by overly complex networks.
[0075] The above description only elaborates in detail on the preferred embodiments and principles of the present invention. For those of ordinary skill in the art, according to the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.
Claims
1. A method for speech emotion recognition, characterized in that, It includes the following steps: S1. Perform parallel channel attention weighting on the input speech feature map to obtain a weighted feature map; S2. Extract features from the weighted feature map through small-scale residual connections and feature convolutions to obtain a deeply fused emotion feature map; Extract shallow features from the weighted feature map through large-scale residual connections to obtain a shallow emotion feature map; S3. Fuse the deeply fused emotion feature map and the shallow emotion feature map to obtain a weighted fused emotion feature map; The step S1 specifically includes the following steps: S11. Use the channel attention mechanism to perform squeezing and excitation operations on the input speech feature map, extract the attention weights of each channel, and obtain the weighted parameters of each channel of the feature map; S12. Reset the output channel dimension parameter of the excitation operation in the channel attention mechanism, and repeat the above step S11 until N a set of weight parameters with different scales ; S13. Take the weights obtained by averaging the weight parameters in X and use them to weight the input speech feature map to obtain a weighted feature map.
2. The method for speech emotion recognition according to claim 1, characterized in that, In step S2, the extraction process of the deeply fused emotion feature map includes: Further extract features from the weighted feature map through six layers of 3×3 convolutions. Every other convolutional layer adjusts the length and width of the output feature map to 1 / 2 of the original, and the number of channels is adjusted to 2 times the original. And insert residual connections into the middle two convolutional layers for small-scale feature fusion. At the same time, use 1×1 convolution for dimension increase operation, and finally obtain the deeply fused emotion feature map in the last layer of convolution.
3. The method for speech emotion recognition according to claim 2, characterized in that, In the step S2, the extraction process of the shallow emotion feature map includes: Insert large-scale residual connections, and perform shallow feature extraction on the weighted feature map through three layers of convolutions: 1×1 convolution, 3×3 convolution, and 1×1 convolution to obtain a shallow emotion feature map; among them, each convolutional operation adjusts the length and width of the output feature map to 1 / 2 of the original, and the number of channels is adjusted to 2 times the original.
4. The method for speech emotion recognition according to claim 2 or 3, characterized in that, After each convolutional operation, batch normalization and activation functions are also passed through.
5. The method for speech emotion recognition according to claim 4, characterized in that, The activation function uses the Leaky ReLU function.
6. The method for speech emotion recognition according to claim 3, characterized in that, The step S3 specifically includes: ; Among them, Z 1 is the deeply fused emotion feature map, Z 2 is the shallow emotion feature map, Z is the weighted fused emotion feature map.
7. A speech emotion recognition system applying the speech emotion recognition method according to any one of claims 1 - 6, characterized in that, The speech emotion recognition system includes: A parallel channel attention weighting module for performing parallel channel attention weighting on the input speech feature map to obtain a weighted feature map; A deep fusion feature extraction module for extracting features from the weighted feature map through small-scale residual connections and feature convolutions to obtain a deeply fused emotion feature map; A shallow feature extraction module for extracting shallow features from the weighted feature map through large-scale residual connections to obtain a shallow emotion feature map; A feature fusion module for fusing the deeply fused emotion feature map and the shallow emotion feature map to obtain a weighted fused emotion feature map.
8. The speech emotion recognition system according to claim 7, characterized in that, The extraction process of the deeply fused emotion feature map by the deep fusion feature extraction module includes: further extracting features from the weighted feature map through six layers of 3×3 convolutions. Every other convolutional layer adjusts the length and width of the output feature map to 1 / 2 of the original, and the number of channels is adjusted to 2 times the original. And insert residual connections into the middle two convolutional layers for small-scale feature fusion. At the same time, use 1×1 convolution for dimension increase operation, and finally obtain the deeply fused emotion feature map in the last layer of convolution; The process of extracting the shallow - layer sentiment feature map by the shallow - layer feature extraction module includes: inserting a large - scale residual connection, and performing shallow - layer feature extraction on the weighted feature map through three - layer convolution: 1×1 convolution, 3×3 convolution, and 1×1 convolution to obtain the shallow - layer sentiment feature map; among them, each layer of convolution operation adjusts the length and width of the output feature map to 1 / 2 of the original, and the number of channels is adjusted to 2 times the original.
9. The speech emotion recognition system according to claim 8, characterized in that, The deep - layer fusion feature extraction module and the shallow - layer feature extraction module form a double - nested residual structure for fusing sentiment feature extraction.
Citation Information
Patent Citations
Speech emotion recognition method
CN109036465A
Residual network expression recognition method integrated with attention
CN112541409A