Speech emotion recognition method and system based on multi-dimensional information perception strategy
By constructing multi-dimensional information perception and cross-dimensional interleaving modules, combining WavLM-Large and MFCC features, the shortcomings of existing speech emotion recognition methods in feature extraction and fusion are solved, and the accuracy and robustness of the model's emotion recognition in complex environments are improved.
Patent Information
- Application Number
- CN202510741490.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing speech emotion recognition methods have shortcomings in feature extraction and multi-source information fusion, and it is difficult to effectively capture emotional features across scales and frequency ranges, resulting in poor accuracy and robustness of the model in complex environments.
A speech emotion recognition method based on multidimensional information perception strategy is adopted, and a multidimensional information perception and cross-dimensional interleaving module is constructed through WavLM-Large and MFCC feature extraction, combining Transformer layer, MDIP layer, CDI layer and convolutional layer, multidimensional information perception and cross-dimensional interleaving module are constructed to enhance the representation ability of time and frequency features, and feature fusion and classification are performed through SENet.
It improves the accuracy, robustness and versatility of the speech emotion recognition model in multiple data sets and multi-situations, enhances the emotion recognition ability, and adapts to the emotion recognition effect at different speech speeds, speaking styles and signal-to-noise ratios.
Smart Images

Figure CN120279950A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech emotion recognition, and in particular, to a speech emotion recognition method and system based on a multi-dimensional information perception strategy. Background Art
[0002] Speech is one of the basic means of human communication and can effectively convey the emotions of the speaker through acoustic signals. With the rapid development of human-computer interaction systems, in-depth analysis of speech signals has become one of the key technologies for improving the quality of human-computer interaction. Speech not only contains rich information but also carries complex emotional data and can express emotional reactions to objects, scenes, or events. The process of automatically identifying a person's emotions from speech signals is called Speech Emotion Recognition (SER), and this technology has become one of the most important research and development fields in the past few decades. At present, speech emotion recognition technology has been widely applied in many fields such as education, medical care, and services and has continued to receive extensive attention. The key technologies of SER lie in the extraction of emotional features and the construction of emotion recognition models. In the process of implementing the present invention, the applicant found that: currently, most studies still mainly focus on the fusion between traditional features, and there is still a lack of in-depth exploration of the correlation and complementarity between traditional features and self-supervised pre-trained features. Most methods still directly use MFCC as the network input, failing to fully extract the spatio-temporal features in speech, ignoring the balance between time-domain and frequency-domain features, and it is difficult to effectively construct a complete emotion representation. In addition, traditional feature extraction networks are usually modeled based on a fixed scale, while emotional information often has a cross-scale distribution, and single-scale modeling is difficult to comprehensively capture emotional features. Existing speech emotion recognition methods still have obvious deficiencies in feature extraction and multi-source information fusion, which limit their application and popularization in complex environments. Emotional information has complex distribution characteristics in speech signals and often exists simultaneously in different time scales and frequency ranges. Traditional feature extraction methods are limited by a fixed receptive field or shallow perception ability and cannot fully capture the deep emotional features hidden in speech, resulting in poor performance of the model in recognizing strong, delicate, or rapidly changing emotional states. Most existing methods only use manually constructed acoustic features such as MFCC, ignoring the complementary relationship with the deep semantic features extracted by self-supervised learning models, resulting in limited expressive and discriminative abilities of the model and difficulty in adapting to diverse and complex emotional expression scenarios. In addition, most current speech emotion recognition models adopt a single-scale, single-path feature extraction network structure, lacking a flexible modeling mechanism for multi-scale features and being difficult to perform stably under different speech rates, speaking styles, and signal-to-noise ratios, with poor generalization performance.
[0003] Therefore, how to improve the accuracy, robustness, and generality of the speech emotion recognition model in multi-datasets and multi-scenarios has become a technical problem to be solved urgently. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art or related technologies, and discloses a speech emotion recognition method and system based on a multi-dimensional information perception strategy, which can extract more comprehensive and prominent emotion information from speech signals, improve the emotion recognition ability of the model, and effectively improve the accuracy, robustness, and generality of the emotion recognition model in multi-datasets and multi-scenarios.
[0005] The first aspect of the present invention discloses a speech emotion recognition method based on a multi-dimensional information perception strategy, including: extracting WavLM features: extracting features of the speech to be recognized through the WavLM-Large model to obtain WavLM features; extracting MFCC features: extracting Mel Frequency Cepstral Coefficients of the speech to be recognized through an audio feature extraction tool to obtain MFCC features; constructing a multi-dimensional information perception and cross-dimensional interleaving module: the multi-dimensional information perception and cross-dimensional interleaving module includes a Transformer layer, an MDIP layer, a linear transformation layer (linear), a CDI layer, and a convolutional layer (conv) connected in sequence; wherein, the Transformer layer processes the WavLM features or MFCC features to generate a frequency feature map and a time feature map; the MDIP layer obtains frequency features through frequency multi-dimensional information perception operations and obtains time features through time multi-dimensional information perception operations; the CDI layer takes the time features and frequency features from the MDIP layer as inputs and enhances the representation capabilities of the time features and frequency features in a feature interleaving manner; processing WavLM features: processing the WavLM features through multiple layers of multi-dimensional information perception and cross-dimensional interleaving modules in sequence to obtain a first feature map; processing MFCC features: processing the MFCC features through multiple layers of multi-dimensional information perception and cross-dimensional interleaving modules in sequence to obtain a second feature map; feature fusion: inputting the first feature map and the second feature map into a SENet for feature fusion, and outputting fused features through fully connected operations and batch normalization operations; emotion classification: classifying the fused features through a classifier to predict the emotion of the speech to be recognized.
[0006] In this technical solution, WavLM-Large is a self-supervised pre-training model based on Transformer proposed by Microsoft, which has powerful context modeling capabilities and the ability to extract deep representations of audio signals. The speech embeddings extracted by it are used for general speech understanding tasks. Mel-Frequency Cepstral Coefficients (MFCC) is a commonly used feature extraction method in audio signal processing. MFCC converts audio signals into a series of cepstral coefficients by simulating the auditory characteristics of the human ear, and these coefficients can capture the important features of the sound. The present invention adopts T×39-dimensional MFCC (composed of its static features T×13 dimensions and its first and second order differences, where T is the number of frames). The MDIP layer can extract more comprehensive and prominent emotional information from speech signals by adaptively perceiving multi-granularity emotional information features in different dimensions of time and frequency, improving the emotion recognition ability of the model. The CDI layer can interactively calibrate time-domain and frequency-domain features, model the complementary relationship between different-dimensional feature information, so as to more richly describe emotional information and enhance the emotional representation ability of features.
[0007] According to the speech emotion recognition method based on the multi-dimensional information perception strategy disclosed by the present invention, preferably, a multi-scale perception attention mechanism is introduced between the Transformer layer and the MDIP layer to model multi-scale time-frequency features respectively, so as to obtain multi-scale time feature maps and frequency feature maps.
[0008] According to the speech emotion recognition method based on the multi-dimensional information perception strategy disclosed by the present invention, preferably, the calculation process of the frequency multi-dimensional information perception operation specifically includes: applying a sliding window on the frequency dimension of each time frame to extract local frequency domain features, so as to select specific frequency points to capture the emotional correlation between different frequency bands in the frequency domain.
[0009] According to the speech emotion recognition method based on the multi-dimensional information perception strategy disclosed by the present invention, preferably, the calculation process of the time multi-dimensional information perception operation specifically includes: applying a sliding window along the time dimension on each frequency point to extract local time features to capture the short-term and long-term dynamic changes in the time domain, where the short-term dynamic change is the instantaneous change of emotion, and the long-term dynamic change refers to the dependence trend between emotion changes.
[0010] According to the speech emotion recognition method based on the multi-dimensional information perception strategy disclosed by the present invention, preferably, the calculation process of the CDI layer specifically includes:
[0011] Information enhancement:
[0012] Receiving the frequency features and time features Generate enhanced features as input:
[0013]
[0014]
[0015]
[0016]
[0017] Among them, IEM(x) represents the information enhancement operation, x represents the input feature, The enhanced feature representing the time feature, The enhanced feature representing the frequency feature, AP() represents the average pooling operation, Contrast represents the contrast enhancement operation, Conv represents the convolution operation, and s represents the Sigmoid activation function;
[0018] Cross-dimensional interleaving:
[0019] First, receive , , and obtain two different attention maps through the spatial attention operation:
[0020]
[0021]
[0022] Among them, W time represents the time feature attention map, W freq represents the frequency feature attention map, and SA represents the spatial attention operation;
[0023] Next, perform cross-dimensional feature complementarity between and :
[0024]
[0025]
[0026] Among them, represents the weighted time feature, represents the weighted frequency feature;
[0027] Subsequently, splice the weighted features and fuse them through two convolutional layers to generate the set weight map W:
[0028]
[0029] Among them, Sigmoid() represents the Sigmoid activation function, conv represents the convolution operation, and ○ represents the concatenation operation by channel;
[0030] Finally, use the set weight map W for and to perform recalibration and fuse through the convolutional layer to generate the output features of multi-dimensional information:
[0031]
[0032] Among them, X MDI represents the multi-dimensional information feature map.
[0033] According to the speech emotion recognition method based on the multi-dimensional information perception strategy disclosed in the present invention, preferably, the Transformer layer is a Vanilla Transformer Encoder.
[0034] The second aspect of the present invention discloses a speech emotion recognition system based on the multi-dimensional information perception strategy, including: a memory for storing program instructions; a processor for calling the program instructions stored in the memory to implement the speech emotion recognition method based on the multi-dimensional information perception strategy as described in any of the above technical solutions.
[0035] Compared with the prior art, the beneficial effects of the present invention at least include: MDIP solves the problem of large distribution differences of emotional information in the frequency domain and time domain in speech signals, and the model can adaptively obtain emotional feature information of different scales. The CDI proposed by the present invention is mainly used to fuse emotional information in different dimensions of the time domain and frequency domain, reduce redundant information in speech emotion embedding, and model the complementary relationship between feature information in different dimensions to enhance the emotional representation ability of speech embedding. In addition, the present invention adopts a two-channel network structure and uses the combination of traditional handcrafted features MFCC and self-supervised pre-trained features WavLM for speech emotion recognition, which has brought certain development to the speech emotion recognition technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Shows a schematic flow chart of a speech emotion recognition method based on the multi-dimensional information perception strategy according to an embodiment of the present invention.
[0037] Figure 2 Shows a schematic block diagram of a speech emotion recognition system based on the multi-dimensional information perception strategy according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] To better understand the above objects, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the present invention is not limited by the limitations of the specific embodiments disclosed below.
[0039] As Figure 1 shown, according to an embodiment of the invention, a speech emotion recognition method based on a multi-dimensional information perception strategy is disclosed, including:
[0040] Step 1, extract WavLM features: Use the WavLM-Large model to extract features from the speech to be recognized to obtain WavLM features;
[0041] Step 2, extract MFCC features: Use an audio feature extraction tool to extract Mel-frequency cepstral coefficients from the speech to be recognized to obtain MFCC features;
[0042] Step 3, construct a multi-dimensional information perception and cross-dimensional interleaving module: The multi-dimensional information perception and cross-dimensional interleaving module includes a Transformer layer, an MDIP layer, a linear transformation layer, a CDI layer, and a convolutional layer connected in sequence; among them, the Transformer layer (Vanilla Transformer Encoder) processes the WavLM features or MFCC features to generate a frequency feature map and a time feature map; the MDIP layer obtains frequency features through frequency multi-dimensional information perception operations and obtains time features through time multi-dimensional information perception operations; the CDI layer takes the time features and frequency features from the MDIP layer as inputs and enhances the representation capabilities of the time features and frequency features in a feature interleaving manner;
[0043] Step 4, process WavLM features: Process the WavLM features through multiple layers of multi-dimensional information perception and cross-dimensional interleaving modules in sequence to obtain a first feature map;
[0044] Step 5, process MFCC features: Process the MFCC features through multiple layers of multi-dimensional information perception and cross-dimensional interleaving modules in sequence to obtain a second feature map;
[0045] Step 6, feature fusion: Input the first feature map and the second feature map into the SENet for feature fusion, and output the fused features through fully connected operations and batch normalization operations;
[0046] Step 7, emotion classification: Classify the fused features through a classifier to predict the emotion of the speech to be recognized.
[0047] Furthermore, a multi-scale perception attention mechanism is introduced between the Transformer layer and the MDIP layer to model the multi-scale time-frequency features respectively, so as to obtain the multi-scale time feature map and frequency feature map.
[0048] Furthermore, the calculation process of the frequency multi-dimensional information perception operation specifically includes: applying a sliding window on the frequency dimension of each time frame to extract local frequency domain features, thereby selecting specific frequency points to capture the emotional correlation between different frequency bands in the frequency domain. The calculation process of the time multi-dimensional information perception operation specifically includes: applying a sliding window along the time dimension at each frequency point to extract local time features to capture the short-term and long-term dynamic changes in the time domain.
[0049] Furthermore, the calculation process of the CDI layer specifically includes:
[0050] Information enhancement:
[0051] Receiving the frequency features and time features generated by the MDIP layer as inputs to generate enhanced features:
[0052]
[0053]
[0054]
[0055]
[0056] Among them, IEM(x) represents the information enhancement operation, x represents the input feature, represents the enhanced feature of the time feature, represents the enhanced feature of the frequency feature, AP() represents the average pooling operation, Contrast represents the contrast enhancement operation, Conv represents the convolution operation, and s represents
[0057] Cross-dimensional interleaving:
[0058] First, receive , , and obtain two different attention maps through the spatial attention operation:
[0059]
[0060]
[0061] Among them, W time represents the time feature attention map, W freq represents the frequency feature attention map, and SA represents the spatial attention operation;
[0062] Next, cross-dimensional feature complementarity is performed between and :
[0063]
[0064]
[0065] Among them, represents the weighted time feature, and represents the weighted frequency feature;
[0066] Subsequently, the weighted features are concatenated and fused through two convolutional layers to generate the set weight map W:
[0067]
[0068] Among them, Sigmoid() represents the Sigmoid activation function, and conv represents the convolution operation; represents the concatenation operation by channel.
[0069] Finally, the set weight map W is used to and for recalibration and fused through a convolutional layer to generate the output features of multi-dimensional information:
[0070]
[0071] Among them, X MDI represents the multi-dimensional information feature map.
[0072] According to another embodiment of the present invention, the specific application process and principle of the above-mentioned speech emotion recognition method based on the multi-dimensional information perception strategy are also disclosed: corresponding to the speech emotion recognition method based on the multi-dimensional information perception strategy disclosed in the above embodiment, a network structure of a speech emotion recognition model based on multi-dimensional information perception and interleaving strategy (named MDIPI-Net by the present invention) is built. In this model, the WavLM feature and the MFCC feature are respectively processed through two branches based on the MDIPI module (Multi-Dimensional Information Perception and Cross-Dimensional Interleaving Module), and then the feature maps of the two branches are fused through SENet (Squeeze-and-Excitation Networks), and after full connection operation and batch normalization operation (FC+BN), they are input into the classifier for emotion prediction. SENet can enhance the perception ability of the convolutional neural network for different features by adaptively adjusting the weights of the feature maps of each channel, thereby improving the classification performance of the model. In the present invention, SENet adaptively assigns weights to different channels through global information, so that the fused features have stronger emotion discrimination ability.
[0073] The operation steps of MDIPI-Net specifically include:
[0074] S1. Data preprocessing:
[0075] The WavLM feature is a deep feature extracted by the self-supervised pre-trained model WavLM-Large. WavLM-Large is a self-supervised pre-trained model based on Transformer proposed by Microsoft, with powerful context modeling capabilities and the ability to extract deep representations of audio signals. The speech embeddings it extracts are used for general speech understanding tasks.
[0076] The MFCC feature is a traditional handcrafted feature Mel-Frequency Cepstral Coefficients (MFCC) extracted using an audio feature extraction tool, adopting T×39-dimensional MFCC (composed of its static feature T×13 dimensions and its first and second order differences, where T is the number of frames).
[0077] S2. The present invention adopts a two-channel network structure to extract emotion features, aiming to extract rich time-domain information and frequency-domain emotion information:
[0078] S21: In the field of speech emotion recognition, emotion information is usually distributed in different time-frequency ranges. Due to problems such as limited receptive fields or vanishing gradients in traditional neural networks, it is difficult to capture long-range dependencies. To enhance the model's perception ability of emotion information, an MDIP module (MDIP layer) based on the dilated convolution strategy is proposed.
[0079] The network structure of MDIP: To extract the frequency-domain information of emotion features, the Frequency Multi-Dimensional Perception (FMDP) extracts local frequency-domain features in a sliding window manner on the frequency dimension of each time frame. By adjusting the dilation rate r f , FMDP sparsely samples the sliding window, thereby selecting specific frequency points to capture the emotional correlation between different frequency bands in the frequency domain. The formal description of FMDP is as follows:
[0080]
[0081] Among them, represents the query at the frequency domain position on, and are respectively the key and value selected from the frequency domain; is the frequency expansion rate for adjusting the sparsity of the selected frequency points. The positions of the selected frequency points within the sliding window are defined as follows:
[0082]
[0083] where represents the size of the frequency domain window.
[0084] Similarly, to extract the emotional features in the time domain, the Temporal Multi-Dimensional Perception (TMDP) applies a sliding window along the time dimension at each frequency point to extract local time features. The dilated sliding window selects time frames according to the expansion rate, aiming to capture the short-term and long-term dynamic changes in the time domain. The formal description of TMDP is as follows:
[0085]
[0086] where represents the query at the time domain position and and are the key and value selected from the time domain respectively; is the time expansion rate for adjusting the sparsity of the selected time frames. The positions of the selected time frames within the sliding window are defined as follows:
[0087]
[0088] where represents the size of the time window.
[0089] To capture multi-scale frequency and time features, the present invention introduces a Multi-Scale Perceptual Attention (MSPA) mechanism on the basis of the extracted source features. MSPA models the multi-scale time-frequency features respectively. Specifically, for a given feature map X, the corresponding query (Q), key (K), and value (V) are obtained through linear projection. Subsequently, the channels are divided into n different heads, and multi-scale information perception (MDP) operations are performed at different expansion rates. The calculation formula of MSPA is as follows:
[0090]
[0091]
[0092] where represents the emotional information feature obtained under the i-th head, Represents the expansion coefficient of the i-th head, h1, h2, …… h n Belongs to h i , Q i , K i , V i Represents the feature map of the i-th head, Concat represents the concatenation operation, and Linear represents the fully connected operation.
[0093] To capture the importance differences of multi-scale feature information, the present invention introduces a weight calculation unit (WeightCalculation Unit, WCU) for calculating the importance of features according to the attention weights obtained from the Transformer. The formal description of the WCU is as follows:
[0094]
[0095]
[0096]
[0097]
[0098] Among them, Represents the attention weight from the i-th head of the MSPA, Represents the weight The element in the r-th row and c-th column, R and C are the row length and height of the matrix aw. Concat represents the concatenation operation, and Softmax represents the Softmax activation function; Represents the feature weight of the i-th head of the frequency feature, , , ……, Belongs to , Represents the feature weight of the i-th head of the time domain feature (temporal feature), , , ……, Belongs to ; Represents the weight coefficient of the frequency feature in the sentiment information, Represents the weight coefficient of the time domain feature (temporal feature) in the sentiment information.
[0099] S22, In the field of SER, the interaction between features of different dimensions plays a vital role. Temporal features contain more dynamic information about how the signal changes over time, while frequency features reflect more characteristics related to frequency distribution. Therefore, in order to more effectively coordinate the properties of features of different dimensions, the present invention proposes a cross-dimensional interweaving (CDI) module. The CDI module (CDI layer) fuses temporal and frequency features and enhances the sentiment discriminative feature representation in the output.
[0100] The CDI module takes the temporal features and frequency features from MDIP as input and enhances the representation capabilities of these two types of features in an interleaved manner. CDI is divided into two stages:
[0101] Information Enhancement Module (IEM);
[0102] Cross-Dimensional Interweaving Unit (CDIU).
[0103] Specifically:
[0104] The IEM module receives the frequency signature generated by the MDIP and time characteristics As input, the mutated edge-enhanced features are generated:
[0105]
[0106]
[0107] Specifically, the fusion of features of different dimensions usually involves upsampling or downsampling operations, and is achieved by element-by-element addition. However, these operations may dilute boundary information in important areas, thereby causing feature redundancy. In speech signals, emotionally strong speech is often accompanied by violent energy fluctuations. By extracting and enhancing these boundary features, the core emotional features can be captured more accurately.
[0108] Therefore, the present invention adopts the IEM module to effectively enhance the boundary area in the feature map while retaining the original feature information. Its formal description is as follows:
[0109]
[0110]
[0111] Where AP() represents the average pooling operation.
[0112] In the CDIU, the input features come from the enhanced feature information of the previous layer. First, two different attention maps are obtained through spatial attention operations:
[0113]
[0114]
[0115] Among them, SA represents the spatial attention mechanism, which is realized through global average pooling and max pooling operations, and then processed by a convolutional layer and a Sigmoid activation function.
[0116] Next, in order to obtain a more comprehensive feature representation, cross-dimensional feature complementarity is performed between the frequency features and the time features, and the weighted features can be obtained:
[0117]
[0118]
[0119] In order to further explore the interaction information between these two types of features, the present invention splices the weighted features and fuses them through two convolutional layers to generate a set weight map , and its calculation method is as follows:
[0120]
[0121] Finally, the set weight map is used to and for recalibration, and they are fused through a convolutional layer to generate output features of multi-dimensional information, as follows:
[0122]
[0123] In this way, features from different dimensions are utilized, and an adaptive attention map is generated by combining the attention mechanism, while calibrating the information of features in different dimensions. In addition, during the process of generating the set weight map, this module also re-optimizes the input features using a convolutional layer to generate more refined and more adaptable output features .
[0124] S3, input the from different source features into the SENet, adaptively assign weights to different channels through global information, so that the fused features have stronger emotion discrimination ability, and input the final emotion features into the classifier to predict the emotion expressed by the current utterance.
[0125] Such as Figure 2As shown, according to another embodiment of the present invention, a voice emotion recognition system 500 based on a multi-dimensional information perception strategy is also disclosed, including: a memory 501 for storing program instructions; a processor 502 for calling the program instructions stored in the memory to implement the voice emotion recognition method based on the multi-dimensional information perception strategy as in the above embodiments.
[0126] All or part of the steps in the various methods of the above embodiments can be completed by a program controlling related hardware. The program can be stored in a readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other readable medium capable of carrying or storing data.
[0127] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A speech emotion recognition method based on a multi-dimensional information perception strategy, characterized in that Including: Extract WavLM features: Use the WavLM-Large model to extract features from the speech to be recognized, and obtain WavLM features; Extract MFCC features: Use an audio feature extraction tool to extract Mel Frequency Cepstral Coefficients from the speech to be recognized, and obtain MFCC features; Construct a multi-dimensional information perception and cross-dimensional interleaving module: The multi-dimensional information perception and cross-dimensional interleaving module includes a Transformer layer, an MDIP layer, a linear transformation layer, a CDI layer, and a convolutional layer connected in sequence; among them, the Transformer layer processes the WavLM features or the MFCC features to generate a frequency feature map and a time feature map; the MDIP layer obtains frequency features through frequency multi-dimensional information perception operations and obtains time features through time multi-dimensional information perception operations; the CDI layer uses the time features and frequency features from the MDIP layer as inputs to enhance the representation capabilities of the time features and frequency features in a feature interleaving manner; Process WavLM features: Process the WavLM features through multiple layers of the multi-dimensional information perception and cross-dimensional interleaving module in sequence to obtain a first feature map; Process MFCC features: Process the MFCC features through multiple layers of the multi-dimensional information perception and cross-dimensional interleaving module in sequence to obtain a second feature map; Feature fusion: Input the first feature map and the second feature map into the SENet for feature fusion, and output the fused features through fully connected operations and batch normalization operations; Emotion classification: Classify the fused features through a classifier to predict the emotion of the speech to be recognized.
2. The speech emotion recognition method based on the multi-dimensional information perception strategy according to claim 1, wherein, Introduce a multi-scale perception attention mechanism between the Transformer layer and the MDIP layer to model the multi-scale time-frequency features respectively to obtain multi-scale time feature maps and frequency feature maps.
3. The speech emotion recognition method based on the multi-dimensional information perception strategy according to claim 1, wherein The calculation process of the frequency multi-dimensional information perception operation specifically includes: Apply a sliding window on the frequency dimension of each time frame to extract local frequency domain features, thereby selecting specific frequency points to capture the emotional correlation between different frequency bands in the frequency domain.
4. The speech emotion recognition method based on the multi-dimensional information perception strategy according to claim 1, wherein, The calculation process of the time multi-dimensional information perception operation specifically includes: Apply a sliding window along the time dimension on each frequency point to extract local time features to capture the instantaneous changes of emotions in the time domain and the dependence trend between emotional changes.
5. The speech emotion recognition method based on the multi-dimensional information perception strategy according to claim 1, wherein, The calculation process of the CDI layer specifically includes: Information enhancement: Receive the frequency features generated by the MDIP layer and time features as inputs to generate enhanced features: Among them, IEM(x) represents an information enhancement operation, and x represents the input feature. represents the enhanced feature of the time feature. represents the enhanced feature of the frequency feature. AP() represents the average pooling operation, Contrast represents the contrast enhancement operation, Conv represents the convolution operation, and s represents the Sigmoid activation function. Cross-dimensional interleaving: First, receive , , and obtain two different attention maps through spatial attention operations: Among them, W time represents the temporal feature attention map, W freq represents the frequency feature attention map, and SA represents the spatial attention operation; Next, perform cross-dimensional feature complementarity between and : Among them, represents the weighted time feature, represents the weighted frequency feature; Subsequently, splice the weighted features and fuse them through two convolutional layers to generate a set weight map W: Among them, Sigmoid() represents the Sigmoid activation function, conv represents the convolutional operation, and ○ represents the splicing operation by channel; Finally, use the set weight graph W to and perform re-calibration and fuse through the convolutional layer to generate the output features of multi-dimensional information: Among them, X MDI represents a multi-dimensional information feature map.
6. The method for speech emotion recognition based on a multi-dimensional information perception strategy according to claim 1, characterized in that The Transformer layer is a Vanilla Transformer Encoder.
7. A speech emotion recognition system based on a multi-dimensional information perception strategy, characterized in that, Including: A memory for storing program instructions; A processor for calling the program instructions stored in the memory to implement the speech emotion recognition method based on the multi-dimensional information perception strategy according to any one of claims 1 to 6.
Citation Information
Patent Citations
Lightweight speech emotion recognition method and system based on multi-scale attention
CN117711443A
Remote emotion recognition method based on multiple modes
CN118279805A
Speech emotion recognition method and system based on multi-feature attention fusion
CN118447880A
Multi-task speech emotion recognition method and device and storage medium
CN118553271A
KR20240072000A