Multi-dimensional speech emotion recognition method and system based on adaptive cross attention

Through the multidimensional speech emotion recognition method with adaptive cross attention, the spectrum map, Mel frequency cepspectral coefficient spectrum and original audio information are combined, and the problem of insufficient single-dimensional feature extraction in the existing technology is solved, achieving a more efficient speech emotion recognition effect.

CN120340464APending Publication Date: 2025-07-18CHINESE PEOPLES LIBERATION ARMY ARMY BORDER & COASTAL DEFENSE ACAD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510538768.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing speech emotion recognition technology mainly focuses on single-dimensional information extraction, ignoring the multi-dimensional characteristics of speech signals, resulting in insufficient understanding of speech data by the model and being unable to fully capture emotional information. Moreover, interrelationship and independence in the fusion of multi-dimensional features are ignored, leading to conflict or redundant information problems.

Method used

The multidimensional speech emotion recognition method based on adaptive cross attention is adopted. Through the fusion of three types of features, spectrogram, Mel frequency cepspectral coefficient spectrum and original audio information, the adaptive cross attention mechanism is used to fusion of features, the attention weight between features is calculated, key information is highlighted and redundant information is suppressed, and multidimensional speech emotion characteristics are generated.

Benefits of technology

It significantly improves the model's ability to understand speech data, improves feature utilization rate and emotional recognition accuracy, enhances the adaptability and robustness of the model, optimizes the calculation process, and avoids conflicts and redundancy between features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340464A_ABST
    Figure CN120340464A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-dimensional speech emotion recognition method and system based on adaptive cross attention, and the method comprises the steps: obtaining a plurality of pieces of original speech data and corresponding labels, extracting a spectrogram image and a Mel-frequency cepstral coefficient spectrum corresponding to each piece of original speech data, and combining the spectrogram image and the Mel-frequency cepstral coefficient spectrum with the corresponding labels to obtain a training data set; constructing a multi-dimensional speech emotion recognition model based on adaptive cross attention; inputting the training data set into a feature extraction layer to extract acoustic features, and generating multi-dimensional speech emotion features through a feature fusion layer; and inputting the multi-dimensional speech emotion features into an emotion classifier, and constructing a cross entropy loss function to train a multi-dimensional speech emotion recognition model based on adaptive cross attention. The emotion information in the voice signal can be comprehensively captured, the understanding ability of the model on the voice data is improved, and the feature utilization rate is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Voice emotion recognition is an important technical aspect in human-computer interaction, aiming to enable machines to accurately capture and understand human emotions, so as to respond in a more natural and context-appropriate way. This technology helps machines obtain emotion information by analyzing the sound features (such as intonation, frequency, rhythm, etc.) and speech content in the voice signal.

[0003] Voice emotion recognition has a wide range of applications in multiple fields. For example, in telephone customer service, voice emotion recognition technology can monitor customer emotions in real time, help enterprises judge customer satisfaction, quickly respond to potential problems, and improve service quality. In the fields of education and healthcare, voice emotion recognition can also be used to identify the psychological states of students or patients, providing a reference for personalized tutoring or diagnosis and treatment. With the introduction of deep learning and pre-trained models, the performance of voice emotion recognition has been continuously improved, providing important support for building intelligent and emotional interaction systems.

[0004] In the prior art, when performing deep learning for voice emotion recognition, it is usually necessary to perform original voice feature extraction and extraction feature fusion; among them, original voice feature extraction usually adopts methods such as spectrogram, Mel frequency, cepstral coefficient spectrum, and pre-trained models. However, these methods mainly focus on the extraction of single-dimensional information, only paying attention to frequency domain features or time domain features, ignoring the multi-dimensional characteristics in the voice signal, resulting in insufficient understanding of voice data by the model, being unable to comprehensively capture the emotion information in the voice signal, and there are significant differences in the performance of different feature extraction methods in the voice emotion recognition task, making it impossible to efficiently select and combine feature extraction methods. Generally speaking, the single-dimensional feature extraction method limits the performance of the voice emotion recognition model and fails to fully exploit the rich information in the voice signal; while extraction feature fusion usually needs to introduce multi-dimensional feature fusion to make up for the deficiencies of single-dimensional methods. However, the mutual correlation and independence between different features are often ignored during the fusion process, resulting in problems such as conflicts or redundant information between features, and moreover, most of the existing fusion methods adopt simple splicing or weighted average methods, and this processing method cannot fully capture the deep interaction information of multi-dimensional features.

[0005] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.

[0006] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0007] The purpose of the embodiments of the present disclosure is to provide a multi-dimensional speech emotion recognition method and system based on adaptive cross-attention, thereby at least to some extent overcoming one or more problems caused by the limitations and defects of related technologies.

[0008] In a first aspect, the present application provides a multi-dimensional speech emotion recognition method based on adaptive cross-attention, including:

[0009] Obtain multiple pieces of original speech data and corresponding labels, extract the spectrogram images and Mel-frequency cepstral coefficient spectra corresponding to each piece of the original speech data, and combine each piece of the original speech data, the spectrogram images, and the Mel-frequency cepstral coefficient spectra with the corresponding labels to obtain a training data set;

[0010] Construct a multi-dimensional speech emotion recognition model based on adaptive cross-attention, where the multi-dimensional speech emotion recognition model based on adaptive cross-attention includes: an input layer, a feature extraction layer, a feature fusion layer, and an emotion classification layer;

[0011] Input the training data set into the feature extraction layer to extract acoustic features, and generate multi-dimensional speech emotion features through the feature fusion layer;

[0012] Input the multi-dimensional speech emotion features into an emotion classifier, and construct a cross-entropy loss function to train the multi-dimensional speech emotion recognition model based on adaptive cross-attention.

[0013] In a possible implementation manner, the step of obtaining multiple pieces of original speech data and corresponding labels, extracting the spectrogram images and Mel-frequency cepstral coefficient spectra corresponding to each piece of the original speech data, and combining each piece of the original speech data, the spectrogram images, and the Mel-frequency cepstral coefficient spectra with the corresponding labels to obtain a training data set includes:

[0014] Perform discrete Fourier transform on the original speech data, select the discrete Fourier transform points as spectrogram features, and obtain the spectrogram images corresponding to the original speech data;

[0015] Extract the original speech data through the Librosa library to obtain the Mel-frequency cepstral coefficient spectra corresponding to the original speech data;

[0016] Combine the original speech data, the spectrogram images, and the Mel-frequency cepstral coefficient spectra with the corresponding labels.

[0017] In a possible implementation manner, the feature extraction layer includes an encoder and a pre-trained speech model; the feature fusion layer includes an adaptive cross-attention mechanism; the emotion classification layer includes an emotion classifier.

[0018] In a possible implementation, the step of inputting the training data set into the feature extraction layer to extract acoustic features and generating multi-dimensional speech emotion features through the feature fusion layer includes:

[0019] Inputting the training data set into the encoder and the pre-trained speech model to obtain spectrogram acoustic features, Mel-frequency cepstral coefficient spectrogram acoustic features, and time-domain speech deep features;

[0020] Generating multi-dimensional speech emotion features according to the spectrogram acoustic features, the Mel-frequency cepstral coefficient spectrogram acoustic features, and the time-domain speech deep features based on the adaptive cross-attention mechanism.

[0021] In a possible implementation, the encoder includes a ResNet50 encoder and an RNN encoder.

[0022] In a possible implementation, the step of inputting the training data set into the encoder and the pre-trained speech model to obtain spectrogram acoustic features, Mel-frequency cepstral coefficient spectrogram acoustic features, and time-domain speech deep features includes:

[0023] Inputting the spectrogram image into the ResNet50 encoder to obtain the spectrogram acoustic features corresponding to the spectrogram image;

[0024] Inputting the Mel-frequency cepstral coefficient spectrogram into the RNN encoder module to obtain the Mel-frequency cepstral coefficient spectrogram acoustic features corresponding to the Mel-frequency cepstral coefficient spectrogram;

[0025] Inputting the original speech data into the WavLM pre-trained speech model to obtain the time-domain speech deep features corresponding to the original speech data.

[0026] In a possible implementation, the step of generating multi-dimensional speech emotion features according to the spectrogram acoustic features, the Mel-frequency cepstral coefficient spectrogram acoustic features, and the time-domain speech deep features based on the adaptive cross-attention mechanism includes:

[0027] Extracting two-dimensional speech emotion features according to the spectrogram acoustic features and the Mel-frequency cepstral coefficient acoustic features through the adaptive cross-attention mechanism;

[0028] Generating multi-dimensional speech emotion features according to the two-dimensional speech emotion features and the time-domain speech deep features through the adaptive cross-attention mechanism.

[0029] In a possible implementation, the step of inputting the multi-dimensional speech emotion features into the emotion classifier and constructing a cross-entropy loss function to train the multi-dimensional speech emotion recognition model based on the adaptive cross-attention includes:

[0030] Input the multi-dimensional speech emotion features into an emotion classifier to predict the emotion label corresponding to the original speech data;

[0031] Construct a cross-entropy loss function based on the emotion label, and backpropagate to adjust the trainable parameters of the multi-dimensional speech emotion recognition model based on adaptive cross-attention to train the multi-dimensional speech emotion recognition model based on adaptive cross-attention.

[0032] In a second aspect, the present application provides a multi-dimensional speech emotion recognition system based on adaptive cross-attention. The system is used to execute the above-mentioned multi-dimensional speech emotion recognition method based on adaptive cross-attention. The system includes:

[0033] A data acquisition module, configured to acquire multiple pieces of original speech data and corresponding labels, extract the spectrogram image and mel-frequency cepstrum coefficients spectrum corresponding to each piece of the original speech data, and combine each piece of the original speech data, the spectrogram image and the mel-frequency cepstrum coefficients spectrum with the corresponding label to obtain a training dataset;

[0034] A model construction module, configured to construct a multi-dimensional speech emotion recognition model based on adaptive cross-attention. The multi-dimensional speech emotion recognition model based on adaptive cross-attention includes: an input layer, a feature extraction layer, a feature fusion layer, and an emotion classification layer;

[0035] A feature generation module, configured to input the training dataset into the feature extraction layer to extract acoustic features, and generate multi-dimensional speech emotion features through the feature fusion layer;

[0036] A model training module, configured to input the multi-dimensional speech emotion features into an emotion classifier, and construct a cross-entropy loss function to train the multi-dimensional speech emotion recognition model based on adaptive cross-attention.

[0037] In a possible implementation manner, the multi-dimensional speech emotion recognition model based on adaptive cross-attention includes: a ResNet50 encoder module, configured to extract spectrogram acoustic features in the spectrogram image;

[0038] An RNN encoder module, configured to extract mel-frequency cepstrum coefficient acoustic features in the mel-frequency cepstrum coefficients spectrum;

[0039] A WavLM pre-trained speech model, configured to extract deep time-domain speech features in the original speech data;

[0040] An adaptive cross-attention mechanism module is used to integrate the speech emotion information between the spectrogram acoustic features and the mel-frequency cepstral coefficient acoustic features to obtain two-dimensional speech emotion features, and is also used to integrate the speech emotion information between the two-dimensional speech emotion features and the time-domain speech deep features to capture the internal correlation between the speech signals in the frequency domain and the time domain, and obtain multi-dimensional speech emotion features;

[0041] An emotion classifier module is used to predict the emotion label corresponding to the original speech data according to the multi-dimensional speech emotion features.

[0042] The technical solution provided by this application may include the following beneficial effects:

[0043] Through the multi-dimensional speech emotion recognition method and system based on adaptive cross-attention of this application, it is possible to fully mine the multi-dimensional information of speech data by fusing three types of features: spectrogram, mel-frequency cepstral coefficient spectrum, and original audio information, extract features of different dimensions respectively, comprehensively capture the emotion information in the speech signal, improve the model's understanding ability of speech data, and significantly improve the feature utilization rate.

[0044] At the same time, an adaptive cross-attention mechanism is used for feature fusion. The cross-attention model can calculate the attention weights between the input features, highlight the key information and suppress the redundant information; when fusing the spectrogram acoustic features and the mel-frequency cepstral coefficient acoustic features, as well as the two-dimensional speech emotion features and the time-domain speech deep features, it fully mines the complementarity between the features, effectively captures the deep interaction information of multi-dimensional features, avoids the conflict and redundancy between the features; optimizes the calculation process, improves the training efficiency, makes the model more efficient in multi-dimensional feature fusion, enhances the adaptability and robustness of the model to real speech scenarios, and realizes accurate speech emotion recognition.

[0045] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0047] Figure 1 Shows a flowchart of the multi-dimensional speech emotion recognition method based on adaptive cross-attention in an exemplary embodiment of the present disclosure;

[0048] Figure 2Shows a detailed flowchart of step S100 of the multi-dimensional speech emotion recognition method based on adaptive cross-attention in an exemplary embodiment of the present disclosure;

[0049] Figure 3 Shows a detailed flowchart of step S300 of the multi-dimensional speech emotion recognition method based on adaptive cross-attention in an exemplary embodiment of the present disclosure;

[0050] Figure 4 Shows a detailed flowchart of step S310 of the multi-dimensional speech emotion recognition method based on adaptive cross-attention in an exemplary embodiment of the present disclosure;

[0051] Figure 5 Shows a detailed flowchart of step S320 of the multi-dimensional speech emotion recognition method based on adaptive cross-attention in an exemplary embodiment of the present disclosure;

[0052] Figure 6 Shows a detailed flowchart of step S400 of the multi-dimensional speech emotion recognition method based on adaptive cross-attention in an exemplary embodiment of the present disclosure;

[0053] Figure 7 Shows the structural diagram of the multi-dimensional speech emotion recognition model based on adaptive cross-attention in an exemplary embodiment of the present disclosure;

[0054] Figure 8 Shows the structural schematic diagram of the multi-dimensional speech emotion recognition system based on adaptive cross-attention in an exemplary embodiment of the present disclosure. Detailed implementation manners

[0055] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0056] In addition, the drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0057] In this exemplary embodiment, a multi-dimensional speech emotion recognition method based on adaptive cross-attention is first provided. This method can be applied to a terminal device, such as a mobile terminal like a mobile phone, a desktop computer, a personal digital assistant, a laptop computer, a tablet computer, a smart watch, etc.

[0058] Referring Figure 1 as shown in, the method may include the following steps:

[0059] Step S100: Obtain multiple pieces of original speech data and corresponding labels, extract the spectrogram image and Mel-frequency cepstral coefficient spectrum corresponding to each piece of the original speech data, and combine each piece of the original speech data, the spectrogram image, and the Mel-frequency cepstral coefficient spectrum with the corresponding label to obtain a training data set.

[0060] Step S200: Construct a multi-dimensional speech emotion recognition model based on adaptive cross-attention. The multi-dimensional speech emotion recognition model based on adaptive cross-attention includes: an input layer, a feature extraction layer, a feature fusion layer, and an emotion classification layer.

[0061] Step S300: Input the training data set into the feature extraction layer to extract acoustic features, and generate multi-dimensional speech emotion features through the feature fusion layer.

[0062] Step S400: Input the multi-dimensional speech emotion features into an emotion classifier, and construct a cross-entropy loss function to train the multi-dimensional speech emotion recognition model based on adaptive cross-attention.

[0063] The above method can change the limitations of previous single-dimensional feature extraction by fully mining and fusing three types of features in speech data, namely spectrogram, Mel-frequency cepstral coefficient spectrum, and original audio information, comprehensively capture the emotional information in the speech signal, improve the model's understanding of speech data, and thus enhance the feature utilization rate; through the adaptive cross-attention mechanism, it can accurately capture the complementary information between multi-dimensional features, highlight key information and suppress redundant information during the feature fusion process, extract more discriminative emotional features, improve the emotion recognition accuracy of the model, and thus enhance the emotion recognition ability; by making the generated emotion labels consistent with the speech context, the model has better adaptability to real speech scenarios, can stably and accurately recognize speech emotions in different speech environments and application scenarios, has stronger robustness, and thus improves the adaptability and robustness; it overcomes the deficiencies in the design of existing multi-dimensional feature fusion technologies, effectively processes the mutual correlation and independence problems between different features, avoids conflicts or redundant information between features, fully captures the deep interaction information of multi-dimensional features, and at the same time optimizes the computational complexity and training efficiency, thus optimizing the feature fusion method.

[0064] Next, referring Figures 1 to 8A more detailed description of each step of the above method in this exemplary embodiment is given.

[0065] In step S100, multiple pieces of original speech data and corresponding labels are obtained, the spectrogram images and Mel-frequency cepstral coefficient spectra corresponding to each piece of the original speech data are extracted, and each piece of the original speech data, the spectrogram images, and the Mel-frequency cepstral coefficient spectra are combined with the corresponding labels to obtain a training data set.

[0066] It should be noted that, optionally, multiple pieces of original speech data S = {s1, s2, …, s n} and their corresponding labels are obtained, and for each piece of speech data s i multiple methods are used to extract the corresponding spectrogram image spec i ∈R C×H×W , where C represents the number of channels of the spectrogram image, H represents the height of the spectrogram image, W represents the width of the spectrogram image, and the Mel-frequency cepstral coefficient spectrum mfcc i ∈R L×C , where L represents the number of frames of the audio segment, and C represents the feature dimension of each time frame.

[0067] In one embodiment, as Figure 2 shown, step S100 may include the following sub-steps:

[0068] In step S110, the original speech data is subjected to a discrete Fourier transform, and the discrete Fourier transform points are selected as spectrogram features to obtain the spectrogram image corresponding to the original speech data.

[0069] It should be noted that specifically, first, the original speech data s i is subjected to a discrete Fourier transform, and spectrogram features are obtained by selecting discrete Fourier transform points, thereby generating the corresponding spectrogram image spec i .

[0070] In step S120, the original speech data is extracted through the Librosa library to obtain the Mel-frequency cepstral coefficient spectrum corresponding to the original speech data.

[0071] It should be noted that specifically, the Mel-frequency cepstral coefficient spectrum is extracted using the Librosa library to obtain the Mel-frequency cepstral coefficient spectrum mfcc i corresponding to the original speech data.

[0072] In step S130, the original speech data, the spectrogram image, and the Mel-frequency cepstral coefficient spectrum are combined with the corresponding labels.

[0073] It should be noted that specifically, the extracted spectrogram image spec i , Mel-frequency cepstral coefficient spectrum mfcc i and the original speech data s i are used as multi-dimensional speech information inputs, and together with their corresponding labels, a high-quality training dataset is constructed to provide rich multi-dimensional speech information for subsequent model training. Expressed by the formula as:

[0074] spec i = DFT(s i ) ∈ R C×H×W ;

[0075] mfcc i = Librosa(s i ) ∈ R L×C ;

[0076] Finally, the multi-dimensional speech information is combined with the labels to form the input dataset of the model.

[0077] In step S200, a multi-dimensional speech emotion recognition model based on adaptive cross-attention is constructed. The multi-dimensional speech emotion recognition model based on adaptive cross-attention includes: an input layer, a feature extraction layer, a feature fusion layer, and an emotion classification layer.

[0078] Furthermore, the feature extraction layer includes an encoder and a pre-trained speech model; the feature fusion layer includes an adaptive cross-attention mechanism; the emotion classification layer includes an emotion classifier.

[0079] In step S300, the training dataset is input into the feature extraction layer to extract acoustic features, and multi-dimensional speech emotion features are generated through the feature fusion layer.

[0080] In one embodiment, as Figure 2 shown, step S300 may include the following sub-steps:

[0081] In step S310, the training dataset is input into the encoder and the pre-trained speech model to obtain spectrogram acoustic features, Mel-frequency cepstral coefficient spectrum acoustic features, and time-domain speech deep features.

[0082] In step S320, through the spectrogram acoustic features, the Mel-frequency cepstral coefficient spectrum acoustic features, and the time-domain speech deep features, multi-dimensional speech emotion features are generated according to the adaptive cross-attention mechanism.

[0083] It should be noted that the encoder includes a ResNet50 encoder and an RNN encoder.

[0084] In one embodiment, as Figure 1As shown, step S310 may include the following sub-steps:

[0085] In step S311, the spectrogram image is input into a ResNet50 encoder to obtain the spectrogram acoustic features corresponding to the spectrogram image.

[0086] It should be noted that the ResNet50 encoder is a deep convolutional neural network consisting of 50 layers. It adopts a residual learning mechanism and effectively alleviates the gradient vanishing problem in deep networks by introducing jump connections. Here, the ResNet50 model is used as a spectrogram image encoder to extract spectrogram acoustic features from the original speech data.

[0087] The spectrogram image spec i Input to the ResNet50 encoder module to extract the spectrogram acoustic features. The ResNet50 model acts as an encoder for the spectrogram image and is responsible for extracting the spectrogram acoustic features from the raw speech data. Specifically:

[0088] feaspec i =Linear(Flatten(ResNet50(spec i )))∈R L′×C′

[0089] Among them, Flatten(·) represents the flattening operation, Linear(·) represents the linear layer, L′ represents the number of time frames of the spectrogram acoustic feature, and C′ represents the number of channels of the spectrogram acoustic feature.

[0090] In step S312, the Mel-frequency cepstral coefficient spectrum is input into an RNN encoder module to obtain the acoustic features of the Mel-frequency cepstral coefficient spectrum corresponding to the Mel-frequency cepstral coefficient spectrum.

[0091] It should be noted that the RNN encoder is a neural network that is good at processing sequence data and captures contextual information by transferring hidden states between time steps. Here, the RNN model is used as a Mel-frequency cepstral coefficient spectrum encoder to extract the Mel-frequency cepstral coefficient spectrum acoustic features from the original speech data.

[0092] Mel frequency cepstrum coefficient spectrum mfcc i Input into the RNN encoder module to extract the acoustic features of the Mel-frequency cepstral coefficient spectrum corresponding to the Mel-frequency cepstral coefficient spectrum. The RNN model, as a Mel-frequency cepstral coefficient spectrum encoder, extracts the deep acoustic features of the Mel-frequency cepstral coefficient spectrum from the original speech data, thereby providing rich time dimension features for subsequent speech emotion analysis. The process can be expressed as:

[0093] feamfcc i= RNN(mfcc i ) ∈ R L″×C″

[0094] Among them, L″ represents the number of time frames of the mel-frequency cepstral coefficient spectrum acoustic feature, and C″ represents the number of channels of the mel-frequency cepstral coefficient spectrum acoustic feature. In this way, the RNN encoder can extract the temporal features of the mel-frequency cepstral coefficient spectrum, providing important information support for subsequent emotion recognition tasks.

[0095] In step S313, the original speech data is input into the WavLM pre-trained speech model to obtain the deep temporal speech features corresponding to the original speech data.

[0096] It should be noted that the WavLM pre-trained speech model is a pre-trained speech model based on the Transformer architecture. It is trained on a large-scale speech data through self-supervised learning and can effectively extract the deep features of temporal speech. Here, the WavLM pre-trained speech model is used as the original speech file encoder to be responsible for extracting the deep temporal speech features from the original speech data.

[0097] Input the original speech data s i into the WavLM pre-trained speech model to extract the deep temporal speech features corresponding to the original speech data. As the original speech file encoder, the WavLM model extracts the temporal speech features through the processing of the original speech data, providing strong support for subsequent speech emotion recognition tasks. This process can be represented by the following formula:

[0098] featim i = WavLM(s i ) ∈ R L″′×C″′

[0099] Among them, L″′ represents the number of time frames of the deep temporal speech features, and C″′ represents the number of channels of the deep temporal speech features. In this way, the WavLM pre-trained speech model can effectively extract the deep temporal features from the original speech data, providing key speech information support.

[0100] In one embodiment, as Figure 2 shown, step S320 may include the following sub-steps:

[0101] In step S321, according to the spectrogram acoustic feature and the mel-frequency cepstral coefficient acoustic feature, two-dimensional speech emotion features are extracted through an adaptive cross-attention mechanism.

[0102] It should be noted that the cross-attention model is a deep learning model that can effectively capture the relationships between different features. By calculating the attention weights between input features, it highlights key information and suppresses redundant information. In this method, the cross-attention model is used as a feature fusion module, which is responsible for integrating the speech emotion information of spectrogram acoustic features and mel-frequency cepstral coefficient acoustic features, ensuring that the model can fully exploit the complementarity of these two types of features, thereby extracting richer and more discriminative two-dimensional speech emotion features.

[0103] Input the spectrogram acoustic feature feaspec i and the mel-frequency cepstral coefficient acoustic feature feamfcc i into the adaptive cross-attention mechanism module to extract two-dimensional speech emotion features. Specifically, in this implementation step, feaspec i is used as the source input, and feamfcc i is used as the target input. Linear transformations are respectively performed on the source and target inputs to generate the query vector Q i , the key vector K i , and the value vector V i :

[0104]

[0105] where is the learned weight matrix, d K is the dimension of the key and the query, and d V is the dimension of the value.

[0106] Then, the obtained query vector Q i , the key vector K i , and the value vector V i are input into the adaptive cross-attention mechanism module to obtain two-dimensional speech emotion features:

[0107]

[0108] where sigmoid(·) represents the sigmoid activation function.

[0109] In step S322, according to the two-dimensional speech emotion features and the time-domain speech deep features, multi-dimensional speech emotion features are generated through the adaptive cross-attention mechanism.

[0110] It should be noted that the cross-attention model is used as an advanced feature fusion module to integrate the two-dimensional speech emotion features (formed by the fusion of spectrogram acoustic features and Mel-frequency cepstral coefficient acoustic features) with the deep features of time-domain speech. Through this feature fusion, the model can not only further capture the intrinsic correlation between speech signals in the frequency domain and time domain, but also use the adaptive cross-attention mechanism to mine the potential emotional expression patterns in multidimensional features, thereby generating multidimensional speech emotion features with higher semantic information density and emotional discrimination.

[0111] The two-dimensional speech emotion feature feadouble j and deep features of speech in time domain j Input the adaptive cross attention mechanism module to generate multi-dimensional speech emotion features. Specifically, this implementation step will feadouble i As source input, featim i As the target input, the source and target inputs are linearly transformed to generate the query vector Q i , the key vector K i , value vector V i :

[0112]

[0113] in, is the learned weight matrix, d K is the dimension of the key and query, d V is the dimension of the value. Then the query vector Q i , the key vector K i , value vector V i Input into the adaptive cross attention mechanism module to obtain two-dimensional speech emotion features:

[0114]

[0115] In step S400, the multidimensional speech emotion feature is input into the emotion classifier, and a cross entropy loss function is constructed to train the multidimensional speech emotion recognition model based on adaptive cross attention.

[0116] In one embodiment, Figure 2 As shown, step S400 may include the following sub-steps:

[0117] In step S410, the multi-dimensional speech emotion feature is input into an emotion classifier to predict the emotion label corresponding to the original speech data.

[0118] It should be noted that the obtained multi-dimensional speech emotion features are input into the emotion classifier module to predict the emotion label corresponding to the original speech data;

[0119] Emo i = EmoClassifier(feacomplex i )

[0120] Among them, EmoCassifier(·) represents the emotion classification module, and there is:

[0121] EmoClassifier(feacomplex i ) = MLP(Avgpool(feacomplex i ))

[0122] Among them, Avgpool(·) represents the average pooling operation, and MLP(·) represents the multi-layer perceptron.

[0123] In step S420, according to the emotion label, construct a cross-entropy loss function, and backpropagate to adjust the trainable parameters of the multi-dimensional speech emotion recognition model based on adaptive cross-attention, and train the multi-dimensional speech emotion recognition model based on adaptive cross-attention.

[0124] It should be noted that by predicting the difference between the label (i.e., the emotion label Emo i ) and the true label (i.e., the corresponding label), adjust the model parameters to make the model output as close as possible to the true result, and ensure that the model performs well on both the training set and unseen test data.

[0125] It can be understood that the cross-entropy loss function can be:

[0126]

[0127] Among them, N is the batch size, C is the total number of emotion categories, y i,c is the true label, is the probability value predicted by the model. The cross-entropy loss function is applicable to classification tasks and can effectively measure the difference between the predicted probability distribution and the true distribution.

[0128] The method of backpropagating to adjust the parameters can be: starting from the loss function through the chain rule, reversely calculate the gradient of each layer of parameters (the gradient represents the influence degree of the parameters on the loss), and use an optimizer (such as Adam, SGD) to adjust the trainable parameters of the model (such as the weights of ResNet50, the hidden state weights of RNN, the weight matrix of adaptive cross-attention, etc.) according to the gradient, so that the loss gradually decreases.

[0129] Furthermore, in this exemplary embodiment, a multi-dimensional speech emotion recognition system based on adaptive cross-attention is also provided. Referring to Figure 8 as shown, the system may include:

[0130] A data acquisition module, configured to acquire multiple pieces of original speech data and corresponding labels, extract the spectrogram images and Mel frequency cepstral coefficient spectra corresponding to each piece of the original speech data, and combine each piece of the original speech data, the spectrogram images and the Mel frequency cepstral coefficient spectra with the corresponding labels to obtain a training data set;

[0131] A model construction module, configured to construct a multi-dimensional speech emotion recognition model based on adaptive cross-attention. The multi-dimensional speech emotion recognition model based on adaptive cross-attention includes: an input layer, a feature extraction layer, a feature fusion layer and an emotion classification layer;

[0132] A feature generation module, configured to input the training data set into the feature extraction layer to extract acoustic features, and generate multi-dimensional speech emotion features through the feature fusion layer;

[0133] A model training module, configured to input the multi-dimensional speech emotion features into an emotion classifier, and construct a cross-entropy loss function to train the multi-dimensional speech emotion recognition model based on adaptive cross-attention.

[0134] In one embodiment, the multi-dimensional speech emotion recognition model based on adaptive cross-attention includes: a ResNet50 encoder module, configured to extract spectrogram acoustic features in the spectrogram images; an RNN encoder module, configured to extract Mel frequency cepstral coefficient acoustic features in the Mel frequency cepstral coefficient spectra; a WavLM pre-trained speech model, configured to extract deep time-domain speech features in the original speech data; an adaptive cross-attention mechanism module, configured to integrate the speech emotion information between the spectrogram acoustic features and the Mel frequency cepstral coefficient acoustic features to obtain two-dimensional speech emotion features, and to integrate the speech emotion information between the two-dimensional speech emotion features and the deep time-domain speech features to capture the internal correlation between the speech signal in the frequency domain and the time domain to obtain multi-dimensional speech emotion features; an emotion classifier module, configured to predict the emotion label corresponding to the original speech data according to the multi-dimensional speech emotion features.

[0135] Regarding the device in the above embodiment, the specific manners in which each module performs operations have been described in detail in the embodiment related to the method, and will not be elaborated herein.

[0136] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units. A component shown as a module or unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed over multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present disclosure. A person of ordinary skill in the art can understand and implement it without creative work.

[0137] In an exemplary embodiment of the present disclosure, an electronic device is further provided. The electronic device may include a processor and a memory for storing executable instructions of the processor. Wherein, the processor is configured to execute the steps of the multi-dimensional speech emotion recognition method based on adaptive cross-attention in any one of the above embodiments by executing the executable instructions.

[0138] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module" or "system" here.

[0139] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the multi-dimensional speech emotion recognition method based on adaptive cross-attention in the above embodiments of the present disclosure.

[0140] In an exemplary embodiment of the present disclosure, a computer storage medium is further provided, on which a computer program is stored. When the program is executed by, for example, a processor, the steps of the multi-dimensional speech emotion recognition method based on adaptive cross-attention in any one of the above embodiments can be implemented.

[0141] In some possible embodiments, aspects of the present invention can also be implemented in the form of a computer program product, which includes a computer program or instructions. When the computer program product runs on a terminal device, the computer program code or instructions are used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above-mentioned part of the multi-dimensional voice emotion recognition method based on adaptive cross-attention in this specification.

[0142] The above program product can be written in any combination of one or more programming languages for executing the program code of the operations of the present invention. The programming languages include object-oriented programming languages - such as Java, C++, etc., and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0143] The computer software product can be stored in a computer storage medium, which includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.

[0144] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

Claims

1. A multi-dimensional speech emotion recognition method based on adaptive cross-attention, characterized in that, Including: Obtain multiple pieces of original speech data and corresponding labels, extract the spectrogram images and Mel-frequency cepstral coefficient spectra corresponding to each piece of the original speech data, and combine each piece of the original speech data, the spectrogram images, and the Mel-frequency cepstral coefficient spectra with the corresponding labels to obtain a training dataset; Construct a multi-dimensional speech emotion recognition model based on adaptive cross-attention. The multi-dimensional speech emotion recognition model based on adaptive cross-attention includes: an input layer, a feature extraction layer, a feature fusion layer, and an emotion classification layer; Input the training dataset into the feature extraction layer to extract acoustic features, and generate multi-dimensional speech emotion features through the feature fusion layer; Input the multi-dimensional speech emotion features into an emotion classifier, and construct a cross-entropy loss function to train the multi-dimensional speech emotion recognition model based on adaptive cross-attention.

2. The multi-dimensional speech emotion recognition method based on adaptive cross-attention according to claim 1, wherein The step of obtaining multiple pieces of original speech data and corresponding labels, extracting the spectrogram images and Mel-frequency cepstral coefficient spectra corresponding to each piece of the original speech data, and combining each piece of the original speech data, the spectrogram images, and the Mel-frequency cepstral coefficient spectra with the corresponding labels to obtain a training dataset includes: Perform discrete Fourier transform on the original speech data, select the discrete Fourier transform points as spectrogram features, and obtain the spectrogram images corresponding to the original speech data; Extract the original speech data through the Librosa library to obtain the Mel-frequency cepstral coefficient spectra corresponding to the original speech data; Combine the original speech data, the spectrogram images, and the Mel-frequency cepstral coefficient spectra with the corresponding labels.

3. The multi-dimensional speech emotion recognition method based on adaptive cross-attention according to claim 1, wherein The feature extraction layer includes an encoder and a pre-trained speech model; the feature fusion layer includes an adaptive cross-attention mechanism; the emotion classification layer includes an emotion classifier.

4. The multi-dimensional speech emotion recognition method based on adaptive cross-attention according to claim 3, wherein The step of inputting the training dataset into the feature extraction layer to extract acoustic features and generating multi-dimensional speech emotion features through the feature fusion layer includes: Input the training dataset into the encoder and the pre-trained speech model to obtain spectrogram acoustic features, Mel-frequency cepstral coefficient spectrum acoustic features, and time-domain speech deep features; Generate multi-dimensional speech emotion features according to the adaptive cross-attention mechanism through the spectrogram acoustic features, the Mel-frequency cepstral coefficient spectrum acoustic features, and the time-domain speech deep features.

5. The multi-dimensional speech emotion recognition method based on adaptive cross-attention according to claim 4, characterized in that The encoder includes a ResNet50 encoder and an RNN encoder.

6. The multi-dimensional speech emotion recognition method based on adaptive cross-attention according to claim 5, wherein The step of inputting the training dataset into the encoder and the pre-trained speech model to obtain spectrogram acoustic features, Mel-frequency cepstral coefficient spectrum acoustic features, and time-domain speech deep features includes: Input the spectrogram images into the ResNet50 encoder to obtain the spectrogram acoustic features corresponding to the spectrogram images; Input the Mel-frequency cepstral coefficient spectra into the RNN encoder module to obtain the Mel-frequency cepstral coefficient spectrum acoustic features corresponding to the Mel-frequency cepstral coefficient spectra; Input the original speech data into the WavLM pre-trained speech model to obtain the time-domain speech deep features corresponding to the original speech data.

7. The multi-dimensional speech emotion recognition method based on adaptive cross-attention according to claim 4, characterized in that, The step of generating multi-dimensional speech emotion features according to the spectrogram acoustic features, the mel-frequency cepstral coefficient spectrum acoustic features, and the time-domain speech deep features based on the adaptive cross-attention mechanism includes: Extracting two-dimensional speech emotion features from the spectrogram acoustic features and the mel-frequency cepstral coefficient acoustic features based on the adaptive cross-attention mechanism; Generating multi-dimensional speech emotion features from the two-dimensional speech emotion features and the time-domain speech deep features based on the adaptive cross-attention mechanism.

8. The multi-dimensional speech emotion recognition method based on adaptive cross-attention according to claim 3, characterized in that The step of inputting the multi-dimensional speech emotion features into an emotion classifier and constructing a cross-entropy loss function to train the multi-dimensional speech emotion recognition model based on the adaptive cross-attention mechanism includes: Inputting the multi-dimensional speech emotion features into an emotion classifier to predict the emotion label corresponding to the original speech data; Constructing a cross-entropy loss function according to the emotion label, backpropagating to adjust the trainable parameters of the multi-dimensional speech emotion recognition model based on the adaptive cross-attention mechanism, and training the multi-dimensional speech emotion recognition model based on the adaptive cross-attention mechanism.

9. A multi-dimensional speech emotion recognition system based on adaptive cross-attention, characterized in that, The system is used to execute the method according to any one of claims 1 to 8, and the system includes: A data acquisition module, configured to acquire multiple pieces of original speech data and corresponding labels, extract the spectrogram image and the mel-frequency cepstral coefficient spectrum corresponding to each piece of the original speech data, and combine each piece of the original speech data, the spectrogram image, and the mel-frequency cepstral coefficient spectrum with the corresponding label to obtain a training data set; A model construction module, configured to construct a multi-dimensional speech emotion recognition model based on the adaptive cross-attention mechanism. The multi-dimensional speech emotion recognition model based on the adaptive cross-attention mechanism includes: an input layer, a feature extraction layer, a feature fusion layer, and an emotion classification layer; A feature generation module, configured to input the training data set into the feature extraction layer to extract acoustic features, and generate multi-dimensional speech emotion features through the feature fusion layer; A model training module, configured to input the multi-dimensional speech emotion features into an emotion classifier, and construct a cross-entropy loss function to train the multi-dimensional speech emotion recognition model based on the adaptive cross-attention mechanism.

10. The multi-dimensional speech emotion recognition system based on adaptive cross-attention according to claim 9, wherein, The multi-dimensional speech emotion recognition model based on the adaptive cross-attention mechanism includes: a ResNet50 encoder module, configured to extract the spectrogram acoustic features in the spectrogram image; An RNN encoder module, configured to extract the mel-frequency cepstral coefficient spectrum acoustic features in the mel-frequency cepstral coefficient spectrum; A WavLM pre-trained speech model, configured to extract the time-domain speech deep features in the original speech data; An adaptive cross-attention mechanism module, configured to integrate the speech emotion information between the spectrogram acoustic features and the mel-frequency cepstral coefficient acoustic features to obtain two-dimensional speech emotion features, and to integrate the speech emotion information between the two-dimensional speech emotion features and the time-domain speech deep features, and capture the internal correlation between the speech signal in the frequency domain and the time domain to obtain multi-dimensional speech emotion features; An emotion classifier module, configured to predict the emotion label corresponding to the original speech data according to the multi-dimensional speech emotion features.