Emotion recognition method based on brain-like multi-mode hierarchical perception

Through the brain-like multimodal hierarchical perception model and multi-head attention mechanism, the problem of insufficient stability of emotional feature extraction in high-dimensional signals is solved, and more robust emotion analysis and multi-modal data fusion are achieved.

CN120067853APending Publication Date: 2025-05-30SHANGHAI UNIV

Patent Information

Application Number
CN202510118849.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When existing emotion recognition methods process high-dimensional signals, the stability of emotional feature extraction is insufficient, which affects the robustness of emotion analysis.

Method used

The emotion recognition method based on brain-like multimodal hierarchical perception is adopted. By constructing a brain-like computing model, human processing of multimodal input information is simulated, and a two-way gating cyclic unit and multi-head attention mechanism is combined to achieve a deep fusion of speech, vision and text features.

Benefits of technology

It improves the robustness and stability of feature extraction in the emotion recognition process, realizes more robust sentiment analysis, and can better process and integrate multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067853A_ABST
    Figure CN120067853A_ABST
Patent Text Reader

Abstract

The invention relates to an emotion recognition method based on brain-like multi-mode hierarchical perception, and the method employs an emotion recognition model to process high-dimensional signals including facial expressions, voices and texts, thereby achieving the emotion recognition. The emotion recognition model comprises a facial expression feature extraction module, a voice feature extraction module, a text processing module, a feature fusion layer and an emotion classification recognition layer, and each feature extraction module comprises a bidirectional gating circulation unit. Compared with the prior art, the method has the advantages that the defect of insufficient extraction stability of high-dimensional signal processing in the emotion recognition process in the prior art is overcome, and a more stable emotion analysis result can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent perception, and in particular to an emotion recognition method based on brain-inspired multi-modal hierarchical perception. Background Art

[0002] Emotion is a highly generalized series of subjective cognitive experiences, a physiological and psychological state generated by a variety of sensations, thoughts, and behaviors. In actual emotion recognition applications such as human-computer interaction, transportation, and the medical field, it is required that the computer understand and generate emotions like humans. Therefore, the research on emotion recognition methods has important research significance. Traditional emotion recognition methods mainly focus on single information such as images, speech, and text, and there are problems such as insufficient information volume and susceptibility to external influences. For example, acoustic features are extracted from speech signals, emotional states are inferred from facial information, or emotional tendencies are extracted through text analysis. However, these single-modal-based methods may be ambiguous when dealing with complex situations such as sarcasm or slang, and the performance of the model is easily affected by the quality of the dataset. In fact, people often integrate multiple modalities when communicating. Therefore, in existing emotion recognition methods, emotion recognition is performed by processing multi-modal data. For example, Chinese Patent Application "CN115169507A" discloses a brain-inspired multi-modal emotion recognition method, which improves the multi-modal feature fusion process by fusing internal features of the same head and external features of different heads and then performing feature splicing. Although it achieves the robustness and accuracy of emotion recognition results to a certain extent, it does not solve the problem of insufficient stability in emotion feature extraction when dealing with high-dimensional signals, thus affecting the robustness of emotion analysis.

[0003] Therefore, it is a technical problem to be solved to provide a method that can process high-dimensional signals more stably and thus perform more stable emotion analysis. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide an emotion recognition method based on brain-inspired multi-modal hierarchical perception. Starting from the perspective of brain-inspired perception, this method imitates the characteristics of the human brain's hierarchical processing of perceptual information, constructs a brain-inspired computing model that can uniformly process multi-modal emotion information, and can simulate the human processing of multi-modal input information; the emotion recognition model integrates a brain-inspired perception mechanism and machine learning. First, deep feature extraction of speech, vision, and text is completed based on the real human brain perception characteristics; then, information fusion is completed through a designed multi-modal multi-head attention mechanism; finally, emotion analysis is realized using the decision layer.

[0005] The purpose of the present invention can be achieved through the following technical solutions:

[0006] The present invention provides an emotion recognition method based on brain-inspired multi-modal hierarchical perception. The method uses an emotion recognition model to process high-dimensional signals including facial expressions, speech, and text to achieve emotion recognition. The emotion recognition model includes a facial expression feature extraction module, a speech feature extraction module, a text processing module, a feature fusion layer, and an emotion classification and recognition layer. Each feature extraction module includes a bidirectional gated recurrent unit. The steps are as follows:

[0007] Obtain the video to be detected, and separate the facial expression images, speech signals, and text signals therefrom;

[0008] Use the facial expression feature extraction module to perform facial expression detection based on the facial expression images, and extract facial expression features. The facial expression feature extraction module includes a multi-task convolutional neural network and a convolutional expert constrained local network;

[0009] Use the speech feature extraction module to preprocess the obtained speech signals and then extract speech features;

[0010] Use the text processing module to extract text features from the obtained text signals;

[0011] Based on the facial expression features, speech features, and text features, use the feature fusion layer to perform feature fusion;

[0012] Based on the fused features, use the emotion classification and recognition layer to perform emotion classification and recognition.

[0013] As a preferred technical solution, the method for extracting facial expression features is as follows:

[0014] Use the multi-task convolutional neural network to perform preliminary human face facial expression recognition based on the facial expression images;

[0015] Based on the results of the preliminary human face facial expression recognition, use the convolutional expert constrained local network to perform human face feature point recognition, obtain the three-dimensional key point information of the human face, and generate the position information of the eyeballs and pupils;

[0016] Perform similarity transformation based on the position information, and extract the facial appearance and its geometric features;

[0017] Based on the facial appearance and its geometric features, use the bidirectional gated recurrent unit to obtain facial expression features containing context information.

[0018] As a preferred technical solution, the method for obtaining facial expression features containing context information is as follows:

[0019]

[0020] Wherein, Represents the forward GRU calculation; Represents the backward GRU calculation; ⊕ represents the concatenation operation; z it Represents the output operation; x it Represents the geometric features of the t-th facial appearance in video i;

[0021] Among them, the calculation method of the GRU calculation is:

[0022] r t = δ(U r x it + W r h t-1 + b r ),

[0023] z t = δ(U z x it + w z h t-1 + b z ),

[0024]

[0025] Represents the candidate hidden state of the geometric feature t of the facial appearance; h t Represents the hidden layer state of the geometric feature t of the facial appearance; U, W, and b represent weights and biases respectively; δ represents the Sigmoid activation function, and * represents the multiplication of corresponding elements of the matrix.

[0026] As a preferred technical solution, the preprocessing includes:

[0027] Perform pre-emphasis on the high-frequency part of the speech signal using a first-order digital filter for the speech signal, and its expression is:

[0028] H(z) = 1 - ξz -1 ,

[0029] Among them, H(z) represents the transfer function of the pre-emphasis filter; ξ represents the pre-emphasis coefficient; z represents the complex frequency variable;

[0030] Perform windowing on the pre-emphasized speech signal, and its expression is:

[0031] S w (z) = S(z) * W(z),

[0032] W(z) = 0.54 - 0.46cos[2πz / (Z - 1)],

[0033] Among them, S w(z) represents the windowed signal transfer function; S(z) represents the transfer function of the original speech signal; W(z) represents the transfer function of the window added, and Z represents the frame length.

[0034] As a preferred technical solution, the speech features include prosodic features, spectral features, and Mel-frequency cepstral coefficients. The method for obtaining the speech features by using the bidirectional gated recurrent unit based on the preprocessed speech signal includes:

[0035] Calculating the autocorrelation function of the preprocessed speech signal to obtain the prosodic features, and its expression is:

[0036]

[0037] where k represents the time delay; N represents the total frame length of the speech signal; n represents the index of the current frame of the speech signal; x i (n) represents the amplitude of the i-th speech frame at the n-th frame; x i (n + k) represents the amplitude of the i-th speech frame at the n + k-th frame;

[0038] Performing constant-Q transform processing on the preprocessed speech signal to obtain the spectral features, and its expression is:

[0039]

[0040] where N k represents the length of the window function; n represents the index of the discrete signal sample; x(n) represents the amplitude of the preprocessed speech signal at the n-th sample point; represents the window function with length N k ; Q represents the constant factor in the transform.

[0041] Obtaining the actual frequency relationship of the preprocessed speech signal, and calculating the Mel-frequency cepstral coefficients based on the actual frequency relationship, and its expression is:

[0042] Mel(f) = 2595 * lg(1 + f / 700),

[0043] where f represents the actual frequency.

[0044] As a preferred technical solution, the method for extracting text features is:

[0045] Processing the text signal by using the self-attention mechanism and the feed-forward neural network to obtain word vectors;

[0046] Extracting text features based on the word vectors by using the bidirectional gated recurrent unit.

[0047] As a preferred technical solution, the method for feature fusion is as follows:

[0048] Transform the facial expression features, speech features, and text features, and unify them to the same dimension;

[0049] Randomly select two features from the facial expression features, speech features, and text features after dimension unification, and use one feature as the query and the other feature as the key and value for modality fusion; repeat the above modality fusion process so that at least one of the queries and the corresponding features of the key and value selected for each modality fusion is different, and six modality fusion results are obtained;

[0050] Obtain the common features based on the modality fusion results, and perform linear feature transformation on the common features to obtain the fusion features.

[0051] As a preferred technical solution, the method for modality fusion is as follows:

[0052] Calculate the query, key, and value, and their expressions are:

[0053] Q = max(0, H α W α + b α ),

[0054] K = V = max(0, H β W β + b β ),

[0055] where Q represents the query; K represents the key; V represents the value; H α and H β represent the α feature and the β feature, and α, β ∈ {v, a, t}, v represents the facial expression feature, a represents the speech feature, and t represents the text feature; W α and W β represent the weight matrices of the a feature and the β feature; b α and b β represent the bias terms of the α feature and the β feature.

[0056] As a preferred technical solution, the method for obtaining the common features is as follows:

[0057] Use multiple independent attention heads to calculate the similarity between the query and the key through scaled dot-product attention to obtain the attention distribution;

[0058] Apply the attention distribution to the value to extract the common features, and its expression is:

[0059]

[0060] where The parameter matrices for the i-th attention head to be used for query Q, key K, and value V respectively.

[0061] As a preferred technical solution, the emotion classification and recognition layer includes a plurality of fully connected layers and a classifier. The method for performing emotion classification and recognition is as follows:

[0062]

[0063] Among them, represents the feature vector; W t , b t represent the weights and biases of the fully connected layer; W soft , b soft represent the weights and biases of the softmax layer; y i represents the emotion recognition classification result.

[0064] Compared with the prior art, the present invention has the following beneficial effects:

[0065] 1), The present invention provides different feature extraction methods for different modal information, and a bidirectional gated recurrent unit is added to each feature extraction module. Based on the preliminarily processed video data (including facial expression images, speech signals, and text signals), the temporal dependence relationship in the input data is captured, and the information from front to back and from back to front is obtained, which not only improves the robustness of feature extraction in the emotion recognition process, but also enables stable features to be extracted when performing high-dimensional signal processing, thereby realizing robust emotion analysis.

[0066] 2), The present invention refers to the process of the human brain's visual and auditory perception and understanding of emotions, provides a method for brain-like multi-level emotion recognition, realizes the extraction of image expression features, text word vector representation and understanding, and speech emotion feature extraction that conform to brain cognition, and provides a method for fusing multi-modal feature data, enabling the emotion recognition model to better utilize the complementary characteristics between different modalities and better mine the interaction information of multi-modal data in high-dimensional information. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 is the flowchart of the method of the present invention;

[0068] Figure 2 is the system block diagram of the emotion recognition method of the present invention;

[0069] Figure 3 is the system block diagram of the facial expression feature extraction method of the present invention;

[0070] Figure 4 is the system block diagram of the speech feature extraction method of the present invention;

[0071] Figure 5 System block diagram of the text feature extraction method of the present invention;

[0072] Figure 6 System block diagram of the feature fusion method of the present invention;

[0073] Figure 7 Schematic diagram of the four-class emotion confusion matrix of the present invention;

[0074] Figure 8 Schematic diagram of the six-class emotion confusion matrix of the present invention. Specific implementation manners

[0075] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0076] The present invention imitates the characteristics of the human brain to hierarchically process perceptual information, and proposes an emotion analysis method based on hierarchical perception of a brain-like multi-modal module. According to the process of the human brain's visual and auditory perception and understanding of emotions, a feature extraction layer is designed; a cross-attention mechanism is used to learn the interaction representation mechanism between modalities to build a fusion layer. According to the computer principle of the brain-like neural network model in the application of multi-modal information fusion, a brain-like computing model that can uniformly process multi-modal emotion perception information is constructed, and hierarchical processing is performed to establish a learning and decision-making model based on multi-modal perception; based on the process of the human brain's visual and auditory perception and understanding of emotions, image expression feature extraction, text word vector representation and understanding, and speech emotion feature extraction that conform to brain cognition are realized; in order to better mine the interaction information between the visual and auditory modalities, a multi-modal fusion layer based on the cross-attention mechanism is designed.

[0077] The method of the present invention uses an emotion recognition model to process high-dimensional signals including facial expressions, speech, and text to achieve emotion recognition. The process is as follows Figure 1 shown. Specifically, the emotion recognition model includes a perception layer, a feature extraction layer, a feature fusion layer, and an emotion classification and recognition layer as Figure 2 shown, wherein the feature extraction layer includes a facial expression feature extraction module, a speech feature extraction module, and a text processing module, and each feature extraction module includes a bidirectional gated recurrent unit.

[0078] The steps for performing emotion recognition are as follows:

[0079] S1. Obtain the video to be detected, and separate the facial expression image, speech signal, and text signal through the perception layer in the emotion recognition model.

[0080] S2. Feature extraction:

[0081] S21. Facial expression feature extraction:

[0082] Using the facial expression feature extraction module, facial expression detection is performed on the facial expression image, and facial expression features are extracted. The facial expression feature extraction module includes a multi-task convolutional neural network and a convolutional expert constrained local network. This process is as Figure 3 shown.

[0083] Specifically including:

[0084] S211. Using the multi-task convolutional neural network to perform preliminary human face facial expression recognition on the facial expression image.

[0085] Based on the results of the preliminary human face facial expression recognition, using the convolutional expert constrained local network to perform human face feature point recognition, obtaining the three-dimensional key point information of the human face and generating the position information of the eyeballs and pupils as well as the facial landmark information.

[0086] S212. Based on the position information and facial landmark information, perform similarity transformation, and extract the facial appearance and its geometric features. This step is similar to the initial processing of the expression seen by the human eye.

[0087] S213. Based on the facial appearance and its geometric features, use a bidirectional gated recurrent unit (i.e., Bi_GRU) to obtain facial expression features containing context information.

[0088] Among them, the method for obtaining facial expression features containing context information is:

[0089]

[0090] Among them, represents the forward GRU calculation; represents the backward GRU calculation; ⊕ represents the concatenation operation; z it represents the output operation; x it represents the geometric features of the t-th facial appearance in video i;

[0091] Among them, the calculation method of the GRU calculation is:

[0092] r t = δ(U r x it + W r h t-1 + b r ),

[0093] z t = δ(U z x it + W z ht-1 +b z ),

[0094]

[0095] represents the candidate hidden state of the geometric feature t of the facial appearance; h t represents the hidden layer state of the geometric feature t of the facial appearance; U, W, and b represent weights and biases respectively; δ represents the Sigmoid activation function, and * represents element-wise multiplication of matrices.

[0096] S22. After preprocessing the acquired speech signal using the speech feature extraction module, speech feature extraction is performed, as Figure 4 shown.

[0097] S221. First, preprocessing is performed in sequence, mainly including pre-emphasis, framing, and windowing.

[0098] Pre-emphasis is performed on the high-frequency part of the speech signal using a first-order digital filter for the speech signal, and its expression is:

[0099] H(z) = 1 - az -1 ,

[0100] where H(z) represents the transfer function of the pre-emphasis filter; a represents the pre-emphasis coefficient; z represents the complex frequency variable;

[0101] Windowing is performed on the pre-emphasized speech signal, and its expression is:

[0102] S w (z) = S(z) * W(z),

[0103] W(z) = 0.54 - 0.46cos[2πz / (Z - 1)],

[0104] where S w (z) represents the transfer function of the windowed signal; S(z) represents the transfer function of the original speech signal; W(z) represents the transfer function of the window added; Z represents the frame length.

[0105] S222. According to the characteristics of human voices, prosodic features, spectral features, and Mel-frequency cepstral coefficients are used as speech features, and the method for extracting speech features is to extract prosodic features, spectral features, and Mel-frequency cepstral coefficients from the speech signal after preprocessing.

[0106] Among them, the method for obtaining speech features includes:

[0107] The autocorrelation function is calculated for the preprocessed speech signal to obtain prosodic features, and its expression is:

[0108]

[0109] Among them, k represents the time delay; N represents the total frame length of the speech signal; n represents the index of the current frame of the speech information; x i (n) represents the amplitude of the i-th speech frame at the n-th frame; x i (n + k) represents the amplitude of the i-th speech frame at the n + k-th frame, which is the value of the speech signal after being delayed by k time units.

[0110] Perform a constant Q transform on the preprocessed speech signal to obtain spectral features, and its expression is:

[0111]

[0112] Among them, N k represents the length of the window function; x(n) represents the amplitude of the preprocessed speech signal at the n-th sample point; represents a window function with a length of N k ; Q represents the constant factor in the transform.

[0113] Since the Mel-frequency cepstral coefficients are close to the characteristics of the human auditory system and reflect the speech features of the power spectrum, the present invention also selects the Mel-frequency cepstral coefficients as one of the speech features, obtains the actual frequency relationship of the preprocessed speech signal, and calculates the Mel-frequency cepstral coefficients based on the actual frequency relationship. Its expression is:

[0114] Mel(f) = 2595 * lg(1 + f / 700),

[0115] where f represents the actual frequency.

[0116] S23. Use the text processing module to extract text features from the obtained text signal, and its process is as Figure 5 shown.

[0117] S231. In the embodiment, the present invention selects the Bert model to complete the word vector representation of the text modality. The Bert model is divided into an input layer, an encoding layer, and an output layer. Assume that the text input obtained by the perception layer is E{E 1 , E 2 , …, E n}. After being processed by multiple self-attention mechanisms and feed-forward neural networks, the generated word vectors are T{T 1 , T 2 , …, T n}.

[0118] S232. Extract text features based on the word vectors using a bidirectional gated recurrent unit.

[0119] S3. Based on facial expression features, speech features, and text features, use the feature fusion layer to perform feature fusion. The detailed process is as Figure 6 shown.

[0120] The present invention provides a fusion strategy for visual and auditory features based on a cross-attention mechanism. The attention mechanism is used to capture information between modalities. The text, audio, and video are respectively used as the inputs of the query Q, key K, and value V for fusion, and six different modality fusion results are output, obtaining richer modality features.

[0121] Specifically, the method for performing feature fusion includes:

[0122] S31. Transform the facial expression features, speech features, and text features to unify them to the same dimension.

[0123] S32. Randomly select two features from the facial expression features, speech features, and text features with unified dimensions, and use one feature as the query and the other feature as the key and value for modality fusion; repeat the above modality fusion process so that at least one of the query and the corresponding features of the key and value selected for each modality fusion is different, and six modality fusion results are obtained. The expression is:

[0124] Q = max(0, H α W α + b α ),

[0125] K = V = max(0, H β W β + b β ),

[0126] where Q represents the query, and K represents the key; V represents the value, and α, β ∈ {v, a, t}, where v represents facial expression features, a represents speech features, and t represents text features; W α and W β represent the weight matrices of the α feature and the β feature; b α and b β represent the bias terms of the α feature and the β feature; d l , and d n represent the lengths of the video and the text respectively; d in represents the vector dimension.

[0127] Specifically, the combination forms for obtaining the six-modal fusion results include: using the facial expression features as the query Q and the speech features as the key K; using the facial expression features as the query Q and the text features as the key K; using the speech features as the query Q and the facial expression features as the key K; using the speech features as the query Q and the text features as the key K; using the text features as the query Q and the facial expression features as the key K; using the text features as the query Q and the speech features as the key K.

[0128] S33. Obtain the common features based on the modal fusion results, and perform linear feature transformation on the common features to obtain the fusion features.

[0129] S331. Use multiple independent attention heads to calculate the similarity between the query and the key through scaled dot-product attention to obtain the attention distribution.

[0130] S332. Apply the attention distribution to the values to extract the common features, and its expression is:

[0131]

[0132] where are the parameter matrices for the i-th attention head to be used for the query Q, the key K, and the value V respectively.

[0133] S333. Concatenate the time averages of the obtained common features according to the time sequence of the video to complete statistical pooling, and finally output three fusion features.

[0134] S4. Perform emotion classification and recognition based on the fusion features using the emotion classification and recognition layer.

[0135] The present invention uses a trained emotion classification and recognition layer for emotion classification and recognition. This layer includes multiple fully connected layers and a classifier. The method for performing emotion classification and recognition is:

[0136]

[0137] where represents the feature vector; W t , b t represent the weights and biases of the fully connected layer; W soft , b soft represent the weights and biases of the softmax layer; y i represents the emotion recognition classification result.

[0138] Specifically, during the training process of this layer, cross-entropy is selected as the loss function, which represents the gap between the probability of the actual predicted category and the probability of the expected model predicted category. The smaller the value of the cross-entropy, the closer the two category prediction probability distributions are. The calculation formula of the loss function is as follows:

[0139]

[0140] Among them, y i is the probability of the expected model prediction category; S i is the probability of the actual model prediction category; C represents the total number of categories in the set; i represents the index of a specific category.

[0141] In this embodiment, the IECOMAP dataset is also selected to train the emotion recognition model provided by the present invention. Specifically, the present invention respectively considers four and six classification emotions. Among the four classification emotions, there are "anger", "happiness", "neutral", and "sadness", where "happiness" and "excitement" are combined into one category. The number of examples for each category is shown in Table 1, which are 1103, 1636, 1708, and 1084 respectively.

[0142] Table 1 Four-class data table

[0143] Angry Happy Neutral Sad Total Data Distribution 1103 1636 1708 1084 5531

[0144] Among the six classification emotions, there are "anger", "happiness", "neutral", "sadness", "depression", and "excitement". The number of examples for each category is shown in Table 2, which are 1103, 648, 1708, 1084, 1849, and 1041 respectively.

[0145] Table 2 Six-class data table

[0146] Angry Happy Neutral Sad Depressed Excited Total Data Distribution 1103 648 1708 1084 1849 1041 7743

[0147] Initialize the emotion recognition model provided by the present invention and make the following settings:

[0148] i). Use the Adam optimizer to optimize the training network learning parameters;

[0149] ii). Use Dropout to prevent overfitting and set the parameter to 0.3;

[0150] iii). Set the initial learning rate to 0.001, the batch size to 32, and use cross-entropy loss as the loss function.

[0151] iiii). Set the number of hidden units of the audio and video GRU networks to 32 and 64 respectively.

[0152] Set the experimental results to be cross-validated 10 times, and divide the dataset into 10 subsets. 8 subsets are used as the training set for training, and the remaining 2 subsets are used as the validation set and the test set for validation and testing respectively.

[0153] To comprehensively analyze the recognition performance of correct and incorrect classifications of emotion categories, experiments on four-class and six-class classifications were respectively conducted in the embodiments. Figure 7 and Figure 8 show the confusion matrices for the two classification cases. As can be seen from Figure 7 , all four emotion categories can be well distinguished, and the overall performance is good. As can be seen from Figure 8 , "anger", "neutral", "sadness", and "frustration" can be well classified, while the emotion categories of "happiness" and "excitement" are confused. In fact, during the real emotion recognition process, the human brain also has a relatively vague distinction between these two emotions, mainly because the corresponding voice, text, and facial expression features of the two are similar, resulting in easy confusion in their emotion classification. In the four-class classification, happiness and excitement are grouped into one category, so the emotion recognition effect of the four-class classification is good.

[0154] To verify the feasibility of the above method, the present invention also compares and analyzes this method with the most advanced related methods. During the comparative experiment process, all methods for emotion recognition adopt common evaluation metrics in machine learning classification tasks, including the accuracy ACC (Accuracy, Acc) and the F1 value (F1-Measure, F1), to evaluate the performance of the algorithm model. Among them, the definition of Acc is the ratio of the number of samples correctly classified by the model to the total number of samples, and its calculation formula is:

[0155] Acc = (TP + TN) / (TP + FP + FN + TN),

[0156] where TP is the number of samples predicted as positive and actually positive; FP is the number of samples predicted as positive and actually negative; TN is the number of samples predicted as negative and actually negative; FN is the number of samples predicted as negative and actually positive.

[0157] The F1 value is the harmonic mean of the precision rate and the recall rate, and its calculation formula is:

[0158] Pre = TP / (TP + FP),

[0159] Rec = TP / (TP + FN),

[0160] F_1 = (2 * Pre * Rec) / (Pre + Rec),

[0161] where Pre is the precision rate, defined as the proportion of truly positive samples among the samples predicted as positive; Rec is the recall rate, defined as the proportion of samples predicted as positive by the model among all positive sample instances.

[0162] The experimental results and comparison of the four-class emotion discrimination are shown in Table 3.

[0163] Table 3 Experimental results of the four-class comparison

[0164]

[0165] The results of the six-class emotion discrimination experiment and the comparison are shown in Table 4 as follows.

[0166] Table 4 Results of the six-class comparison experiment

[0167]

[0168] It can be seen from the table that the model proposed in this paper is close to the best performance in terms of weighted average accuracy and F1 score, and performs well in the model metrics on the IEMOCAP dataset. The recognition performance is more balanced in these four categories. The relatively uniform recognition accuracy indicates that the brain-inspired perception recognition provided by the invention is more in line with the recognition of real emotions. Therefore, the method provided by the present invention is feasible.

[0169] Furthermore, in this embodiment, an ablation experiment on modal data is also carried out, and the results are shown in Table 5.

[0170] Table 5 Results of the ablation experiment

[0171]

[0172] It can be seen from the table that the more modalities there are, the better the performance. When all three modalities are present, the performance is the best, indicating the effectiveness of brain-inspired multi-modal perception. When there is only one modality, the text modality has better performance than the other two modalities. In summary, in the process of brain-inspired multi-modal perception, using text to focus on speech and visual images can better analyze emotions.

[0173] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An emotion recognition method based on brain-like multimodal hierarchical perception, characterized in that: The method uses an emotion recognition model to process high-dimensional signals including facial expressions, voices and texts to realize emotion recognition. The emotion recognition model includes a facial expression feature extraction module, a voice feature extraction module, a text processing module, a feature fusion layer and an emotion classification and recognition layer, and each feature extraction module includes a bidirectional gated recurrent unit. The steps are: Obtain the video to be detected and separate the facial expression image, voice signal and text signal from it; Utilizing the facial expression feature extraction module to perform facial expression detection based on the facial expression image and extract facial expression features, the facial expression feature extraction module includes a multi-task convolutional neural network and a convolutional expert constrained local network; After preprocessing the acquired voice signal using the voice feature extraction module, voice feature extraction is performed; Using the text processing module to extract text features from the acquired text signal; Based on the facial expression features, speech features and text features, using the feature fusion layer to perform feature fusion; Based on the fusion features, the emotion classification and recognition layer is used to classify and recognize emotions.

2. The emotion recognition method based on brain-like multimodal hierarchical perception according to claim 1 is characterized in that: The method for extracting facial expression features is: Using the multi-task convolutional neural network to perform preliminary facial expression recognition based on the facial expression image; Based on the preliminary facial expression recognition results, the convolutional expert constrained local network is used to recognize facial feature points, obtain the three-dimensional key point information of the face and generate the position information of the eyeball and pupil; Perform similarity changes based on the position information and extract facial appearance and its geometric features; The facial expression features containing context information are obtained by using the bidirectional gated recurrent unit based on the facial appearance and its geometric features.

3. The emotion recognition method based on brain-like multimodal hierarchical perception according to claim 2 is characterized in that: The method for obtaining facial expression features containing context information is as follows: in, Represents forward GRU calculation; Represents backward GRU calculation; Indicates the splicing operation; z it Indicates output operation; x it Represents the geometric features of the t-th facial appearance in video i; Among them, the calculation method of GRU calculation is: r t =δ(U r x it +W r h t-1 +b r ), z t =δ(U z x it +W z h t-1 +b z ), The candidate hidden state of the geometric feature t representing the facial appearance; h t Represents the hidden layer state of the geometric feature t of facial appearance; U, W and b represent weights and biases respectively; δ represents the Sigmoid activation function, and * represents the multiplication of corresponding elements of the matrix.

4. The emotion recognition method based on brain-like multimodal hierarchical perception according to claim 1 is characterized in that: The pretreatment comprises: The first-order digital filter of the speech signal performs pre-emphasis on the high-frequency part, and its expression is: H(z)=1-ξz -1 , Wherein, H(z) represents the transfer function of the pre-emphasis filter; ξ represents the pre-emphasis coefficient; z represents the complex frequency variable; The pre-emphasized speech signal is windowed, and its expression is: S w (z)=S(z)*W(z), W(z)=0.54-0.46cos[2πz / (Z-1)], Among them, S w (z) represents the signal transfer function after windowing; S(z) represents the transfer function of the original speech signal; W(z) represents the windowed transfer function, and Z represents the frame length.

5. The emotion recognition method based on brain-like multimodal hierarchical perception according to claim 1 is characterized in that: The speech features include prosodic features, spectral features and Mel-frequency cepstral coefficients. The method for obtaining the speech features based on the preprocessed speech signal using the bidirectional gated recurrent unit includes: The autocorrelation function is calculated on the preprocessed speech signal to obtain the prosodic feature, which is expressed as follows: Where k represents the time delay; N represents the total frame length of the speech signal; n represents the index of the current frame of the speech signal; x i (n) represents the amplitude of the i-th speech frame in the n-th frame; x i (n+k) represents the amplitude of the i-th speech frame in the n+k-th frame; The pre-processed speech signal is subjected to constant Q transformation to obtain the spectrum feature, which is expressed as follows: Among them, N k represents the length of the window function; x(n) represents the amplitude of the preprocessed speech signal at the nth discrete signal sample point; Indicates length N k The window function of ; Q represents the constant factor in the transformation; The actual frequency relationship of the preprocessed speech signal is obtained, and the Mel-frequency cepstral coefficient is calculated based on the actual frequency relationship, and the expression is: Mel(f)=2595*lg(1+f / 700), Wherein, f represents the actual frequency.

6. The emotion recognition method based on brain-like multimodal hierarchical perception according to claim 1 is characterized in that: The method for extracting text features is as follows: Use the self-attention mechanism and feedforward neural network to process text signals and obtain word vectors; The bidirectional gated recurrent unit is used to extract text features based on the word vector.

7. The emotion recognition method based on brain-like multimodal hierarchical perception according to claim 1 is characterized in that: The method for performing feature fusion is: Transforming the facial expression features, speech features and text features to unify them into the same dimension; Randomly select two features from the facial expression features, voice features, and text features after dimensionality unification, and use one of the features as the query and the other as the key and value to perform modal fusion; repeat the above modal fusion process so that at least one of the query and key and value corresponding features selected for each modal fusion is different, and obtain six modal fusion results; Common features are obtained based on the modal fusion results, and the common features are subjected to linear feature transformation to obtain fusion features.

8. The emotion recognition method based on brain-like multimodal hierarchical perception according to claim 7 is characterized in that: The modal fusion method is: Calculate the query, key and value, the expression is: Q=max(0,H α W α +b α ), K=V=max(0,H β W β +b β ), Among them, Q represents query; K represents key; V represents value; H represents α and H β represents α feature and β feature, and α, β∈{v, a, t}, v represents facial expression feature, a represents speech feature, and t represents text feature; W α and W β Represents the weight matrix of α feature and β feature; b α and b β Represents the bias term of α feature and β feature.

9. The emotion recognition method based on brain-like multimodal hierarchical perception according to claim 7 is characterized in that: The method for obtaining common features is: Use multiple independent attention heads to perform scaled dot product attention to calculate the similarity between the query and the key and obtain the attention distribution; Applying the attention distribution to the value to extract the common features, the expression is: head i =Attention(QW i Q ,KW i K ,VW i V ), Among them, W i Q , W i K , W i V are the parameter matrices of the ith attention head for query Q, key K, and value V respectively.

10. The emotion recognition method based on brain-like multimodal hierarchical perception according to claim 1, characterized in that: The emotion classification and recognition layer includes multiple fully connected layers and classifiers. The method for performing emotion classification and recognition is as follows: in, represents the feature vector; W t , b t represents the weight and bias of the fully connected layer; W soft , b soft represents the weight and bias of the softmax layer; y i Represents the emotion recognition classification results.

Citation Information

Patent Citations

  • Brain-like multi-mode emotion recognition network, brain-like multi-mode emotion recognition method and emotion robot

    CN115169507A

Cited By

  • Micro-expression recognition method and device, electronic equipment and storage medium

    CN121617144A