Multi-modal emotion recognition method and system based on VAFNet
The VAFNet model of filtering video frames through the PLAS algorithm and combining the positive traffic channel and interactive attention mechanism, the problem of insufficient frame selection and feature expression in short video emotion recognition is solved, and efficient and accurate emotion recognition is achieved.
Patent Information
- Application Number
- CN202510177949.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to effectively capture emotional changes in short video emotions recognition. The traditional frame selection method is affected by face rotation and scaling and has high calculation costs. The Mel spectrum feature expression ability is insufficient, and it is easy to lose relying on audio features.
The PLAS algorithm is used to filter video frames, combine the positive traffic channel attention mechanism and interactive attention mechanism of audio mode, and multimodal feature fusion is performed through the VAFNet model, and the classification performance is optimized using balanced cross entropy and modal consistency loss function.
It improves the video frame sampling efficiency, enhances the accuracy and robustness of multi-modal feature fusion, reduces computing costs, and adapts to emotion recognition in multi-lingual environments.
Smart Images

Figure CN120260145A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of emotion recognition, and particularly relates to a multi-modal emotion recognition method and system for audio-visual feature extraction and fusion based on deep learning. Background Art
[0002] In short video emotion recognition, how to extract effective frames from video signals is one of the current research difficulties. Traditional video frame selection methods, such as simple interval sampling, random sampling or methods based on image difference, often fail to effectively capture emotion changes. Although methods such as optical flow method and local binary pattern (LBP) can distinguish the change trend, they are interfered by pixel changes unrelated to emotions, such as rotation, movement and scaling of the face. At the same time, the calculation needs to consider all pixel points of each frame image, resulting in a high calculation cost.
[0003] In addition, the color information of the Mel spectrogram represents the intensity of energy at frequency points. The traditional channel attention mechanism uses global average pooling (GAP) to compress channel information, which will lead to the loss of channel detail information, reducing the feature expression ability. At the same time, the features extracted only relying on the Mel spectrogram are prone to losing the original audio features. Summary of the Invention
[0004] To solve the above problems of the prior art, the present invention proposes an efficient and accurate multi-modal emotion recognition method and system through the PLAS (Procrustes Landmark Alignment Similarity, frame similarity between frames referring to facial key points) frame selection algorithm in the video modality, the orthogonal channel attention mechanism OCA feature extraction strategy and the interactive attention mechanism in the audio modality. The present invention does not rely on language semantic information and is only based on two modalities of audio and video, and can meet the needs of short video emotion recognition for different races and cultural backgrounds around the world, providing a unified solution for emotion analysis in a multilingual environment.
[0005] The present invention adopts the following technical solutions:
[0006] A multi-modal emotion recognition method based on VAFNet (Video-Audio Fusion Net) is carried out according to the following steps:
[0007] S1. After preprocessing the video data set, sampling is performed, and then it is converted into a peak frame sequence;
[0008] S2. The audio signal in the video is extracted as a one-dimensional time series signal and preprocessed to obtain the Mel spectrogram corresponding to each audio;
[0009] S3. Input the video peak frame sequence obtained in step S1 into the pre-trained ResNet18 model to extract spatial features to obtain facial spatial features Denote the facial spatial features of the M-th frame image, and then integrate the time information through one-dimensional convolution to obtain a video feature vector; input the Mel spectrogram and the processed audio signal obtained in step S2 into the OCANet model and the WaveNet respectively to extract spectral features and original waveform features, and obtain the Mel feature vector f mel and the waveform feature vector f wav ;
[0010] S4. Input the extracted Mel feature vector f mel , the waveform feature vector f wav and the facial spatial features into the feature fusion model based on the interactive attention mechanism to obtain the interactive feature vector f interact . Concatenate and fuse the Mel feature vector f mel , the video feature vector f video and the interactive feature vector f interact to obtain a multi-modal feature vector, perform emotion classification, and output the emotion classification result.
[0011] As a preferred solution, in step S1, preprocess the video dataset. Each video is sampled after frame division, face cropping, grayscale conversion, and key point extraction, and then converted into a peak frame sequence; specifically:
[0012] Perform video frame division, face cropping, and grayscale conversion on the videos in the video dataset to obtain a video frame sequence {I1, I2, …, I N}}. Then, extract the facial key point marks (landmarks) for each frame to obtain the facial key point coordinate set of each frame. Denote the key point coordinate set of the n-th frame:
[0013]
[0014] Considering that in the emotion recognition task, facial expression changes are directly reflected in the areas densely surrounded by some key points, such as the mouth, eyes, etc. The key point coordinates do not depend on the brightness value or texture information of the pixels, so they are robust to non-geometric factors such as illumination changes and image noise. In addition, the speaker in the video may move, rotate, move away from, or approach the camera on the lens, and conventional pixel texture algorithms cannot exclude these noises unrelated to emotions.
[0015] Based on the above situation, the present invention introduces the PLAS (Procrustes Landmark Alignment Similarity) algorithm, which performs key-point alignment while considering the geometric similarity of facial key-points between video frames to eliminate the influence of translation, scaling, and rotation. First, for the key-point coordinate set P n perform centering processing to eliminate the influence of translation and obtain P ′ n :
[0016]
[0017] To eliminate the influence of scaling, normalize the centered key-point coordinate set P′ n to obtain P″ n :
[0018]
[0019] where scale n is the scaling factor of the nth frame. After completing the translation and scaling operations, it is still necessary to eliminate the influence of rotation. Using the key-point sets P″1 and P″ of the first frame and the nth frame n and the singular value decomposition (SVD) method to calculate the optimal alignment rotation matrix R between the key-point sets of the first frame and the nth frame:
[0020] H = (P″1) T P″ n = U∑V T
[0021]
[0022] The covariance matrix H is the similarity transformation matrix between the first frame and the nth frame. The rotation matrix R is a second-order orthogonal matrix, and r 11 , r 22 represent the rotation amounts in the x-axis direction and y-axis direction in sequence, and r 12 , r 21 represent the influence of the rotation amount in the y-axis direction on the x-axis direction and the influence of the rotation amount in the x-axis direction on the y-axis direction in sequence. In summary, the key-point spatial similarity s[1, n] between the first frame and the nth frame can be obtained as:
[0023]
[0024] where n = 2, 3, 4…, N. Sort the values of s[1, 2], s[1, 3], … s[1, N] from small to large, select the first M frames with the smallest similarity as the peak frames, and arrange them in chronological order to obtain the peak frame sequence {I1, I2, …, I M} after PLAS processing.
[0025] As a preferred solution, step S2 includes:
[0026] Extract the audio signal from the video. The preprocessing includes pre-emphasizing, framing, amplitude normalizing, and windowing the audio signal after extracting the effective sound segment using voice activity detection (VAD) to obtain the preprocessed signal a[n]. After applying the short-time Fourier transform (STFT) to a[n], it undergoes Mel filtering and logarithmic transformation to obtain the Mel spectrogram that conforms to the human ear's auditory perception.
[0027] As a preferred solution, step S3 includes:
[0028] Input the processed video peak frame sequence {I1, I2, …, I M} with a size of (224, 224, 1, M) into the ResNet18 model pre-trained on the FER dataset to extract the facial spatial features. After each frame extracts the spatial features through ResNet18, the features Then input F s into the 1D CNN model to extract the time information to obtain the 512-dimensional video feature vector f video .
[0029] As a preferred solution, step S3 also includes:
[0030] Input the processed audio signal a[n] into WaveNet to extract the features f wav of the original waveform. WaveNet extracts the time information through causal dilated convolution; input the Mel spectrogram processed in step S2 into the audio feature extraction model OCANet. After the Mel spectrogram is extracted by OCANet, the 512-dimensional audio feature vector f mel is obtained.
[0031] As a preferred solution, in step S4, the loss function for training the VAFNet model combines the balanced cross-entropy loss and the modality consistency loss.
[0032] As a preferred solution, step S4 includes:
[0033] Denote the spatial features without integrated time information extracted in step 3 as the feature matrix F s , and input F s and the audio features [f mel , f wav extracted in step 3 into the feature fusion (Interactive attention) model based on the interactive attention mechanism. The query (Query), key (Key), and value (Value) constructed for its training are as follows:
[0034] Q audio = [f mel , f wav W Q , K video = F s W K
[0035] V video = F s W V
[0036] Among them, W Q , W K , W V are all weight matrices participating in training. The similarity weight matrix A represents the attention weight of audio features to each frame of video features. The weight matrix A and the interaction feature vector f interact are in turn:
[0037]
[0038] f interact = flatten(AV video )
[0039] Among them, flatten represents the feature flattening operation. Since the dimension of AV video is (2, 512) and contains two parts of information, namely the attention degrees of Mel features and waveform features to video frames. Finally, the spectrogram feature f mel , the video feature f video , and the interaction feature f interact need to be concatenated and fused to obtain a 2048-dimensional multimodal vector f multi . After passing through two fully connected layers (before the end of VAFNet and before classification), the softmax function is used to convert the output into the probability distribution of each emotion category, and the emotion classification result is output.
[0040] In the emotion classification task, the balanced cross-entropy loss is commonly used to measure the difference between the emotion category probability distribution and the true distribution. Because in the emotion recognition task, the same model has different recognition abilities for different emotions. For example, it is difficult to recognize neutral samples. The calculation formula of the balanced cross-entropy loss is as follows:
[0041]
[0042] Among them, C is the number of categories, p c is the probability that the sample belongs to category c, and y c is the true category label of the sample. Only when it corresponds to the true category, y c = 1, and for the remaining categories, y c= 0. γ is a hyperparameter used to adjust the attention to difficult samples, (1 - p c ) γ to make low-probability samples obtain greater loss weighting. To further improve the coordination between audio and video modalities, the present invention introduces a modality consistency loss to enhance the recognition effect. Since the Mel features of the audio modality have higher expressive power, the Mel feature f mel and the video feature f video are used to calculate the modality consistency loss:
[0043]
[0044] In summary, combining the classification loss and the modality consistency loss, the total loss of the model for a single sample (each sample contains the processed video frame sequence, Mel spectrogram, and processed audio signal) is:
[0045] L = L CE + βL modality
[0046] The parameter β is used to balance the classification loss and the modality consistency loss.
[0047] The present invention also discloses a multi-modal emotion recognition system based on VAFNet for performing the above method, including the following modules:
[0048] Peak frame screening algorithm module: Preprocess the video dataset. Each video is sampled after frame division, face cropping, and grayscale conversion, and then converted into a peak frame sequence;
[0049] Mel spectrogram acquisition module: Extract the audio signal in the video as a one-dimensional time series signal and perform preprocessing to obtain the Mel spectrogram corresponding to each audio;
[0050] Feature vector extraction module: First input the obtained video peak frame sequence into the pre-trained ResNet18 model to extract spatial features to obtain facial spatial features and then integrate the time information through one-dimensional convolution to obtain the video feature vector; Input the Mel spectrogram and the processed audio signal obtained in step S2 into the OCANet model and the WaveNet respectively to extract spectral features and original waveform features, and obtain the Mel feature vector f mel and the waveform feature vector f wav ; Emotion classification module: Input the extracted Mel feature vector f mel , the waveform feature vector f wav and the facial spatial features into the feature fusion model based on the interactive attention mechanism to obtain the interactive feature vector f interact , and the Mel feature vector f mel, the video feature vector f video and the interaction feature vector f interact are concatenated and fused to obtain a multi-modal feature vector, and then emotion classification is performed to output the emotion classification result.
[0051] Compared with the prior art, the present invention has the following remarkable technical effects:
[0052] Firstly, the present invention improves the video frame sampling efficiency, screens a small number of effective facial emotion peak frames as the model input to reduce the complexity of the model for video signals. Then, the improved Orthogonal Channel Attention (OCA) mechanism is applied to enhance the spectrogram features, combined with facial emotion features, Mel spectrogram features, and original audio features, and the Interactive Attention mechanism is used for multi-modal feature fusion to enhance the information interaction between modalities. The present invention also proposes a new loss function, combining cross-entropy loss and modality consistency loss, which can enhance the feature consistency between different modalities while optimizing the classification performance, thereby improving the robustness of emotion recognition. Description of the Drawings
[0053] Figure 1 is the preprocessing flowchart of step S1 in the preferred embodiment of the present invention;
[0054] Figure 2 is the visualization diagram of facial key points;
[0055] Figure 3 is the VAFNet model diagram proposed in the preferred embodiment of the present invention;
[0056] Figure 4 is the OCA orthogonal channel attention mechanism diagram;
[0057] Figure 5 is the ResNet18 model structure diagram;
[0058] Figure 6 is the interactive attention mechanism model diagram;
[0059] Figure 7 is the block diagram of a multi-modal emotion recognition system based on VAFNet in the preferred embodiment of the present invention. Detailed Embodiments
[0060] The technical solutions of the present invention will be further explained and illustrated through the following preferred embodiments.
[0061] As Figures 1-4 shown, a multi-modal emotion recognition method based on VAFNet in this embodiment includes the following steps:
[0062] S1. For the video datasets CREMA-D and RAVDESS, use Dlib to frame the videos, crop the faces, grayscale them, and extract key points. After preprocessing, the size of the video frames is 224×224, the number of key points extracted is 68, and the visualization of the facial key points is as shown in Figure 2 . Among them, the number of key points K participating in the screening of facial emotion peak frames by PLAS is 51, numbered from 18 to 68, excluding the facial contour key points. The video frame rate is 60fps, and the number of frames screened for the video M = 20. After the PLAS algorithm, each video sample is converted into a peak frame sequence.
[0063] S2. Extract the audio signal in the video as a one-dimensional time series signal, perform pre-emphasis, frame it, and normalize the amplitude to [-1,1] to obtain the processed signal a[n]. Apply the short-time Fourier transform (STFT) to a[n], and then obtain the Mel spectrogram that conforms to the human ear's auditory perception through Mel filtering and logarithmic transformation.
[0064] S3. First, input the video peak frame sequence obtained in step S1 into the pre-trained ResNet18, load the pre-trained weights on the FER dataset, and use this as the initial weight for fine-tuning to achieve the purpose of accelerating convergence. The structure of ResNet18 is as shown in Figure 5 . The model extracts spatial features to obtain facial spatial features . Then, integrate the time information through one-dimensional convolution to obtain the video feature vector. At the same time, input the Mel spectrogram obtained in step S2 into OCANet. Compared with the traditional channel attention mechanism, the orthogonal attention mechanism OCA generates more accurate channel weights and outperforms other attention mechanisms such as SE, CBMA, and ECA in terms of performance. The specific algorithm process is as follows:
[0065] Assume that the size of the input feature X is (H, W, C), where C is the number of channels. OCA first makes the convolution kernels orthogonal to each other by using random orthogonalization. The orthogonal convolution kernel is K h,w,c After convolving each channel of X, apply global average pooling to obtain the channel weight feature vector z(c):
[0066]
[0067] Among them, x h,w,c represents the element at a specific position of the feature X. Then, apply the squeeze-and-excitation operation to the vector z(c) to recalibrate the channel weights and obtain the final attention weight vector z s (c):
[0068] z s (c) = σ(W2δ(W1z s (c)))
[0069] Among them, W1 and W2 represent the learnable weight matrices of z(c) passing through two fully connected layers, and σ and δ represent the sigmoid and ReLU activation functions respectively. Then, z s (c) is multiplied by the original input feature X channel by channel to obtain the output feature X ′ . The specific details of OCANet and the residual network structure are as Figure 4 shown. The Mel spectrogram and the processed audio signal a[n] in step S2 are respectively input into the OCANe model and the WaveNet to extract spectral features and waveform features, obtaining the Mel feature vector f mel and the waveform feature vector f wav ;
[0070] S4. For the extracted audio features f mel , f wav and the video facial spatial features are input into the feature fusion model based on the interactive attention mechanism to obtain the interactive feature vector f interact . The spectrogram feature f mel , the video feature f video , and the interactive feature f interact are then concatenated and fused to obtain the multi-modal feature vector, and emotion classification is performed to output the emotion classification result.
[0071] As Figure 5 shown, this embodiment discloses a multi-modal emotion recognition system based on VAFNet for executing the above method, including the following modules:
[0072] Peak frame screening algorithm module: Preprocess the video dataset. Each video is sampled using the PLAS algorithm after frame division, face cropping, and grayscale conversion, and then converted into a peak frame sequence;
[0073] Mel spectrogram acquisition module: Extract the audio signal in the video as a one-dimensional time series signal and perform preprocessing to obtain the Mel spectrogram corresponding to each audio;
[0074] Feature vector extraction module: First input the obtained video peak frame sequence into the pre-trained ResNet18 model to extract spatial features to obtain the facial spatial features , and then integrate the time information through one-dimensional convolution to obtain the video feature vector; Input the Mel spectrogram and the processed audio signal obtained in step S2 into the OCANet model and the WaveNet respectively to extract spectral features and original waveform features, obtaining the Mel feature vector f mel and the waveform feature vector f wav ;
[0075] Emotion classification module: For the extracted Mel feature vector f mel, waveform feature vector f wav and facial spatial features are input into the feature fusion model based on the interactive attention mechanism to obtain the interactive feature vector f interact . The Mel feature vector f mel , video feature vector f video and interactive feature vector f interact are concatenated and fused to obtain a multi-modal feature vector, and emotion classification is performed to output the emotion classification result.
[0076] For other contents of this embodiment, reference may be made to the above method embodiment.
[0077] The present invention realizes multi-modal emotion recognition based on audio and video, improves the recognition accuracy while reducing the computational cost.
[0078] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A multimodal emotion recognition method based on VAFNet, characterized in that, It includes the following steps: S1. After preprocessing the video dataset, sample it and then convert it into a peak frame sequence; S2. Extract the audio signal in the video as a one-dimensional time series signal and perform preprocessing to obtain the Mel spectrogram corresponding to each audio; S3. Input the video peak frame sequence obtained in step S1 into the pre-trained ResNet18 model to extract spatial features to obtain facial spatial features Denote the facial spatial features of the M-th frame image, and then integrate the temporal information through one-dimensional convolution to obtain a video feature vector; at the same time, input the Mel spectrogram obtained in step S2 and the audio signal processed in step S2 into the OCANe model and the WaveNet respectively to extract spectral features and waveform features, and obtain a Mel feature vector f mel and a waveform feature vector f wav ; S4. For the extracted Mel feature vector f mel , waveform feature vector f wav and facial spatial features , input them into the feature fusion model based on the interactive attention mechanism to obtain the interactive feature vector f interact . Concatenate and fuse the Mel feature vector f mel , video feature vector f video and interactive feature vector f interact to obtain a multi-modal feature vector, perform emotion classification, and output the emotion classification result.
2. The multimodal emotion recognition method according to claim 1, wherein In step S1, the videos in the video dataset are frame-divided, face-cropped, and grayscaled to obtain a video frame sequence {I1, I2, …, I N}, facial key point markers are extracted for each frame to obtain a set of facial key point coordinates for each frame. Denote the set of key point coordinates of the nth frame as:
3. The multimodal emotion recognition method according to claim 2, wherein In step S1, the PLAS algorithm is used for sampling. For the PLAS algorithm, first, the key point coordinate set P n is centralized to obtain P' n : The centralized key point coordinate set P' n is normalized to obtain P'' n : where scale n is the scaling factor of the nth frame. After completing the translation and scaling operations, the key point sets P″1 and P″ of the first frame and the nth frame are used n and the singular value decomposition method to calculate the optimal alignment rotation matrix R between the key point sets of the first frame and the nth frame: H = (P″1) T P″ n = U∑V T The covariance matrix H is the similarity transformation matrix between the first frame and the nth frame. The rotation matrix R is a second-order orthogonal matrix, and r 11 , r 22 represent the rotation amounts in the x-axis direction and y-axis direction respectively, and r 12 , r 21 represent the influence of the rotation amount in the y-axis direction on the x-axis direction and the influence of the rotation amount in the x-axis direction on the y-axis direction respectively. The spatial similarity s[1, n] of the key points between the first frame and the nth frame is obtained as follows: Among them, n = 2, 3, 4, …, N. Sort the values of s[1, 2], s[1, 3], …, s[1, N] in ascending order, select the first M frames with the smallest similarity as peak frames, and arrange them in chronological order to obtain a peak frame sequence {I1, I2, …, I M}.
4. The multimodal emotion recognition method according to claim 3, characterized in that In step S2, the preprocessing includes using endpoint activity detection to extract the effective sound segment, and then pre-emphasizing, framing, amplitude normalizing, and windowing the audio signal to obtain the preprocessed signal a[n]. After performing short-time Fourier transform on the signal a[n], it is passed through Mel filtering and logarithmic transformation to obtain the Mel spectrogram.
5. The multimodal emotion recognition method according to claim 3 or 4, characterized in that In step S3: The video peak frame sequence {I1, I2, …, I M} is first passed through the ResNet18 model pre-trained on the FER dataset to extract facial spatial features. After each frame extracts spatial features through subsequent layers, the feature Then, the F s is input into the 1D CNN model to extract temporal information to obtain the video feature vector f video .
6. The multimodal emotion recognition method according to claim 5, characterized in that In step S3: Input the audio signal a[n] into WaveNet to extract the waveform feature vector f of the original waveform wav , WaveNet extracts temporal information through causal dilated convolutions; input the Mel spectrogram obtained in step S2 into the audio feature extraction model OCANet, and after extraction by OCANet, the Mel spectrogram obtains the Mel feature vector f mel .
7. The multimodal emotion recognition method according to claim 6, characterized in that Step S4 is specifically as follows: The spatial features of the unintegrated time information extracted in step 3 are denoted as the feature matrix F s , and F s is input together with the audio features [f mel , f wav extracted in step 3 into the feature fusion model of the interactive attention mechanism, and the Query, Key, and Value for training are constructed as follows: Q audio = [f mel , f wav W Q , K video = F s W K V video = F s W V Among them, W Q , W K , W V are all weight matrices participating in the training. The similarity weight matrix A represents the attention weight of audio features to each frame of video features. The weight matrix A and the interaction feature vector f interact are in sequence as follows: f interact = flatten(AV video ) Among them, flatten represents the feature flattening operation, AV vidro The dimension is (2, 512) and contains two parts of information, namely the attention degrees of Mel features and wav features to video frames; the spectrogram features f mel , video features f video , interaction features f interact are concatenated and fused to obtain a 2048-dimensional multimodal vector f multi . After passing through two fully connected layers, the softmax function is used to convert the output into the probability distribution of each emotion category, and the emotion classification result is output.
8. The multimodal emotion recognition method according to claim 7, wherein: In the model classification task, the balanced cross-entropy loss is used to measure the difference between the class probability distribution and the true distribution. The calculation formula of the balanced cross-entropy loss is as follows: where C is the number of classes, and p c is the probability that the sample belongs to class c, y c is the true class label of the sample, and γ is a hyperparameter.
9. The multimodal emotion recognition method according to claim 8, wherein Introduce the modal consistency loss and select the Mel feature f mel and the video feature f video to calculate the modal consistency loss: Combining the classification loss and the modality consistency loss, the total loss for a single sample is: L = L CE + βL modality Among them, the parameter β is used to balance the classification loss and the modality consistency loss.
10. A multimodal emotion recognition system based on VAFNet for performing the method according to any one of claims 1-9, characterized in that, It includes the following modules: Peak frame screening algorithm module: Preprocess the video dataset. Each video is sampled after being framed, face cropped, and grayscaled, and then converted into a peak frame sequence; Mel spectrogram obtaining module: Extract the audio signal in the video as a one-dimensional time series signal and perform preprocessing to obtain the Mel spectrogram corresponding to each audio; Feature vector extraction module: First, input the obtained video peak frame sequence into the pre-trained ResNet18 model to extract spatial features and obtain facial spatial features Then, integrate temporal information through one-dimensional convolution to obtain video feature vectors; Input the Mel spectrogram and the processed audio signal obtained in step S2 into the OCANet model and the WaveNet respectively to extract spectral features and original waveform features, and obtain the Mel feature vector f mel and the waveform feature vector f wav ; Emotion classification module: For the extracted Mel feature vector f mel , waveform feature vector f wav and facial spatial features are input into the feature fusion model based on the interactive attention mechanism to obtain the interactive feature vector f interact . The Mel feature vector f mel , video feature vector f video and the interactive feature vector f interact are concatenated and fused to obtain a multi-modal feature vector for emotion classification, and the emotion classification result is output.