An Audio-Visual Emotion Classification Method Based on Improved ConvMixer Network and Dynamic Focal Loss
Through the improved ConvMixer network and ResNet34 network combined with dynamic focus loss function, the problem of insufficient local feature extraction and low recognition rate of difficult-to-separate samples in audio-visual emotion classification is solved, and better feature extraction and classification effects are achieved.
Patent Information
- Application Number
- CN202211015781.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-08-24
AI Technical Summary
In the existing audio-visual emotion classification methods, dual-modal emotion recognition has the problem of extracting local features of video images while ignoring global features, simple dual-modal fusion method, and difficulty in separating samples from loss functions.
The improved ConvMixer network is used to extract visual features in combination with the adjacency matrix, and the ResNet34 network extracts auditory features, and features fusion is performed through the cross-modal time attention module. The model is optimized using the dynamic focus loss function to improve the recognition rate of difficult-to-separate samples.
Effectively extract the global and local spatial temporal features of the image sequence, capture the cross-modal temporal correlation, and improve the generalization ability of the model and the recognition rate of difficult samples.
Smart Images

Figure CN115346261B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of audio-visual emotion classification, and specifically relates to an audio-visual emotion classification method based on an improved ConvMixer network and dynamic focal loss. Background Art
[0002] With the popularization of the Internet and computers, human-computer interaction behaviors occur more frequently. The audio-visual emotion classification method measures and analyzes the external manifestations of humans and calculates the influence on emotions. Applied in modern human-computer interaction systems, it can make the interaction mode more natural and friendly, and improve the human-computer interaction experience.
[0003] The audio-visual emotion classification method based on deep learning does not require manual feature extraction based on professional knowledge, and also shows better performance than traditional methods, achieving good results in audio-visual emotion classification.
[0004] In the paper "End-to-End Multimodal Emotion Recognition Using Deep Neural Networks" published by Trigeorgis et al. in the IEEE Journal of Selected Topics in Signal Processing in 2017, one-dimensional and two-dimensional Convolutional Neural Networks (CNNs) were used to extract features from audio signals and video signals respectively. After concatenating the audio and video features, they were input into a Recurrent Neural Network (RNN) for emotion analysis. However, the simple concatenation of audio and video features could not fuse the temporal correlation between the two, resulting in the RNN being unable to fully utilize its own advantages and characteristics in extracting temporal features. In the paper "Learning Affective Features with a Hybrid Deep Model for Audio-Visual Emotion Recognition" published by Zhang et al. in the IEEE Transactions on Circuits and Systems for Video Technology in 2018, a two-stream network was constructed using three-dimensional CNN and two-dimensional CNN to extract features from the video frame sequences and audio spectrograms segmented by a certain time length. The corresponding video and audio features for each segment were fused through a Deep Belief Nets (DBN), and finally, the features for all time periods were averaged through average pooling to obtain global features for classification. However, the video and audio features corresponding to all time periods were extracted through the same two-stream network, making it impossible to perform targeted feature extraction for each time period. Moreover, using average pooling to fuse the features for all time periods could not prominently reflect the temporal information of the original video.
[0005] CN114724224A discloses a multi-modal emotion recognition method for a medical care robot, which extracts facial expression self-attention emotion features and action self-attention emotion features according to the video information, and extracts speech self-attention emotion features and text self-attention emotion features according to the audio information; the four self-attention emotion features are fused based on a mutual attention mechanism to obtain complete multi-modal emotion features; the multi-modal emotion features are used to extract context emotion features based on a graph convolutional neural network to obtain multi-modal emotion features containing context information; the multi-modal emotion features containing context information are used for emotion classification recognition to obtain an emotion label result. In this method, there is an operation of first concatenating the four self-attention emotion features into a multi-modal feature and then performing mutual attention mechanism emotion feature fusion. The concatenation operation will increase the computational complexity of the subsequent mutual attention mechanism, and the operation of concatenating the four features to increase the feature dimension will ignore the time correlation of the four modalities. CN114582372A discloses a multi-modal emotion feature recognition method for driver emotion judgment, which inputs visual information and audio information into a visual facial expression feature recognition model and a speech emotion feature recognition model respectively to obtain a visual feature vector and a speech feature vector, and inputs the two features into a bi-modal emotion feature recognition model to obtain an emotion recognition result of decision-level fusion. The speech feature vector extracted in this method is a statistical feature extracted based on the mel cepstrum coefficient spectrogram of the audio, ignoring the energy change information related to emotion in the audio. CN113989893A discloses a child emotion recognition algorithm based on facial expression and speech bi-modalities, which constructs a semantic feature space using the emotion label information of speech features and facial expression features, extracts local features and global features of audio and video through a multi-scale feature extraction method, projects the local features and global features of audio and video into the semantic feature space, and the semantic feature space selects important features that contribute to emotion classification for emotion judgment and recognition. In this method, there is an operation of performing object detection and precise positioning on the video feature map, which is vulnerable to the influence of the face deflection angle. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the technical problem to be solved by the present invention is: to provide an audiovisual emotion classification method based on an improved ConvMixer network and dynamic focal loss. The visual features are extracted from the image sequence by the ConvMixer network combined with the adjacency matrix, the ResNet34 network extracts the auditory features from the mel spectrogram, the cross-modal temporal attention module of the feature fusion and classification network fuses the visual features and auditory features, and the fused features are used to judge the emotion category of the video; the network model is trained by the Focal Loss function that fuses dynamic weights, and the model parameters are optimized to improve the recognition rate of difficult-to-separate samples. The present invention overcomes the problems of the existing audiovisual bimodal emotion recognition method that focuses on extracting local features of the video frame while ignoring global features, the bimodal fusion method is simple, and the loss function cannot make the model focus on difficult-to-separate samples.
[0007] The technical solution adopted by the present invention to solve this technical problem is as follows:
[0008] An audiovisual emotion classification method based on an improved ConvMixer network and dynamic focal loss, comprising the following steps:
[0009] In the first step, collect videos involving the human face area expressing emotions, extract the image sequence and audio signal from the videos, and convert the audio signal into a mel spectrogram;
[0010] In the second step, construct a ConvMixer network combined with an adjacency matrix, including three parts of operations, namely, block embedding operation, Layer module operation, and average pooling operation; input the image sequence into the ConvMixer network combined with the adjacency matrix to extract visual features, and obtain the feature map F;
[0011] Step (2.1), block embedding operation:
[0012] The image sequence is subjected to block embedding operation through a convolutional layer, an activation function layer, and a normalization layer in sequence to obtain the feature map F output by the block embedding operation 2.1 ;
[0013] Step (2.2), Layer module operation, including four cascaded Layer modules;
[0014] Input the feature map F 2.1 into the first Layer module, and the first Layer module constructs a two-dimensional spatial coordinate matrix of each image block in the feature map F 2.1 according to the spatial size of the feature map F 2.1 and copies and splices the two-dimensional spatial coordinate matrix according to the temporal size of the feature map F 2.1 to obtain a matrix that is the same size as the feature map F 2.1Spatial position encoding with the same size; the feature map F 2.1 is concatenated with the spatial position encoding, and then passes through a linear layer to obtain the feature map According to the feature map F 2.1 a spatial adjacency matrix is randomly generated according to the spatial size of, and the feature map is multiplied by the spatial adjacency matrix, and then passes through an activation function layer and a normalization layer to obtain the feature map F s ; the feature map is superimposed with F s to obtain the feature map F s ';
[0015] According to the temporal size of the feature map F s ', a one-dimensional temporal coordinate matrix of each image patch in the feature map F s ' is constructed. According to the spatial size of the feature map F s ', the one-dimensional temporal coordinate matrix is copied and concatenated to obtain a temporal position encoding with the same size as the feature map F s '; the feature map F s ' is concatenated with the temporal position encoding, and then passes through a linear layer to obtain the feature map According to the temporal size of the feature map F s ', a temporal adjacency matrix is randomly generated, and the feature map is multiplied by the temporal adjacency matrix, and then passes through an activation function layer and a normalization layer to obtain the feature map F t ; the feature map is superimposed with F t to obtain the feature map F t '; the feature map F t ' sequentially passes through a pointwise convolutional layer, an activation function layer, and a normalization layer to obtain the feature map output by the first Layer module;
[0016] Step (2.3), average pooling operation:
[0017] The feature map output by the fourth Layer module is subjected to spatial dimension average pooling operation through an average pooling layer to obtain the feature map F;
[0018] In the third step, the ResNet34 network is used to extract auditory features from the mel cepstrum coefficient spectrogram to obtain the feature map M;
[0019] In the fourth step, a feature fusion and classification network is constructed to fuse visual features and auditory features, and perform sentiment classification on each video according to the fused features; the feature fusion and classification network includes two cross-modal temporal attention modules, pooling and concatenation operations, and classification operations;
[0020] Step (4.1), the first cross-modal temporal attention module:
[0021] Input the feature map F and the feature map M into the first cross-modal temporal attention module. The feature map F passes through a linear layer and a normalization layer to obtain the feature Q1; the feature map M passes through two independent linear layers and two independent normalization layers to obtain the features K1 and V1 respectively.
[0022] Generate a learnable intermediate matrix LIM1 according to the temporal dimension sizes of the feature Q1 and K1, and perform random parameter initialization on the intermediate matrix LIM1; multiply the feature Q1 with the transpose of the initialized intermediate matrix LIM1 and the feature K1, then divide by the square root of the number of channels of the feature K1, and then input it into the softmax layer to obtain the normalized weight; multiply the normalized weight with the feature V1 and then add it to the feature Q1 to obtain the cross-modal attention feature F based on the image sequence. att ; The cross-modal attention feature F based on the image sequence att Pass through a pointwise convolutional layer to obtain the cross-modal feature F based on the image sequence. cm ;
[0023] Step (4.2), the second cross-modal temporal attention module:
[0024] Input the feature map F and the feature map M into the second cross-modal temporal attention module; the feature map M passes through a linear layer and a normalization layer to obtain the feature Q2; the feature map F passes through two independent linear layers and two independent normalization layers to obtain the features K2 and V2.
[0025] Generate a learnable intermediate matrix LIM2 according to the temporal dimension sizes of the feature Q2 and K2, and perform random parameter initialization on the intermediate matrix LIM2; multiply the feature Q2 with the intermediate matrix LIM2 and the transpose of the feature K2, then divide by the square root of the number of channels of the feature K2, and then input it into the softmax layer to obtain the normalized weight; multiply the normalized weight with the feature V2 and then add it to the feature Q2 to obtain the cross-modal attention feature M based on the mel cepstrum coefficient spectrogram. att ; The cross-modal attention feature M based on the mel cepstrum coefficient spectrogram att Pass through a pointwise convolutional layer to obtain the cross-modal feature M based on the mel cepstrum coefficient spectrogram. cm ;
[0026] Step (4.3), pooling and concatenation operations:
[0027] Perform average pooling on the cross-modal feature F based on the image sequence cm and the cross-modal feature M based on the mel cepstrum coefficient spectrogram cm respectively, and then perform concatenation to obtain the feature f. FM ;
[0028] Step (4.4), classification operation:
[0029] Input the feature f FM into the linear layer and then pass it through the softmax layer to obtain the predicted probability distribution P{Y1, Y2,..., Y i ,..., Y q} for E types of emotion categories. Y i represents the predicted probability distribution of the i-th video for E types of emotion categories, denoted as Y i {y i1 ,…, y ie ,…, y iE}, where y ie represents the predicted probability that the i-th video belongs to the e-th emotion category, and q represents the number of videos;
[0030] In the fifth step, train the ConvMixer network, ResNet34 network, feature fusion and classification network combined with the adjacency matrix, and calculate the training loss through the focal loss function that fuses dynamic weights; use the trained ConvMixer network combined with the adjacency matrix to extract visual features from the image sequence, use the trained ResNet34 network to extract auditory features from the mel spectrogram, then fuse the visual features and auditory features through the trained feature fusion and classification network, and perform emotion classification based on the fused features to predict the corresponding emotion category of the video.
[0031] Furthermore, the focal loss function that fuses dynamic weights is:
[0032]
[0033] In formula (21), α represents the dynamic weight, and log represents the logarithmic function with base 2. represents the predicted probability distribution of the true emotion category to which the i-th video belongs;
[0034] The calculation formula for the dynamic weight α is as follows:
[0035]
[0036] In formula (22), C represents the confusion matrix of the previous training cycle. represents the number of times that the video with the true emotion category is predicted as category t in the previous training cycle. represents the number of videos with the true emotion category in the previous training cycle. represents predicting the video with the true emotion category as category The number of times.
[0037] Furthermore, the construction process of the confusion matrix C is as follows: Before the start of each training cycle, a zero matrix of size E×E is generated, where the number of rows and columns ranges from 0 to E-1; According to the prediction of each sample during training, when a sample with a true sentiment category of 0 is predicted as category 2, the element at the 0th row and 2nd column of the matrix is incremented by 1, and the same applies to the other samples, updating the zero matrix to obtain the confusion matrix C; When the training cycle is completed, the confusion matrix of the current cycle is used to calculate the dynamic weight of the loss function in the next training cycle.
[0038] Compared with the prior art, the outstanding substantive features and significant progress of the present invention are as follows:
[0039] (1) The ConvMixer (Adjacent Matrix-based ConvMixer, hereinafter referred to as ACAM) proposed in the present invention combines an adjacency matrix. The operation of the Layer module therein expands the receptive field of the network by virtue of the spatial adjacency matrix and the temporal adjacency matrix, enabling the network to extract global and local spatial and temporal features of the image sequence simultaneously. In order to capture the temporal correlation of the emotional changes in the image sequence and the mel spectrogram, the Cross-Modal Temporal Attention Module (hereinafter referred to as CMTAM) proposed in the present invention connects the temporal features of different time scales of video and audio through a learnable intermediate matrix. CMTAM learns the ability to capture cross-modal temporal correlation through the intermediate matrix. The Focal Loss function with fused dynamic weights proposed in the present invention can adjust the loss value of training samples through dynamic weights, enabling the model to pay more attention to misclassified samples during the optimization process and improving the generalization ability of the model.
[0040] (2) CN114694076A discloses a multi-modal sentiment analysis method based on multi-task learning and cascaded cross-modal fusion, which splices and fuses the extracted single-modal features and uses the cross-entropy loss function as the optimization objective of the model. Compared with CN114694076A, the present invention captures cross-modal temporal correlation through the cross-modal temporal attention module with a learnable intermediate matrix, extracts cross-modal temporal attention features, and improves the recognition effect of the model.
[0041] (3)CN114582372A discloses a multi-modal emotion feature recognition method for driver emotion judgment. The visual feature vector and the speech feature vector are input into a bimodal emotion feature recognition model to obtain an emotion recognition result of decision-level fusion. Compared with CN114582372A, the present invention adopts an end-to-end training method, and the feature extraction networks of audio and video are trained simultaneously, saving the model training time; moreover, the present invention not only extracts the spatial information of the video image sequence, but also extracts the time features, ensuring the applicability of the algorithm in different scenarios.
[0042] (4) The method of the present invention adopts the idea of deep learning. Traditional detection methods only extract low-level spatial features frame by frame from the video image sequence and cannot extract the time features of the sequence. Deep learning can extract high-level semantic features and can better express the images. Description of the Drawings
[0043] Figure 1 is the flowchart of the training stage of the present invention;
[0044] Figure 2 is the flowchart of the classification stage of the present invention;
[0045] Figure 3 is a schematic diagram of the block embedding operation in the process of constructing the ConvMixer network combined with the adjacency matrix of the present invention;
[0046] Figure 4 is a schematic diagram of the Layer module operation and average pooling operation in the process of constructing the ConvMixer network combined with the adjacency matrix of the present invention;
[0047] Figure 5 is a schematic diagram of the shallow feature extraction operation in the ResNet34 network of the present invention;
[0048] Figure 6 is a schematic diagram of the first to third residual modules, fifth to seventh residual modules, ninth to thirteenth residual modules, and fifteenth to sixteenth residual modules in the ResNet34 network;
[0049] Figure 7 is a schematic diagram of the fourth, eighth, and fourteenth residual modules in the ResNet34 network;
[0050] Figure 8 is a schematic diagram of the first cross-modal temporal attention module in the process of constructing the feature fusion and classification network;
[0051] Figure 9 is a schematic diagram of the second cross-modal temporal attention module in the process of constructing the feature fusion and classification network;
[0052] Figure 10It is a schematic diagram of pooling, concatenation, and classification operations in the process of constructing a feature fusion and classification network;
[0053] Figure 11 It is a schematic diagram of the present invention for constructing a Focal Loss loss function that combines dynamic weights. Specific implementation manners
[0054] The technical solution of the present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation manners, but the protection scope of this application is not limited thereby.
[0055] The present invention is an audiovisual emotion classification method based on an improved ConvMixer network and dynamic focal loss (hereinafter referred to as the method, see Figures 1 to 11 ). The ConvMixer network combined with an adjacency matrix is used to extract visual features from an image sequence, and the ResNet34 network is used to extract auditory features from an audio mel spectrogram. The audiovisual features are fused through a cross-modal temporal attention module, and a Focal Loss loss function with dynamic weights is obtained in combination with a confusion matrix as the optimization objective function of the network model. The specific steps are as follows:
[0056] In the first step, videos involving the human face area expressing emotions are collected, an image sequence and an audio signal are extracted from the videos, and the audio signal is converted into a mel spectrogram.
[0057] In step (1.1), images are extracted from the videos, and the videos are converted into an image sequence.
[0058] A group of video sequences consists of multiple videos. OpenCV software is used to extract images from the videos. The images extracted from each video form an image sequence. Therefore, the image dataset is a collection of multiple image sequences, denoted as T{V1, V2,..., V i ,..., V q}, where V i represents the image sequence corresponding to the i-th video, and q represents the number of videos; each image sequence contains N frames of images. For example, N frames of images are extracted from the i-th video, that is, the i-th image sequence is represented as V i {v i1 , v i2 ,..., v id ,..., v iN}, N is 64, and v id represents the d-th frame of the image sequence extracted from the i-th video. The obtained image sequence is normalized, and the size of each frame of the image is adjusted to 112×112 pixels. Therefore, the size of each image sequence is 64×112×112;
[0059] Step (1.2): Separate the audio signal from the video and convert the audio signal into a mel spectrogram;
[0060] Use Librosa software to separate the audio signal from the video and extract a mel spectrogram with 32 frequency domain features; the mel spectrogram corresponding to the \(i\)-th video is denoted as \(A\) i \(\{a\) i1 , a\) i2 , \(\cdots, a\) id , \(\cdots, a\) iN \}\), where \(a\) id represents the mel coefficient of the \(d\)-th time segment of the mel spectrogram extracted from the \(i\)-th video, and the set of mel spectrograms corresponding to the entire dataset is \(M = \{A_1, A_2, \cdots, A\) i , \(\cdots, A\) q \}\);
[0061] Step 2: Construct a ConvMixer network combined with an adjacency matrix, including three operations: patch embedding operation, Layer module operation, and average pooling operation; input the image sequence into the ConvMixer network combined with the adjacency matrix to extract visual features and obtain the feature map \(F\);
[0062] Step (2.1): Patch embedding operation:
[0063] Apply the patch embedding operation to the image sequence obtained in Step (1.1) with a size of \(64\times112\times112\) and 3 channels by passing it through a convolutional layer, an activation function layer, and a normalization layer in sequence, to obtain a feature map \(F\) with a size of \(16\times16\times16\) and 512 channels 2.1 , see Figure 3 ; the patch embedding operation is shown in Equation (1):
[0064]
[0065] In Equation (1), \(F\) in represents the input of the patch embedding operation, represents a convolutional layer with a stride and kernel size of \(4\times7\times7\), \(c\) in and \(h\) are the number of input channels and output channels of the convolutional layer respectively, GELU represents the activation function layer, and BN represents the normalization layer;
[0066] Step (2.2): Layer module operation, including four cascaded Layer modules;
[0067] Input the feature map \(F\) 2.1 obtained in Step (2.1) above into the first Layer module, and the first Layer module processes the feature map \(F\) 2.1Construct the feature map F according to the spatial dimension size 2.1 The two-dimensional spatial coordinate matrix of each image patch in 2.1 According to the time dimension size of the feature map F, copy and splice the two-dimensional spatial coordinate matrix to obtain a spatial position encoding with the same size as the feature map F 2.1 Splice the feature map F 2.1 With the spatial position encoding, and then pass through a linear layer to obtain the feature map According to the feature map F 2.1 Randomly generate a spatial adjacency matrix according to the spatial dimension size, multiply the feature map With the spatial adjacency matrix, and then pass through an activation function layer and a normalization layer to obtain the feature map F s ; Splice the feature map With F s To obtain the feature map F s ';
[0068] According to the feature map F s 's time dimension size, construct the one-dimensional time coordinate matrix of each image patch in the feature map F s ', according to the spatial dimension size of the feature map F s Copy and splice the one-dimensional time coordinate matrix to obtain a time position encoding with the same size as the feature map F s '; Splice the feature map F s ' with the time position encoding, and then pass through a linear layer to obtain the feature map According to the feature map F s 's time dimension size, randomly generate a time adjacency matrix, multiply the feature map With the time adjacency matrix, and then pass through an activation function layer and a normalization layer to obtain the feature map F t ; Splice the feature map With F t To obtain the feature map F t '; The feature map F t ' passes through a pointwise convolution layer, an activation function layer, and a normalization layer in sequence to obtain the feature map F 2.2.1 ; The feature map F 2.2.1 Is input into the second Layer module, and the feature map F 2.2.2 Is output; The feature map F 2.2.2 Is input into the third Layer module, and the feature map F 2.2.3 Is output; The feature map F 2.2.3 Is input into the fourth Layer module, and the feature map F 2.2.4 Is output, see Figure 4 ;
[0069] Each Layer module is shown in formulas (2) to (6):
[0070]
[0071] In formula (2), F0 represents the input of the Layer module, SPC represents the spatial position encoding, Concat represents the concatenation operation, and Linear represents the linear layer;
[0072]
[0073] In formula (3), SAM represents the spatial adjacency matrix;
[0074]
[0075] In formula (4), TPC represents the temporal position encoding;
[0076]
[0077] In formula (5), TAM represents the temporal adjacency matrix;
[0078] F out = BN(GELU(Conv pw (F t '))) (6)
[0080] In formula (6), Conv pw represents a pointwise convolutional layer with a kernel size of 1×1×1;
[0081] Step (2.3), average pooling operation:
[0082] Perform spatial dimensional average pooling operation on the feature map F with a size of 16×16×16 and 512 channels obtained in the above step (2.2) through the average pooling layer to obtain a feature map F with a size of 16×1×1 and 512 channels; 2.2.4 Thus, a ConvMixer network combined with an adjacency matrix is constructed;
[0083] Step 3, construct a ResNet34 network, including a shallow feature extraction layer and sixteen residual modules, and the output of the previous residual module is the input of the next residual module; input the mel cepstrum coefficient spectrogram into the ResNet34 network to extract auditory features and obtain a feature map M;
[0084] Step (3.1), shallow feature extraction operation:
[0085]
[0086] The mel cepstral coefficient spectrogram obtained in step (1.2) is input into the ResNet34 network. The size of the mel cepstral coefficient spectrogram is 32×590, and the number of channels is 1. The mel cepstral coefficient spectrogram is input into a shallow feature extraction layer composed of a convolutional layer, a normalization layer, and an activation function layer to obtain a feature map M with a size of 8×148 and 64 channels. 3.1 , see Figure 5 ; The shallow feature extraction operation is shown in formula (7):
[0087] M 3.1 = RELU(BN(Conv 1,64,7,2,3 (M in ))) (7)
[0089] In formula (7), M in represents the input of the ResNet34 network, and M 3.1 represents the output of the shallow feature extraction layer. Conv 3,64,7,2,3 represents a convolutional layer with 1 input channel, 64 output channels, a convolutional kernel size of 7, a stride of 2, and a padding size of 2 on the edges;
[0090] Step (3.2), deep feature extraction operation:
[0091] Input the feature map M obtained in step (3.1) above 3.1 into the first residual module. The output channels of the first to third residual modules are all 64, the output channels of the fourth to seventh residual modules are all 128, the output channels of the eighth to thirteenth residual modules are all 256, and the output channels of the fourteenth to sixteenth residual modules are all 512. The sixteenth residual module outputs a feature map M with a size of 1×19 and 512 channels;
[0092] The operations of the first to third residual modules, the fifth to seventh residual modules, the ninth to thirteenth residual modules, and the fifteenth to sixteenth residual modules are shown in formula (8):
[0093] M l = BN2(Conv2(RELU(BN1(Conv1(M l-1 ))))) + M l-1 (8)
[0095] In formula (8), M l-1 represents the input of the l-th residual module, and M l represents the output of the l-th residual module. Conv1 and Conv2 represent two independent convolutional layers with a convolutional kernel size of 3, a stride of 1, and a padding size of 1 on the edges. BN1 and BN2 represent two independent normalization layers;
[0096] The operations of the fourth, eighth, and fourteenth residual modules are shown in Equation (9):
[0097] M l = BN2(Conv4(RELU(BN1(Conv3(M l-1 )))))+BN3(Conv5(M l-1 )) (9)
[0099] In Equation (9), Conv3 represents a convolutional layer with a kernel size of 3, a stride of 2, and a padding size of 1, Conv4 represents a convolutional layer with a kernel size of 3, a stride of 1, and a padding size of 1, Conv5 represents a convolutional layer with a kernel size of 1, a stride of 2, and a padding size of 0, and BN1, BN2, and BN3 represent three independent normalization layers;
[0100] Step 4: Construct a feature fusion and classification network for fusing visual features and auditory features and performing sentiment classification on each video according to the fused features; the feature fusion and classification network includes two cross-modal temporal attention modules, pooling and concatenation operations, and classification operations;
[0101] Step (4.1): The first cross-modal temporal attention module:
[0102] Input the feature map F obtained in step (2.3) and the feature map M obtained in step (3.2) into the first cross-modal temporal attention module. The feature map F passes through a linear layer and a normalization layer to obtain the feature Q1; the feature map M passes through two independent linear layers and two independent normalization layers to obtain the features K1 and V1 respectively. The operations for obtaining the features Q1, K1, and V1 are shown in Equations (10), (11), and (12):
[0103] Q1 = BN(Linear(F)) (10)
[0105] K1 = BN(Linear(M)) (11)
[0107] V1 = BN(Linear(M)) (12)
[0109] Generate a learnable intermediate matrix LIM1 according to the temporal dimension sizes of the features Q1 and K1, and perform random parameter initialization on the intermediate matrix LIM1; multiply the feature Q1 by the initialized intermediate matrix LIM1 and the transpose of the feature K1, and then divide by the square root of the number of channels of the feature K1 Then, it enters the softmax layer to obtain the normalized weights. After multiplying the normalized weights by the feature V1 and then adding the result to the feature Q1, the cross-modal attention feature F based on the image sequence is obtained. att ; Calculate F att The operation is shown in Equation (13):
[0110]
[0111] In Equation (13), T represents matrix transpose;
[0112] The cross-modal attention feature F based on the image sequence att Passes through a pointwise convolutional layer to obtain the cross-modal feature F based on the image sequence cm , as shown in Figure 8 ; Calculate F cm The operation is shown in Equation (14):
[0113] F cm = Conv pw (F att ) (14)
[0115] In Equation (14), Conv pw represents a pointwise convolutional layer with a kernel size of 1;
[0116] Step (4.2), the second cross-modal temporal attention module:
[0117] Input the feature map F obtained in step (2.3) and the feature map M obtained in step (3.2) into the second cross-modal temporal attention module; the feature map M passes through a linear layer and a normalization layer to obtain the feature Q2; the feature map F passes through two independent linear layers and two independent normalization layers to obtain the features K2 and V2; the operations for calculating Q2, K2, and V2 are shown in Equations (15), (16), and (17):
[0118] Q2 = BN(Linear(M)) (15)
[0120] K2 = BN(Linear(F)) (16)
[0122] V2 = BN(Linear(F)) (17)
[0124] Generate a learnable intermediate matrix LIM2 based on the temporal dimension sizes of the features Q2 and K2, and perform random parameter initialization on the intermediate matrix LIM2; multiply the feature Q2 by the intermediate matrix LIM2 and the transpose of the feature K2, and then divide by the square root of the number of channels of the feature K2 Then, it enters the softmax layer to obtain the normalized weights; multiply the normalized weights by the feature V2 and then add the result to the feature Q2 to obtain the cross-modal attention feature M based on the mel-frequency cepstral coefficient spectrogram att ; Calculate M att The operation is shown in Equation (18):
[0125]
[0126] The cross-modal attention feature M based on the mel-frequency cepstral coefficient spectrogram att passes through a pointwise convolutional layer to obtain the cross-modal feature M based on the mel-frequency cepstral coefficient spectrogram cm , see Figure 9 ; Calculate M cm The operation is shown in Equation (19):
[0127] M cm = Conv pw (M att ) (19)
[0129] Step (4.3), pooling and concatenation operations:
[0130] The cross-modal feature F based on the image sequence obtained in step (4.1) cm and the cross-modal feature M based on the mel-frequency cepstral coefficient spectrogram obtained in step (4.2) cm are respectively subjected to average pooling and then concatenated to obtain a feature f of size 1×1 and channel number 1024 FM ; The pooling and concatenation operations are shown in Equation (20):
[0131] f FM = Concat(AvgPool(F cm ), AvgPool(M cm )) (20)
[0133] In Equation (20), AvgPool represents the average pooling operation;
[0134] Step (4.4), classification operation:
[0135] The feature f obtained in step (4.3) FM is input into a linear layer and then passes through a softmax layer to obtain the predicted probability distribution P{Y1, Y2,..., Y i ,..., Y q} for E emotion categories, where Y i represents the predicted probability distribution of the i-th video for the E emotion categories, denoted as Y i {y i1,…,y ie ,…,y iE} where y ie represents the predicted probability that the i-th video belongs to the e-th emotion category;
[0136] Step 5: Construct a Focal Loss function that fuses dynamic weights to train the ConvMixer network combined with the adjacency matrix, the ResNet34 network, and the feature fusion and classification network. Calculate the training loss through the Focal Loss function that fuses dynamic weights; Use the trained ConvMixer network combined with the adjacency matrix to extract visual features from the image sequence, use the trained ResNet34 network to extract auditory features from the mel spectrogram, and then fuse the visual features and auditory features through the trained feature fusion and classification network, and perform emotion classification based on the fused features to predict the emotion category corresponding to the video;
[0137] Calculate the loss between the predicted probability distribution output in step (4.4) and the true emotion category according to formula (21);
[0138]
[0139] In formula (21), α represents the dynamic weight, and log represents the logarithmic function with base 2. represents the true emotion category to which the i-th video belongs of the predicted probability distribution;
[0140] The calculation formula for the dynamic weight α is as follows:
[0141]
[0142] In formula (22), C represents the confusion matrix of the previous training cycle. represents the number of times in the previous training cycle that the video with the true emotion category is predicted as category t. represents the number of videos with the true emotion category in the previous training cycle. represents the number of times of predicting the video with the true emotion category as category ;
[0143] Construction of the confusion matrix C in formula (22): Before the start of each training cycle, a zero matrix of size E×E is generated, where the number of rows and columns ranges from 1 to E; according to the prediction of each sample during training, the zero matrix is updated to obtain the confusion matrix C; when a sample of the true sentiment category 1 is predicted as category 2, the element at the first row and second column of the matrix is incremented by 1; when the training cycle is completed, the confusion matrix C of the current cycle is used to calculate the dynamic weight α of the loss function in the next training cycle.
[0144] In the fifth step, the batch size is 8, the number of iterations is set to 120, the Adam optimizer is used, the initial learning rate is 0.0001, the momentum factor is 0.9, and the learning rate is reduced by 90% every 30 iterations.
[0145] Where the present invention is not described shall be applicable to the prior art.
Claims
1. An audiovisual emotion classification method based on an improved ConvMixer network and dynamic focal loss, characterized in that It includes the following steps: In the first step, collect videos of the human face area expressing emotions, extract image sequences and audio signals from the videos, and convert the audio signals into mel spectrograms; In the second step, construct a ConvMixer network combined with an adjacency matrix, including three parts of operations, namely, patch embedding operation, Layer module operation, and average pooling operation in sequence; input the image sequence into the ConvMixer network combined with the adjacency matrix to extract visual features and obtain a feature map F; Step (2.1), patch embedding operation: The image sequence is sequentially subjected to a convolutional layer, an activation function layer, and a normalization layer for a block embedding operation to obtain a feature map F output by the block embedding operation 2.1 ; Step (2.2), Layer module operation, including four cascaded Layer modules; Input the feature map F 2.1 into the first Layer module. The first Layer module constructs a two-dimensional spatial coordinate matrix for each image patch in the feature map F 2.1 according to the spatial size of the feature map F 2.1 . It duplicates and concatenates the two-dimensional spatial coordinate matrix according to the temporal size of the feature map F 2.1 to obtain a spatial position encoding with the same size as the feature map F 2.1 . Concatenate the feature map F 2.1 with the spatial position encoding, and then pass through a linear layer to obtain the feature map Randomly generate a spatial adjacency matrix according to the spatial size of the feature map F 2.1 . Multiply the feature map by the spatial adjacency matrix, and then pass through an activation function layer and a normalization layer to obtain the feature map F s ; Stack the feature map with F s to obtain the feature map F s '. According to the feature map F s 'The time size of the feature map F s 'The one-dimensional time coordinate matrix of each image block in the feature map F s 'The spatial size of the one-dimensional time coordinate matrix is copied and spliced to obtain the feature map F s 'Time position encoding of the same size; the feature map F s 'Concatenate with the time position code, and then pass through the linear layer to get the feature map According to the feature map F s 'The time size randomly generates the time adjacency matrix and converts the feature map Multiply it with the time adjacency matrix, and then pass through the activation function layer and normalization layer to get the feature map F t ; The feature map With F t Superimpose to obtain the feature map F t ';Feature map F t 'After passing through the point-by-point convolution layer, activation function layer and normalization layer in sequence, the feature map output by the first Layer module is obtained; Step (2.3), average pooling operation: Perform spatial dimension average pooling operation on the feature map output by the fourth Layer module through an average pooling layer to obtain the feature map F; In the third step, use the ResNet34 network to extract auditory features from the mel spectrogram to obtain a feature map M; In the fourth step, construct a feature fusion and classification network for fusing visual features and auditory features and performing emotion classification on each video according to the fused features; the feature fusion and classification network includes two cross-modal temporal attention modules, pooling and concatenation operations, and classification operations; Step (4.1), the first cross-modal temporal attention module: Input the feature map F and the feature map M into the first cross-modal temporal attention module. The feature map F passes through a linear layer and a normalization layer to obtain a feature Q1; the feature map M passes through two independent linear layers and two independent normalization layers to obtain features K1 and V1 respectively; Generate a learnable intermediate matrix LIM1 based on the temporal dimension sizes of features Q1 and K1, and perform random parameter initialization on the intermediate matrix LIM1; multiply feature Q1 by the transposed initialized intermediate matrix LIM1 and feature K1, then divide by the square root of the number of channels of feature K1, and then input it into the softmax layer to obtain the normalized weights; multiply the normalized weights by feature V1 and then add it to feature Q1 to obtain the cross-modal attention feature F based on the image sequence att ; The cross-modal attention feature F based on the image sequence att Pass through the pointwise convolutional layer to obtain the cross-modal feature F based on the image sequence cm ; Step (4.2), the second cross-modal temporal attention module: Input the feature map F and the feature map M into the second cross-modal temporal attention module; the feature map M passes through a linear layer and a normalization layer to obtain a feature Q2; the feature map F passes through two independent linear layers and two independent normalization layers to obtain features K2 and V2; Generate a learnable intermediate matrix LIM2 based on the temporal dimension sizes of features Q2 and K2, and perform random parameter initialization on the intermediate matrix LIM2; multiply feature Q2 by the transpose of the intermediate matrix LIM2 and feature K2, divide by the square root of the number of channels of feature K2, and then input it into the softmax layer to obtain the normalized weights; multiply the normalized weights by feature V2 and add it to feature Q2 to obtain the cross-modal attention feature M based on the mel cepstrum coefficient spectrogram att ; The cross-modal attention feature M based on the mel cepstrum coefficient spectrogram att Pass through a pointwise convolutional layer to obtain the cross-modal feature M based on the mel cepstrum coefficient spectrogram cm ; Step (4.3), pooling and concatenation operations: The cross-modal feature F based on the image sequence cm and the cross-modal feature M based on the mel spectrogram cm are respectively subjected to average pooling and then concatenated to obtain the feature f FM ; Step (4.4), classification operation: Input the feature f FM into the linear layer and then pass through the softmax layer to obtain the predicted probability distribution P{Y1, Y2,..., Y i ,..., Y q} for E emotion categories. Y i represents the predicted probability distribution for the i-th video with respect to E emotion categories, denoted as Y i {y i1 , …, y ie , …, y iE}, where y ie represents the predicted probability that the i-th video belongs to the e-th emotion category, and q represents the number of videos; In the fifth step, train the ConvMixer network combined with the adjacency matrix, the ResNet34 network, and the feature fusion and classification network, and calculate the training loss through a focal loss function that fuses dynamic weights; Use the trained ConvMixer network combined with the adjacency matrix to extract visual features from the image sequence, use the trained ResNet34 network to extract auditory features from the mel spectrogram, then fuse the visual features and auditory features through the trained feature fusion and classification network, and perform emotion classification according to the fused features to predict the emotion category corresponding to the video.
2. The audiovisual emotion classification method based on the improved ConvMixer network and dynamic focal loss according to claim 1, characterized in that, The focal loss function that fuses dynamic weights is: In formula (21), α represents the dynamic weight, and log represents the logarithmic function with base 2. represents the true emotion category to which the i-th video belongs is the predicted probability distribution; The calculation formula of the dynamic weight α is as follows: In formula (22), C represents the confusion matrix of the previous training cycle, represents the number of times in the previous training cycle that a video with the true emotion category is predicted as category t, represents the number of videos with the true emotion category in the previous training cycle, represents the number of times that a video with the true emotion category is predicted as category .
3. The audiovisual emotion classification method based on the improved ConvMixer network and dynamic focal loss according to claim 2, wherein The construction process of the confusion matrix C is as follows: Before the start of each training cycle, a zero matrix of size E×E is generated, where the number of rows and columns ranges from 1 to E; According to the prediction of each sample during training, when a sample with the true sentiment category 1 is predicted as category 2, the element at the first row and second column of the matrix is incremented by 1, and the same applies to the rest of the samples. The zero matrix is updated to obtain the confusion matrix C; When the training cycle is completed, the confusion matrix of the current cycle is used to calculate the dynamic weight of the loss function in the next training cycle.
Citation Information
Patent Citations
Expression and voice bimodal-based child emotion recognition algorithm
CN113989893A
Multi-modal driver emotion feature recognition method and system
CN114582372A
Multi-modal sentiment analysis method based on multi-task learning and stacked cross-modal fusion
CN114694076A
Image-text matching method and device, storage medium and equipment
CN110147457A
Method and system for automatic chromosome classification
CN110689036A