A Multimodal Correlation-based Emotion Recognition Method and System Using Vision and Speech
Through the C3D network and sliding window self-attention network combined with COVAREP acoustic analysis framework, the problem of large model parameters and insufficient fusion of multimodal in video emotion classification is solved, and efficient multimodal emotion recognition is achieved.
Patent Information
- Application Number
- CN202310167361.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-02-27
AI Technical Summary
In the video emotion classification, the prior art has problems such as large number of model parameters, serious overfitting, difficulty in effectively extracting the time and spatial dimension information of multi-frame images, and insufficient fusion of multimodal data.
The C3D network is used to extract the spatiotemporal characteristics of video data with self-attention network with sliding windows, and the acoustic characteristics of speech data are extracted in combination with the COVAREP acoustic analysis framework. The emotional classification is performed through feature-level fusion and Softmax classifier, and the cross entropy loss function and gradient descent optimization model is used.
Effectively extracting the space-time and spatial emotional information of video data, improving the accuracy and efficiency of emotional recognition, fully integrating visual and speech modal information, reducing the amount of model parameters and avoiding overfitting.
Smart Images

Figure CN116167014B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer emotion computing, and in particular relates to a multi-modal association-type emotion recognition method and system based on vision and speech. Background Art
[0002] With the rapid development of the Internet, smooth and natural human-computer interaction systems have become a research hotspot. This undoubtedly requires that human-computer interaction should be like interpersonal communication. Machines should be able to understand people's emotions and true intentions and make corresponding responses. Emotional computing research is an attempt to create a computing system that can perceive, recognize and understand human emotions, and make targeted intelligent, sensitive and friendly responses. In general, it is to enable computers to have the same observation, understanding and expression capabilities as humans, so that computers can interact with users with emotions like humans. To achieve the above content, it is necessary to do a good job in the research of the two main tasks of emotional computing: identifying user emotions and generating emotional responses. This patent mainly completes the task of identifying user emotions.
[0003] Traditional methods generally use the method of manually designing features for sentiment recognition. After years of development, certain achievements have been made. However, manually designing features often requires a large amount of work, and it is difficult to break through the bottleneck in recognition performance. With the booming development of deep learning, convolutional neural networks have been widely applied to sentiment recognition tasks. Generally speaking, by stacking various complex network models, a relatively high recognition rate can be achieved, but usually a large amount of computing resources are consumed. Multi-head self-attention has achieved great success in the field of natural language processing in recent years. Recently, many works have also applied multi-head self-attention networks to the field of computer vision, attempting to develop a backbone network model that is universal in both the field of computer vision and natural language processing. Along with the rapid growth of computer computing power, such works have also achieved remarkable results comparable to convolutional neural networks. The literature [Dosovitskiy, Alexey, et al. "An image is worth 16x16 words: Transformers for image recognition at scale." 2020.] can be said to be the pioneering work that applies multi-head self-attention to the visual field. This model divides the image into multiple image patches according to the specified size, and each image patch is linearly mapped into a one-dimensional vector, thus meeting the input requirements of the multi-head self-attention model and achieving extremely advanced results in various tasks in the visual field. [Liu.Z, Lin.Y, Cao.Y, Hu.H, Wei.Y, and Zhang.Z, "Swin transformer: hierarchical vision transformer using shifted windows." 2021] proposed a self-attention model with sliding windows. The original data gradually reduces the data size through the model, enabling the model to have a receptive field similar to that of a CNN, thereby improving the performance of extracting multi-scale information. The sliding window cleverly realizes the extraction of global information and greatly reduces the number of parameters compared to the original multi-head self-attention model.
[0004] The Chinese patent application "A method and system for video sentiment recognition based on visual and language annotation association" (patent application number CN202210511572.0, publication number CN114882412A) proposes to divide the image patches into 9 equal parts and then use C3D for temporal feature extraction, use CNN for spatial feature extraction, and then send them into a multi-head self-attention neural network respectively to further extract the sentiment features in the spatial and temporal dimensions and cascade them. Finally, sentiment classification is performed by combining the text sentiment features. This method extracts visual information through the form of two groups of convolutional neural networks plus a multi-head attention model, resulting in a huge number of parameters, making the model difficult to train and prone to overfitting.
[0005] Although the multi - head self - attention model network has achieved great success in the field of natural language processing and there have been many attempts in the field of computer vision, there are still many challenges in the field of video emotion classification. First, most current work focuses on single - frame image tasks. Video data consists of multiple consecutive images, and there is a great correlation between the front and back. It is very important to design a suitable network to extract the visual information contained in the video data. Second, video data in real life often consists of multiple modalities of data such as image frames, text captions, and speech data at the same time. It is necessary to effectively extract and fuse the emotional information contained in different modalities and classify them. Third, speech data is usually a continuous audio data, and the extraction of the basic features of speech data usually requires certain prior knowledge. Summary of the Invention
[0006] In view of the deficiencies of the prior art, the present invention provides a multi - modal associated emotion recognition method and system based on vision and speech, which while fully extracting the emotional information in the time dimension and space dimension of video data, fuses the emotional information contained in speech data, and realizes the emotion classification of short - video data.
[0007] To achieve the above functions, the present invention adopts the following technical solutions:
[0008] A multi - modal associated emotion recognition method based on vision and speech, comprising the following steps:
[0009] S1. Pre - process the video stream of the short - video sample, segment a specified number of image frames, and uniformly adjust the resolution of the image frames to a specified size.
[0010] S2. Use the C3D network (Convolution 3D, 3D convolutional neural network) to extract the temporal features of the image frames in step S1 to obtain a feature map of a specified size; input the feature map into a self - attention neural network with a sliding window to further extract the spatial - dimension information on the basis of the time dimension, and obtain a visual deep - layer emotion feature vector with spatio - temporal feature information.
[0011] S3. For the speech data corresponding to the short - video content, use the COVAREP acoustic analysis framework to extract the acoustic features of the speech data, and then use the self - attention network to further extract the deep - layer emotion feature vector of the speech data.
[0012] S4. Feature - level fuse the visual deep - layer emotion feature vector and the speech deep - layer emotion feature vector extracted in steps S2 and S3 in a concatenated form, and then pass the fused feature vector through a fully - connected network, and further use Softmax as a classifier to classify the emotion to obtain a complete emotion recognition model;
[0013] S5. Based on the complete emotion recognition model obtained in step S4, calculate the cross-entropy loss function using the output emotion probability distribution matrix, and use the gradient descent method as the optimization method. Continuously iterate and train the network through backpropagation to obtain the complete trained network model.
[0014] S6. Input the short video to be recognized into the network model obtained in step S5 for emotion classification recognition.
[0015] Furthermore, the specific content of step S1 is as follows: Extract F image frames at equal intervals starting from the first frame of the short video sample video stream. When there are less than F frames, fill the last frame using the oversampling method; uniformly adjust the resolution of the obtained image frames to M×M.
[0016] Furthermore, the specific steps of step S2 are as follows:
[0017] S201. Input the F M×M image frames extracted in step S1 into a 3D convolutional neural network to extract temporal features, and the output is a feature map of a specified size.
[0018] S202. Input the feature map into a self-attention network with a sliding window, perform original self-attention calculation under a window of a specified size, slide the window to the right and down, and the sliding distance is half of the window width. Perform self-attention calculation after window sliding, and then set another window of a different size to perform original self-attention calculation. Slide the window to the right and down again, and the sliding distance is half of the window width. Perform self-attention calculation after window sliding, extract the information in the spatial dimension, and output a feature map with a specified size of N×N×C. The self-attention calculation formula is as follows:
[0019]
[0020]
[0021] where Q, K, and V represent the query, key, and value matrices respectively, X is the input sequence of the self-attention network, W Q , W K , W V are obtained through training, d represents the dimension of the query vector, and B represents the relative position bias matrix.
[0022] S203. Perform global average pooling operation on the feature map output in step S202 to obtain a C×1-dimensional feature vector with spatio-temporal feature information.
[0023] Furthermore, the specific steps of step S3 are as follows:
[0024] S301. Use the COVAREP acoustic analysis framework to extract acoustic features in three aspects of speech data rhythm, voice quality, and spectrum, and obtain Among them, T a represents the segmented frame number of the audio, and A i represents the acoustic feature vector of the i-th frame. d is the dimension of the acoustic feature vector extracted from each frame of audio data.
[0025] S302. The extracted acoustic feature dimension is (T a , d). Embed the position information in the extracted acoustic features and add a class label vector with a dimension of (1, d) to form a feature sequence with a dimension of (T a +1, d) and input it into the self-attention network to calculate the deep emotional feature vector of the speech data.
[0026] Furthermore, the specific steps of step S4 are as follows:
[0027] S401. Directly splice the visual deep emotional feature vector Feature v extracted in step S2 and the speech deep emotional feature vector Feature a extracted in S3 to obtain a fused feature vector F with a specific dimension va :
[0028]
[0029]
[0030] Among them, represents the value of the i-th dimension of the visual feature vector, represents the value of the j-th dimension of the speech feature vector. The sizes of V and A respectively represent the dimension sizes of the visual and speech feature vectors.
[0031] S402. Input the fused feature vector F va into the fully connected layer network, and further use the Softmax classifier to classify the emotion:
[0032]
[0033] Among them, J is the emotion category; Score i is the predicted score of the i-th emotion, i = 1, 2,..., J; x i is the value of the i-th dimension of the classifier input vector x.
[0034] The Softmax classifier calculates the scores of various emotions through the form of vector exponential normalization, and obtains the emotion distribution probability matrix as P = [Socre1, Socre2,..., Socre J .
[0035] S403. Select the category corresponding to the subscript of the Score with the highest probability according to the emotion distribution probability matrix as the final result.
[0036] Furthermore, the specific formula for using the cross-entropy loss function in step S5 is as follows:
[0037]
[0038] Among them, J represents the emotion category; Score i is the predicted score of the i-th type of emotion; y i represents the true label of the sample data, and the value of y is 1 when the category is correct and 0 otherwise.
[0039] Furthermore, the present invention also provides a multi-modal associated emotion recognition system based on vision and speech, including:
[0040] A video stream segmentation module, configured to segment the video stream of the video data to obtain a specified number of image frames, and adjust the resolution of these image frames to a unified specified size.
[0041] A visual feature extraction module, configured to extract spatio-temporal feature information of the video data to obtain a deep emotion feature vector of the video data.
[0042] A speech feature extraction module, configured to extract an emotion feature vector from the speech data corresponding to the video data.
[0043] A fused feature emotion score calculation module, configured to perform feature fusion on the visual emotion feature vector and the speech emotion feature vector in a concatenated form, then input the fused vector into a fully connected layer network, use Softmax as a classifier to calculate the scores of each emotion, obtain an emotion distribution probability matrix, and take the emotion with the highest score as the final classification result.
[0044] A visual and speech network model training module, configured to calculate the cross-entropy loss function value for the complete network model according to the emotion distribution probability matrix, use the gradient descent method as an optimization method, and continuously iterate and train the network through backpropagation to obtain a trained complete network model.
[0045] Furthermore, the visual feature extraction module includes a temporal feature extraction module unit and a spatial feature extraction module unit, where:
[0046] The temporal feature extraction module unit is configured to perform the following actions: extract the temporal features of the selected image frames using a 3D convolutional neural network to obtain a feature map of a specific size.
[0047] The spatial feature extraction module unit is configured to perform the following actions: input the feature map output by the temporal feature extraction module into a self-attention neural network with a sliding window, and extract spatial features through the original self-attention calculation and the self-attention calculation after window sliding to obtain a feature map of a specific size.
[0048] Furthermore, the speech feature extraction module includes an acoustic feature extraction module unit and a speech emotion feature extraction module unit, where:
[0049] The acoustic feature extraction module unit is configured to perform the following actions: use the COVAREP acoustic analysis framework to extract the acoustic features in three aspects of prosody, voice quality, and spectrum of the speech data, and obtain A = {A1, A2,..., A i ,..., A Ta}, where T a represents the number of segmented frames of the audio, and A i represents the acoustic feature vector of the i-th frame. d is the dimension of the acoustic feature vector extracted from each frame of audio data.
[0050] The speech emotion feature extraction module unit is configured to perform the following actions: embed the position information into the acoustic features extracted by the acoustic feature extraction module unit and add a class label vector with a dimension of (1, d) to form a feature sequence with a dimension of (T a + 1, d), and input it into the self-attention network to extract the deep emotion feature vector of the speech data.
[0051] Furthermore, the present invention also provides an electronic device, which is characterized by including a computing device including a memory and a processor, and there is a readable storage medium in the computer. The program stored in the readable storage medium can run on the processor. When the computer program is executed by the processor, it implements the steps of the multi-modal associated emotion recognition method based on vision and speech described above.
[0052] The present invention adopts the above technical solutions. Compared with the prior art, the remarkable technical effects are as follows:
[0053] (1) The present invention uses C3D combined with a self-attention network with a sliding window to extract the deep visual emotion feature information of video data, which can effectively extract the emotion information from the time dimension and the space dimension. The attention model with a sliding window can efficiently extract local and global spatial information, making the receptive field more friendly to multi-scale data and reducing the number of model parameters;
[0054] (2) The present invention uses the COVAREP acoustic analysis framework to extract acoustic features in three aspects of prosody, voice quality, and spectrum of speech data, improving the feature extraction efficiency. Furthermore, an attention model is used to extract deep emotional feature information of speech data, improving the accuracy and efficiency of emotion recognition.
[0055] (3) The present invention integrates a visual and speech emotion feature extraction module, extracts visual features and speech features of data samples, fully integrates visual emotion information and speech emotion information. The integration of these two modal information fills a certain information gap and realizes the full utilization of multi-modal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is the overall step flow chart of the present invention.
[0057] Figure 2 is the structure diagram of the emotion recognition system of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0058] The following further describes the specific implementation technical solutions of the present invention with reference to the accompanying drawings:
[0059] As Figure 1 shown, an embodiment of the present invention discloses a multi-modal associated emotion recognition method based on vision and speech, which specifically includes the following steps:
[0060] S1. Extract 16 image frames at equal intervals starting from the first frame of the short video sample video stream. When there are less than 16 frames, oversampling is used to fill the last frame. The resolution of the obtained image frames is uniformly adjusted to 224×224. In this example, the CMU-MOSEI dataset is used as the data source.
[0061] S2. Use the C3D network to extract temporal features of the image frames in step S1 to obtain a feature map of a specified size; input this feature map into a self-attention neural network with a sliding window to further extract spatial dimension information on the basis of the time dimension, and obtain a visual deep emotional feature vector with spatio-temporal feature information. The specific steps are as follows:
[0062] S201. Send the 16 224×224 image frames extracted in step S1 into the C3D network for temporal feature extraction, and output a feature map with a size of 56×56×4. The C3D network is composed of three convolutional layers and three pooling layers connected alternately. The size of the convolutional kernel is 3×3×3. The four convolutional layers use 2, 4, and 4 convolutional kernels in sequence, and all use the Relu function for activation. The sizes of the first two pooling kernels are 2×2×2, and the last one is 1×1×4, and the maximum pooling strategy is adopted.
[0063] S202. Input the feature map into the self-attention network with a sliding window. Perform the original self-attention calculation with a window size of 4×4. Slide the window 2 units to the right and 4 units down, and perform the self-attention calculation after window sliding. Subsequently, perform the original self-attention calculation with a window size of 8×8. Slide the window 2 units to the right and 4 units down again, and perform the self-attention calculation after window sliding. Extract the information in the spatial dimension and output a feature map with a size of 7×7×256. The self-attention calculation formula is as follows:
[0064]
[0065]
[0066] Among them, Q, K, and V represent the query, key, and value matrices respectively. X is the input sequence of the self-attention network. The input sequences of the two operation units are 196 64-dimensional vectors and 49 256-dimensional vectors respectively. W Q ,W K ,W V are obtained through training. d represents the dimension of the query vector, and B represents the relative position bias matrix. The values of d and B are determined by each specific calculation.
[0067] S203. Perform global average pooling operation on the feature map output in step S202 to obtain a 256×1 feature vector with spatio-temporal feature information.
[0068] S3. For the voice data in wav format corresponding to the short video content, use the COVAREP acoustic analysis framework to extract the acoustic features of the voice data, and then use the self-attention network to further extract the deep emotional feature vector of the voice data, with a dimension of 74×1. The specific steps are as follows:
[0069] S301. Use the COVAREP acoustic analysis framework to extract the acoustic features in three aspects of prosody, voice quality, and spectrum of the voice data, and obtain A = {A1, A2, …, A i , …, A 128}, where A i represents the acoustic feature vector of the i-th frame. The number of segmented frames of the audio data is 128, and the dimension of the extracted acoustic feature vector is 74.
[0070] S302. The dimension of the extracted acoustic features is (128, 74). Embed the position information in the extracted acoustic features and add a class label vector with a dimension of (1, 74) to form a feature sequence with a dimension of (129, 74) and input it into the self-attention network to calculate the deep emotional feature vector of the voice data.
[0071] S4. Feature-level fusion is performed on the visual deep emotional feature vectors and speech deep emotional feature vectors extracted in steps S2 and S3 respectively in a concatenated form, and then the fused feature vectors are passed through a fully connected network. Further, Softmax is used as a classifier to classify emotions, obtaining a complete emotion recognition model. The specific steps are as follows:
[0072] S401. Concatenate directly the visual deep emotional feature vector Feature v extracted in step S2 and the speech deep emotional feature vector Feature a extracted in step S3 to obtain a fused feature vector F va :
[0073]
[0074]
[0075] wherein, represents the value of the i-th dimension of the visual feature vector, represents the value of the j-th dimension of the speech feature vector, and the sizes of V and A respectively represent the dimension sizes of the visual and speech feature vectors.
[0076] S402. Input the fused feature vector F va into the fully connected layer network. The first layer of the fully connected layer contains 1024 nodes, and the second layer contains 256 nodes. Use the Relu activation function and further use the Softmax classifier to classify emotions:
[0077]
[0078] where J is the emotion category, including six emotions, namely happy, sad, angry, fearful, disgusted, and surprised; Score i is the predicted score of the i-th type of emotion, i = 1, 2,..., 6; x i is the value of the i-th dimension of the input vector x of the classifier.
[0079] The Softmax classifier calculates the scores of various emotions in the form of vector exponential normalization, obtaining an emotion distribution probability matrix P = [Socre1, Socre2,..., Socre6]. The specific situation is shown in Table 1:
[0080] Table 1 Emotion categories corresponding to different subscripts
[0081]
[0082] S403. According to the emotion distribution probability matrix, select the category corresponding to the subscript of the maximum probability Score as the final result.
[0083] S5. According to the complete emotion recognition model obtained in step S4, calculate the cross-entropy loss function using the output emotion probability distribution matrix, and use the gradient descent method as the optimization method. Continuously iterate and train the network through backpropagation to obtain the trained complete visual-audio bimodal emotion recognition model. The specific formula of the cross-entropy loss function is as follows:
[0084]
[0085] Among them, J represents emotion categories, including six emotions: happy, sad, angry, fearful, disgusted, and surprised; Score i is the predicted score of the i-th type of emotion; y i represents the true label of the sample data. When the category is correct, the value of y is 1, and the rest are 0.
[0086] S6. Input the short video to be recognized into the network model obtained in step S5 for emotion classification recognition.
[0087] As Figure 2 shown, an embodiment of the present invention also proposes a multimodal correlation-based emotion recognition system based on vision and speech, including: a video stream segmentation module, a visual feature extraction module, a speech feature extraction module, a fused feature emotion score calculation module, and a visual and speech network model training module.
[0088] It should be noted that each module in the above system corresponds to the specific steps of the method provided by the embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiment of the present invention.
[0089] For the above-mentioned multimodal correlation-based emotion recognition method and system embodiments based on vision and speech, their technical principles, the technical problems solved, and the technical effects produced are similar to those of the method embodiments, belonging to the same inventive concept. For specific implementation details and related descriptions, reference can be made to the corresponding processes in the foregoing multimodal correlation-based emotion recognition method embodiments based on vision and speech, which will not be elaborated here.
[0090] Those skilled in the art can understand that the modules in the embodiments can be adaptively changed and set in one or more systems different from this embodiment. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components.
[0091] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, including a computing device including a memory and a processor, and a computer-readable storage medium. When a program that can be run on the processor is stored in the readable storage medium, the multi-modal association-based emotion recognition method based on vision and voice described above is implemented.
[0092] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention fall within the protection scope of the present invention.
Claims
1. A multi-modal associative emotion recognition method based on vision and speech, characterized in that, Including: S1. Preprocess the video stream of the short video sample, segment a specified number of image frames, and uniformly adjust the resolution of the image frames to a specified size; S2. Use a 3D convolutional neural network to extract temporal features from the image frames in step S1 to obtain a feature map of a specified size; input the feature map into a self-attention neural network with a sliding window to further extract spatial dimension information on the basis of the time dimension, and obtain a visual deep emotional feature vector with spatio-temporal feature information; S3. For the speech data corresponding to the short video content, use the COVAREP acoustic analysis framework to extract the acoustic features of the speech data, and then use the self-attention network to further extract the deep emotional feature vector of the speech data; S4. Feature-level fuse the visual deep emotional feature vector and the speech deep emotional feature vector extracted in steps S2 and S3 in a concatenated form, and then pass the fused feature vector through a fully connected network, and further classify the emotion using Softmax to obtain a complete emotion recognition model; S5. According to the complete emotion recognition model obtained in step S4, calculate the cross-entropy loss function using the output emotion probability distribution matrix, and use the gradient descent method as the optimization method to continuously iterate and train the network through backpropagation to obtain a complete trained network model; S6. Input the short video to be recognized into the network model obtained in step S5 for emotion classification recognition.
2. The multimodal associated emotion recognition method based on vision and speech according to claim 1, wherein The specific content of step S1 is: equally spaced extract F image frames from the first frame of the short video sample video stream. When there are less than F frames, fill the last frame using the oversampling method; uniformly adjust the resolution of the obtained image frames to M×M.
3. The multimodal correlation-based emotion recognition method based on vision and speech according to claim 1, wherein The specific steps of step S2 are: S201. Send the F M×M image frames extracted in step S1 into a 3D convolutional neural network to extract temporal features, and the output is a feature map of a specified size; S202. Input the feature map into a self-attention network with a sliding window, perform the original self-attention calculation under a window of a specified size, slide the window to the right and down, and the sliding distance is half of the window width. Perform the self-attention calculation after window sliding, and then set another size of window to perform the original self-attention calculation. Slide the window to the right and down again, and the sliding distance is half of the window width. Perform the self-attention calculation after window sliding to extract the information of the spatial dimension, and output a feature map with a specified size of N×N×C; where the self-attention calculation formula is as follows: Among them, Q, K, and V respectively represent the query, key, and value matrices, X is the input sequence of the self-attention network, W Q , W K , W V are obtained through training, d represents the dimension of the query vector, and B represents the relative position bias matrix; S203. Perform global average pooling operation on the feature map output in step S202 to obtain a C×1-dimensional feature vector with spatio-temporal feature information.
4. The multimodal correlation-based emotion recognition method based on vision and speech according to claim 1, wherein The specific steps of step S3 are: S301. Extract the acoustic features in three aspects of prosody, voice quality and spectrum from the speech data by using the COVAREP acoustic analysis framework, and obtain \(A = \{A_1, A_2, \cdots, A_{T}\}\), where \(T\) i , \cdots, A_{ Ta}\), where \(T\) a represents the number of segmented frames of the audio, and \(A_{i}\) i represents the acoustic feature vector of the \(i\)-th frame. \(d\) is the dimension of the acoustic feature vector extracted from each frame of audio data. S302. The extracted acoustic feature dimension is (T a , d). Embed the position information in the extracted acoustic features and add a class label vector with a dimension of (1, d) to form a feature sequence with a dimension of (T a + 1, d) and input it into the self-attention network to calculate the deep emotional feature vector of the speech data.
5. The multimodal associated emotion recognition method based on vision and speech according to claim 1, wherein The specific steps of step S4 are: S401. Directly splice the visual deep emotional feature vector Feature extracted in step S2 v and the speech deep emotional feature vector Feature extracted in S3 a to obtain a fused feature vector F of a specific dimension va : Among them, f i v represents the value of the i-th dimension of the visual feature vector, represents the value of the j-th dimension of the speech feature vector, and the sizes of V and A respectively represent the dimensionality sizes of the visual and speech feature vectors; S402. Input the fusion feature vector F va into the fully connected layer network, and further classify the sentiment using the Softmax classifier: Among them, J is the emotion category; Score i is the predicted emotion score of the i-th category, where i = 1, 2,..., J; x i is the value of the i-th dimension of the classifier input vector x; The Softmax classifier calculates the scores of various emotions in the form of exponential normalization of vectors, and obtains the emotion distribution probability matrix as P = [Socre1, Socre2, …, Socre J ; S403. According to the emotion distribution probability matrix, select the category corresponding to the subscript of the maximum probability Score as the final result.
6. The multimodal correlation-based emotion recognition method based on vision and speech according to claim 1, characterized in that, The specific formula for using the cross-entropy loss function in step S5 is as follows: Among them, J represents the emotion category; Score i is the predicted emotion score of the i-th category; y i represents the true label of the sample data, where the value of y is 1 when the category is correct and 0 otherwise.
7. A multimodal associated emotion recognition system based on vision and speech, characterized in that, Including: A video stream segmentation module, which is used to segment the video stream of video data to obtain a specified number of image frames, and adjust the resolution of these image frames to a unified specified size; A visual feature extraction module, which is used to extract the spatio-temporal feature information of video data to obtain the deep emotional feature vector of the video data; A speech feature extraction module, which is used to extract the emotional feature vector in the speech data corresponding to the video data; A fused feature emotion score calculation module, which is used to fuse the visual emotion feature vector and the speech emotion feature vector in a concatenated form, then input the fused vector into a fully connected layer network, use Softmax as a classifier to calculate the scores of each emotion, obtain an emotion distribution probability matrix, and take the emotion with the highest score as the final classification result; A visual and speech network model training module, which is used to calculate the cross-entropy loss function value for the complete network model according to the emotion distribution probability matrix, use the gradient descent method as the optimization method, and continuously iteratively train the network through backpropagation to obtain the trained complete network model.
8. The multimodal associated emotion recognition system based on vision and speech according to claim 7, wherein The visual feature extraction module includes a temporal feature extraction module unit and a spatial feature extraction module unit, where: The temporal feature extraction module unit is configured to perform the following actions: use a 3D convolutional neural network to extract the temporal features of the selected image frames to obtain a feature map of a specific size; The spatial feature extraction module unit is configured to perform the following actions: input the feature map output by the temporal feature extraction module into a self-attention neural network with a sliding window, and propose spatial features through the original self-attention calculation and the self-attention calculation after window sliding to obtain a feature map of a specific size.
9. The multi-modal associated emotion recognition system based on vision and speech according to claim 7, wherein, The speech feature extraction module includes an acoustic feature extraction module unit and a speech emotion feature extraction module unit, where: An acoustic feature extraction module unit is configured to perform the following actions: use the COVAREP acoustic analysis framework to extract acoustic features in three aspects of prosody, voice quality, and spectrum of speech data, resulting in \(A = \{A_1, A_2, \cdots, A\) i , \cdots, A Ta \}, where \(T\) a represents the number of segmented frames of the audio, and \(A\) i represents the acoustic feature vector of the \(i\)-th frame, and \(d\) is the dimension of the acoustic feature vector extracted from each frame of audio data; The voice emotion feature extraction module unit is configured to perform the following actions: embed the position information of the acoustic features extracted by the acoustic feature extraction module unit and add a class token vector with a dimension of (1, d) to form a feature sequence with a dimension of (T a + 1, d), and input it into the self-attention network to extract the deep emotion feature vector of the voice data.
10. An electronic device, characterized in that, It includes a computing device containing a memory and a processor, and the computer has a readable storage medium, and a program that can run on the processor is stored in the readable storage medium. When the computer program is executed by the processor, it implements the multi-modal correlation-based emotion recognition method according to any one of claims 1-5.
Citation Information
Patent Citations
Vision and language-based annotation association type short video emotion recognition method and system
CN114882412A
Annotated and associated short video emotion recognition method and system based on vision and language
CN114882412B
GIF short video emotion recognition method and system fusing text information
CN109145712A
Multi-modal emotion recognition method based on situational attention neural network
CN112348075A