Voice visual emotion recognition method, device and system, and storage medium
By designing adaptive alignment and time-agging networks in speech and visual emotion recognition technology, the problems of redundant and time-varying feature capture in feature extraction are solved, and the emotion recognition effect with high accuracy and low computational complexity is achieved.
Patent Information
- Application Number
- CN202510298761.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
Existing video-based speech visual emotion recognition technology ignores heterogeneity when extracting speech and visual modal features, resulting in feature redundancy; when video-level feature fusion, it is difficult to capture time-varying features of emotions, and the computational complexity is high.
Adaptive alignment and time aggregation networks are designed, including low-redundant voice vision adaptive alignment modules and time-adaptive aggregation modules, which extract low-redundant features through adaptive alignment and adaptively aggregate features over a long period of time to reduce computational complexity.
It improves the accuracy of emotion recognition, reduces the computational complexity, and effectively captures the time-varying characteristics of emotions.
Smart Images

Figure CN120148558A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of emotion recognition, and particularly relates to a method and device, system, and storage medium for speech-visual emotion recognition. Background Art
[0002] The video-based speech-visual emotion recognition technology aims to recognize the emotional states of humans by fusing speech-visual features extracted from emotional videos. As a core component of human-computer interaction, this technology has great application potential in various intelligent systems, such as intelligent driving, e-education, intelligent robots, etc. With the rapid development of artificial intelligence, deep learning-based methods have increasingly dominated video-based speech-visual emotion recognition, driving significant progress in emotion recognition technology.
[0003] Video-based speech-visual emotion recognition can be divided into two steps: segment-level speech-visual feature extraction and video-level speech-visual feature fusion. Current mainstream segment-level speech-visual feature extraction methods usually use neural networks to extract features of speech and visual modalities separately. This ignores the heterogeneity between the speech and visual modalities, resulting in redundant speech and visual features being proposed. For video-level speech-visual feature fusion, average pooling operations are usually used to fuse, which is difficult to capture the time-varying features of emotions. Some works use multi-head attention to adaptively aggregate segment-level speech-visual features. However, the time aggregation method based on multi-head attention does not fully consider the time-varying features of emotions at the video level, resulting in higher computational complexity. This is because the temporal correlation of emotions in emotional videos often decreases over time. For example, at the beginning and end of an emotional video, the temporal correlation of emotions is weak. If multi-head attention is directly used to process the speech-visual features of each short time period, it is difficult to effectively capture this temporal correlation. In addition, the computational complexity of this operation increases quadratically with the increase in the number of segment levels, further exacerbating the problem. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and device, system, and storage medium for speech-visual emotion recognition.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for speech-visual emotion recognition, comprising:
[0007] Step S1, extracting speech-visual features at the segment level;
[0008] Step S2, inputting the extracted speech-visual features at the segment level into an adaptive alignment and time aggregation network to obtain emotional video features of multiple short time segments;
[0009] Step S3: Perform emotion recognition based on the emotion video features of multiple short-time segments.
[0010] Preferably, the adaptive alignment and temporal aggregation network includes: a low-redundancy speech-visual adaptive alignment module for obtaining low-redundancy alignment features for each short time period, and a temporal adaptive aggregation module for adaptively aggregating the aligned features over a long time period.
[0011] Preferably, in step S1, let X = [x 1 , …, x n , …, x N represent an emotion video containing N segments, where x n represents the nth segment, and each segment x n is divided into a speech modality and a visual modality Input the speech modality and the visual modality into the speech feature extractor and the visual feature extractor respectively obtain segment-level speech features and visual features That is:
[0012]
[0013] Among them,
[0014] The present invention also provides a speech-visual emotion recognition device, including:
[0015] A feature extraction module for extracting speech-visual features at the segment level;
[0016] A processing module for inputting the speech-visual features extracted at the segment level into an adaptive alignment and temporal aggregation network; obtaining emotion video features of multiple short-time segments;
[0017] A recognition module for performing emotion recognition based on the emotion video features of multiple short-time segments.
[0018] Preferably, the adaptive alignment and temporal aggregation network includes: a low-redundancy speech-visual adaptive alignment module for obtaining low-redundancy alignment features for each short time period, and a temporal adaptive aggregation module for adaptively aggregating the aligned features over a long time period.
[0019] Preferably, the processing process of the feature extraction module is as follows: Let X = [x 1 , …, x n , …, x N represent an emotion video containing N segments, where xn Denotes the nth segment, and each segment x n Is divided into the speech modality And the visual modality Input the speech modality And the visual modality Into the speech feature extractor And the visual feature extractor Respectively obtain segment-level speech features And visual features That is:
[0020]
[0021] Wherein,
[0022] The present invention also provides a speech-visual emotion recognition system, including: a memory and a processor, wherein a computer program run by the processor is stored on the memory, and the computer program executes a speech-visual emotion recognition method when being run by the processor.
[0023] The present invention also provides a storage medium, on which a computer program is stored, and the computer program executes a speech-visual emotion recognition method when running.
[0024] The present invention first designs a Low Redundancy Speech-Visual Adaptive Alignment (LRSVAA) module to obtain low redundancy alignment features of speech-visual modalities; then designs a Computationally Efficient Time-Adaptive Aggregation (CETAA) module to adaptively aggregate the aligned features in a long time period. By adopting the technical solution of the present invention, low redundancy speech-visual features are extracted in an adaptive alignment manner at the segment level; and time-adaptive emotion aggregation with low computational complexity is realized at the video level, which can improve the accuracy of emotion recognition. Description of the Drawings
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0026] Figure 1 Is the flowchart of the speech-visual emotion recognition method according to the embodiment of the present invention;
[0027] Figure 2 Schematic diagram of data processing for the voice-visual emotion recognition method according to an embodiment of the present invention;
[0028] Figure 3 Schematic diagram of the LRSVAA structure;
[0029] Figure 4 Schematic diagram of the CETAA structure. Detailed implementation manners
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0031] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0032] Embodiment 1:
[0033] As Figure 1 , 2 shown, an embodiment of the present invention provides a voice-visual emotion recognition method, including:
[0034] Step S1, extract voice-visual features at the segment level;
[0035] Step S2, input the voice-visual features extracted at the segment level into an Adaptive Alignment and Time Aggregation Network (AataNet) to obtain emotion video features of multiple short-time segments;
[0036] Step S3, perform emotion recognition according to the emotion video features of multiple short-time segments.
[0037] As an implementation manner of an embodiment of the present invention, as Figure 2 shown, AataNet includes: a Low Redundancy Speech-Visual Adaptive Alignment (LRSVAA) module for obtaining low-redundancy alignment features of each short time period and a Computationally Efficient Time-Adaptive Aggregation (CETAA) module for adaptively aggregating the aligned features in a long time period.
[0038] As an implementation manner of an embodiment of the present invention, in step S1, let X = [x 1 ,…,x n ,…,x N represent an emotional video containing N segments, where x n represents the nth segment. Each segment x n can be further divided into a speech modality and a visual modality which are input into a speech feature extractor and a visual feature extractor to respectively obtain segment-level speech features and visual features which are represented as follows:
[0039]
[0040] Among them, Select the wav2vec2.0 model that is often used in the speech field. Select the Spatio-Temporal Vision Transformer (ST-VT) because it effectively extracts spatio-temporal features of the visual modality by combining spatial and temporal attention. In addition, a fully connected layer is added before the classification layers of the wav2vec2.0 model and the ST-ViT to obtain and
[0041] As an implementation manner of an embodiment of the present invention, in step S2, to reduce the redundancy between and , AataNet designs LRSVAA to adaptively align the speech-visual features of each segment. At the video level, AataNet also designs CETAA to adaptively aggregate the aligned features with lower computational complexity.
[0042] Further, as shown in Figure 3 , LRSVAA first linearly projects and into queries, keys, and values:
[0043]
[0044] Among them, represents and corresponding weight matrices for queries, keys, and values, all of which can be automatically learned during the training phase.
[0045] Then, to reduce the redundancy between and , LRSVAA designed cross-modal multi-head attention to project between and to study the perceptual information. Specifically, to explore the perception of on , LRSVAA designed a Speech Cross-Modal Multi-Head Attention (SCM-HA), where the query comes from while the key and value come from and Similarly, the Visual Cross-Modal Multi-Head Attention (VCM-HA) also aims to explore the situation where is perceived by . Similarly, the query, key, and value of VCM-HA are and Mathematically, this is expressed as:
[0046]
[0047] where SCM-MHA(·) and VCM-MHA(·) are similar to MHA. represents the information perceived by by , represents the information perceived by by . Based on and , the speech-visual adaptive weight calculation formula can be expressed as
[0048]
[0049] If and are well-aligned in semantic information, will have a higher weight, otherwise a lower weight. Similarly, and are well-aligned in semantic information, will have a higher weight, and vice versa. Finally, and for being perceived by Adaptive weighting is performed to obtain low-redundancy alignment features. Similar to Transformer, a Feed-Forward (FF) layer, Layer Normalization (LN), and Residual Connections (RC) are also applied, which are denoted as
[0050]
[0051] where FF(·), LN(·), and RC(·) represent the FF, LN, and RC operations respectively. Denote and the aligned features.
[0052] Furthermore, through LRSVAA, low-redundancy alignment features at the segment level are obtained A simple method is to directly use Transformer for adaptive aggregation at the video level where each can be regarded as a patch. However, this requires N 2 attention calculations, resulting in a serious computational burden. For this reason, AataNet designed CETAA, as Figure 4 shown. Specifically, and are linearly projected into queries, keys, and values:
[0053]
[0054] where represents the output of the (n - 1)-th CETAA, note that denote and the weight matrices corresponding to queries, keys, and values in the linear projection, which are learned during the training phase.
[0055] Based on CETAA adaptively calculates the temporal clustering information of adjacent short-time segments using Cross-Time Multi-Head Attention (CT-MHA), and the clustered temporal information is then aggregated into the next short time period. This is because adjacent segments have a high temporal correlation in terms of emotion, which is more in line with the time-varying characteristics of emotion. In addition, CETAA only requires N - 1 attention calculations for N short-time segments, which reduces the complexity from N 2 to N - 1. Mathematically, this is expressed as:
[0056]
[0057] Among them, CT-MHA(·) is similar to SCM-MHA(·) and VCM-MHA(·). and denotes and The cross-temporal clustering features of. These features are then processed by the FF layer, LN, and RC to obtain and
[0058]
[0059] On the other hand, different segments contribute differently to speech-visual emotion recognition. Therefore, an adaptive attention mechanism is applied as follows:
[0060]
[0061]
[0062] where q represents the attention vector. and respectively denote and The weight matrices of. represents the output of the nth CETAA. Note that represents the emotion video features of N short-time segments processed by AataNet, which are input into the classification layer to achieve downstream emotion classification. Among them, the classification layer uses a three-layer multi-layer perceptron (MLP), and the softmax function is selected for the last layer to output the emotion classification result. In addition, the classification layer uses the classic cross-entropy loss function, which is defined as:
[0063]
[0064] where, represents the sample set of emotion videos, i represents the ith emotion video sample in, y i represents the label of the ith emotion video sample, represents the predicted label of the classification layer for the ith emotion video sample.
[0065] The RAVDESS, RML, and eNTERFACE05 datasets are selected to evaluate the performance of the proposed method. Table 1 shows the results of the comparative experiments. Comparative method 1 means and are directly summed and aligned. Comparative method 2 means using the average pooling operation to achieve Aggregation. Comparative method 3 involves a Transformer-based method, where and are directly summed and then processed through FF, LN, and RC to obtain Comparative method 4 uses a Long Short-Term Memory (LSTM) network to aggregate Comparing the proposed method with comparative methods 1 and 3 shows that the designed LRSVAA is superior to direct summation and the Transformer-based method. This is because LRSVAA adaptively calculates and at the short-time segment level to align speech-visual features, thus reducing the redundancy of the extracted speech-visual features. In addition, the comparison with comparative methods 2 and 4 shows that the proposed CETAA is superior to average pooling and LSTM network-based aggregation. This is because CETAA designs CT-MHA to adaptively calculate the temporal aggregation information from adjacent short time periods, effectively simulating the time-varying characteristics of emotions.
[0066] Table 1
[0067]
[0068]
[0069] Example 2:
[0070] The embodiment of the present invention also provides a speech-visual emotion recognition device, including:
[0071] A feature extraction module for extracting speech-visual features at the segment level;
[0072] A processing module for inputting the speech-visual features extracted at the segment level into an adaptive alignment and temporal aggregation network; obtaining emotion video features of multiple short time segments;
[0073] An identification module for performing emotion recognition based on the emotion video features of multiple short time segments.
[0074] As an implementation manner of the embodiment of the present invention, the adaptive alignment and temporal aggregation network includes: a low-redundancy speech-visual adaptive alignment module for obtaining low-redundancy alignment features for each short time period and a temporal adaptive aggregation module for adaptively aggregating the aligned features within a long time period.
[0075] As an implementation manner of the embodiment of the present invention, the processing process of the feature extraction module is as follows: Let X = [x 1 , …, x n , …, x Nrepresents an emotional video containing N segments, where x n represents the nth segment, and each segment x n is divided into a speech modality and a visual modality Input the speech modality and the visual modality into a speech feature extractor and a visual feature extractor respectively obtain segment-level speech features and visual features That is:
[0076]
[0077] Among them,
[0078] Embodiment 3:
[0079] An embodiment of the present invention also provides a speech-visual emotion recognition system, including: a memory and a processor, where a computer program run by the processor is stored on the memory, and the computer program executes a speech-visual emotion recognition method when run by the processor.
[0080] Embodiment 4:
[0081] An embodiment of the present invention also provides a storage medium, where a computer program is stored on the storage medium, and the computer program executes a speech-visual emotion recognition method when running.
[0082] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention should fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for speech and visual emotion recognition, characterized in that: include: Step S1, extracting segment-level speech visual features; Step S2, inputting the extracted segment-level speech and visual features into an adaptive alignment and time aggregation network to obtain emotional video features of multiple short-time segments; Step S3: performing emotion recognition based on the emotion video features of multiple short time segments.
2. The method for speech and visual emotion recognition according to claim 1, characterized in that: The adaptive alignment and time aggregation network includes: a low-redundancy speech-vision adaptive alignment module for obtaining low-redundancy alignment features for each short period of time and a time-adaptive aggregation module for adaptively aggregating the aligned features in a long period of time.
3. The method for speech and visual emotion recognition according to claim 1, wherein: In step S1, let X = [x1,…,x n ,…,x N ] represents an emotional video containing N clips, where x n represents the nth fragment, each fragment x n Speech Mode and visual modality Voice Mode and visual modality Input to speech feature extractor and visual feature extractor Get segment-level speech features and visual features Right now: in, 4. A speech and visual emotion recognition device, characterized in that: include: Feature extraction module, used to extract segment-level speech visual features; A processing module is used to input the extracted segment-level speech and visual features into an adaptive alignment and time aggregation network to obtain emotional video features of multiple short-time segments; The recognition module is used to perform emotion recognition based on the emotion video features of multiple short time segments.
5. The speech-visual emotion recognition device according to claim 4, characterized in that: The adaptive alignment and time aggregation network includes: a low-redundancy speech-vision adaptive alignment module for obtaining low-redundancy alignment features for each short period of time and a time-adaptive aggregation module for adaptively aggregating the aligned features in a long period of time.
6. The speech and visual emotion recognition device according to claim 5, characterized in that: The processing of the feature extraction module is as follows: Let X = [x1,…,x n ,…,x N ] represents an emotional video containing N clips, where x n represents the nth fragment, each fragment x n Speech Mode and visual modality Voice Mode and visual modality Input to speech feature extractor and visual feature extractor Get segment-level speech features and visual features Right now: in, 7. A speech and visual emotion recognition system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the method for speech-visual emotion recognition as described in any one of claims 1 to 3 is executed.
8. A storage medium, characterized in that: The storage medium stores a computer program, which, when running, executes the speech-visual emotion recognition method as described in any one of claims 1 to 3.