Multi-element emotion recognition method and device based on electroencephalogram spatiotemporal dynamic characterization
Patent Information
- Application Number
- CN202511170600.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-08-20
AI Technical Summary
第一,情绪模型过于简单(即情绪维度单一化)
第一,针对现有情绪识别方法中的个体差异鲁棒性差的问题,本申请通过确定用户在多个情绪亚群类型中对应的目标情绪亚群类型,并获取目标情绪亚群类型对应的目标解码策略的方式进行解决。现有技术中通常将个体间的系统性差异视为噪声并抹平,而本申请利用这种差异,即人群中存在情绪反应模式显著不同的亚群(如高情绪多样性组与高情绪颗粒度组),通过为不同亚群匹配专属的解码策略,能够进行精细化、个性化的解码,显著提升情绪识别模型在不同用户间的泛化能力和鲁棒性。
Smart Images

Figure CN121265045B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the interdisciplinary fields of biomedical engineering and artificial intelligence, and in particular to a method and device for multi-emotion recognition based on spatiotemporal dynamic representation of brain waves. Background Technology
[0002] Emotion recognition is a key technology at the intersection of artificial intelligence and biomedical engineering. It aims to enable machines to understand human emotional states by analyzing users' physiological or behavioral signals. This technology is crucial for building more natural and empathetic human-computer interactions and shows great potential for applications in areas such as mental health assessment, intelligent education, and assisted diagnosis of mental illnesses.
[0003] Existing technologies for emotion recognition via EEG face at least three core bottlenecks: First, emotion models are too simplistic (i.e., they are too narrow in their emotional dimensions). Traditional methods can usually only identify single, discrete emotions such as "joy," "anger," and "sadness," or locate them on a two-dimensional valence-arousal model. This is seriously inconsistent with the complex state of multiple emotions that people experience in reality (for example, when watching an angry video, one experiences anger, sadness, and disgust at the same time), resulting in the loss of a large amount of emotional information.
[0004] Second, the model exhibits poor robustness across individuals (poor robustness to individual differences). EEG signals are highly individual-specific. Current technologies attempt to homogenize data from different individuals through simple linear transformations, but they ignore the systematic differences in emotional response patterns within a population, treating these fundamental differences as noise. This results in poor model generalization ability and low accuracy.
[0005] Third, the extraction of neural features is static (i.e., insufficient utilization of dynamic features). Traditional methods rely heavily on static features such as energy in specific frequency bands (e.g., power spectral density), failing to capture the millisecond-level, rapidly changing dynamics of brain networks during emotional processing. Studies using EEG microstates have demonstrated that brain states exhibit unique, rapidly shifting patterns across different frequency bands. This dynamic information is crucial for understanding emotions but is overlooked by current technologies. Summary of the Invention
[0006] This application provides a method and device for multi-emotion recognition based on spatiotemporal dynamic representation of EEG, which addresses three shortcomings in existing emotion recognition technologies and enables accurate recognition of users' multi-emotions.
[0007] This application provides a multi-emotion recognition method based on spatiotemporal dynamic representation of brainwaves, including: Acquire the brainwave signals generated by the user while watching the target video; Frequency decomposition of the EEG signal yields a time-domain feature map representing the temporal dynamic characteristics of the EEG signal at multiple frequency scales. Based on the time-domain feature map, a spatiotemporal dynamic representation of the EEG signal is obtained, which represents the dynamic evolution pattern of the EEG signal in the whole brain space. Determine the target emotion subgroup type corresponding to the user among multiple emotion subgroup types, and obtain the target decoding strategy corresponding to the target emotion subgroup type; The target decoding strategy is used to decode the spatiotemporal dynamic representation of EEG to obtain the user's multiple emotions.
[0008] This application also provides a multi-emotion recognition device based on spatiotemporal dynamic representation of brainwaves, characterized in that it includes: The first acquisition module is used to acquire the brain signals generated by the user while watching the target video; The frequency decomposition module is used to perform frequency decomposition on the EEG signal to obtain a time-domain feature map representing the temporal dynamic characteristics of the EEG signal at multiple frequency scales. The second acquisition module is used to acquire, based on the time-domain feature map, a spatiotemporal dynamic representation of the EEG signal representing the dynamic evolution pattern of the EEG signal in the whole brain space. The determination module is used to determine the target emotion subgroup type corresponding to the user among multiple emotion subgroup types, and to obtain the target decoding strategy corresponding to the target emotion subgroup type; The decoding module is used to decode the spatiotemporal dynamic representation of the EEG using the target decoding strategy to obtain the user's multiple emotions.
[0009] The multi-emotion recognition method based on spatiotemporal dynamic representation of EEG proposed in this application has at least the following technical advantages: First, addressing the issue of poor robustness to individual differences in existing emotion recognition methods, this application solves the problem by identifying the target emotion subgroup type corresponding to a user among multiple emotion subgroup types and obtaining the target decoding strategy corresponding to the target emotion subgroup type. Existing technologies typically treat systematic differences between individuals as noise and smooth them out. However, this application utilizes this difference—that is, the existence of subgroups with significantly different emotional response patterns within the population (such as a high emotion diversity group and a high emotion granularity group)—by matching exclusive decoding strategies to different subgroups, enabling refined and personalized decoding, significantly improving the generalization ability and robustness of the emotion recognition model across different users.
[0010] Secondly, addressing the issue of insufficient utilization of dynamic features in existing emotion recognition methods, this application solves the problem through two consecutive feature extraction steps: frequency decomposition of EEG signals to obtain a time-domain feature map, and acquisition of a spatiotemporal dynamic representation of EEG signals representing their dynamic evolution patterns across the whole brain. Unlike existing technologies that rely on static frequency domain features or fixed microstate templates, the dynamic evolution pattern proposed in this application can directly model the continuous changes in the whole-brain potential topography, effectively capturing key dynamic information such as EEG microstate transitions, thereby revealing the dynamic nature of emotion processing.
[0011] Third, addressing the issue of limited emotional dimensions in existing emotion recognition methods, this application ultimately obtains multiple emotional dimensions of the user, resolving the drawback of existing methods that forcibly categorize complex emotions into a single label, resulting in significant information loss. This application no longer outputs isolated labels such as "happy" or "angry," but instead outputs a result containing multiple emotional dimensions (such as anger, sadness, disgust, etc.), which can realistically and comprehensively represent the complex state of multiple coexisting emotions felt by the user in a specific context. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart illustrating a multi-emotion recognition method based on spatiotemporal dynamic representation of brainwaves, as shown in one embodiment of this application.
[0014] Figure 2 This is a schematic diagram illustrating a variety of emotional experiences across different emotional subgroups, as shown in one embodiment of this application.
[0015] Figure 3 This is a schematic diagram illustrating the process of obtaining positive and negative sample pairs in one embodiment of this application.
[0016] Figure 4 This is a schematic diagram of an encoder shown in one embodiment of this application.
[0017] Figure 5 This is a schematic diagram of a projector shown in one embodiment of this application.
[0018] Figure 6 This is a schematic diagram illustrating the training process of a multi-emotion recognition model according to an embodiment of this application.
[0019] Figure 7This is a structural block diagram of a multi-emotion recognition device based on spatiotemporal dynamic representation of brainwaves, as shown in one embodiment of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] To address the problems existing in the prior art, this application provides a method for multi-emotion recognition based on spatiotemporal dynamic representation of electroencephalogram (EEG). The method in this application is executed by any electronic device with data processing capabilities. This electronic device is equipped with a pre-trained multi-emotion recognition model. After the collected electroencephalogram (EEG) signals of the user are input into the multi-emotion recognition model, the model can output the recognized multi-emotions of the user.
[0022] Figure 1 This is a flowchart illustrating a multi-emotion recognition method based on spatiotemporal dynamic representation of electroencephalogram (EEG) according to an embodiment of this application. (Refer to...) Figure 1 The multi-emotion recognition method of this application may include the following steps: Step 101: Obtain the EEG signals generated by the user while watching the target video.
[0023] In step 101, the user can watch a target video whose multiple emotions are to be predicted, and the EEG signals generated by the user are recorded simultaneously.
[0024] In practice, the EEG cap can be used to collect the user's brain signals in real time while the user is watching the target video.
[0025] After acquiring EEG signals, the continuous EEG signal stream can be segmented into short windows of fixed length, for example, in units of 1 second, as the basic unit for subsequent processing. In this application, the EEG signal of one segmented data window is referred to as an EEG signal unit X. In this application, after acquiring the initial EEG signal, the initial EEG signal can also be preprocessed to improve the quality of the EEG signal. This application does not impose specific limitations on the preprocessing strategy.
[0026] Step 102: Perform frequency decomposition on the EEG signal to obtain a time-domain feature map representing the temporal dynamic characteristics of the EEG signal at multiple frequency scales.
[0027] In this application, multiple frequency scales include five frequency bands: Delta (1-4 Hz), Theta (4-8 Hz), Alpha (8-13 Hz), Beta (13-30 Hz), and Gamma (30-47 Hz). Therefore, by performing step 102 to perform frequency decomposition on the EEG signal, a temporal feature map representing the temporal dynamic characteristics of the EEG signal in five different frequency bands can be obtained.
[0028] The temporal feature map is the result of deep, adaptive frequency domain decomposition and dynamic feature extraction of the acquired EEG signals. It contains multi-scale frequency information corresponding to different frequency bands. The temporal feature map preserves complete time series information, describing the intensity fluctuations and changes of neural activity in different frequency bands at each time point, i.e., the temporal dynamic characteristics.
[0029] In this application, a temporal feature map can be obtained for each EEG signal unit X. Here, X is a two-dimensional matrix with dimensions... , This represents the number of brainwave channels (e.g., 30). This represents the number of sampling points (i.e., time points) within that time window (e.g., for a 1-second window at a 250Hz sampling rate). (250).
[0030] In this application, X can be processed using temporal convolution operations to obtain a temporal feature map. The dimension of a temporal feature map is... ,in, This is the total number of temporal convolution kernels (e.g., 160). , The meaning is the same as before. The temporal convolution process will be described in detail later.
[0031] Step 103: Based on the time-domain feature map, obtain the spatiotemporal dynamic representation of EEG that represents the dynamic evolution pattern of EEG signals in the whole brain space.
[0032] In this application, the spatiotemporal dynamic representation of EEG is the result of deep modeling of the continuous change process of the spatial distribution of brain potentials based on the time-domain feature map.
[0033] The spatiotemporal dynamic representation of EEG encompasses dynamic evolution patterns across the entire brain. This application no longer analyzes static EEG topography but instead directly captures millisecond-level features of transitions from one whole-brain potential distribution pattern to another through specific spatial dynamic convolution operations. Physiologically, this corresponds to the transition process of EEG microstates. Furthermore, these spatial dynamic evolution patterns are deeply integrated with the frequency information extracted in step 102. Therefore, the spatiotemporal dynamic representation of EEG in this application is a feature tensor containing complete information across three dimensions: frequency, space, and time. Each value represents the intensity of a specific dynamic change (e.g., a shift from prefrontal activation to parietal activation) of a specific frequency component (e.g., a beta wave) across the entire brain at a specific moment. The spatial dynamic convolution operation will be described in detail later.
[0034] In this application, a spatiotemporal dynamic representation of EEG can be obtained for each temporal feature map.
[0035] Step 104: Determine the target emotion subgroup type corresponding to the user among multiple emotion subgroup types, and obtain the target decoding strategy corresponding to the target emotion subgroup type.
[0036] In this application, before implementing step 101, multiple emotion subgroup types can be predetermined, as well as the target emotion subgroup type to which the user's emotion type belongs among the multiple emotion subgroup types.
[0037] In one implementation, determining the target emotion subgroup type to which a user belongs among multiple emotion subgroup types may include: After a user watches multiple different videos, obtain the user's rating for each video across multiple preset emotional dimensions; Based on the scores, the diversity of negative emotions, the diversity of positive emotions, the correlation between positive and negative emotions, the co-occurrence of positive and negative emotions, the granularity of negative emotions, and the granularity of positive emotions were determined. By inputting the diversity of negative emotions, the diversity of positive emotions, the correlation between positive and negative emotions, the co-occurrence of positive and negative emotions, the granularity of negative emotions, and the granularity of positive emotions into a pre-trained latent profile analysis model, the target emotion subgroup type is obtained.
[0038] The pre-defined emotional dimensions include four negative emotional dimensions and four positive emotional dimensions. The four negative emotions are anger, disgust, fear, and sadness, while the four positive emotions are pleasure, motivation, happiness, and warmth.
[0039] After watching each video, users need to provide an 8-dimensional rating vector for each video based on the above 8 emotional dimensions.
[0040] In addition, this application proposes six indicators of emotional complexity: negative emotion diversity, positive emotion diversity, correlation between positive and negative emotions, co-occurrence of positive and negative emotions, granularity of negative emotions, and granularity of positive emotions. These indicators together constitute a six-dimensional vector that reflects the user's comprehensive emotional response characteristics.
[0041] Each emotional complexity indicator is calculated based on an 8-dimensional rating vector provided by the user. For negative / positive emotion diversity: this indicator is calculated based on four negative emotions and four positive emotions respectively. A higher value indicates a greater richness and evenness in the individual's experience of negative / positive emotions. For the correlation between positive and negative emotions: this indicator is derived by calculating the Pearson correlation coefficient between positive and negative emotion ratings across all videos. A higher value indicates a stronger ability for the individual to tolerate both positive and negative emotions. For the co-occurrence of positive and negative emotions: this indicator calculates the frequency of co-occurrence of positive and negative emotions. A higher value indicates a higher frequency of co-occurrence of positive and negative emotions in the individual's emotional experience. For the granularity of negative / positive emotions: this indicator is calculated based on the intra-class correlation (ICC) algorithm. A higher value indicates a higher level of finesse in the individual's differentiation between different emotions when experiencing negative or positive emotions.
[0042] In this application, multiple emotion subgroups include a high emotion diversity group (Cluster A), a high emotion granularity group (Cluster B), and a moderate-level control group (Cluster C). The following is a combination of... Figure 2 This section provides a brief explanation of the meaning of each emotion subgroup type. Figure 2 This is a schematic diagram illustrating a variety of emotional experiences across different emotional subgroups, as shown in one embodiment of this application.
[0043] High Emotional Diversity Group (Cluster A): This subgroup is characterized by the richness and complexity of their emotional experiences. On the emotional complexity index, participants in this group exhibited high levels of positive and negative emotional diversity, as well as high levels of co-occurrence of positive and negative emotions; conversely, their emotional granularity level was lower. From the perspective of specific emotion rating models (such as...) Figure 2 As shown in the figure, when watching videos, even if they can feel the dominant emotion, they tend to experience multiple strong non-dominant emotions at the same time. This experience is not limited to similar emotions (e.g., multiple negative emotions co-occurring), but also includes cross-valence emotion mixing (e.g., experiencing warmth when sad).
[0044] High Emotional Granularity Group (Cluster B): The core characteristic of this subgroup is the precision and differentiation of their emotional experiences. Their emotional complexity indicators, in contrast to the high emotional diversity group, show high levels of granularity in both positive and negative emotions, while emotional diversity and co-occurrence levels are lower. Regarding emotion rating patterns (such as...) Figure 2 As shown in the image, the participants' experience was highly focused, with the dominant emotion score evoked by each video being significantly higher than the scores for all other non-dominant emotions. This indicates that they were able to clearly distinguish and experience the dominant emotion, while being less affected by other emotions, resulting in a low level of emotional co-occurrence.
[0045] The moderate-level control group (Cluster C) performed moderately across all emotional response patterns and can serve as a baseline for the other two groups. On the six indicators of emotional complexity, this group's scores fell between the high emotional diversity group and the high emotional granularity group. Their emotional rating patterns (such as...) Figure 2 The data (shown) also reflects this moderate level of characteristics: they can experience clear dominant emotions, but at the same time they also feel a certain degree of non-dominant emotions. The degree of emotional mixing is not as rich as that of the high emotional diversity group, but it is not as highly differentiated and singular as that of the high emotional granularity group.
[0046] Dominant emotion refers to the strongest emotion experienced by a user while watching a specific video. Non-dominant emotion refers to all other emotions experienced simultaneously with the strongest dominant emotion.
[0047] In this application, a large number of experiments can be conducted in advance to divide the group into a high emotional diversity group (Cluster A), a high emotional granularity group (Cluster B), and a moderate level control group (Cluster C). This application does not impose specific restrictions on the division method.
[0048] The process of determining the target emotional subgroup type for a user will be described in detail below, including the following steps: Step 1: Data Collection and Calibration. Have users watch several pre-set video clips that evoke different emotions (e.g., select 8 videos covering 8 dominant emotions from the THU-EP or FACED datasets). After watching, users are required to rate their feelings about each video on 8 emotional dimensions (anger, disgust, fear, sadness, pleasure, encouragement, happiness, warmth).
[0049] Step 2: Calculation of Emotional Complexity Indicators. Based on user rating data, calculate six indicators representing their personal emotional response characteristics, including: negative emotion diversity, positive emotion diversity, correlation between positive and negative emotions, co-occurrence of positive and negative emotions, granularity of negative emotions, and granularity of positive emotions.
[0050] Step 3: Emotion Subgroup Classification. The calculated 6-dimensional emotion complexity index vector is input into a pre-trained Latent Profile Analysis (LPA) model. The model will output the user's most likely emotion subgroup classification.
[0051] In this application, each emotion subgroup type has its own exclusive decoding strategy (decoder), that is, the high emotion diversity group (Cluster A), the high emotion granularity group (Cluster B), and the medium level control group (Cluster C) each have their own exclusive decoding strategy.
[0052] Step 105: Decode the spatiotemporal dynamic representation of EEG using a target decoding strategy to obtain the user's diverse emotions.
[0053] In step 105, after determining the target decoding strategy applicable to the user, the spatiotemporal dynamic representation of the user's EEG is decoded through the target decoding strategy, thereby obtaining the user's multiple emotions.
[0054] In this application, multiple emotions are not a single emotion (such as happiness or anger), but an 8-dimensional scoring vector containing the aforementioned 8 preset emotional dimensions. Each element in this vector represents the intensity of a specific emotion experienced. For example, when outputting a user's multiple emotions, instead of simply outputting anger, a graph similar to [Anger: 5.02, Sadness: 4.27, Nausea: 2.74, ...] is output. According to this result, it can be seen that the user experienced strong anger, accompanied by relatively high intensity of sadness and moderate intensity of nausea.
[0055] The multi-emotion recognition method proposed in this application can effectively solve the problems in the prior art: First, addressing the issue of poor robustness to individual differences, this application solves the problem by identifying the target emotion subgroup type corresponding to a user among multiple emotion subgroup types and obtaining the target decoding strategy corresponding to the target emotion subgroup type. Existing technologies typically treat systematic differences between individuals as noise and smooth them out. However, this application leverages these differences—that is, the existence of subgroups with significantly different emotional response patterns within the population (such as a high emotion diversity group and a high emotion granularity group)—by matching exclusive decoding strategies to different subgroups. This enables refined and personalized decoding, significantly improving the generalization ability and robustness of the multi-emotion recognition model across different users.
[0056] Secondly, addressing the issue of insufficient utilization of dynamic features, this application solves the problem through two consecutive feature extraction steps: frequency decomposition of EEG signals to obtain a time-domain feature map, and acquisition of a spatiotemporal dynamic representation of EEG signals representing their dynamic evolution patterns across the whole brain. Unlike existing technologies that rely on static frequency domain features or fixed microstate templates, the dynamic evolution pattern proposed in this application can directly model the continuous changes in the whole-brain potential topography, effectively capturing key dynamic information such as EEG microstate transitions, thereby revealing the dynamic nature of emotional processing.
[0057] Third, addressing the issue of singular emotional dimensions, this application ultimately obtains the user's diverse emotions, resolving the drawback of existing methods that forcibly categorize complex emotions into a single label, resulting in significant information loss. This application no longer outputs isolated labels such as happiness or anger, but instead outputs a result encompassing multiple emotional dimensions (e.g., anger, sadness, disgust), which can realistically and comprehensively represent the complex state of multiple coexisting emotions felt by the user in a specific context.
[0058] In conjunction with the above embodiments, in one implementation, step 102 may include: Step 1021: Input the EEG signal into a multi-scale temporal convolution unit. Through multiple sets of parallel one-dimensional temporal convolution kernels in the multi-scale temporal convolution unit, perform one-dimensional convolution operation on the EEG signal along the time dimension to obtain the preliminary temporal feature map output by each set of one-dimensional temporal convolution kernels. Step 1022: Integrate the various preliminary time-domain feature maps to obtain the time-domain feature map; Among them, different one-dimensional time-domain convolution kernels correspond to different frequency bands, namely Delta, Theta, Alpha, Beta and Gamma, each of which corresponds to a set of one-dimensional time-domain convolution kernels.
[0059] In this application, the multi-scale temporal convolutional unit is located within the encoder. The multi-scale temporal convolutional unit comprises five sets of parallel one-dimensional temporal convolutional kernels. Each set of one-dimensional temporal convolutional kernels contains multiple (e.g., 32) one-dimensional convolutional kernels. The size of each convolutional kernel is... ,in, The number of time points covered by the convolutional kernel is denoted as follows: In this application, the parameters of these convolutional kernels are denoted as... .
[0060] During processing, the EEG signal unit X (with dimensions of...) Simultaneously, five sets of parallel one-dimensional temporal convolution kernels are input, and one-dimensional convolution operations are performed independently on the time series of each EEG channel to extract temporal variation patterns under different frequency bands. For each EEG channel, the convolution kernel is used along the time axis ( Perform sliding calculations on the dimension, i.e., execute The operation, in which This represents the convolution operation.
[0061] In this application, before performing convolution, zero-padding can be applied to both ends of the time dimension of X to ensure that the length of the output temporal feature map in the time dimension is the same as that of the input. It remains unchanged. After the convolution operation, it can be processed by a non-linear activation function (such as ReLU) to increase the non-linear expressive power.
[0062] Since the five sets of one-dimensional temporal convolutional kernels are parallel, X is simultaneously input into all five sets of kernels, and each set of kernels independently outputs features corresponding to its frequency band. Finally, the outputs (preliminary temporal feature maps) of the five parallel convolutional kernels are integrated to form a higher-dimensional temporal feature map containing multi-band, multi-channel temporal dynamic information. Specifically, since each convolutional kernel is used in all... A preliminary temporal feature map is generated on each channel. Assuming each group of temporal convolutional kernels has 32 kernels, then a total of 5 × 32 = 160 preliminary temporal feature maps are generated across the five groups. These 160 preliminary temporal feature maps are stacked along the new feature dimension, ultimately outputting a three-dimensional tensor, i.e., the temporal feature map, with a dimension of . . This is the total number of temporal convolution kernels (e.g., 160). It refers to the number of brainwave channels (e.g., 30). This is the number of time points (e.g., 250). Compared to the original EEG signal, this temporal feature map contains dynamic information at different frequency scales after adaptive filtering by the model.
[0063] In this application, by employing multiple sets of parallel one-dimensional temporal convolution kernels corresponding to different frequency bands, the problem of frequency domain information loss caused by the fixed frequency band division in traditional Fourier transform methods can be solved. Specifically, this application does not use fixed mathematical filtering, but rather, through end-to-end training, allows the multi-emotion recognition model to automatically learn the optimal filtering parameters most relevant to the emotion task. This adaptive frequency band decomposition method can more flexibly and accurately capture neural rhythm information that plays a crucial role in multi-emotion recognition.
[0064] In conjunction with the above embodiments, in one implementation, step 1022 may include: Each preliminary temporal feature map is input into the attention network to obtain the weight coefficients corresponding to each frequency band; Based on each weight coefficient, the preliminary time-domain feature maps are weighted and fused to obtain the fusion result, which is a time-domain feature map.
[0065] In this application, when integrating the outputs of five sets of one-dimensional temporal convolutional kernels to form a temporal feature map, the integration method used is not equal-weighted stacking, but weighted fusion based on an attention mechanism. The multi-scale temporal convolutional unit automatically learns which frequency band features should be given more attention when decoding emotion and assigns it higher weight.
[0066] In this application, an attention module is added to the network of multi-scale temporal convolutional units. This module can dynamically calculate the weight coefficients assigned to each frequency band based on the input features. The higher the weight of a frequency band, the greater its influence on subsequent calculations. Specifically, the encoder inputs the preliminary temporal feature maps output from five sets of one-dimensional temporal convolutional kernels into a small, learnable attention network, which outputs a set of fusion coefficients (i.e., weights), for example... , , , , The sum of these weights is 1. Next, the calculated weights are multiplied by the corresponding frequency band feature maps. The five attention-weighted feature maps are then integrated (e.g., by addition or stacking) to form a single, weighted temporal feature map that incorporates attention information. In subsequent processing, this weighted temporal feature map is input into a spatial dynamic convolutional unit for further processing.
[0067] In this application, the attention network, trained end-to-end, automatically learns and calculates the contribution of each frequency band feature to the final emotion recognition task, and assigns different weight coefficients accordingly. The attention network autonomously identifies and assigns higher weights to frequency bands with stronger relevance in emotion representation (e.g., the Beta band), thereby focusing on the most information-rich neural rhythms. This dynamic weighted fusion mechanism enables the attention network to adaptively integrate specific emotion representations distributed across different frequency bands and brain regions based on the characteristics of the input signal. For example, the attention network can learn to co-weight occipital region activity in the Alpha band with temporal lobe activity in the Gamma band to more accurately represent complex, multi-faceted emotional states, thereby significantly improving the representational power and decoding accuracy of the features.
[0068] In conjunction with the above embodiments, in one implementation, step 103 may include: Step 1031: Input the temporal feature map into the spatial dynamic convolution unit. Perform a two-dimensional convolution operation on the temporal feature map through the two-dimensional spatial dynamic convolution kernel in the spatial dynamic convolution unit to obtain the spatiotemporal dynamic representation of EEG.
[0069] Among them, the size of the two-dimensional spatial dynamic convolution kernel covers all EEG channels when collecting EEG signals in the spatial dimension, and covers multiple consecutive time points in the temporal dimension.
[0070] In this application, the encoder also includes a spatial dynamic convolution unit, and the output of the multi-scale temporal convolution unit serves as the input of the spatial dynamic convolution unit. The spatial dynamic convolution unit employs a set of two-dimensional spatial dynamic convolution kernels, the number of which is... (For example, 8).
[0071] In this set of spatial dynamic convolution kernels, the size of each convolution kernel is... . (For example, 30) is the height of the convolutional kernel, which is equal to the total number of EEG channels. This means that when performing calculations, the convolutional kernel can instantly see the potential distribution map of the entire brain, rather than just looking at a local area as in traditional image processing. (For example, 2) represents the width of the convolutional kernel. A value of 2 means that it can process data from two consecutive time points simultaneously (at a sampling rate of 250Hz, one time point is approximately 4 milliseconds, and two time points are approximately 8 milliseconds). Therefore, this set of spatial dynamic convolutional kernels does not learn a static EEG, but rather data from time points. arrive The changing pattern. This application denotes the parameters of this group of dynamic convolutional kernels as... .
[0072] In this application, the obtained time-domain feature maps (using...) (This is represented as) sequentially inputting spatial dynamic convolutional units, and performing sliding calculations on the input temporal feature map for each spatial dynamic convolutional kernel to identify specific, continuous dynamic change patterns occurring throughout the whole brain. Each temporal feature map can yield a corresponding spatiotemporal dynamic representation of EEG.
[0073] Specifically, for the input For each of the temporal feature maps, the spatial dynamic convolutional unit uses... Two-dimensional convolutional kernels along the time axis ( Perform sliding calculations on the dimension, i.e., execute... operate.
[0074] At each sliding position, the convolutional kernel calculates how well the current small region (e.g., covering all 30 channels and two consecutive time points) matches the dynamic pattern it represents. For example, a convolutional kernel might be trained to specifically recognize a dynamic pattern where "alpha wave activity in the frontal lobe decreases within 8 milliseconds, while activity in the parietal lobe increases." When a similar change occurs in the input data, the output of this convolutional kernel will produce a high activation value.
[0075] Finally, the outputs of all spatial dynamic convolution kernels are integrated to form a high-dimensional EEG representation containing complete information in the three dimensions of frequency, space, and time.
[0076] Finally, the spatial dynamic convolution unit outputs a three-dimensional tensor (using... (representation), namely, the spatiotemporal dynamic representation of EEG, its dimensions are . It is the number of spatial dynamic patterns learned (e.g., 8). It is the number of inherited time-domain (frequency) patterns (e.g., 160). It represents the number of time points (e.g., 250).
[0077] The final output spatiotemporal dynamic representation of EEG It is a highly condensed and abstract representation of the original electroencephalogram (EEG) signals. Each value represents the intensity of a specific dynamic change (such as the transition from mode A to mode D) that occurs in the whole brain at a specific moment in a specific frequency component (such as a beta wave).
[0078] In this application, a two-dimensional spatial dynamic convolutional kernel is designed to extend the traditional analysis of static EEG spatial patterns to the direct modeling of dynamic transition sequences. Specifically, because this two-dimensional spatial dynamic convolutional kernel covers all EEG channels spatially and multiple consecutive time points temporally, the model can directly capture the continuous changes in the whole-brain potential topography at the millisecond level, thereby achieving quantitative representation of key neural dynamics such as EEG microstate transitions at the algorithm level. This design can transform the specific spatial patterns of microstates in different frequency bands into learnable parameters, significantly improving the representational power of features. Experimental results demonstrate that when the number of covered time points is 2, this method achieves optimal performance in multi-emotion recognition tasks.
[0079] In conjunction with the above embodiments, in one implementation, step 105 may include: Step 1051: Perform differential entropy calculation and temporal smoothing on the spatiotemporal dynamic representation of EEG to obtain the smoothed feature vector after dimensionality reduction; Step 1052: Using a target decoding strategy, decode the smoothed feature vector after dimensionality reduction to obtain the user's diverse emotions.
[0080] In this application, before decoding the spatiotemporal dynamic representation of EEG, it is necessary to perform feature compression and smoothing on the spatiotemporal dynamic representation of EEG.
[0081] Steps 1051 and 1052 above are also executed by the encoder.
[0082] First, obtain the spatiotemporal dynamic representation of EEG from the output of the spatial dynamic convolutional unit, representing a single time window (i.e., a single EEG signal unit). , It is a three-dimensional tensor with dimension . .
[0083] Next, the differential entropy (DE) is calculated to compress the temporal information, reducing the duration of each feature channel to 1 second. The dynamic signal (at several time points) is compressed into a single value by calculating differential entropy. This value represents the overall activity or logarithmic energy spectrum of the feature within that second. Since EEG signals of a certain length can approximately follow a Gaussian distribution, the complex calculation of differential entropy can be simplified, as shown in the formula: ,in Let V be the variance of the input signal. For the input tensor... Each feature channel in (total) Each of the given variables (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 16, 17, 18, 19 ... , dimension Thus far, the time series dimension... It was successfully compressed.
[0084] Next, to enhance the stability of the features and remove potential noise, a temporal smoothing algorithm is used to smooth the calculated features. Further processing is then performed. This application employs the Linear Dynamical Systems (LDS) smoothing algorithm for denoising. The LDS algorithm comprehensively considers the DE eigenvalues of the current time window and its adjacent time windows to generate an optimal estimate of the current window's features, thereby filtering out abrupt, non-physiological fluctuations. This step outputs a smoothed and denoised two-dimensional matrix, with dimensions still being [dimension O(n].] . Next, the compressed and smoothed two-dimensional feature matrix is flattened into a one-dimensional long vector to meet the input requirements of the final decoding model (fully connected neural network). Specifically, this is done through standard matrix transformation operations to flatten the two-dimensional... The matrix is flattened into a one-dimensional vector in sequence, and the final output is a one-dimensional eigenvector. , length is .
[0085] In this application, the high-dimensional dynamic representation of temporal sequence is compressed into stable features with high information density through differential entropy calculation. Then, irrelevant noise in the signal is effectively removed through temporal smoothing processing, which can provide high-quality and robust input feature vectors for subsequent decoding.
[0086] Regarding encoder computation optimization, this application employs a hierarchical parameter sharing strategy to reduce the complexity of the encoder network. The parameters of multi-scale temporal convolutional units are shared in the spatial dynamic convolutional unit section, maintaining the independence of each frequency band while controlling the growth of the number of parameters. This design allows the encoder to maintain its performance advantages while significantly reducing the number of parameters compared to traditional stacked architectures, meeting the requirements of real-time processing. Furthermore, this application can alleviate the gradient vanishing problem in deep network training by introducing residual connections and layer normalization.
[0087] Compared to traditional methods, the features extracted by the encoder in this application can more accurately locate brain regions related to emotion processing, such as the anterior insula and posterior cingulate cortex, core nodes of the salience network. Particularly when dealing with co-occurring emotion scenarios, the system successfully captures changes in Gamma band activity caused by the ablation of fear through dynamic convolutional kernels—a feature that existing static feature extraction methods cannot achieve at all.
[0088] In conjunction with the above embodiments, in one implementation, the encoder is pre-trained through the following steps: Positive and negative sample pairs were constructed. A positive sample pair consisted of EEG signal samples generated by two different subjects watching the same video at the same time. A negative sample pair consisted of an EEG signal sample from a single subject and any EEG signal sample that could not be paired with it to form a positive sample pair. Positive and negative sample pairs are input into the encoder in sequence to obtain the corresponding spatiotemporal dynamic representation samples of EEG. The spatiotemporal dynamic representation samples of EEG output from the encoder are input into the projector to obtain the hidden layer representation vector. The projector is used to nonlinearly map the features extracted by the encoder into the low-dimensional hidden space. Based on the spatiotemporal dynamic representation samples of EEG output and the hidden layer representation vectors, a loss value is determined. The contrastive loss function used to determine the loss value is constructed with the goal of maximizing the similarity of the representation vectors of all positive sample pairs and minimizing the similarity of the representation vectors of all negative sample pairs. Based on the loss value, the network parameters in the encoder and projector are updated through backpropagation.
[0089] In this application, to address the commonalities amidst individual differences, a powerful encoder needs to be trained to map the highly differentiated raw EEG signals of any individual to a common, consistent feature space shared by all individuals. The entire training process involves a data sampler, an encoder, and a projector.
[0090] This application employs a contrastive learning method to train the encoder. The training process specifically includes the following steps: Step 1: The data sampler generates multiple pairs of EEG signal samples for training.
[0091] The data sampler generates batches containing multiple pairs of EEG signal samples. Specifically, data sampling for each batch is conducted with two subjects: EEG signal samples from the same time period are randomly selected from each video segment (assuming there are 24 videos) from both subjects, and represented as a set. , The dimension is In this batch, for a given sample Only with samples Forming positive sample pairs, while forming negative sample pairs with other samples (e.g.) Figure 3 (As shown). One training cycle will iterate through all pairwise pairings of subjects in the training set, therefore for The number of participants can obtain Group subjects were paired. This application uses a 5-second time window to divide samples, with adjacent samples overlapping by 3 seconds; therefore, approximately 13 EEG signal samples can be obtained from each video segment (30 seconds). Figure 3 In the middle, using EEG signal samples For example, in its batch, only samples They form positive sample pairs, while they form negative sample pairs with other samples. Figure 3 This is a schematic diagram illustrating the process of obtaining positive and negative sample pairs in one embodiment of this application.
[0092] Step 2: Extract spatiotemporal dynamic representation samples of EEG from positive and negative sample pairs by the encoder.
[0093] The encoder acquires the spatiotemporal dynamic representation of EEG, representing a single time window, from the output of the spatial dynamic convolution step. , dimension ,express The spatiotemporal variation pattern of EEG shows the activity intensity at each time point.
[0094] Step 3: The projector maps the spatiotemporal dynamic representation of EEG obtained in Step 2 to another latent space.
[0095] The projector consists of a time-dimension average pooling layer and two time-domain convolutional layers, employing... Activation function. Specifically, for the encoder output. First, an average pooling layer is used to extract the average value of the signal over a period of time. The pooling layer size is 𝑆, and the output is... , dimension , where ⌊. ⌋ represents rounding down. Then, two layers of temporal convolution are used to further extract the temporal variation pattern, using the following formula: .in, , (i = 1, 2, ...) ), where are the parameters of the two temporal convolution layers, respectively. The number of convolution kernels, This covers the number of time points. Finally, the projector output... Transform into a long vector This is used for subsequent comparison loss calculation.
[0096] Step 4: Calculate the contrast loss.
[0097] The contrastive loss is defined based on the similarity of sample pairs. Specifically, the output of a batch of projectors is represented as a set. Then the similarity between two samples is defined as their hidden layer representations. Cosine similarity: Since the contrastive learning strategy maximizes the similarity of EEG representations corresponding to the same stimulus (positive sample pairs) while minimizing the similarity of EEG representations corresponding to different stimuli (negative sample pairs), this application, referencing existing research, uses normalized cross-entropy loss to define the contrastive loss for each sample. For example, the loss values are as follows: Therefore, the contrastive loss function for this batch is: Step 5: Optimize model parameters.
[0098] Using optimizers like Adam, gradients are backpropagated based on the calculated contrastive loss, and the network parameters of the encoder and projector are updated. This process is repeated on all paired subjects to obtain a well-trained encoder. This encoder is generalizable.
[0099] Among them, the encoder is as follows Figure 4 As shown, the projector is as follows Figure 5 As shown. Figure 4 This is a schematic diagram of an encoder shown in one embodiment of this application. Figure 5 This is a schematic diagram illustrating a projector according to an embodiment of this application. Figure 5 middle, It calculates the new length of the data in the time dimension after the data has passed through the average pooling layer. It calculates the length of the data in the time dimension after the data passes through the first layer of temporal convolution. It calculates the final length of the data in the time dimension after the data undergoes the second layer of temporal convolution. It is the total number of feature channels after the data passes through the first temporal convolution layer in the projector. It is the final total number of feature channels after the data passes through the second temporal convolution layer in the projector.
[0100] In conjunction with the above embodiments, in one implementation, step 105 may include: The spatiotemporal dynamic representation of brainwaves is input into a decoder containing a target decoding strategy. The decoder decodes the spatiotemporal dynamic representation of brainwaves to obtain the user's multiple emotions. The decoder is obtained by performing regression training on a multi-layer fully connected neural network, with the input being the spatiotemporal dynamic representation samples of EEG of users belonging to the target emotion subgroup in the training set, and the output being the multivariate emotion samples corresponding to each spatiotemporal dynamic representation sample.
[0101] In this application, the training process of the decoder includes: The weights of the loss values corresponding to various emotions in the multiple emotions are determined based on the emotion co-occurrence matrix. The emotion co-occurrence matrix is a matrix that quantifies the intensity of multiple different emotions appearing simultaneously and being related to each other in the experience. The elements of the emotion co-occurrence matrix are the similarity between the ratings of two different emotions. Based on the loss values corresponding to various emotions and the weights of the loss values corresponding to various emotions, the predicted output corresponding to the spatiotemporal dynamic representation sample of EEG is determined, and the total loss value between it and the multi-emotion sample is determined. Based on the total loss value, the network parameters in the multi-layer fully connected neural network are updated to obtain the decoder. In this application, the training method for the decoder corresponding to each emotion subgroup type is as follows: Step 1: Divide the users in the training set into subgroups.
[0102] For all subjects in the training set, six emotional complexity indicators were calculated based on their emotional rating reports. The six-dimensional emotional complexity indicator vectors of all subjects were used as input to run the LPA algorithm, which clustered them into multiple subgroups with different emotional response patterns, resulting in Cluster A, Cluster B, and Cluster C.
[0103] Step 2: Prepare dedicated training data for each emotion subgroup type.
[0104] The entire training set is divided into multiple independent sub-training sets based on the partitioning result of step 1. For example, the EEG data and emotion scores of all subjects belonging to Cluster A constitute the Cluster A training set. The EEG data of each sub-training set are input into a pre-trained encoder for feature extraction, resulting in compressed and smoothed feature vectors. .
[0105] Step 3: Independently train the dedicated decoder for each emotional subgroup.
[0106] Training process for the decoder for the high emotion diversity group (Cluster A): 1) Prepare dedicated training data: Select all participants classified as high emotional diversity group (Cluster A) from the overall training set. Then, extract the spatiotemporal dynamic representations (i.e., feature vectors) of these participants' EEG data. Together with their corresponding multi-dimensional emotion rating labels, they form a dedicated training dataset for use only by Cluster A.
[0107] 2) Calculate a specific emotion co-occurrence matrix: Using only the multivariate emotion rating data of the Cluster A participant group, calculate a matrix according to the following formula. The emotion co-occurrence matrix is called . It accurately reflects the unique emotional co-occurrence patterns of people with high emotional diversity (usually, multiple emotions co-occur to a high degree).
[0108] 3) Training a dedicated decoder: Initialize a brand new three-layer fully connected neural network (FCN) decoder specifically for Cluster A; process the feature vectors from the Cluster A training set. The FCN decoder is input to obtain predicted multivariate sentiment scores. A weighted mean squared error loss function is used to calculate the loss between the predicted scores and the true score labels. When calculating the loss for each training sample, the true dominant sentiment of that sample is considered. From the exclusive matrix Identify the dominant emotion and other non-dominant emotions within the context of the analysis. co-present value and count it backwards ( This weight is multiplied by the prediction error corresponding to the non-dominant emotion. The network parameters of the FCN decoder are updated using the calculated weighted total loss via backpropagation. This process is repeated until the model converges, resulting in a fully trained decoder specifically adapted to a population with high emotional diversity. The weighted mean squared error loss function is as follows: Training process for the decoder for the high-sensory-granularity group (Cluster B): 1) Prepare dedicated training data: Select all participants classified into the high-emotion-granularity group (Cluster B) from the overall training set. Then, extract the spatiotemporal dynamic representations (i.e., feature vectors) of these participants' EEG data. Together with their corresponding multi-dimensional emotion rating labels, they form a dedicated training dataset for use only by Cluster B.
[0109] 2) Calculate a proprietary emotion co-occurrence matrix: using only the multivariate emotion rating data from the Cluster B participant group, based on... The calculation formula yields an emotion co-occurrence matrix, called... . It accurately reflects the unique emotional co-occurrence patterns of people with high emotional granularity (who usually have a low degree of co-occurrence between emotions).
[0110] 3) Train a dedicated decoder. Initialize a brand new, three-layer fully connected neural network (FCN) decoder specifically for Cluster B. This involves processing the feature vectors from the Cluster B training set. Inputting this FCN decoder yields predicted multivariate sentiment scores. A weighted mean squared error loss function (as above) is used to calculate the loss between the predicted score and the true score label. When calculating the loss for each training sample, the system considers the true dominant sentiment of that sample. From the exclusive matrix Identify the dominant emotion and other non-dominant emotions within the context of the analysis. co-present value and count it backwards ( This weight is multiplied by the prediction error of the corresponding non-dominant emotion. The network parameters of the FCN decoder are updated based on the calculated weighted total loss using the backpropagation algorithm. This process is repeated until the model converges, ultimately resulting in a fully trained decoder specifically adapted to high-sensory-granularity audiences.
[0111] Training process of the decoder for the intermediate-level control group (Cluster C): 1) Prepare dedicated training data: Select all subjects who were assigned to the intermediate-level control group (Cluster C) from the overall training set. Then, extract the spatiotemporal dynamic representations of their EEG data (i.e., feature vectors). Together with their corresponding multi-dimensional emotion rating labels, they form a dedicated training dataset for use only by Cluster C.
[0112] 2) Calculate a proprietary emotion co-occurrence matrix: using only multivariate emotion rating data from the Cluster C participant group, based on... The calculation formula yields an emotion co-occurrence matrix, called... . It accurately reflects the average emotional co-occurrence pattern of the average population at a moderate level.
[0113] 3) Train a dedicated decoder: Initialize a brand new, three-layer fully connected neural network (FCN) decoder specifically for Cluster C. This involves processing the feature vectors from the Cluster C training set. The input to the FCN decoder yields predicted multivariate sentiment scores. A weighted mean squared error loss function (as above) is used to calculate the loss between the predicted scores and the true score labels. When calculating the loss for each training sample, the true dominant sentiment d for that sample is derived from the dedicated matrix. Find the co-occurrence value of the dominant emotion with each of the other non-dominant emotions i. and count it backwards ( This weight is multiplied by the prediction error corresponding to the non-dominant sentiment. The network parameters of the FCN decoder are updated based on the calculated weighted total loss using a backpropagation algorithm. This process is repeated until the model converges, ultimately resulting in a fully trained decoder specifically adapted to a moderate-level control population.
[0114] In this application, the training process for the encoder and decoder can be referred to Figure 6 As shown. Figure 6 This is a schematic diagram illustrating the training process of a multi-emotion recognition model according to an embodiment of this application.
[0115] In existing technologies, systematic differences between individuals are usually treated as noise and smoothed out. However, this application utilizes such differences, namely, the existence of subgroups with significantly different emotional response patterns in the population (such as a high emotional diversity group and a high emotional granularity group). By matching exclusive decoding strategies to different subgroups, refined and personalized decoding can be performed, which significantly improves the generalization ability and robustness of the emotion recognition model among different users.
[0116] This application also provides a dynamic adaptation mechanism. This mechanism innovatively integrates individual differences in two dimensions: long-term traits and short-term states, enabling continuous personalized optimization of the multi-emotion recognition model. On one hand, the device can dynamically fine-tune the parameters of the pre-trained multi-emotion recognition model based on the user's long-term stable emotional complexity index (i.e., traits). For example, when a high co-occurrence index of positive and negative emotions is detected, the device automatically enhances the weight coefficients of specific neural features such as micro-state C, making the multi-emotion recognition model more adaptable to the user's personal response style. On the other hand, the device can also adapt to the user's short-term emotional response fluctuations (i.e., states) by updating the emotion co-occurrence covariance matrix online. This dual adaptation mechanism, combining traits and states, allows the multi-emotion recognition model to evolve from a static pre-trained model into an adaptive model that grows alongside the user, thus maintaining high recognition stability when facing special individuals with high emotional diversity.
[0117] In this application, a strategy combining shared and dedicated methods is adopted during the training phase. All emotion subgroups share a common encoder to extract common features, but each emotion subgroup is trained with a dedicated projector and decoder. This design significantly improves recognition accuracy (e.g., the MAE of non-dominant emotion recognition in Cluster B is reduced by approximately 12.8%) with only a small increase in the number of parameters, achieving a balance between efficiency and performance.
[0118] Furthermore, the method in this application incorporates a dynamic emotion subgroup assignment algorithm in its practical application. When a new user's EEG characteristics differ significantly from the centers of all known subgroups, a new emotion subgroup generation mechanism is automatically triggered. (The model employs a dynamic algorithm to handle new users who do not fully conform to any known emotion subgroup (A, B, C). It determines their affiliation by calculating the distance between the new user's EEG characteristics and the centers of existing emotion subgroups. If all distances exceed a preset threshold, the model automatically triggers a new emotion subgroup generation mechanism.) This enables the model to recognize and adapt to new emotional response patterns, evolving from a static classification system into a dynamically evolving adaptive system.
[0119] The features extracted in this application possess high biological interpretability, revealing the neural basis behind different emotional response pattern subgroups. Source localization analysis validates that the subgroup-specific representations extracted by the model accurately reflect different emotion processing strategies. For example, individuals with high emotional diversity (Cluster A) rely on prefrontal-temporal co-activation, while individuals with high granularity (Cluster B) exhibit stronger parietal cortex activation. This interpretable differential modeling allows the model to maintain high recognition accuracy even when faced with individual differences, providing an effective new paradigm for solving the cold-start problem (i.e., inaccurate recognition of new users) common in the field of emotion-based brain-computer interfaces.
[0120] In this application, an emotion co-occurrence matrix is calculated to quantify the mixing strength and direction of different emotions. The model calculates the cosine similarity between each pair of eight emotion ratings, resulting in a matrix that reveals complex correlation patterns. This quantitative analysis provides a scientific basis for subsequent dynamic weight allocation. For example, experimental data shows that the co-occurrence strength of the high emotion diversity group (Cluster A) (0.38) is significantly higher than that of the high granularity group (Cluster B) (0.12). Ultimately, this application can adaptively handle mixed emotion scenarios, successfully transforming relationships such as the co-occurrence pattern of anger and sadness (r=0.45) into collaborative constraints within the model, thereby significantly improving the accuracy and realism of multi-emotion recognition.
[0121] This application optimizes model training through a dual attention mechanism: on the one hand, a fixed weight is applied to the dominant emotion to ensure basic recognition accuracy; on the other hand, the loss weight is dynamically adjusted based on the co-occurrence intensity of non-dominant emotions and the dominant emotion, forcing the model to learn to predict those accompanying emotions that are more difficult to capture. This application also draws on neuroscience findings from ablation experiments, specifically strengthening the weight coefficients of highly arousing emotions such as fear. This ingenious loss function design significantly reduces the mean absolute error (MAE) of the model on multi-emotion recognition tasks compared to traditional methods, greatly improving the overall predictive performance of the multi-emotion recognition model.
[0122] In this application, the system internally employs an innovative hierarchical emotion relation graph network to finely model different emotional features and their interactions. This network consists of two layers: the lower layer handles the basic neural representations of a single emotion, such as associating alpha-band occipital region activity with positive emotions, while the upper layer further models the dynamic interactions between different emotion categories through a cross-attention mechanism. SHAP analysis provides evidence for this, indicating that the model successfully learns the key contributions of specific brain activities, such as high-frequency prefrontal cortex activity, to cross-valence emotion mixing. Ultimately, the system utilizes a learnable 8 The 8-correlation matrix transforms this complex neural interaction pattern into computable weights, enabling a highly integrated end-to-end mapping from EEG dynamic features to multivariate emotion scores.
[0123] Secondly, this application employs two refined strategies to optimize computational efficiency and training stability. First, a phased training strategy is used: the first phase pre-trains the base network using all emotion dimension data, and the second phase performs targeted fine-tuning based on the emotion co-occurrence matrix. This strategy successfully controls the number of model parameters within the limits of traditional ensemble regression chain methods while ensuring high performance in dominant emotion recognition, achieving a balance between efficiency and accuracy. Secondly, this application also sets differentiated convergence thresholds for emotion subgroups with varying emotional complexity. For example, a 15% loss tolerance is relaxed for the high emotion diversity group (Cluster A), which exhibits higher emotional volatility. This adaptive training control makes the training process of each subgroup model more stable and efficient, avoiding overfitting to noise.
[0124] This application can effectively solve three problems in the prior art.
[0125] First, regarding the problem of overly simplistic emotion models.
[0126] This application addresses this problem by constructing a multi-objective regression architecture. It abandons the traditional single-emotion classification framework, which forcibly categorizes complex emotional experiences into a single label. Instead, it defines the emotion recognition task as simultaneously predicting the intensity of eight different emotion dimensions. To accurately model the complex co-occurrence relationships among these emotions, an emotion co-occurrence weighted mean square error (MSE) algorithm is employed during training. This algorithm dynamically adjusts the prediction error weights of each non-dominant emotion in the loss function using an emotion co-occurrence matrix calculated from the training data. This design forces the model to learn and output a multivariate emotion vector containing eight-dimensional ratings that reflects the coexistence of multiple emotions, thereby fully preserving the user's authentic and mixed emotional experience information and solving the serious information loss problem in traditional methods.
[0127] Second, it addresses the issue of poor robustness in cross-individual decoding.
[0128] This application employs a dual strategy of first addressing commonalities and then individual differences to tackle the problem of individual variations. First, to extract common neural representations among individuals, a contrastive learning framework is used to train the feature extraction encoder. By constructing positive and negative sample pairs and optimizing using a contrastive loss function, the encoder is forced to learn a universal representation that maps the EEG signals of different individuals to a unified feature space. Second, to handle systematic individual variations, the application uses Latent Profile Analysis (LPA) to divide users into multiple subgroups with different response patterns (such as a high emotional diversity group) based on their emotional complexity index. Then, a dedicated decoder is trained for each subgroup, and the emotional co-occurrence matrix used by this decoder is calculated only based on the data of its corresponding subgroup. Through this strategy combining universal feature extraction and subgroup-specific decoding, the model can effectively address individual variations and significantly improve the robustness and accuracy of cross-individual recognition.
[0129] Third, addressing the issue of staticizing neural feature extraction.
[0130] This application replaces traditional static feature extraction methods with a multi-scale temporal convolutional unit. The network first uses multi-scale temporal convolution, employing multiple parallel convolutional kernels corresponding to different physiological frequency bands, to adaptively decompose the EEG signal into frequency bands, capturing the temporal dynamic characteristics within each band. Secondly, subsequent spatial dynamic convolutional layers use special two-dimensional convolutional kernels, whose size spatially covers the entire EEG channel and temporally covers multiple consecutive time points. This design allows the model to directly model the continuous evolution of the whole-brain potential topography, rather than analyzing static EEG spatial features. This enables the model to effectively capture crucial neural dynamic information such as millisecond-level EEG microstate transitions, thus solving the problem that traditional static features cannot characterize the dynamic nature of emotional processing.
[0131] The following describes the multi-emotion recognition device based on spatiotemporal dynamic representation of EEG provided in this application. The multi-emotion recognition device described below can be referred to in correspondence with the multi-emotion recognition method described above.
[0132] Figure 7 This is a structural block diagram of a multi-emotion recognition device based on spatiotemporal dynamic representation of electroencephalography (EEG) according to an embodiment of this application. (Refer to...) Figure 7 The multi-emotion recognition device 700 based on spatiotemporal dynamic representation of brainwaves in this application may include: The first acquisition module 701 is used to acquire the brain signals generated by the user while watching the target video; The frequency decomposition module 702 is used to perform frequency decomposition on the EEG signal to obtain a time-domain feature map representing the temporal dynamic characteristics of the EEG signal at multiple frequency scales. The second acquisition module 703 is used to acquire, based on the time-domain feature map, a spatiotemporal dynamic representation of the EEG signal representing the dynamic evolution pattern of the EEG signal in the whole brain space. The determination module 704 is used to determine the target emotion subgroup type to which the user belongs among multiple emotion subgroup types, and to obtain the target decoding strategy corresponding to the target emotion subgroup type; The decoding module 705 is used to decode the spatiotemporal dynamic representation of the EEG using the target decoding strategy to obtain the user's multiple emotions.
[0133] According to the multi-emotion recognition device 700 provided in this application, the frequency decomposition module 702 is specifically used for: inputting the EEG signal into a multi-scale temporal convolution unit; performing a one-dimensional convolution operation on the EEG signal along the time dimension through multiple sets of parallel one-dimensional temporal convolution kernels in the multi-scale temporal convolution unit to obtain preliminary temporal feature maps output by each set of one-dimensional temporal convolution kernels; integrating each of the preliminary temporal feature maps to obtain the temporal feature map; wherein, different one-dimensional temporal convolution kernels correspond to different frequency bands.
[0134] According to the multi-emotion recognition device 700 provided in this application, the frequency decomposition module 702 is further configured to: input each of the preliminary temporal feature maps into an attention network to obtain the weight coefficients corresponding to each frequency band; and perform weighted fusion on each of the preliminary temporal feature maps according to each of the weight coefficients to obtain a fusion result, wherein the fusion result is the temporal feature map.
[0135] According to the multi-emotion recognition device 700 provided in this application, the second acquisition module 703 is specifically used to: input the temporal feature map into the spatial dynamic convolution unit, and perform a two-dimensional convolution operation on the temporal feature map through the two-dimensional spatial dynamic convolution kernel in the spatial dynamic convolution unit to obtain the spatiotemporal dynamic representation of the EEG; wherein, the size of the two-dimensional spatial dynamic convolution kernel covers all EEG channels when the EEG signal is acquired in the spatial dimension, and covers multiple consecutive time points in the temporal dimension.
[0136] According to the multi-emotion recognition device 700 provided in this application, the decoding module 705 is specifically used for: performing differential entropy calculation and temporal smoothing processing on the spatiotemporal dynamic representation of the EEG to obtain a dimensionality-reduced smooth feature vector; and decoding the dimensionality-reduced smooth feature vector through the target decoding strategy to obtain the user's multi-emotion.
[0137] According to the multi-emotion recognition device 700 provided in this application, the decoding module 705 is further configured to: input the EEG spatiotemporal dynamic representation to a decoder containing the target decoding strategy, and decode the EEG spatiotemporal dynamic representation through the decoder to obtain the user's multi-emotion; wherein, the decoder is obtained by performing regression training on a multi-layer fully connected neural network with EEG spatiotemporal dynamic representation samples of users belonging to the target emotion subgroup type in the training set as input and multi-emotion samples corresponding to each EEG spatiotemporal dynamic representation sample as output.
[0138] According to the multi-emotion recognition device 700 provided in this application, the training process of the decoder includes: determining the weights of the loss values corresponding to various emotions in the multi-emotion recognition based on the emotion co-occurrence matrix, wherein the emotion co-occurrence matrix is a matrix that quantifies the intensity of multiple different emotions appearing simultaneously and being interrelated in an experience, and the elements of the emotion co-occurrence matrix are the similarity between the ratings of two different emotions; determining the total loss value between the predicted output corresponding to the EEG spatiotemporal dynamic representation sample and the multi-emotion sample based on the loss values corresponding to the various emotions and the weights of the loss values corresponding to the various emotions; and updating the network parameters in the multi-layer fully connected neural network based on the total loss value to obtain the decoder.
[0139] According to the multi-emotion recognition device 700 provided in this application, the multi-scale temporal convolutional unit and the spatial dynamic convolutional unit are located in the encoder, which is pre-trained through the following steps: Construct positive and negative sample pairs. A positive sample pair is a sample pair consisting of EEG signal samples generated by two different subjects watching the same video at the same time. A negative sample pair is a sample pair consisting of an EEG signal sample of a single subject and any EEG signal sample that cannot be paired with it to form a positive sample pair. The positive sample pairs and the negative sample pairs are sequentially input into the encoder to obtain the corresponding spatiotemporal dynamic representation samples of EEG. The various spatiotemporal dynamic representation samples of EEG output by the encoder are nonlinearly mapped to a low-dimensional latent space to obtain the latent layer representation vector. Based on the spatiotemporal dynamic representation samples of EEG output by the encoder and the hidden layer representation vector, a loss value is determined. The contrastive loss function used to determine the loss value is constructed with the goal of maximizing the similarity of the representation vectors of all positive sample pairs and minimizing the similarity of the representation vectors of all negative sample pairs. Based on the loss value, the network parameters in the encoder and the projector are updated through backpropagation.
[0140] According to the multi-emotion recognition device 700 provided in this application, the determining module 704 is specifically used for: after a user watches multiple different videos, obtaining the user's ratings for each video on multiple preset emotion dimensions; determining the diversity of negative emotions, the diversity of positive emotions, the correlation between positive and negative emotions, the co-occurrence of positive and negative emotions, the granularity of negative emotions, and the granularity of positive emotions based on the ratings; and inputting the diversity of negative emotions, the diversity of positive emotions, the correlation between positive and negative emotions, the co-occurrence of positive and negative emotions, the granularity of negative emotions, and the granularity of positive emotions into a pre-trained latent profile analysis model to obtain the target emotion subgroup type.
[0141] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0142] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A multi-emotion recognition method based on spatiotemporal dynamic representation of brainwaves, characterized in that, include: Acquire the brainwave signals generated by the user while watching the target video; Frequency decomposition of the EEG signal yields a time-domain feature map representing the temporal dynamic characteristics of the EEG signal at multiple frequency scales. Based on the time-domain feature map, a spatiotemporal dynamic representation of the EEG signal is obtained, which represents the dynamic evolution pattern of the EEG signal in the whole brain space. Determining the target emotion subgroup type corresponding to the user among multiple emotion subgroup types includes: after the user watches multiple different videos, obtaining the user's ratings for each video on multiple preset emotion dimensions; based on the ratings, determining the diversity of negative emotions, the diversity of positive emotions, the correlation between positive and negative emotions, the co-occurrence of positive and negative emotions, the granularity of negative emotions, and the granularity of positive emotions; inputting the diversity of negative emotions, the diversity of positive emotions, the correlation between positive and negative emotions, the co-occurrence of positive and negative emotions, the granularity of negative emotions, and the granularity of positive emotions into a pre-trained latent profile analysis model to obtain the target emotion subgroup type, and obtaining the target decoding strategy corresponding to the target emotion subgroup type; The target decoding strategy is used to decode the spatiotemporal dynamic representation of the EEG to obtain the user's multiple emotions. This includes: inputting the spatiotemporal dynamic representation of the EEG into a decoder containing the target decoding strategy, and decoding the spatiotemporal dynamic representation of the EEG through the decoder to obtain the user's multiple emotions; wherein the decoder is obtained by performing regression training on a multilayer fully connected neural network with EEG spatiotemporal dynamic representation samples of users belonging to the target emotion subgroup type in the training set as input and the multiple emotion samples corresponding to each EEG spatiotemporal dynamic representation sample as output. The training process of the decoder includes: determining the weights of the loss values corresponding to various emotions in the multi-emotion spectrum based on the emotion co-occurrence matrix, wherein the emotion co-occurrence matrix is a matrix that quantifies the intensity of multiple different emotions occurring simultaneously and being interrelated in an experience, and the elements of the emotion co-occurrence matrix are the similarity between the ratings of two different emotions; determining the total loss value between the predicted output corresponding to the EEG spatiotemporal dynamic representation sample and the multi-emotion sample based on the loss values corresponding to the various emotions and the weights of the loss values corresponding to the various emotions; and updating the network parameters in the multi-layer fully connected neural network based on the total loss value to obtain the decoder.
2. The multi-emotion recognition method according to claim 1, characterized in that, The step of performing frequency decomposition on the EEG signal to obtain a time-domain feature map representing the temporal dynamic characteristics of the EEG signal at multiple frequency scales includes: The EEG signal is input into a multi-scale temporal convolutional unit. Through multiple sets of parallel one-dimensional temporal convolutional kernels in the multi-scale temporal convolutional unit, a one-dimensional convolution operation is performed on the EEG signal along the time dimension to obtain the preliminary temporal feature map output by each set of one-dimensional temporal convolutional kernels. For each of the aforementioned preliminary time-domain features The graphs are integrated to obtain the time-domain feature map; Different one-dimensional time-domain convolution kernels correspond to different frequency bands.
3. The multi-emotion recognition method according to claim 2, characterized in that, The process of integrating the preliminary time-domain feature map to obtain the time-domain feature map includes: Each of the preliminary temporal feature maps is input into the attention network to obtain the weight coefficients corresponding to each frequency band; Based on the respective weight coefficients, the respective preliminary temporal feature maps are weighted and fused to obtain a fusion result, which is the temporal feature map.
4. The multi-emotion recognition method according to claim 2, characterized in that, The step of obtaining a spatiotemporal dynamic representation of the EEG signal representing its dynamic evolution pattern in the whole brain space based on the time-domain feature map includes: The temporal feature map is input into the spatial dynamic convolution unit. The two-dimensional convolution operation is performed on the temporal feature map through the two-dimensional spatial dynamic convolution kernel in the spatial dynamic convolution unit to obtain the spatiotemporal dynamic representation of EEG. The size of the two-dimensional spatial dynamic convolution kernel covers all EEG channels when the EEG signals are acquired in the spatial dimension, and covers multiple continuous sampling points in the temporal dimension.
5. The multi-emotion recognition method according to any one of claims 1-3, characterized in that, The process of decoding the spatiotemporal dynamic representation of the EEG using the target decoding strategy to obtain the user's multiple emotions includes: Differential entropy calculation and temporal smoothing are performed on the spatiotemporal dynamic representation of EEG to obtain a smoothed feature vector after dimensionality reduction; The target decoding strategy is used to decode the dimensionality-reduced smooth feature vector to obtain the user's multiple emotions.
6. The multi-emotion recognition method according to claim 4, characterized in that, The multi-scale temporal convolutional unit and the spatial dynamic convolutional unit are located in the encoder, which is pre-trained through the following steps: Construct positive and negative sample pairs. A positive sample pair is a sample pair consisting of EEG signal samples generated by two different subjects watching the same video at the same time. A negative sample pair is a sample pair consisting of an EEG signal sample of a single subject and any EEG signal sample that cannot be paired with it to form a positive sample pair. The positive sample pairs and the negative sample pairs are sequentially input into the encoder to obtain the corresponding spatiotemporal dynamic representation samples of EEG. The various spatiotemporal dynamic representation samples of EEG output by the encoder are nonlinearly mapped to a low-dimensional latent space to obtain the latent layer representation vector. Based on the spatiotemporal dynamic representation samples of EEG output by the encoder and the hidden layer representation vector, a loss value is determined. The contrastive loss function used to determine the loss value is constructed with the goal of maximizing the similarity of the representation vectors of all positive sample pairs and minimizing the similarity of the representation vectors of all negative sample pairs. Based on the loss value, the network parameters in the encoder are updated through backpropagation to obtain the encoder.
7. A multi-emotion recognition device for implementing the multi-emotion recognition method based on spatiotemporal dynamic representation of EEG as described in any one of claims 1-6, characterized in that, include: The first acquisition module is used to acquire the brain signals generated by the user while watching the target video; The frequency decomposition module is used to perform frequency decomposition on the EEG signal to obtain a time-domain feature map representing the temporal dynamic characteristics of the EEG signal at multiple frequency scales. The second acquisition module is used to acquire, based on the time-domain feature map, a spatiotemporal dynamic representation of the EEG signal representing the dynamic evolution pattern of the EEG signal in the whole brain space. A determination module is used to determine the target emotion subgroup type corresponding to the user among multiple emotion subgroup types, and to obtain the target decoding strategy corresponding to the target emotion subgroup type. Specifically, the determination module is used to: after the user watches multiple different videos, obtain the user's ratings for each video on multiple preset emotion dimensions; based on the ratings, determine the diversity of negative emotions, the diversity of positive emotions, the correlation between positive and negative emotions, the co-occurrence of positive and negative emotions, the granularity of negative emotions, and the granularity of positive emotions; input the diversity of negative emotions, the diversity of positive emotions, the correlation between positive and negative emotions, the co-occurrence of positive and negative emotions, the granularity of negative emotions, and the granularity of positive emotions into a pre-trained latent profile analysis model to obtain the target emotion subgroup type. A decoding module is used to decode the spatiotemporal dynamic representation of the EEG using the target decoding strategy to obtain the user's multiple emotions. Specifically, the decoding module is used to: input the spatiotemporal dynamic representation of the EEG into a decoder containing the target decoding strategy, and decode the spatiotemporal dynamic representation of the EEG using the decoder to obtain the user's multiple emotions; wherein, the decoder is obtained by performing regression training on a multilayer fully connected neural network with EEG spatiotemporal dynamic representation samples of users belonging to the target emotion subgroup type in the training set as input and the multiple emotion samples corresponding to each EEG spatiotemporal dynamic representation sample as output; The training process of the decoder includes: determining the weights of the loss values corresponding to various emotions in the multi-emotion spectrum based on the emotion co-occurrence matrix, wherein the emotion co-occurrence matrix is a matrix that quantifies the intensity of multiple different emotions occurring simultaneously and being interrelated in an experience, and the elements of the emotion co-occurrence matrix are the similarity between the ratings of two different emotions; determining the total loss value between the predicted output corresponding to the EEG spatiotemporal dynamic representation sample and the multi-emotion sample based on the loss values corresponding to the various emotions and the weights of the loss values corresponding to the various emotions; and updating the network parameters in the multi-layer fully connected neural network based on the total loss value to obtain the decoder.