Emotion recognition method based on multi-view contrast learning of brain activation region

By employing a multi-view contrastive learning method based on brain activation regions, data augmentation and feature decoupling are performed on EEG signals, addressing the issues of existing methods' dependence on labeled data and insufficient feature utilization, thus achieving more efficient emotion recognition results.

CN119498849BActive Publication Date: 2026-01-06XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411577096.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2026-01-06
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

Existing EEG emotion recognition methods rely on large amounts of labeled data. Data augmentation strategies fail to effectively protect key emotional information and lack spatiotemporal feature decoupling methods, affecting the model's generalization ability and recognition accuracy.

Method used

We employ a multi-view contrastive learning method based on brain activation regions. By pre-training a dual-path decoupled multi-view encoder, we utilize high-confidence positive samples for data augmentation, implement specific data augmentation strategies for different brain regions, and combine cross-view knowledge exchange of temporal and spatial features to construct an emotion recognition model.

Benefits of technology

It improves the accuracy and reliability of emotion recognition, enhances the diversity of data samples and the generalization ability of the model, and ensures the protection of key information and the effective extraction of features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119498849B_ABST
    Figure CN119498849B_ABST
Patent Text Reader

Abstract

The emotion recognition method based on brain activation area multi-view contrast learning provided by the application comprises the following steps: obtaining an emotion signal to be recognized; performing space-time transformation on the emotion signal to be recognized to obtain a space-time signal to be recognized; inputting the space-time signal to be recognized into a pre-trained dual-path decoupling multi-view model to obtain an emotion recognition result; by setting two encoders for extracting time features and space features in the final dual-path decoupling multi-view encoder, multi-dimensional information in the space-time signal to be recognized can be fully mined and utilized, and the decoupling of corresponding space-time features is realized; by dividing the space-time signal sample into a high activation area and a low activation area according to the activation degree of the brain area, a targeted data enhancement strategy can be performed according to the different activation degrees of the brain area, and the pre-trained dual-path decoupling multi-view model is trained by using the data-enhanced sample, thereby improving the accuracy and reliability of the pre-trained dual-path decoupling multi-view model for emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of biosignal processing and artificial intelligence, specifically to an emotion recognition method based on multi-view comparative learning of brain activation areas. Background Technology

[0002] Emotion, as a crucial component of complex human physiological and psychological states, occupies a central position in cognition, decision-making, behavior, and interpersonal interaction. In recent years, with technological advancements, emotion recognition technology has demonstrated broad application potential in various fields such as human-computer interaction, mental health assessment, and affective computing. Compared to overt emotional signals such as facial expressions and voice tone, electroencephalography (EEG), as an intrinsic physiological indicator with high temporal and spatial resolution, can more directly and objectively reflect an individual's true emotional state, thus attracting significant attention in emotion recognition research. EEG signals capture brain electrical activity through multiple electrodes, possessing rich spatiotemporal characteristics and providing unique and complementary information for a deeper understanding of brain function.

[0003] However, most current EEG emotion recognition methods rely heavily on supervised learning paradigms. While these methods offer excellent performance, their success is highly dependent on large amounts of labeled training data. Unlike image and text data, EEG signals lack intuitive, human-recognizable patterns, making the labeling process complex and time-consuming, significantly increasing data preparation costs. To address this issue, some researchers have begun exploring self-supervised learning methods such as contrastive learning, hoping to reduce reliance on labeled data by designing effective data augmentation and contrastive tasks. For example, some researchers have applied contrastive learning to EEG emotion recognition, employing data augmentation techniques such as temporal masking, linear scaling, and Gaussian noise to learn the representation of EEG signals by maximizing the consistency of the same sample after different augmentations. Furthermore, some researchers have proposed graph-based contrastive learning methods that utilize frequency and spatial mosaics to generate augmented data, aiming to better capture the spatial and frequency features of EEG signals.

[0004] While contrastive learning methods alleviate the reliance on labeled data to some extent, several shortcomings remain. First, existing data augmentation strategies often treat all brain regions equally, resulting in a failure to effectively protect key emotional information. Second, strong data augmentation may destroy key semantic information in EEG signals, while weak augmentation struggles to provide sufficient sample diversity, affecting the model's generalization ability. Finally, existing methods lack effective decoupling mechanisms when processing the spatiotemporal features of EEG signals, failing to fully mine and utilize the multidimensional information within the signals, thus limiting the accuracy and reliability of emotion recognition. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, this invention provides an emotion recognition method based on multi-view comparative learning of brain activation areas.

[0006] The technical problem to be solved by this invention is achieved through the following technical solution:

[0007] This invention provides an emotion recognition method based on multi-view contrastive learning of brain activation areas, comprising:

[0008] Acquire the EEG emotion signals to be identified;

[0009] Spatiotemporal transformation is performed on the EEG emotion signal to be identified to obtain the spatiotemporal signal to be identified;

[0010] The spatiotemporal signal to be identified is input into a pre-trained dual-path decoupled multi-view model to obtain the emotion recognition result. The pre-trained dual-path decoupled multi-view model includes a final dual-path decoupled multi-view encoder and a pre-trained classifier. The final dual-path decoupled multi-view encoder has two encoders, which are used to identify the temporal and spatial features of the spatiotemporal signal to be identified, respectively. The final dual-path decoupled multi-view encoder is constructed based on contrastive learning of high-confidence positive samples and under the cross-view knowledge exchange strategy of temporal and spatial features. Among them, the high-confidence positive samples of contrastive learning are obtained by performing corresponding data augmentation strategies on spatiotemporal signal samples according to the activation level of preset brain regions.

[0011] Optionally, the EEG emotion signal to be identified is subjected to spatiotemporal transformation to obtain the spatiotemporal signal to be identified, including:

[0012] The EEG emotion signal to be identified is segmented into segments along the time dimension using a sliding window method to obtain multiple segments of the signal to be identified.

[0013] Each segment signal to be identified is segmented using a sliding window to obtain multiple sub-segments to be identified containing time segments;

[0014] Multiple sub-segments to be identified are subjected to cosine similarity calculation according to electrode channels, and spatiotemporal signals to be identified are generated based on the results of the cosine similarity calculation.

[0015] Optionally, multiple sub-segments to be identified are subjected to cosine similarity calculation according to electrode channels, and a spatiotemporal signal to be identified is generated based on the result of the cosine similarity calculation, including:

[0016] Cosine similarity is calculated for multiple sub-segments to be identified according to electrode channels to obtain the first cosine similarity result;

[0017] Edges are established between the sub-segments to be identified corresponding to the first cosine similarity results that are greater than the similarity threshold, so as to obtain the spatiotemporal signal to be identified.

[0018] Optionally, the final dual-channel decoupled multi-view encoder is obtained by fine-tuning a pre-trained dual-channel decoupled multi-view encoder using spatiotemporal label samples; the training process of the pre-trained dual-channel decoupled multi-view encoder includes:

[0019] Acquire brainwave emotion signal samples;

[0020] Spatiotemporal transformation of EEG emotion signal samples is performed to obtain spatiotemporal signal samples;

[0021] Calculate the class activation map corresponding to the spatiotemporal signal sample in the preset brain region, and divide the spatiotemporal signal to be identified into high activation region and low activation region according to the class activation map;

[0022] Data augmentation and signal fusion are performed sequentially on high-activation and low-activation regions to obtain fused signal samples; data augmentation includes at least one of the following: spatial mosaicking, temporal mosaicking, frequency band mosaicking, random occlusion, noise injection, or scaling transformation;

[0023] The fused signal samples are input into the initial dual-channel decoupled multi-view encoder and trained through contrastive learning;

[0024] The initial dual-path decoupled multi-view encoder that meets the preset iteration stopping condition is used as the pre-trained dual-path decoupled multi-view encoder.

[0025] The presupposed brain regions include: the prefrontal lobe, frontal lobe, left frontal lobe, right frontal lobe, left temporal lobe, right temporal lobe, central region, left parietal lobe, right parietal lobe, and occipital lobe.

[0026] Optionally, the class activation map corresponding to the spatiotemporal signal sample in a preset brain region is calculated, and the spatiotemporal signal to be identified is divided into high-activation regions and low-activation regions according to the class activation map, including:

[0027] Global average pooling and linear projection are performed sequentially on both the spatiotemporal signal samples and the static benchmark to obtain the sample global embedding and the static global embedding, respectively; the static benchmark is the mean of all spatiotemporal signal samples.

[0028] Calculate the cosine similarity between the global embedding of each sample and the static global embedding to obtain the second cosine similarity result;

[0029] The gradient information of the second cosine similarity result is calculated by backpropagation to obtain the cosine similarity gradient, and the cosine similarity gradient is then subjected to global average pooling to obtain the gradient pooling result.

[0030] The gradient pooling results are averaged according to the preset brain regions to obtain the class activation map;

[0031] The class activation map is compared with a preset activation threshold, and the spatiotemporal signal to be identified is divided into high activation region and low activation region based on the result of the threshold comparison.

[0032] Optionally, the class activation map is compared with a preset activation threshold, and the spatiotemporal signal to be identified is divided into high-activation regions and low-activation regions based on the threshold comparison result, including:

[0033] Compare the class activation graph with a preset activation threshold;

[0034] The brain regions corresponding to activation maps that exceed the preset activation threshold are designated as high-activation regions.

[0035] Brain regions corresponding to activation maps that are below a preset activation threshold are designated as low-activation regions.

[0036] Optionally, data augmentation and signal fusion processing are performed sequentially on the high-activation region and the low-activation region to obtain a fused signal sample, including:

[0037] Weak data augmentation is performed on highly active regions to obtain weakly augmented data;

[0038] Perform strong data augmentation on low-activation regions to obtain strongly augmented data;

[0039] Data fusion processing is performed on the weakly enhanced data and the strongly enhanced data to obtain fused signal samples.

[0040] Optionally, the initial dual-channel decoupled multi-view encoder has the same structure as the pre-trained dual-channel decoupled multi-view encoder, which includes a spatial encoder and a temporal encoder.

[0041] The spatial encoder and the temporal encoder are connected in parallel; the spatial encoder is equipped with the following components connected in series: a first unit, a spatial transformation module, a first word embedding module, a first Transformer encoding module, and a first pooling module;

[0042] The time encoder is configured with the following components connected in series: a second unit, a time conversion module, a second word embedding module, a second Transformer encoding module, and a second pooling module.

[0043] The first unit is provided with a first branch unit and a second branch unit connected in parallel; the first branch unit is provided with a first linear projection module; the second branch unit is provided with a graph encoding module and a second linear projection module connected in series.

[0044] The first and second units have the same structure.

[0045] Optionally, the fused signal samples are input into the initial dual-path decoupled multi-view encoder and trained through contrastive learning, including:

[0046] The fused signal sample is input into the initial dual-channel decoupled multi-view encoder to obtain temporal embedding, global embedding and spatial embedding; temporal embedding is the output of the temporal encoder, spatial embedding is the output of the spatial encoder; global embedding is the concatenation result of temporal embedding and spatial embedding.

[0047] Calculate the similarity between the temporal embedding, global embedding, and spatial embedding and all embeddings in the corresponding queue to obtain the temporal similarity result, global similarity result, and spatial similarity result, respectively.

[0048] The time similarity results are sorted in a first descending order, and the top K times in the first descending order are embedded as the first guidance information.

[0049] The spatial similarity results are sorted in a second descending order, and the top K spatial embeddings in the second descending order are used as the second guidance information.

[0050] The global similarity results are sorted in a third descending order, and the top K global embeddings in the third descending order are used as the third guidance information.

[0051] The first, second, and third guidance information were used as high-confidence positive samples.

[0052] The initial dual-path decoupled multi-view encoder is trained using high-confidence positive samples and contrastive learning to obtain a pre-trained dual-path decoupled multi-view encoder.

[0053] Optionally, the generation process of the pre-trained dual-path decoupled multi-view model includes:

[0054] Acquire pre-trained dual-path decoupled multi-view encoder and labeled spatiotemporal samples;

[0055] The labeled spatiotemporal samples are input into the pre-trained dual-path decoupled multi-view encoder and the initial classifier for iterative fine-tuning training.

[0056] The pre-trained dual-channel decoupled multi-view encoder corresponding to the stopping condition is used as the final dual-channel decoupled multi-view encoder.

[0057] The initial classifier that is reached when the stopping condition is used as the pre-trained classifier;

[0058] The final dual-path decoupled multi-view encoder and the pre-trained classifier together constitute the pre-trained dual-path decoupled multi-view model.

[0059] This invention provides an emotion recognition method based on multi-view contrastive learning of brain activation areas, comprising: acquiring an EEG emotion signal to be recognized; performing a spatiotemporal transformation on the EEG emotion signal to be recognized to obtain a spatiotemporal signal to be recognized; inputting the spatiotemporal signal to be recognized into a pre-trained dual-path decoupled multi-view model to obtain an emotion recognition result; the pre-trained dual-path decoupled multi-view model includes: a final dual-path decoupled multi-view encoder and a pre-trained classifier; the final dual-path decoupled multi-view encoder is equipped with two encoders, which are used to recognize the temporal and spatial features of the spatiotemporal signal to be recognized, respectively; the final dual-path decoupled multi-view encoder is constructed based on contrastive learning of high-confidence positive samples, under a cross-view knowledge exchange strategy for temporal and spatial features; wherein, the high-confidence positive samples for contrastive learning are obtained by performing corresponding data augmentation strategies on spatiotemporal signal samples according to the activation level of preset brain regions. In this invention, by setting two encoders for extracting temporal and spatial features in the final dual-channel decoupled multi-view encoder, the multidimensional information in the spatiotemporal signal to be identified can be fully explored and utilized, achieving decoupling of the corresponding spatiotemporal features. In addition, by dividing the spatiotemporal signal samples into high-activation and low-activation regions according to the activation level of brain regions, targeted data augmentation strategies can be executed according to the different activation levels of brain regions, thereby protecting the key information and enhancing the diversity of spatiotemporal signal samples. Based on this, a pre-trained dual-channel decoupled multi-view model is obtained by training the data-augmented samples, which improves the generalization ability of the pre-trained dual-channel decoupled multi-view model as well as the accuracy and reliability of emotion recognition.

[0060] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0061] Figure 1 This is a flowchart illustrating an emotion recognition method based on multi-view comparative learning of brain activation areas, provided in an embodiment of the present invention.

[0062] Figure 2 This is a schematic diagram of the structure of the initial dual-channel decoupled multi-view encoder provided in an embodiment of the present invention;

[0063] Figure 3 This invention provides an overall architecture for comparative learning based on an initial dual-path decoupled multi-view encoder, as provided in the embodiments of the present invention.

[0064] Figure 4 This is a schematic diagram illustrating the training process of the pre-trained dual-path decoupled multi-view model provided in an embodiment of the present invention. Detailed Implementation

[0065] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0066] To improve the accuracy and reliability of emotion recognition, this invention provides an emotion recognition method based on multi-view comparative learning of brain activation areas. Figure 1 This is a flowchart illustrating an emotion recognition method based on multi-view comparative learning of brain activation areas, provided as an embodiment of the present invention. Figure 1 As shown, it includes:

[0067] S101. Acquire the EEG emotion signal to be identified.

[0068] The emotional signals to be identified are those collected via brain electrodes. In this embodiment of the invention, the emotional signals include, but are not limited to: anger, disgust, fear, sadness, pleasure, joy, inspiration, tenderness, arousal, valence, familiarity, and liking.

[0069] S102. Perform spatiotemporal transformation on the EEG emotion signal to be identified to obtain the spatiotemporal signal to be identified.

[0070] Optionally, S102 may specifically include:

[0071] The EEG emotion signal to be identified is segmented into segments along the time dimension using a sliding window method to obtain multiple segments of the signal to be identified.

[0072] Each segment signal to be identified is segmented using a sliding window to obtain multiple sub-segments to be identified containing time segments;

[0073] Multiple sub-segments to be identified are subjected to cosine similarity calculation according to electrode channels, and spatiotemporal signals to be identified are generated based on the results of the cosine similarity calculation.

[0074] Optionally, multiple sub-segments to be identified are subjected to cosine similarity calculation according to electrode channels, and a spatiotemporal signal to be identified is generated based on the result of the cosine similarity calculation, including:

[0075] Cosine similarity is calculated for multiple sub-segments to be identified according to electrode channels to obtain the first cosine similarity result;

[0076] Edges are established between the sub-segments to be identified corresponding to the first cosine similarity results that are greater than the similarity threshold, so as to obtain the spatiotemporal signal to be identified.

[0077] S103. Input the spatiotemporal signal to be identified into the pre-trained dual-path decoupled multi-view model to obtain the emotion recognition result.

[0078] The pre-trained dual-path decoupled multi-view model includes: a final dual-path decoupled multi-view encoder and a pre-trained classifier; the final dual-path decoupled multi-view encoder has two encoders, which are used to identify the temporal and spatial features of the spatiotemporal signal to be identified, respectively; the final dual-path decoupled multi-view encoder is constructed based on contrastive learning of high-confidence positive samples, under the cross-view knowledge exchange strategy of temporal and spatial features; wherein, the high-confidence positive samples of contrastive learning are obtained by performing corresponding data augmentation strategies on spatiotemporal signal samples according to the activation level of preset brain regions.

[0079] This invention provides an emotion recognition method based on multi-view contrastive learning of brain activation regions. By setting two encoders in the final dual-channel decoupled multi-view encoder for extracting temporal and spatial features, the multi-dimensional information in the spatiotemporal signal to be recognized can be fully explored and utilized, achieving decoupling of corresponding spatiotemporal features. In addition, by dividing the spatiotemporal signal samples into high-activation and low-activation regions according to the activation level of brain regions, targeted data augmentation strategies can be executed according to the different activation levels of brain regions, thereby protecting key information and enhancing the diversity of spatiotemporal signal samples. Based on this, a pre-trained dual-channel decoupled multi-view model is obtained by training the data-augmented samples, improving the generalization ability of the pre-trained dual-channel decoupled multi-view model as well as the accuracy and reliability of emotion recognition.

[0080] Optionally, the final dual-channel decoupled multi-view encoder is obtained by fine-tuning a pre-trained dual-channel decoupled multi-view encoder using spatiotemporal label samples; the training process of the pre-trained dual-channel decoupled multi-view encoder includes:

[0081] Acquire brainwave emotion signal samples;

[0082] Spatiotemporal transformation of EEG emotion signal samples is performed to obtain spatiotemporal signal samples;

[0083] Calculate the class activation map corresponding to the spatiotemporal signal sample in the preset brain region, and divide the spatiotemporal signal to be identified into high activation region and low activation region according to the class activation map;

[0084] Data augmentation and signal fusion are performed sequentially on high-activation and low-activation regions to obtain fused signal samples; data augmentation includes at least one of the following: spatial mosaicking, temporal mosaicking, frequency band mosaicking, random occlusion, noise injection, or scaling transformation;

[0085] The fused signal samples are input into the initial dual-channel decoupled multi-view encoder and trained through contrastive learning;

[0086] The initial dual-path decoupled multi-view encoder that meets the preset iteration stopping condition is used as the pre-trained dual-path decoupled multi-view encoder.

[0087] The presupposed brain regions include: the prefrontal lobe, frontal lobe, left frontal lobe, right frontal lobe, left temporal lobe, right temporal lobe, central region, left parietal lobe, right parietal lobe, and occipital lobe.

[0088] In this embodiment of the invention, the EEG emotion signal samples were obtained through publicly available EEG emotion recognition datasets, such as SEED, SEED-IV, and FACED. These datasets consist of EEG signals recorded by subjects while watching emotion-evoked videos, covering different emotional states and subject groups. The SEED dataset contains EEG data from 15 subjects, collected by 62 electrodes. The SEED-IV dataset also consists of data from 15 subjects, each collected through 62 channels (electrodes). The FACED dataset contains data from 123 subjects, each collected through 30 channels. Furthermore, necessary preprocessing steps, such as filtering, denoising, and standardization, were performed on the EEG emotion signals from the above three datasets to improve data quality. In this embodiment, electrode localization was performed according to the international 10-20 EEG system localization method. Table 1 shows the relationship between brain regions and corresponding electrodes provided in this embodiment of the invention.

[0089] Table 1 Brain regions and corresponding electrodes

[0090]

[0091] Optionally, the class activation map corresponding to the spatiotemporal signal sample in a preset brain region is calculated, and the spatiotemporal signal to be identified is divided into high-activation regions and low-activation regions according to the class activation map, including:

[0092] Global average pooling and linear projection are performed sequentially on both the spatiotemporal signal samples and the static benchmark to obtain the sample global embedding and the static global embedding, respectively; the static benchmark is the mean of all spatiotemporal signal samples.

[0093] Calculate the cosine similarity between the global embedding of each sample and the static global embedding to obtain the second cosine similarity result;

[0094] The gradient information of the second cosine similarity result is calculated by backpropagation to obtain the cosine similarity gradient, and the cosine similarity gradient is then subjected to global average pooling to obtain the gradient pooling result.

[0095] The gradient pooling results are averaged according to the preset brain regions to obtain the class activation map;

[0096] The class activation map is compared with a preset activation threshold, and the spatiotemporal signal to be identified is divided into high activation region and low activation region based on the result of the threshold comparison.

[0097] Specifically, in this embodiment of the invention, the gradient information of the second cosine similarity result is calculated by backpropagation and inversion.

[0098] Optionally, the class activation map is compared with a preset activation threshold, and the spatiotemporal signal to be identified is divided into high-activation regions and low-activation regions based on the threshold comparison result, including:

[0099] Compare the class activation graph with a preset activation threshold;

[0100] The brain regions corresponding to activation maps that exceed the preset activation threshold are designated as high-activation regions.

[0101] Brain regions corresponding to activation maps that are below a preset activation threshold are designated as low-activation regions.

[0102] Optionally, data augmentation and signal fusion processing are performed sequentially on the high-activation region and the low-activation region to obtain a fused signal sample, including:

[0103] Weak data augmentation is performed on highly active regions to obtain weakly augmented data;

[0104] Perform strong data augmentation on low-activation regions to obtain strongly augmented data;

[0105] Data fusion processing is performed on the weakly enhanced data and the strongly enhanced data to obtain fused signal samples.

[0106] In this invention, data augmentation includes at least one of the following: spatial mosaicking, temporal mosaicking, frequency band mosaicking, random occlusion, noise injection, or scaling transformation processing.

[0107] Specifically, the data augmentation process using random occlusion in this embodiment of the invention is as follows:

[0108] EEG samples (high-activation or low-activation regions) are randomly masked according to the feature dimension. Each EEG sample is independently sampled according to a binomial distribution with a masking ratio r, thereby generating a feature masking matrix Mask. C,v The channel retention matrix Mask is generated by preserving all information from a single channel by setting all elements in each row to 1. V,v A probability matrix U∈[0,1] is randomly generated, with its elements sampled from a uniform distribution. A probability threshold σ is set to control the weights of the two masking strategies. For each channel v, the corresponding channel probability value U in the probability matrix U is used... v Choose a masking strategy:

[0109]

[0110] Where, x i,vThis represents the v-th channel (region) of the i-th spatiotemporal signal sample. For x i,v The corresponding masked samples (weakly enhanced data or strongly enhanced data), ⊙ represents the element-wise multiplication operation.

[0111] Furthermore, in this embodiment, weak data augmentation can be applied to highly active regions, i.e., a higher masking ratio r = 0.7 and a probability threshold σ = 0.7 can be set to reduce the masking of highly active regions, retain more original features, and obtain weakly augmented data. Strong data augmentation can be applied to low-active regions, i.e., a lower masking ratio r = 0.3 and a probability threshold σ = 0.3 can be set to obtain strongly augmented data to increase data diversity.

[0112] Furthermore, the data augmentation strategies illustrated above are merely illustrative. Any method that can adaptively perform matching augmentation on low-activation and high-activation regions can be considered as an alternative to the data augmentation methods described herein.

[0113] Optionally, the initial dual-channel decoupled multi-view encoder has the same structure as the pre-trained dual-channel decoupled multi-view encoder, which includes a spatial encoder and a temporal encoder.

[0114] Figure 2 This is a schematic diagram of the initial dual-channel decoupled multi-view encoder provided in an embodiment of the present invention. Figure 2 As shown, the spatial encoder and the temporal encoder are connected in parallel; the spatial encoder is equipped with the following components connected in series: a first unit, a spatial transformation module, a first word embedding module, a first Transformer encoding module, and a first pooling module;

[0115] The time encoder is configured with the following components connected in series: a second unit, a time conversion module, a second word embedding module, a second Transformer encoding module, and a second pooling module.

[0116] The first unit is provided with a first branch unit and a second branch unit connected in parallel; the first branch unit is provided with a first linear projection module; the second branch unit is provided with a graph encoding module and a second linear projection module connected in series.

[0117] The first and second units have the same structure.

[0118] In this embodiment of the invention, the spatial conversion module is used to convert the input signal to a spatial dimension. The time conversion module is used to convert the input signal to a time dimension.

[0119] Optionally, the fused signal samples are input into the initial dual-path decoupled multi-view encoder and trained through contrastive learning, including:

[0120] The fused signal sample is input into the initial dual-channel decoupled multi-view encoder to obtain temporal embedding, global embedding and spatial embedding; temporal embedding is the output of the temporal encoder, spatial embedding is the output of the spatial encoder; global embedding is the concatenation result of temporal embedding and spatial embedding.

[0121] Calculate the similarity between the temporal embedding, global embedding, and spatial embedding and all embeddings in the corresponding queue to obtain the temporal similarity result, global similarity result, and spatial similarity result, respectively.

[0122] The time similarity results are sorted in a first descending order, and the top K times in the first descending order are embedded as the first guidance information.

[0123] The spatial similarity results are sorted in a second descending order, and the top K spatial embeddings in the second descending order are used as the second guidance information.

[0124] The global similarity results are sorted in a third descending order, and the top K global embeddings in the third descending order are used as the third guidance information.

[0125] The first, second, and third guidance information were used as high-confidence positive samples.

[0126] The initial dual-path decoupled multi-view encoder is trained using high-confidence positive samples and contrastive learning to obtain a pre-trained dual-path decoupled multi-view encoder.

[0127] Figure 3 This invention provides an overall architecture for comparative learning based on an initial dual-path decoupled multi-view encoder, as described in embodiments of the invention. Figure 3As shown, under the contrastive learning framework, spatiotemporal signal samples undergo simultaneous training in both the training and contrastive channels. First, the spatiotemporal signal samples are divided into high-activation and low-activation brain regions according to a predefined brain region classification, and corresponding data augmentation is performed based on the activation level. The fused signal samples obtained after augmentation are input into the initial dual-path decoupled multi-view encoder for training, yielding temporal embeddings, spatial embeddings, and global embeddings. These, combined with the temporal, spatial, and global embeddings obtained from the contrastive channel, constitute positive samples. Finally, the positive samples obtained through the training channel and the corresponding negative samples in the queue learn from each other based on contrastive loss, resulting in a pre-trained dual-path decoupled multi-view encoder. The temporal, global, and spatial queues are used to store the temporal, spatial, and global embeddings output by the initial dual-path decoupled multi-view encoder from previous batches, respectively, and serve as negative samples for the current batch. Furthermore, after obtaining the positive samples, high-confidence positive samples can be determined based on the similarity results calculated using cosine similarity. The initial dual-path decoupled multi-view encoder is then trained based on these high-confidence positive samples, negative samples, and contrastive learning, resulting in the pre-trained dual-path decoupled multi-view encoder. The process for determining high-confidence positive samples has been described in the above embodiments, and will not be repeated in this embodiment.

[0128] Furthermore, in this embodiment of the invention, the contrastive learning uses contrastive loss to determine when to stop the iteration. That is, the contrastive learning training process terminates when the number of iterations exceeds a preset iteration threshold or the value of the contrastive loss remains below a preset loss threshold.

[0129] Optionally, the generation process of the pre-trained dual-path decoupled multi-view model includes:

[0130] Acquire pre-trained dual-path decoupled multi-view encoder and labeled spatiotemporal samples;

[0131] The labeled spatiotemporal samples are input into the pre-trained dual-path decoupled multi-view encoder and the initial classifier for iterative fine-tuning training.

[0132] The pre-trained dual-channel decoupled multi-view encoder corresponding to the stopping condition is used as the final dual-channel decoupled multi-view encoder.

[0133] The initial classifier that is reached when the stopping condition is used as the pre-trained classifier;

[0134] The final dual-path decoupled multi-view encoder and the pre-trained classifier together constitute the pre-trained dual-path decoupled multi-view model.

[0135] After obtaining the pre-trained dual-channel decoupled multi-view encoder, a classification head is added for a specific emotion recognition task. Then, the entire model (pre-trained dual-channel decoupled multi-view encoder + initial classifier) ​​is fine-tuned using a small number of labeled spatiotemporal samples. During fine-tuning, the parameters of the pre-trained dual-channel decoupled multi-view encoder are subtly adjusted through end-to-end training to better adapt to the specific needs of emotion recognition, while retaining the rich and general feature representations obtained from the contrastive learning stage. This allows the pre-trained dual-channel decoupled multi-view model to exhibit good generalization ability even with limited labeled data. The cross-entropy loss function is used during the training of the pre-trained dual-channel decoupled multi-view model, and the PyTorch deep learning framework is employed.

[0136] Figure 4 This is a schematic diagram illustrating the training process of a pre-trained dual-path decoupled multi-view model provided in an embodiment of the present invention. Figure 4 As shown, in the first stage, a pre-trained dual-path decoupled multi-view encoder is obtained through unsupervised training. Specifically, the pre-trained dual-path decoupled multi-view encoder is trained through contrastive learning using spatiotemporal signal samples under corresponding data augmentation strategies. In the second stage, a small number of labeled samples (labeled spatiotemporal samples) are used to simultaneously fine-tune the pre-trained dual-path decoupled multi-view encoder and the initial classifier, ultimately yielding the pre-trained dual-path decoupled multi-view model.

[0137] In summary, the emotion recognition method based on multi-view contrastive learning of brain activation areas provided in this invention optimizes data augmentation strategies, adaptively designing strong or weak augmentation strategies for low-activation and high-activation regions respectively. This effectively protects key information in high-activation regions while introducing sufficient diversity to low-activation regions, avoiding the loss of semantic information and excessive similarity of data samples. Furthermore, the adaptive data augmentation method based on activation levels makes data processing more flexible and intelligent, dynamically adjusting the augmentation strategy according to the activation status of different brain regions, thus improving the overall model's flexibility and adaptability. Through adaptive data transformation and multi-view consistency learning, the model exhibits stronger generalization ability and broader adaptability under different data distributions and practical application scenarios. Specifically, the dual-path multi-view encoder fully utilizes the temporal and spatial features in EEG signals, ensuring independent extraction and fusion of temporal and spatial features, improving the accuracy and comprehensiveness of feature representation. Finally, by exchanging high-confidence positive samples (temporal, spatial, and global), the consistency of the contrastive context under different perspectives is maintained, promoting stable and meaningful knowledge mining and utilization.

[0138] The method provided in this embodiment of the invention can be applied to electronic devices. Specifically, the electronic device can be a desktop computer, a portable computer, a smart mobile terminal, a server, etc., and this embodiment of the invention does not limit the application to such devices.

[0139] It should be noted that the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention.

[0140] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0141] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings and the disclosure, will understand and implement other variations of the disclosed embodiments in carrying out the claimed invention. In the description of the invention, the word "comprising" does not exclude other components or steps, "a" or "an" does not exclude a plurality, and "a plurality" means two or more, unless otherwise explicitly specified. Furthermore, while different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce good results.

[0142] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the inventive concept, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. An emotion recognition method based on multi-view contrastive learning of brain activation regions, characterized in that, The method comprises the following steps: obtaining an electroencephalogram emotion signal to be recognized; spatiotemporal transformation is performed on the electroencephalogram emotion signal to be recognized to obtain a to-be-recognized spatiotemporal signal; inputting the to-be-recognized spatiotemporal signal into a pre-trained dual-path decoupling multi-view model to obtain an emotion recognition result; the pre-trained dual-path decoupling multi-view model comprises a final dual-path decoupling multi-view encoder and a pre-trained classifier; two encoders are arranged in the final dual-path decoupling multi-view encoder, which are respectively used for recognizing the time feature and the space feature of the to-be-recognized spatiotemporal signal; the final dual-path decoupling multi-view encoder is jointly constructed based on contrast learning of high-confidence positive samples under a cross-view knowledge exchange strategy of the time feature and the space feature; wherein the high-confidence positive samples of the contrast learning are obtained by performing corresponding data enhancement strategies on the spatiotemporal signal samples according to the activation degree of a preset brain region.

2. The emotion recognition method based on multi-view contrastive learning of brain activation areas according to claim 1, characterized in that, The spatiotemporal transformation performed on the electroencephalogram emotion signal to be recognized to obtain a to-be-recognized spatiotemporal signal comprises: the electroencephalogram emotion signal to be recognized is segmented in the time dimension by using a sliding window to obtain a plurality of to-be-recognized segment signals; each to-be-recognized segment signal is segmented by using a sliding window to obtain a plurality of to-be-recognized sub-segments containing time segments; cosine similarity calculation is performed on the plurality of to-be-recognized sub-segments according to electrode channels, and the to-be-recognized spatiotemporal signal is generated according to the results of the cosine similarity calculation.

3. The emotion recognition method based on multi-view contrastive learning of brain activation areas according to claim 2, characterized in that, The cosine similarity calculation performed on the plurality of to-be-recognized sub-segments according to electrode channels, and the to-be-recognized spatiotemporal signal generated according to the results of the cosine similarity calculation, comprises: cosine similarity calculation is performed on the plurality of to-be-recognized sub-segments according to electrode channels to obtain first cosine similarity results; edges are established between to-be-recognized sub-segments corresponding to the first cosine similarity results greater than a similarity threshold to obtain the to-be-recognized spatiotemporal signal.

4. The emotion recognition method based on multi-view contrastive learning of brain activation areas according to claim 1, characterized in that, The final dual-path decoupling multi-view encoder is obtained by fine-tuning a pre-trained dual-path decoupling multi-view encoder based on spatiotemporal label samples; The training process of the pre-trained dual-path decoupling multi-view encoder comprises: obtaining electroencephalogram emotion signal samples; spatiotemporal transformation is performed on the electroencephalogram emotion signal samples to obtain spatiotemporal signal samples; calculating the class activation map corresponding to the preset brain region of the spatiotemporal signal sample, and dividing the to-be-recognized spatiotemporal signal into a high activation region and a low activation region according to the class activation map; data enhancement and signal fusion processing are sequentially performed on the high activation region and the low activation region to obtain a fusion signal sample; the data enhancement comprises at least one of spatial puzzle, time puzzle, frequency band puzzle, random masking, noise injection or scaling transformation processing; the fusion signal sample is input into an initial dual-path decoupling multi-view encoder, and training is performed through contrast learning; the initial dual-path decoupling multi-view encoder meeting a preset iteration stop condition is taken as the pre-trained dual-path decoupling multi-view encoder; wherein, the preset brain region comprises: frontal lobe, frontal lobe, left frontal lobe, right frontal lobe, left temporal lobe, right temporal lobe, central region, left parietal lobe, right parietal lobe and occipital lobe.

5. The emotion recognition method based on multi-view contrastive learning of brain activation areas according to claim 4, characterized in that, The calculating the class activation map corresponding to the preset brain region of the spatiotemporal signal sample comprises: The spatiotemporal signal sample and the static reference are sequentially subjected to global average pooling and linear projection, and sample global embedding and static global embedding are correspondingly obtained; the static reference is the mean of all spatiotemporal signal samples; The cosine similarity between each sample global embedding and the static global embedding is calculated to obtain a second cosine similarity result; Gradient information of the second cosine similarity result is calculated by back propagation to obtain a cosine similarity gradient, and the cosine similarity gradient is subjected to global average pooling processing to obtain a gradient pooling result; The gradient pooling result is averaged according to the preset brain region to obtain the class activation map; The class activation map is compared with a preset activation threshold, and the to-be-recognized spatiotemporal signal is divided into the high-activation region and the low-activation region according to a result of the threshold comparison.

6. The emotion recognition method based on multi-view contrastive learning of brain activation areas according to claim 5, characterized in that, The class activation map is compared with the preset activation threshold, and the to-be-recognized spatiotemporal signal is divided into the high-activation region and the low-activation region according to a result of the threshold comparison. The class activation map is compared with the preset activation threshold, and the to-be-recognized spatiotemporal signal is divided into the high-activation region and the low-activation region according to a result of the threshold comparison. The high-activation region and the low-activation region are sequentially subjected to data enhancement and signal fusion processing to obtain a fusion signal sample, comprising: Weak data enhancement processing is performed on the high-activation region to obtain weak enhancement data; 7. The emotion recognition method based on multi-view contrastive learning of brain activation areas according to claim 4, characterized in that, Strong data enhancement processing is performed on the low-activation region to obtain strong enhancement data; Data fusion processing is performed on the weak enhancement data and the strong enhancement data to obtain the fusion signal sample. The initial dual-path decoupled multi-view encoder has the same structure as the pre-trained dual-path decoupled multi-view encoder, and the initial dual-path decoupled multi-view encoder comprises a spatial encoder and a temporal encoder; The spatial encoder and the temporal encoder are connected in parallel; the spatial encoder is provided with, which are connected in series, a first unit, a spatial conversion module, a first word embedding module, a first Transformer encoding module, and a first pooling module; 8. The emotion recognition method based on multi-view contrastive learning of brain activation areas according to claim 4, characterized in that, The temporal encoder is provided with, which are connected in series, a second unit, a time conversion module, a second word embedding module, a second Transformer encoding module, and a second pooling module; The first unit is provided with a first branch unit and a second branch unit connected in parallel; the first branch unit is provided with a first linear projection module; and the second branch unit is provided with a graph encoding module and a second linear projection module connected in series; The first unit and the second unit have the same structure. The fusion signal sample is input into an initial dual-path decoupled multi-view encoder, and training is performed through contrastive learning, comprising: ​ 9. The emotion recognition method based on multi-view contrastive learning of brain activation areas according to claim 4, characterized in that, ​ The fusion signal sample is input into the initial two-path decoupling multi-view encoder to obtain a time embedding, a global embedding, and a spatial embedding; the time embedding is an output result of a time encoder, the spatial embedding is an output result of a spatial encoder, and the global embedding is a splicing result of the time embedding and the spatial embedding; Similarities of the time embedding, the global embedding, and the spatial embedding with all embeddings in a corresponding pair queue are calculated, and time similarity results, global similarity results, and spatial similarity results are obtained correspondingly; The time similarity results are sorted in a first descending order, and the first K time embeddings in a result of the first descending order sorting are taken as first guide information; The spatial similarity results are sorted in a second descending order, and the first K spatial embeddings in a result of the second descending order sorting are taken as second guide information; The global similarity results are sorted in a third descending order, and the first K global embeddings in a result of the third descending order sorting are taken as third guide information; The first guide information, the second guide information, and the third guide information are taken as the high-confidence positive samples; The initial two-path decoupling multi-view encoder is trained based on the high-confidence positive samples and the contrast learning to obtain the pre-trained two-path decoupling multi-view encoder.

10. The emotion recognition method based on multi-view contrastive learning of brain activation areas according to claim 9, characterized in that, The generation process of the pre-trained two-path decoupling multi-view model includes: The pre-trained two-path decoupling multi-view encoder and a labeled spatio-temporal sample are obtained; The labeled spatio-temporal sample is input into the pre-trained two-path decoupling multi-view encoder and an initial classifier for iterative fine-tuning training; The pre-trained two-path decoupling multi-view encoder corresponding to when a stop condition is reached is taken as the final two-path decoupling multi-view encoder; The initial classifier corresponding to when the stop condition is reached is taken as the pre-trained classifier; The final two-path decoupling multi-view encoder and the pre-trained classifier jointly constitute the pre-trained two-path decoupling multi-view model.

Citation Information

Patent Citations

  • Multi-channel electroencephalogram emotion recognition method based on space-time fusion feature network

    CN114224342A

  • Electroencephalogram emotion recognition method based on mixed enhancement and time sequence comparative learning

    CN116746929A