Brain electrical emotion recognition system based on comparative learning and parallel multi-source domain adaptation

By introducing contrast learning and parallel multi-source domain adaptation technology into the EEG emotion recognition system, combined with the self-attention mechanism, simulating the coordinated operation of the human brain visual cortex, the challenges of the existing system in terms of accuracy and individual differences are solved, and more efficient emotion recognition effects are achieved.

CN120046032APending Publication Date: 2025-05-27EAST CHINA UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510038590.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing EEG sentiment recognition systems have challenges in improving accuracy, especially in dealing with individual differences and frequency band characteristics.

Method used

The EEG emotion recognition system based on contrast learning and parallel multi-source domain adaptation is adopted, and the ventral and dorsal flow in the human brain visual cortex is simulated to work together through contrast learning, and the common and characteristic visual perception features are extracted.

Benefits of technology

The accuracy of EEG emotion recognition is improved, and by simulating the neuron transmission mechanism of the human brain's visual cortex, it captures the commonalities and characteristics of different subjects in the perception of the visual cortex, enhancing the performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046032A_ABST
    Figure CN120046032A_ABST
Patent Text Reader

Abstract

An electroencephalogram emotion recognition system based on comparative learning and parallel multi-source domain adaptation comprises the following steps that firstly, a data sampler is used for collecting positive and negative electroencephalogram samples under the same or different stimulations for comparative learning; secondly, mapping data distribution of the source domain and the target domain to the same feature distribution space as much as possible in a two-step alignment mode; finally, the feature alignment network maps characterization to potential spaces under different parallel branches, and the similarity of positive and negative sample pairs of a source domain and a target domain in feature space and prediction results is improved by minimizing parallel contrast loss between subjects at the same time. Due to the fact that different subjects have generality and characteristics on a brain visual cortex perception mechanism, a learning mode similar to self-supervision of a human brain is simulated by utilizing a comparative learning framework and a parallel multi-source domain adaptive network, so that rich physiological information contained in brain electrical emotion signals is obtained, and the accuracy of brain electrical emotion recognition is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electroencephalogram (EEG) emotion recognition. Specifically, the present invention relates to an EEG emotion recognition system for emotion classification of EEG signals based on contrast learning and parallel multi-source domain adaptation. Background Art

[0002] Emotion, as an important physiological information, is the core of the quality of human daily communication. The human-computer interaction system aims to establish a user-friendly interaction interface through machine decision-making, which is crucial for situation assessment and information acquisition affected by emotions. Among them, the emotion brain-computer interface, as a bridge between the emotions extracted from the brain and the computer, has shown great prospects and potential in the fields of medicine and human-computer interaction. To quantify emotions, most researchers focus on using traditional methods, such as classifying emotions with facial expressions or language. In recent years, due to the non-invasive brain-computer interface technology based on electroencephalogram (EEG) having the advantages of reliability, easy acquisition, and high precision, it has been widely used in intelligent emotion recognition systems.

[0003] The human brain is the center and core of emotion and cognitive processing, logical reasoning, enabling humans to quickly generate and feedback emotions from external emotional stimuli, which is difficult for machines. With the in-depth study of the brain nerves and cortical anatomical structure, researchers have begun to try to use mathematical models to simulate the generation of emotions and cognition or seemingly abstract problems such as object discrimination. Although the structure and principle of the brain's complex emotion and perception system still need to be fully understood, by summarizing and generalizing extensive experimental data, researchers can establish a biomathematical model to characterize the activity mechanisms of different neurons and cerebral cortex. Therefore, starting from the brain's emotion and cognitive system, exploring the process of emotion stimulation generation and information processing, so as to construct a more human-cognition-law-compliant intelligent human-computer interaction system.

[0004] Due to the differences in the individual experiences and personality characteristics of the subjects, it is relatively difficult to establish a general EEG emotion recognition model. Considering that most of the current emotion datasets based on physiological signals use videos, games, or pictures as emotional stimuli, the exploration of the human brain's visual perception mechanism and its connection with EEG signals is worthy of in-depth excavation. The present invention regards each subject as a separate parallel domain, extracts the common visual perception features of all subjects through multi-source domain adaptive learning, and then adopts a contrast learning method between each parallel domain to simulate the characteristics of the dorsal stream and ventral stream communicating with each other under parallel operation, so as to better explain the information received by the human brain's visual cortex. Obtain the rich physiological information contained in the EEG emotion signals, and further improve the accuracy of EEG emotion recognition. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a more effective electroencephalogram (EEG) emotion recognition system. Through this EEG emotion recognition system, the accuracy of EEG emotion recognition can be further improved. Inspired by the operation mechanism of bilateral stream neurons in the visual cortex of the brain, the object of the present invention is to propose an EEG emotion recognition system based on contrast learning and parallel multi-source domain adaptation, construct a neurology-inspired multi-source domain contrast learning method, and simulate the perception mechanism of the transmission of visual stimulus information through the parallel branches of the ventral stream and dorsal stream in the visual cortex of the human brain by introducing contrast loss in each parallel domain, so as to learn the commonalities and characteristics of different subjects in visual cortex perception. A new method for EEG emotion recognition simulation calculation is given, so as to further improve the performance of EEG emotion recognition.

[0006] 1. An EEG emotion recognition system based on contrast learning and parallel multi-source domain adaptation, characterized by comprising the following steps:

[0007] S1. Use a data sampler to collect EEG data of different subjects under the same and different stimuli to form positive and negative samples, and then generate a batch of positive and negative EEG sample training pairs for contrast learning;

[0008] S2. For the EEG sample pairs obtained in step S1, use an EEG feature extraction network based on the self-attention mechanism to learn the characteristics of each frequency band's EEG emotion response, and extract a common frequency band weight representation;

[0009] S3. For the common features extracted in step S2, adopt a two-step alignment method to make the data distributions of the source domain and the target domain be mapped to the same feature distribution space as much as possible, so that their outputs are as similar as possible;

[0010] S4. The feature alignment network maps the representations obtained in step S3 to the latent space under different parallel branches, and improves the similarity of the positive and negative sample pairs of the source domain and the target domain in the feature space and the prediction results by simultaneously minimizing the parallel contrast loss of the subjects;

[0011] S5. In the test step, preprocess the original test EEG signal to extract differential entropy features, then input them into the trained multi-source domain adaptation network, and then recognize the emotional state of the EEG signal through a fully connected layer;

[0012] 2. The EEG emotion recognition system based on contrast learning and parallel multi-source domain adaptation according to claim 1, characterized in that: in S1, a data sampler for contrast learning is designed to collect positive and negative sample pairs. In the dataset, the EEG data of a subject includes N trials, and each trial contains the EEG signals recorded when the subject watches an emotional stimulus video. To obtain a batch of positive and negative sample data, first randomly extract an EEG sample N times from each trial of subject A to obtain N samples, denoted as ( where M is the number of electrodes and T is the number of sampling points of a segment of EEG sample), and then similarly obtain the same N samples from another subject B, denoted as sample and correspond to the same moment of the i-th trial, forming a sample set In a batch, given a sample and sample form a positive pair, and the remaining samples are combined with to form 2(N - 1) negative pairs.

[0013] 3. The EEG emotion recognition system based on contrast learning and parallel multi-source domain adaptation according to claim 1, characterized in that: in S2, in order to capture the weight influence of different frequency bands on the EEG signals of the same emotional stimulus in each subject, a self-attention based EEG feature extraction network is proposed to extract the frequency band weighted features of all subjects. After a series of preprocessing operations on the original EEG signals, use the Butterworth filter to decompose them into five frequency bands, and extract their differential entropy features as the features before network input. Then apply the self-attention mechanism to each frequency band to calculate the weighted output of the frequency band and splice it into a frequency band weighted feature vector of a segment of EEG signal, so as to learn the self-attention weights within the frequency band and the complementary information between frequency bands.

[0014] 4. An EEG emotion recognition system based on contrastive learning and parallel multi-source domain adaptation according to claim 1, characterized in that: the multi-source domain adaptation based on two-step alignment in S3 converts data from a single subject to between-subject aligned representations and prediction labels by aligning the feature distributions of each domain and the outputs of each domain classifier respectively. The proposed two-step alignment multi-source domain adaptation network consists of a common feature extractor (CFE) and a domain-specific feature extractor (DSFE). First, the CFE extracts the common representations of all subject domains, mapping the data of different subjects from the original feature space to the common feature space. Second, the DSFE extracts domain-specific features for each domain. In order to map each pair of source domain and target domain to the same specific feature space, the distribution difference between the source domain and the target domain is minimized by using the Wasserstein distance to learn domain-invariant features, thus achieving alignment at the feature level. Then, a softmax classifier is established for each domain by DSC to predict the classification outputs of their respective domains, and cross-entropy loss is used to calculate the classification loss specific to each domain. However, the classifiers are trained on different source domains, so their predictions for target domain samples, especially samples near the decision boundary, may vary. In order to make the prediction outputs of different classifiers on the target domain as consistent as possible, absolute value loss is used to measure the prediction differences of each source domain classifier on the target domain samples, thus achieving alignment at the decision level.

[0015] 5. An EEG emotion recognition system based on contrastive learning and parallel multi-source domain adaptation according to claim 1, characterized in that: using contrastive loss in the parallel structure of the multi-source domain can maximize the similarity between two positive EEG signals. The original input sample set After passing through the feature extractor in the multi-source domain adaptation network, it is converted into a set of latent features of the input samples Thus, the input sample and The similarity of Is calculated from the cosine distance between the latent features of A and B and As shown in the following formula:

[0016] Similar to the SimCLR framework, this method uses a normalized temperature scale to calculate the cross-entropy loss of the input samples and The cross-entropy loss of and The sum of the two is the contrastive loss L of this sample paircon , as follows: Here, I [j≠i] ∈ {0, 1} is an indicator function that is set to 1 when j ≠ i. τ is the temperature scale coefficient.

[0017] The beneficial effects of the present invention are as follows: The EEG emotion recognition system based on contrastive learning and parallel multi-source domain adaptation of the present invention, in order to simulate the neuron representation mechanism of the human brain similar to self-supervised learning, adds a parallel contrastive loss in the parallel branch structure to simulate the cooperative operation mode of the ventral stream and the dorsal stream. In order to minimize the data distribution difference between the source subject and the target subject, a multi-source domain adaptation method is introduced, and the extracted differential entropy features are synchronously aligned to the feature space and the decision space by using a two-step alignment method, and then the feature space distance and classification loss between each source domain and the target domain are minimized, so that the spatial distributions of each source domain and the target domain can be close to each other. In addition, due to the differences in the EEG emotional responses of different frequency bands, a feature extractor based on the self-attention mechanism is proposed, and the EEG signals are weighted in different frequency bands during the two-step alignment stage, so as to capture the commonality and uniqueness of the frequency band emotion weights of each subject under the same emotional stimulus. Therefore, compared with other methods that only focus on mathematical models or statistical models, this method has stronger persuasiveness in biology and improves the performance of EEG emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is the flowchart of the overall framework in the present invention;

[0019] Figure 2 is the structural diagram of the EEG feature extraction network based on self-attention in the present invention;

[0020] Figure 3 is the architecture diagram of the multi-source domain adaptation based on two-step alignment in the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0021] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments: The method of the present invention is divided into three parts in total.

[0022] The first part: EEG feature extraction network based on self-attention

[0023] In order to capture the weight influence of EEG signals in different frequency bands on the same emotional stimulus for each subject, the present invention proposes an EEG feature extraction network based on self-attention to extract the frequency band weighted features of all subjects, and the overall module is as Figure 2 shown. Among them, Figure 2 (a) represents the process of the EEG signal from frequency band decomposition to obtaining the frequency band weighted output after passing through the self-attention module,Figure 2 (b) describes the specific details of the self-attention module.

[0024] Specifically, first, after a series of preprocessing operations such as downsampling, noise reduction, and artifact removal on the original EEG signal, it is decomposed into five frequency bands using a Butterworth filter, namely the Delta band, Theta band, Alpha band, Beta band, and Gamma band. Since DE features have been proven to be very effective in previous EEG emotion recognition studies, this method also uses DE features as the features before network input. DE is an extended form of Shannon information entropy in the range of continuous variables, which is equivalent to the logarithm of the energy spectrum within a specific frequency band. It is usually assumed that the EEG signal follows a specific Gaussian distribution N(μ,σ 2 ), then the specific calculation form can be expressed as: where e is the Euler's constant, μ is the mean value of the EEG signal, and σ is the standard deviation of the EEG signal.

[0025] Then, the EEG signal is segmented using a Hamming window with a length of 1 second on each frequency band, and the DE features of m electrode channels are extracted from each segment, thus forming a 5m-dimensional feature vector, denoted as where represents the dimension of the EEG feature vector of all frequency bands, b ∈ {δ,θ,α,β,γ}, and t x represents the number of EEG signal segments. Thus, the EEG signals on the five frequency bands are respectively expressed as where represents the dimension of the EEG feature vector on each frequency band, equal to 1×m.

[0026] For the EEG input x at the i-th electrode position in the b-th frequency band i b , the importance weight value α i b at the i-th electrode position is calculated as follows: where σ(·) represents the activation function, W s and b s are the parameter matrix and bias of the activation layer respectively, and u i b is the output of the activation layer at the i-th electrode position in the b-th frequency band.

[0027] Finally, multiply α i b by the original input x ib Obtain the output on frequency band b Subsequently, the weighted outputs on all frequency bands are concatenated into a one-dimensional vector as the frequency band weighted feature vector of a segment of EEG signal to learn the self-attention weights within the frequency band and the complementary information z between frequency bands. The specific calculation process is as follows: z = concat(o δ , o θ , o α , o β , o γ ) (9)

[0028] Part Two: Multi-source domain adaptation based on two-step alignment

[0029] Multi-source domain adaptation based on two-step alignment converts data from individual subjects into between-subject aligned representations and prediction labels by aligning the feature distributions of each domain and the outputs of each domain classifier respectively. The two-step alignment is divided into feature-level alignment and decision-level alignment. Feature-level alignment refers to the alignment of the source subject domain and the target subject domain in the common feature space, and decision-level alignment refers to the alignment of the source subject domain and the target subject domain in the decision space, that is, the alignment of the classifier outputs.

[0030] As Figure 3 shown, the proposed two-step alignment-based multi-source domain adaptation network consists of CFE, DSFE, and DSC. First, CFE extracts the common representation of all subject domains, mapping the data of different subjects from the original feature space to the common feature space. Second, DSFE extracts specific features for each domain. To map each pair of source domain and target domain to the same specific feature space, the domain-invariant features are learned by minimizing the distribution difference between the source domain and the target domain using the Wasserstein distance, thus achieving alignment at the feature level. Then, DSC establishes a softmax classifier for each domain to predict the classification output of its respective domain, and the cross-entropy loss is used to calculate the classification loss specific to each domain. However, the classifiers are trained on different source domains, so their predictions for target domain samples, especially those near the decision boundary, may be inconsistent. To make the prediction outputs of different classifiers on the target domain as consistent as possible, the absolute value loss is used to measure the prediction differences of each source domain classifier on the target domain samples, thus achieving alignment at the decision level.

[0031] Specifically, the input of the network is N independent source domain data and a target domain data X T . In the feature-level alignment stage, these data are input into the common feature extraction module to obtain the domain-invariant features and Q T, and then each DSFE extracts common features and Q T specific domain features of and Next is the decision-level alignment stage, where the target domain features and all the source domain features extracted in the last step are used to obtain the corresponding classification prediction results through DSC output and Finally, the classification loss is calculated based on the prediction results of the source domain. Since the target domain is input into all source domain classifiers, multiple prediction results for the target domain will be generated. The model uses these prediction results to calculate the Discrepancy Loss, and finally takes the average of these target domain predictions as the final output.

[0032] In the feature-level alignment stage, EEG data first passes through a CFE to map the source domain data and the target domain data from the original feature space to a common latent space, thereby extracting the common representations of all domains and helping to capture some underlying domain-invariant features. After obtaining the common features of all domains, DSFEs corresponding to N source domains and the target domain are established, and each pair of source domain and target domain data is mapped to a latent feature space specific to its own domain to obtain the domain-specific features of each branch. To make the outputs of the source domain and the target domain data in the latent feature space closer, this method uses the Wasserstein distance loss L was to calculate the distance between different domains instead of the Maximum Mean Discrepancy (MMD) loss, as shown in Equation (10). During training, the gap between the source domain and the target domain is reduced by minimizing the Wasserstein distance, which helps to better predict the target domain. Compared with the MMD loss, the Wasserstein distance can measure the minimum of the average distance that data needs to move from one distribution to another. Even if the overlap between these two distributions is very small, it can still reflect the distance between the two distributions. Therefore, the Wasserstein distance is suitable for EEG signals that follow the same distribution but are quite different from each other. where P s , P t are the probability distributions of the source domain and the target domain respectively, and Π(P s , P t ) represents the set of joint distributions composed of the marginal distributions of P s , P t , and represents sampling (x w from the joint distribution γ w, y w ) to γ to obtain the sample x w , y w And calculate the distance ||x of this pair of samples under this joint distribution w - y w The lower bound of the P expectation value, that is, the minimum path cost under the joint distribution γ w under.

[0033] In the decision-level alignment stage, each source domain and target domain DSC calculates the classification loss L of each classifier by establishing N softmax classifiers and using cross-entropy cls , as shown in formula (11). However, if the average value of the prediction results of these N classifiers is simply used as the final result, it may lead to a high variance due to the large differences in the prediction results. Especially when the target domain samples are near the decision boundary, it will have a great negative impact on the prediction results. To reduce the possibility of this error occurring, this method introduces the L1 loss L disc as the difference loss metric to make the prediction results of these N classifiers converge, as shown in formula (12).

[0034] Part Three: Multi-source domain sentiment classification based on the contrast learning framework

[0035] In the present invention, this contrast learning framework assumes that when the subjects receive the same part of the emotional stimuli, their neural activities will be in a similar state. Based on this basic idea, the present invention learns the common and unique information of the visual perception mechanisms among different subjects by aligning the similar emotional stimulus activity features. Specifically, the present invention includes two stages, namely the domain alignment stage and the contrast learning stage. The domain alignment stage aligns the features of multiple source subject domains and target subject domains into the same representation space and decision space through a two-step alignment method, so that the features and prediction labels of different subjects are as consistent as possible. In the contrast learning stage, inspired by the commonality and uniqueness of the human visual perception mechanism, a parallel contrast learning structure for multi-source domains is proposed. By minimizing the parallel contrast loss between the representation spaces and decision spaces of different subjects, the similarity of signals under the same emotional stimulus (positive pair) response is maximized, while the similarity between the corresponding signals under different stimuli (negative pair) is minimized, as Figure 3 shown. Finally, the total loss L on one batch is calculated as: L = L cls + α 1 L was + β 1 L disc + η 1 L con (13) Among them, α 1 , β 1 , η 1 are the parameters for controlling the Wasserstein distance, the divergence loss, and the contrast loss, respectively. The method trains the entire network and updates the parameters by minimizing the above loss functions. For these four losses, minimizing the Wasserstein distance can capture the domain-invariant features of each pair of source domain and target domain; minimizing the classification loss can make the predictions of the source domain classifier more accurate; minimizing the divergence loss will make it easier for multiple classifiers to converge; minimizing the contrast loss will increase the similarity between and , rather than other sample pairs that may involve . The basic steps of multi-source domain sentiment classification based on the contrast learning framework are as follows: 1: Take samples and labels from the source domain Take samples from the target domain 2: Sample two subjects A and B from the source domain and the target domain; 3: Use the data sampler to sample 2N pieces of data from A and B 4: For each frequency band b, use the self-attention mechanism to obtain the final weighted output o of this frequency band b ; 5: Concatenate the weighted outputs of each frequency band to obtain the frequency band weighted features of subjects A and B; 6: Calculate each loss and sum them up to obtain the total loss; 7: Minimize the total loss to update the model parameters; 8: After all data pairs are enumerated, the steps end; Experimental design

[0036] Experimental dataset: The SEED dataset used in the experiment is an electroencephalogram (EEG) emotion dataset led and released by the team of Shanghai Jiao Tong University. It contains the EEG and eye movement data of 12 subjects and the EEG data of another 3 subjects. The dataset selected 15 movie clips from the material library as the emotional stimuli for the experiment. Each movie clip has a duration of about 4 minutes, and it is ensured that the corresponding emotions can be continuously and prominently triggered. These movie clips contain three different emotions: positive, neutral, and negative. After each subject watches each movie clip, they will immediately be asked to complete a questionnaire to report their emotional reactions to each movie clip. During the acquisition process, the electrodes were placed using the 62-channel international 10-20 system, and the sampling frequency was 1000 Hz. In the preprocessed version, the EEG signal data was downsampled to 200 Hz, and at the same time, the EEG signal was filtered using a band-pass filter of 0 - 75 Hz. Experimental results

[0037] Experimental results under different classification models: To verify the effectiveness of the present invention compared with other domain adaptation methods, some typical deep learning domain adaptation methods are compared with the present invention. These methods include four single-source domain adaptation methods, namely: Deep Adaptation Network (DAN), Domain-Adversarial Neural Network (DANN), Deep Domain Confusion (DDC), Deep Correlation Alignment (DCORAL), and a multi-source domain adaptation method, Two-Way Multi-source Domain Adaptation (TWMDA). In all comparison models, the network parameters and hyperparameters used are the same. The experimental comparison results of different domain adaptation methods on the SEED and DEAP datasets are listed in Table 1. It can be clearly seen from Table 1 that the performance of the multi-source domain adaptation method TWMDA and the present invention both exceed that of the single-source domain adaptation methods (DAN, DANN, DDC, DCORAL). This shows that retaining the unique visual perception mechanisms of multiple subjects is crucial for predicting other subjects, and helps the perceptual features of the target subject to be closer to those of other subjects. The comparison between the present invention and TWMDA mainly proves that the introduction of the parallel contrast loss and the self-attention mechanism successfully simulates the similar visual information processing and perception mechanisms of different subjects, which are lacking in other domain adaptation methods. Table 1 Accuracy comparison of different domain adaptation methods on SEED Table 1 Accuracy comparative results on SEED and DEAP with different DA methods.

[0038] Experimental results of different innovative modules: The present invention has three innovative parts: the addition of multiple source domains to supplement specific difference information, the introduction of contrast loss in the parallel branch structure, and the improvement of the feature extraction network by the self-attention mechanism. To evaluate the contribution of each innovative part, three variant models with some innovative points removed are designed as follows:

[0039] M1: The base model, where all source domains are mixed together and regarded as one domain;

[0040] M2: Based on M1, the source domain is considered as multiple domains, and the features of each source domain and target domain are aligned using a two-step alignment method;

[0041] M3: Improves the feature extractor based on M2 and introduces a self-attention mechanism on multiple frequency bands;

[0042] M4 (i.e. the proposed model): introduces parallel contrast loss to the overall framework based on M3;

[0043] Table 2 shows the classification accuracy comparison between the three variant models and the proposed method in SEED to intuitively reflect the contribution of different innovative parts. It can be clearly seen from the table that the multi-source domain adaptation model M2 has the greatest improvement among all variant models, which shows that the use of multiple source domains can retain the perceptual mechanism characteristics of each subject at the feature level and prediction level at the same time, so that the target domain can find a source domain closer to it; secondly, the introduction of parallel contrast loss (belonging to M4) improves the accuracy of the proposed model, proving that the human brain has a similar operating mechanism when receiving and processing the same emotional stimuli, which helps to improve the model's extraction of common emotional perception information. In addition, the introduction of the self-attention mechanism also slightly improves the recognition accuracy of the model, indicating that there is a commonality in the weight differences of each frequency band of different subjects to the EEG emotional response, which is helpful for the extraction of common emotional features of all subjects. Table 2 Comparison of the contribution of different innovations to model performance improvement Table 2 Contribution comparison of different novel parts in improving performance.

Claims

1. An EEG emotion recognition system based on contrastive learning and parallel multi-source domain adaptation, characterized in that: The steps include: S1. Use a data sampler to collect EEG data of different subjects under the same and different stimuli to form positive samples and negative samples, and then generate a batch of positive and negative EEG sample training pairs for comparative learning; S2, for the EEG sample pairs obtained in step S1, using an EEG feature extraction network based on a self-attention mechanism to learn the characteristics of each frequency band's response to EEG emotions and extract a common frequency band weight representation; S3, for the common features extracted in step S2, a two-step alignment method is used to map the data distribution of the source domain and the target domain to the same feature distribution space as much as possible, so that their outputs are as close as possible; S4, the feature alignment network maps the representation obtained in step S3 to the latent space under different parallel branches, and improves the similarity of the positive and negative sample pairs in the source domain and the target domain in the feature space and the prediction results by minimizing the parallel contrast loss of the subjects at the same time; In S5, the test step, the original test EEG signal is preprocessed to extract the differential entropy features, which are then input into the trained multi-source domain adaptation network, and then the emotion state of the EEG signal is recognized through the fully connected layer.

2. The EEG emotion recognition system based on contrastive learning and parallel multi-source domain adaptation according to claim 1, characterized in that: S1 describes a data sampler for contrastive learning designed to collect positive and negative sample pairs; in the dataset, the EEG data of a subject includes N trials, each trial contains an EEG signal recorded when the subject watches an emotional stimulus video; in order to obtain a batch of positive and negative sample data, first randomly extract an EEG sample N times from each trial of subject A, and obtain N samples, which are expressed as ( M is the number of electrodes, T is the number of sampling points of a segment of EEG sample), and then similarly obtain the same N samples from another subject B, expressed as sample and Corresponding to the same moment of the i-th trial, the sample set In a batch, given a sample and samples Form a positive pair, and the remaining samples s∈{A,B}} is equal to 2(N-1) negative pairs are formed.

3. The EEG emotion recognition system based on contrastive learning and parallel multi-source domain adaptation according to claim 1, characterized in that: S2 described that in order to capture the weighted influence of different frequency bands in each subject on the EEG signal of the same emotional stimulus, a self-attention-based EEG feature extraction network is proposed to extract the frequency band weighted features of all subjects; after a series of preprocessing operations, the original EEG signal is decomposed into five frequency bands using a Buttonworth filter, and its differential entropy features are extracted as the features before network input, and then the self-attention mechanism is applied to each frequency band to calculate the weighted output of the frequency band and splice it into a frequency band weighted feature vector of an EEG signal, thereby learning the self-attention weights within the frequency band and the complementary information between frequency bands.

4. The EEG emotion recognition system based on contrastive learning and parallel multi-source domain adaptation according to claim 1, characterized in that: The multi-source domain adaptation based on two-step alignment described in S3 converts the data from a single subject into aligned representations and predicted labels between subjects by aligning the feature distribution of each domain and the output of each domain classifier respectively; the proposed two-step alignment multi-source domain adaptation network consists of a common feature extractor (CFE) and a domain-specific feature extractor (DSFE); first, CFE extracts the common representation of all subject domains and maps the data of different subjects from the original feature space to the common feature space; secondly, DSFE extracts specific features for each domain, and in order to enable each pair of source and target domains to be mapped to the same specific feature space, the Wasserstein distance is used to minimize the distribution difference between the source and target domains to learn domain-invariant features, thereby achieving feature-level alignment; Then, DSC builds a softmax classifier for each domain to predict the classification output of the respective domain, and uses cross entropy loss to calculate the classification loss specific to each domain; however, the classifiers are trained on different source domains, so their predictions for target domain samples, especially those close to the decision boundary, may differ; in order to make the prediction outputs of different classifiers on the target domain as consistent as possible, the absolute value loss is used to measure the prediction differences of each source domain classifier on the target domain samples, thereby achieving alignment at the decision level.

5. The EEG emotion recognition system based on contrastive learning and parallel multi-source domain adaptation according to claim 1, characterized in that: The use of contrast loss in the parallel structure of multi-source domains as described in S4 can maximize the similarity of two EEG signals facing each other; after passing through the feature extractor in the multi-source domain adaptive network, the original input sample set After passing through the feature extractor in the multi-source domain adaptation network, it is converted into a potential feature set of the input sample Therefore, the input sample and Similarity By calculating the latent features of A and B and The cosine distance between them is obtained as shown in the following formula: Similar to the SimCLR framework, this method uses a normalized temperature scale to calculate the input samples and The cross entropy loss and The sum of the two is the contrast loss of the sample pair as follows: here, is an indicator function, which is set to 1 when j≠i; τ is the temperature scale coefficient.