A cross-subject electroencephalogram emotion recognition method based on double-path attention and global-local collaborative integrated alignment
Patent Information
- Application Number
- CN202610960289.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明的目的在于提供一种基于双路径注意力与全局局部协同集成对齐的跨被试脑电情感识别方法,解决现有技术缺乏在输入端自适应抑制个体异质性的机制、域适应易引发过度对齐且训练不稳定、以及缺乏显式语义聚类约束导致鉴别边界坍塌的问题,实现了跨被试情感识别准确性、泛化性及鲁棒性的显著提升
1)本发明通过引入三阶段级联校正框架,有效克服了现有技术中跨被试信号分布差异大、模型适应性差的局限;
Smart Images

Figure CN122827679A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of emotion computing and brain-computer interface technology, and in particular relates to a cross-subject EEG emotion recognition method based on dual-path attention and global-local collaborative integration alignment. Background Technology
[0002] Automated emotion recognition technology is an important research direction in the fields of affective computing and brain-computer interfaces. It aims to achieve an objective and convenient assessment of psychological states by analyzing an individual's physiological signals. Among numerous signal sources, electroencephalography (EEG) data has attracted widespread attention due to its ability to directly record the intrinsic cortical neural activity behind emotional processing, offering unique physiological interpretability and technological advantages.
[0003] Existing research indicates significant differences in emotional expression among individuals' EEG signals. The transformation of EEG emotion recognition technology from controlled laboratory environments to large-scale real-world deployment has long been hampered by a bottleneck resulting in significant degradation in cross-subject generalization performance. This degradation primarily stems from the inherent physiological heterogeneity of individual brains. Specifically, structural and functional differences between individuals create specific spatiotemporal physiological feature spaces, leading to extremely complex cascading effects. First, at the source level, different individuals exhibit significantly different neural response patterns to the same emotional stimuli, with emotion-related signals often being masked by physiological noise specific to each subject. Second, at the transmission level, these response differences propagate to the extracted features, causing severe physiological distribution shifts within the feature space. Finally, at the terminal level, these distribution shifts ultimately lead to semantic misalignment of emotions; even after alignment, different categories of emotional neural patterns may still experience feature overlap or incorrect clustering. Because these biases are amplified throughout the entire process, optimizing only a single stage cannot solve the cross-subject generalization problem.
[0004] However, current cross-subject EEG emotion recognition methods still face many challenges. First, regarding feature enhancement and processing, existing feature reconstruction and decoupling processes are mostly performed in high-level feature spaces, where source-level response differences have already been propagated and amplified. Existing solutions lack mechanisms for adaptively suppressing heterogeneity at the input end. Second, in mitigating feature distribution differences, mainstream domain adaptation methods often employ rigid statistical alignment or adversarial learning. This approach is highly susceptible to over-alignment problems, where blind distribution matching suppresses discriminative features related to emotion. Furthermore, these methods often encounter training instability and convergence difficulties when facing scenarios with high individual heterogeneity. Third, in handling emotional semantic misalignment, existing methods often treat semantic consistency as a byproduct of global distribution alignment, lacking specific end-to-end explicit constraints for cross-subject feature clustering. This leads to class overlap and collapse of discriminative boundaries after the model eliminates inter-domain differences.
[0005] Therefore, how to overcome the amplification effect of neurophysiological heterogeneity in the conduction chain and provide a full-link correction scheme that can suppress individual noise from the source, coordinately align distribution offsets, and ensure consistency of the emotional semantic space is a technical problem that urgently needs to be solved in the field of cross-subject EEG emotion recognition. Summary of the Invention
[0006] The purpose of this invention is to provide a cross-subject EEG emotion recognition method based on dual-path attention and global-local collaborative integration alignment. This method addresses the problems of existing technologies, such as the lack of a mechanism to adaptively suppress individual heterogeneity at the input end, the tendency of domain adaptation to lead to over-alignment and unstable training, and the lack of explicit semantic clustering constraints that cause the identification boundary to collapse. This invention achieves a significant improvement in the accuracy, generalization and robustness of cross-subject emotion recognition.
[0007] To address the aforementioned technical problems, the technical solution adopted in this invention is a cross-subject EEG emotion recognition method based on dual-path attention and global-local collaborative integration alignment, comprising the following steps: S1: Acquire EEG signals from the source and target domains, extract multi-band, multi-channel EEG differential entropy feature matrices as raw inputs, and standardize the raw inputs to obtain the standardized initial feature matrix; S2: Construct a dual-path attention correction module for frequency band feature branches and channel feature branches; S3: Construct a deep feature extraction backbone network consisting of a cascaded dynamic graph convolution module, a multi-head self-attention module, and a feature mapping module; S4: Construct a global-local cooperative alignment module to perform distribution offset correction on the high-level feature representations of the source domain and the target domain, and calculate the global-local cooperative joint loss; S5: Construct a deep clustering module with learnable centers. Based on the high-level feature representations of the source domain and the target domain, calculate the supervised contrastive loss and the learnable center loss, and construct the joint loss for deep clustering based on the supervised contrastive loss and the learnable center loss. S6: Based on the loss function constructed in S4-S5, the cascaded correction network is trained. The cascaded correction network consists of a dual-path attention correction module, a deep feature extraction backbone network, and a global-local collaborative alignment module and a deep clustering module with learnable centers connected in parallel after the deep feature extraction backbone network. The trained network is then used to identify the emotional state of the EEG signals of the target domain subjects.
[0008] Furthermore, the specific steps of S1 are as follows: S11: Using a multi-channel EEG acquisition device, raw EEG signals from the scalp of different subjects under a preset emotional evoked paradigm were obtained, and power frequency filtering, bandpass filtering, and independent component analysis were performed to remove artifacts and obtain pure EEG signals. S12: Framing, windowing, and baseline correction are performed on the pure EEG signal to obtain a baseline-corrected multi-channel EEG timing signal; S13: The baseline-corrected multichannel EEG time-series signal is decomposed into several frequency bands, and the differential entropy characteristics of each frequency band are calculated, as shown in Equation (1): (1) in: EEG signals in a specific frequency band The eigenvalues of the differential entropy, Let be the probability density function of a time series signal in a specific frequency band. A time series of EEG signals in a specific frequency band. For the instantaneous values of the EEG signal time series, The variance of the time-series signal in this specific frequency band. It is a natural constant. A time series of EEG signals in a specific frequency band. It is the natural logarithm. Pi; S14: Concatenate and stack the differential entropy eigenvalues obtained within a single time window to reconstruct a two-dimensional feature matrix, and then perform a standardization operation on the two-dimensional feature matrix to obtain a standardized initial feature matrix. .
[0009] Furthermore, the specific steps of S2 are as follows: S21: Construct a frequency band feature branch, transpose, pool, and perform nonlinear mapping on the standardized initial feature matrix to obtain the frequency band attention weight vector, construct a diagonal matrix using the frequency band attention weight vector, and perform right multiplication calibration on the standardized initial feature matrix to obtain the frequency band calibration feature matrix. S22: Construct channel feature branches, perform pooling and nonlinear mapping on the standardized initial feature matrix to obtain channel attention weight vectors, construct diagonal matrices using channel attention weight vectors, and perform left multiplication on the standardized initial feature matrix to obtain channel space calibration feature matrices; S23: The frequency band calibration feature matrix and the channel space calibration feature matrix are nonlinearly recombined using the Hadamard product to obtain the correction feature matrix.
[0010] Furthermore, the specific steps of S3 are as follows: S31: The scalp EEG electrode channels are abstracted as vertices of a mathematical graph model, and a spatial adjacency matrix is constructed using electrode spatial coordinates. The Laplace matrix is calculated, and the graph convolution kernel is truncated based on Chebyshev polynomials. The correction feature matrix is then graph convolved to obtain the spatial node fusion feature matrix. S32: Using three sets of learnable weight matrices, the spatial node fusion feature matrix is projected into a query matrix, a key matrix, and a value matrix. The feature is divided into multiple parallel attention heads along the feature dimension, and the context-weighted features of a single attention head are calculated separately. The local weighted features of all attention heads are concatenated along the channel dimension and then fused to reduce the dimension, thus obtaining the attention feature matrix. S33: Map the attention feature matrix to a high-dimensional latent feature space to obtain the high-level feature representations of the source domain and the target domain, respectively.
[0011] Furthermore, the specific steps of S4 are as follows: S41: Calculate the difference between the high-level feature representation of the source domain and the high-level feature representation of the target domain in the first-order statistic, and construct the global macro-distribution alignment loss, as shown in Equation (2).
[0012] in: For global macroscopic distribution alignment loss, It is the L2 norm. Let be the center vector of the first-order statistics of the source domain. Let the center vector of the first-order statistic of the target domain be . The dimension of the high-dimensional latent feature space; S42: Calculate the difference between the high-level feature representation of the source domain and the high-level feature representation of the target domain in the second-order statistic, and construct the local alignment loss as shown in Equation (3):
[0013] in: This is an empirical normalization scaling factor derived based on the norm bounds of high-dimensional matrices. For local alignment loss, For the Frobenius norm, Let be the unbiased covariance matrix of the source domain. Let be the unbiased covariance matrix of the target domain features; S43: Superimpose the global macro-distribution alignment loss with the local alignment loss to construct a global-local collaborative joint loss.
[0014] Furthermore, the specific steps of S5 are as follows: S51: Perform L2 norm normalization on each sample feature vector in the high-level feature representation matrix of the current batch, and calculate the temperature scaling cosine similarity between each pair of samples in the batch based on the normalized feature vectors to obtain the calibrated spatial similarity, as shown in Equation (4):
[0015] in: For any two samples within a batch and The calibrated spatial similarity between them For each independent sample, the high-level feature vector The new eigenvectors obtained by normalization For temperature overparameter, For each independent sample, the high-level feature vector The new feature vector obtained by normalization; S52: Using the sample sentiment category label as the supervision signal, other samples with the same label as the anchor sample within the batch are defined as the positive sample set, and samples with different labels are defined as the negative sample set. A supervised contrastive loss function is constructed based on the calibrated spatial similarity, as shown in Equation (5):
[0016] in: To supervise the contrastive loss function, This represents the total number of samples in the current training batch. This is the index of the current anchor sample in the batch. For anchor point samples A set of indices of other samples with the same sentiment label. Positive sample set The number of elements in the middle, for Index in For the sample With sample The calibrated spatial similarity between them To iterate over all that are not equal to Sample index, For the sample With sample The calibrated spatial similarity between them It is an exponential function. It is the natural logarithm function; S53: Initialize a learnable class center vector for each sentiment category in the latent space, calculate the mean squared error between the high-level feature vector of each sample in the batch and its corresponding class center vector, and construct the center loss function based on the mean squared error, as shown in Equation (6):
[0017] in: For the center loss function, For the first The original high-level feature vector of each sample, To be consistent with the sample Authentic Labels The corresponding learnable category center vector; S54: The supervised contrast loss function and the center loss function are weighted and combined to construct a deep clustering joint loss.
[0018] Furthermore, the specific steps of S6 are as follows: S61: Input the source domain high-level feature representation and the target domain high-level feature representation output from S3 into a linear sentiment classifier to obtain the predicted probability distribution and construct the total target loss function; S62: Backpropagate the error of the network based on the total objective loss function, and use the optimization algorithm to perform end-to-end synchronous iterative updates of the network weights until the total objective loss function converges, thus obtaining a cascaded correction network with solidified weights. S63: Collect EEG signals from new subjects and extract differential entropy feature matrix. After processing by a cascaded correction network with fixed weights, output high-level feature representation, which is then input into a linear emotion classifier and passed through a Soft max layer to obtain the emotion category prediction result.
[0019] The beneficial effects of this invention are: 1) This invention effectively overcomes the limitations of existing technologies, such as large differences in the distribution of subject signals and poor model adaptability, by introducing a three-stage cascaded correction framework; 2) This invention utilizes a dual-path attention correction module to perform early calibration at the feature input end, effectively decoupling and enhancing emotion-related neurophysiological features, and suppressing individual-specific noise from the physical source; 3) This invention unifies the statistical alignment of global centroid and local covariance through a global-local collaborative alignment mechanism, which not only alleviates the inter-domain distribution shift but also protects the inherent discriminative cues of emotion and avoids over-alignment. 4) This invention introduces a deep clustering mechanism in the feature space, which effectively corrects the misalignment of emotional semantics across individual features through the dual constraints of intra-class compactness and inter-class separation; 5) This invention effectively improves the accuracy of extracting key features of emotional state and cross-subject stability, thereby achieving more robust automated emotion assessment in complex real-world scenarios. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of the overall architecture of the cross-subject EEG emotion recognition method based on dual-path attention and global-local collaborative alignment, which is a schematic diagram of the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] like Figure 1 As shown, this invention provides a cross-subject EEG emotion recognition method based on dual-path attention and global-local collaborative integration alignment, comprising the following steps: S1: Acquire EEG signals from the source and target domains, extract multi-band, multi-channel EEG differential entropy feature matrices as raw inputs, and standardize the raw inputs to obtain standardized initial feature matrices; In a specific embodiment of the present invention, step S1 is as follows: S11: Using a multi-channel EEG acquisition device, raw scalp EEG signals from different subjects under a preset emotional evoked paradigm are obtained, and then subjected to power frequency filtering, bandpass filtering, and ICA (Independent Component Analysis) artifact removal to obtain pure EEG signals. In a specific embodiment of the present invention, firstly, a multi-channel EEG acquisition device is used to acquire raw scalp EEG signals from different subjects under a preset emotional evoked paradigm. The dataset collected from known subjects and bearing clear emotional category labels is defined as the source domain dataset. As shown in equation (1): (1) in: For the source domain dataset, For the sample index of the source domain dataset, For the first Source domain EEG features corresponding to each source domain sample For the first The sentiment category label corresponding to each source domain sample , The total number of emotion categories, This represents the total number of samples in the source domain.
[0024] Then, an unlabeled dataset from unknown subjects who did not participate in the joint training of the model is obtained and defined as the target domain dataset. As shown in equation (2): (2) in: For the target domain dataset, For the first in the target domain j EEG characteristics corresponding to each sample The total number of samples in the target domain dataset. j This refers to the sample index in the target domain dataset.
[0025] Due to the complexity of the actual acquisition environment and the physiological state of the subjects, various artifacts inevitably interfere with the raw scalp EEG signals. To ensure that the extracted neural response patterns can truly and purely reflect the subject's emotional and cognitive processing state, multi-level signal quality control processing needs to be performed before feature extraction. The specific implementation method is as follows: A notch filter with a stopband set to 50Hz was used to remove source domain datasets. and target domain dataset To address the power frequency interference noise introduced during the data acquisition process, a bandpass filter with a lower cutoff frequency of 1Hz and an upper cutoff frequency of 75Hz was used to filter out source domain datasets. and target domain dataset The study investigated low-frequency baseline drift caused by the subject's breathing, eliminated high-frequency environmental white noise, and finally used an independent component analysis algorithm to separate the blind source of the multi-channel mixed signal into multiple statistically independent components. By comparing spatial distribution characteristics with temporal waveform characteristics, artifact components highly correlated with the subject's blinking, eye movement, and muscle occlusion were identified and removed. Subsequently, the remaining effective neural components were projected and reconstructed into a pure multi-channel EEG temporal signal.
[0026] S12: Perform frame-segmentation, windowing, and baseline correction on the pure multi-channel EEG timing signal to obtain the baseline-corrected multi-channel EEG timing signal; In a specific embodiment of the present invention, S12 is specifically implemented by setting a fixed-length observation time window for the pure multi-channel EEG time-series signal stream, without time overlap between adjacent windows, and rigorously decomposing the long-term continuous signal into multiple independent short-term domain analysis frames. The device sampling rate of the EEG signal is set to... Each independent time window contains While there are discrete sampling points, significant physiological heterogeneity exists among different subjects in terms of scalp impedance, skull thickness, and baseline resting neuronal firing state, resulting in inherent individual baseline shifts in the amplitude and DC component of the EEG signal. To eliminate the interference of this baseline shift on the distribution of emotional characteristics across subjects, this application performs baseline correction on each time-domain frame sample of the multi-channel EEG time-series signal. A fixed-duration resting-state EEG signal without task is extracted before the presentation of the emotion-evoked stimulus as an individual baseline reference scale. The time average of this baseline segment on each channel is then calculated, and the baseline average of the corresponding channel is subtracted from each sampling point of the subsequent evoked state analysis frame. Through this baseline correction process of subtracting the mean, background neural electrical activity unrelated to emotion processing is effectively suppressed, highlighting the characteristic intensity of the emotion-evoked neural response components.
[0027] S13: Decompose the baseline-corrected multichannel EEG time-series signal into several frequency bands and calculate the differential entropy characteristics of each frequency band; In a specific embodiment of this invention, the baseline-corrected multi-channel EEG time-series signal undergoes signal domain transformation and frequency band separation. However, due to the highly nonlinear and cross-band distribution of emotional information carried in the EEG signal, this invention utilizes a Butterworth bandpass filter bank to perform frequency domain filtering and separation on the signal within each discrete time window, specifically five sub-frequency bands with definite cortical neurophysiological correlations, namely... Frequency band (1-3Hz, usually associated with unconscious states, but can be helpful in identification under certain strong emotions). Frequency band (4-7Hz, highly correlated with cognitive control and emotional inhibition) Frequency band (8-13Hz, negatively correlated with resting arousal and relaxation state of the brain) Frequency band (14-30Hz, directly mapping active cognitive processing, alertness, and emotional dimensions such as tension), and The frequency band (31-50Hz, representing complex multimodal information integration and intense emotional fluctuations) yielded a multi-channel, multi-frequency time series. Considering that the amplitude distribution of the filtered EEG signal within a specific finite frequency band approximately follows a Gaussian distribution... ,in, It follows a normal distribution. The mean, Given the variance, and considering that differential entropy can effectively quantify the logarithmic energy and system complexity of this continuous random variable, calculate the neural signal sequence for each sub-band. Differential entropy characteristics As shown in equation (3): (3) in: EEG signals in a specific frequency band The eigenvalues of the differential entropy, Let be the probability density function of a time series signal in a specific frequency band. For a specific frequency band of EEG signal time series, For the instantaneous values of the EEG signal time series, The variance of the time-series signal in this specific frequency band. It is a natural constant. For a specific frequency band of EEG signal time series, It is the natural logarithm. Pi is the mathematical constant of a circle.
[0028] Compared to traditional time-domain amplitude or frequency-domain power spectral density, the differential entropy feature used in this invention has superior low-frequency and high-frequency energy resolution, which can effectively highlight weak abnormal emotional fluctuations in large-scale background electrical activity, thereby achieving a highly discriminative expression of complex emotional states.
[0029] S14: Concatenate and stack the differential entropy eigenvalues obtained within a single independent time window to reconstruct a two-dimensional feature matrix. ,in, For the set of real numbers, This represents the total number of physiological sub-bands. To determine the total number of scalp electrode channels, a Z-score operation is performed on the two-dimensional feature matrix to obtain a standardized initial feature matrix. .
[0030] After extracting the differential entropy for all observation windows, the extracted discrete features need to be structurally reorganized according to the mapping relationship between the brain's physical spatial topological location and frequency depth, traversing all... Individual scalp electrode channels and the above-described decomposition Each physiological sub-band is used to concatenate and compactly stack all differential entropy eigenvalues obtained within a single independent time window, reconstructing a two-dimensional feature matrix with clear physical meaning. In this matrix structure, row vectors represent the overall energy state of specific cortical spatial nodes in the brain, while column vectors characterize the topological energy distribution of the whole brain at specific oscillation frequencies.
[0031] The two-dimensional feature matrix is normalized using Z-score, as shown in equation (4):
[0032] in: The initial feature matrix after standardization. It is a two-dimensional feature matrix. This represents the statistical mean of all samples from a single subject across a specific channel and frequency band. The mean and standard deviation of all samples from a single subject across a specific channel and frequency band dimension.
[0033] In a specific embodiment of this invention, the purpose of performing Z-score operation on the two-dimensional feature matrix is that, due to irreversible physiological and physical heterogeneity among individuals, such as the folding degree of the cerebral sulci cortex, skull conductivity, and electrode contact impedance, the extracted differential entropy features will still have systematic biases in the global amplitude scaling scale and data boundary distribution. Therefore, within the independent dataset of each subject, individualized Z-score standardization operation is performed on the feature matrix extracted from all time window samples, dimension by dimension.
[0034] Through the above elimination operation, the original physiological characteristics of all heterogeneous individuals are forcibly aligned into a unified low-order statistical distribution space with zero mean and unit variance. The calibrated and standardized initial feature matrix is output, which serves as the standard input base for the subsequent Bilateral Cross-Domain Attention (BCDA) correction module, providing a stable data foundation for feature decoupling across subjects.
[0035] S2: Construct a dual-path attention correction module to perform frequency band and channel dimension attention calibration on the standardized initial feature matrix respectively, and use Hadamard product to nonlinearly recombine the two calibration results to obtain the correction feature matrix; In a specific embodiment of the present invention, in order to suppress the deviation of neural response patterns caused by the heterogeneity of the physiological structure of the subject's brain at the physical source and prevent individual-specific noise from propagating into the deep feature space, the standardized initial feature matrix output in S1 is... A dual-path attention correction module is constructed, which decouples the weight quantization of the input matrix through parallel frequency band branch and channel branch.
[0036] Since there are significant individual differences in the dominant frequency bands of cerebral cortex neural activity when different subjects receive the same emotional stimuli, this step independently extracts the frequency domain dimension feature expression and assigns adaptive attention weights to it.
[0037] In a specific embodiment of the present invention, S2 specifically includes the following steps: S21: Construct a frequency band feature branch, transpose, pool, and perform nonlinear mapping on the normalized initial feature matrix to obtain the frequency band attention weight vector, and use the frequency band attention weight vector to calibrate the normalized initial feature matrix to obtain the frequency band calibration feature matrix; In a specific embodiment of the present invention, S21 specifically includes the following sub-steps: S211: Obtain the standardized initial feature matrix And transpose, and then normalize the initial feature matrix after transposition. Pooling is performed to obtain a one-dimensional frequency band aggregated feature vector; In a specific implementation of this application, S211 is specifically implemented as follows: First, obtain the standardized initial feature matrix. The matrix transpose operation yields , The initial feature matrix after standardization The matrix is transposed so that its row dimension is changed to the physiological frequency band dimension. Then, a global average pooling operation is applied to the transposed standardized feature matrix along the physical channel dimension to generate a matrix of length [length missing]. The middle one-dimensional frequency band aggregated feature vector, the physical essence of the pooling operation is to aggregate the energy of each electrode node distributed throughout the entire scalp space without discrimination, eliminate the interference of differences in spatial topology, and thus extract a macroscopic statistical scalar that characterizes the global overall energy activity level of each physiological sub-band.
[0038] S212: Input the one-dimensional frequency band aggregated feature vector into the frequency band multilayer perceptron to obtain the frequency band attention weight vector. Using frequency band attention weight vectors Each diagonal element constructs a frequency band diagonal matrix against a normalized initial feature matrix. Perform a right multiplication to obtain the frequency band calibration feature matrix. ; In a specific implementation of this application, S212 is implemented as follows: First, a one-dimensional frequency band aggregated feature vector is input into a frequency band multilayer perceptron composed of a linear layer network. The linear network performs nonlinear feature projection and transformation on the input vector in a high-dimensional space to deeply explore the potential mapping dependency between energy fluctuations in different sub-frequency bands and complex discrete emotional states.
[0039] After completing the nonlinear projection, the Sigmoid activation function is used. The output of the frequency band multilayer perceptron (fully connected network) is normalized, and the unbounded projection values are forcibly compressed and mapped to a continuous real number interval of (0,1) to obtain the frequency band attention weight vector. The calculation process is shown in equation (5):
[0040] in: For frequency band attention weight vectors, It is the Sigmoid activation function. For a fully connected network, This is a global average pooling operation along the physical channel dimension.
[0041] Frequency band attention weight vector Each component precisely quantifies the importance of the corresponding physiological sub-band in the current emotion classification task.
[0042] Finally, using the frequency band attention weight vector Each component constructs a diagonal element The frequency band diagonal matrix will be used to standardize the initial feature matrix. By performing a right multiplication operation with the frequency band diagonal matrix, the frequency band calibration feature matrix is obtained. The calculation process is shown in equation (6):
[0043] in: For frequency band calibration feature matrix, Frequency band attention weight vector The constructed diagonal matrix.
[0044] This step enables adaptive signal amplification of highly emotionally sensitive frequency bands within the frequency space, while effectively suppressing irrelevant frequency bands containing a large amount of individual background noise.
[0045] S22: Construct channel feature branches, perform pooling and nonlinear mapping on the standardized initial feature matrix to obtain channel attention weight vectors, and use the channel attention weight vectors to calibrate the standardized initial feature matrix to obtain the channel space calibration feature matrix.
[0046] Parallel to the computational logic of the frequency band feature branch, this step aims to independently extract the feature representation of the physical space dimension and dynamically assign spatial attention weights.
[0047] S22 specifically includes the following sub-steps: S221: The initial feature matrix normalized along the physiological sub-band dimension Perform global average pooling to obtain a length of One-dimensional channel aggregated feature vector; The physical essence of S221 is to aggregate the energy response information of a specific spatial electrode node across all physiological frequency bands without discrimination, thus eliminating the interference of frequency-specific fluctuations. This allows for the extraction of a macroscopic statistical scalar characterizing the overall activity of each physical electrode's local spatial location, generating a scalar of length [missing information]. The middle one-dimensional channel aggregates the feature vector.
[0048] S222: Input the one-dimensional channel aggregated feature vector into the channel multilayer perceptron to obtain the channel attention weight vector. Then use the channel attention weight vector Each diagonal element constructs a channel diagonal matrix, which is then used to normalize the initial feature matrix. Left multiplication with the channel diagonal matrix yields the channel space calibration feature matrix. ; In a specific embodiment of this application, S222 is implemented as follows: First, the one-dimensional channel aggregated feature vector is input into an independently constructed channel multilayer perceptron. This linear network performs nonlinear feature mapping on the input vector in the physical space dimension, aiming to explore the intrinsic spatial topological relationship between the activation state of specific functional brain regions and the specific emotion induction process.
[0049] Then, the Sigmoid activation function is used. The output of the fully connected network of this channel is normalized, and the spatial projection values are forcibly constrained to... Within a continuous real number interval, the channel attention weight vector is calculated. As shown in equation (7):
[0050] in: Here is the channel attention weight vector. This is a global average pooling operation along the physiological sub-band dimension.
[0051] Channel attention weight vector It can dynamically assign higher activation weights to cortical regions that are highly related to emotion processing, while adaptively suppressing spatial channels with low signal-to-noise ratios, poor electrode contact, or severe individual-specific physiological background noise.
[0052] Based on the channel attention weight vector Each component, construct a The channel diagonal matrix is then compared with the normalized initial characteristic matrix. Perform a left multiplication operation, as shown in equation (8), to obtain the channel space calibration feature matrix. .
[0053]
[0054] in: To calibrate the feature matrix for the channel space, For channel attention weight vectors The constructed diagonal matrix.
[0055] S23: The Hadamard product is used to perform nonlinear feature reconstruction on the frequency band calibration feature matrix and the channel space calibration feature matrix to obtain the source-level correction feature matrix. .
[0056] In the specific embodiments of this application, frequency band calibration feature matrices for amplification of specific frequency bands are obtained respectively. Channel spatial calibration feature matrix for cortical spatial filtering Subsequently, the dual-path features need to be synergistically fused in orthogonal dimensions to integrate calibration information from the frequency domain and physical space domain. To preserve the fine-grained interaction information and local neural activation patterns of the frequency and spatial domain features to the greatest extent possible during the fusion process, this step excludes linear additive fusion, which may cause information smoothing, and obtains the frequency band calibration feature matrix. With channel space calibration feature matrix The Hadamard product is used to perform nonlinear feature recombination on the two to generate a source-level correction feature matrix. As shown in equation (9):
[0057] in: To correct the feature matrix, For Hadama accumulation.
[0058] By completing the decoupling and reconstruction of the above dual paths, this invention successfully filters out the response pattern bias introduced by the heterogeneity of the subject's brain physiological structure at the input end before the low-order features are transmitted to the high-order deep network. This source-level correction mechanism greatly reduces the global optimization difficulty of the subsequent domain adaptation module, ensuring that the subsequent distribution alignment and deep clustering operations in the cascaded architecture can be built on a pure feature base with high emotional discriminative power and maximum suppression of individual-specific physiological noise.
[0059] S3: Construct a deep feature extraction backbone network consisting of a cascaded dynamic graph convolution module, a multi-head self-attention module, and a feature mapping module. Input the correction feature matrix into the three modules to obtain high-level feature representations of the source domain and the target domain. In a specific embodiment of the present invention, the correction feature matrix output by S2 is... , To further explore the dynamic spatial collaboration network and deep topological associations among different regions of the cerebral cortex during emotional cognitive processing, a deep feature extraction backbone network was constructed for high-level mapping.
[0060] S3 specifically includes the following sub-steps: S31: The scalp EEG electrode channels are abstracted as vertices in a mathematical graph model. A spatial adjacency matrix is constructed using the electrode spatial coordinates, and the Laplace matrix is calculated. The graph convolution kernel is truncated based on Chebyshev polynomials. The calibration feature matrix is then subjected to graph convolution to obtain the calibration feature matrix. Spatial characteristics; S3.1 includes the following sub-steps: S3.1.1: Abstract the topological structure of the human brain into a mathematical graph model, a set of vertices. Corresponding to Each scalp electroencephalographic electrode channel (i.e. ), edge set To represent the functional connectivity between different brain regions, Euclidean distances were calculated using the three-dimensional spatial coordinates of various scalp EEG electrodes within the International 10-20 system. A normalized spatial adjacency matrix was then constructed using a Gaussian kernel function. To quantify the spatial dependence strength between different electrode nodes, the Laplace matrix is calculated based on graph theory as shown in equation (10): (10) in: It is the identity matrix. For degree matrix, It is a spatial adjacency matrix.
[0061] S3.1.2: To reduce the enormous computational complexity of eigenvalue decomposition of the feature matrix during graph convolution, this invention utilizes Chebyshev polynomials to truncate the graph convolution kernel, obtaining the spatial node fusion feature matrix of the graph convolution output. As shown in equation (11).
[0062]
[0063] In the formula: The spatial node fusion feature matrix is the output of graph convolution. Let be the truncation order of the Chebyshev polynomial. This is the scaled Laplace matrix. For the first Chebyshev polynomials To correct the feature matrix, For the first The order-learnable graph convolution spatial weight matrix, where k is the order index.
[0064] The scaled Laplacian matrix in the spatial graph convolution operation is shown in equation (12): (12) In the formula: The largest eigenvalue, For the original Laplace matrix No. Chebyshev polynomial The functional relationships shown in equations (13)-(14) are satisfied: (13) in: It is a Chebyshev polynomial of order 0. It is an identity matrix.
[0065] (14) in: It is a first-order Chebyshev polynomial.
[0066] This invention is achieved by performing By performing graph convolution operations of order 3, the backbone network can effectively aggregate the topological spatial energy distribution features of each EEG channel within its third-order neighborhood, thereby fusing isolated node features into a graph node feature matrix with global brain region collaborative perception capabilities. This step, while retaining the advantages of noise correction at the input end, realizes deep feature encoding of EEG signals in the non-Euclidean physical space domain.
[0067] S32: Using three sets of learnable weight matrices, the spatial features are projected into a query matrix, a key matrix, and a value matrix. The features are then divided into multiple parallel attention heads along the feature dimension, and the context-weighted features of each attention head are calculated to obtain the attention feature matrix. In obtaining the spatial node fusion feature matrix of the graph convolution output Furthermore, considering that the cognitive processing of human emotions is not only a local spatial activation, but also a globally distributed processing process with a high context dependence, S32 further introduces a multi-head self-attention mechanism to perform deep semantic reconstruction of features.
[0068] S32 includes the following sub-steps: S321: Using three sets of independent learnable linear transformation weight matrices, the spatial node fusion feature matrix is used. Projected into query matrices respectively Key matrix Value matrix Query matrix Key matrix Value matrix Divide along the feature dimension Each attention head has a parallel attention head, and the context-weighted features of each attention head are calculated.
[0069] In a specific embodiment of the present invention, the input spatial graph features are transformed using three sets of independent learnable linear transformation weight matrices. Projected onto the query matrix respectively. ), key matrix (Key, ) and value matrix (Value, ).
[0070] In a specific embodiment of the present invention, in order to enable the model to jointly focus on emotional semantic information of different dimensions from multiple orthogonal representation subspaces, the following measures are taken: , , The matrix is divided equally along the feature dimension. Each attention head has a parallel attention head. Within each independent attention head, a scaled dot product attention computation is performed to compute the query matrix. AND key matrix The transposed dot product is divided by a scaling factor (i.e., the square root of the single-head feature dimension) to prevent gradient vanishing or exploding during backpropagation. Then, a soft max normalization function is applied to transform it into an attention weight matrix. This weight matrix precisely quantifies the degree of attention each node assigns to all other nodes in the feature space. Finally, this weight matrix is compared with the value matrix... Perform matrix multiplication to obtain the context-weighted features of a single head output.
[0071] S322: Concatenate the local weighted features of all attention heads along the channel dimension, input them into a fully connected linear projection layer for fusion and dimensionality reduction, and then use a random inactivation mechanism with an inactivation rate of 0.5 to suppress overfitting, resulting in an attention feature matrix rich in global emotional semantic interaction information. .
[0072] In the specific embodiments of this application, all After parallel computation of each attention head, the local weighted features of each head are concatenated along the channel dimension, and then fused and dimensionality reduced through a fully connected linear projection layer. In this computational process, to further enhance the model's generalization ability, a random deactivation mechanism with a deactivation rate of 0.5 is introduced. This means that during the training phase, the activation transmission of neurons is randomly blocked with a 50% probability, effectively suppressing the high-dimensional overfitting tendency that deep, complex networks are prone to in cross-subject feature learning. After this step, an attention feature matrix rich in global emotional semantic interaction information is output.
[0073] S33: Attention Feature Matrix Mapping to a high-dimensional latent feature space yields high-level feature representations of the source domain and the target domain.
[0074] In a specific embodiment of this application, in order to make the attention feature matrix To perfectly adapt to and support the distance metric calculation of subsequent conduction-level and terminal-level corrections, the attention feature matrix needs to be... It is mapped to a unified, high-dimensional latent feature space with strong representational capabilities.
[0075] The specific implementation method is as follows: After being flattened, the input is fed into a feature mapping multilayer perceptron module consisting of a nonlinear fully connected layer and a ReLU activation function. Through parameterized matrix affine transformation and nonlinear mapping, the matrix features are reshaped into dense feature vectors of fixed dimensions, realizing the mapping of features to a high-dimensional latent space. To achieve the optimal balance between the information capacity of feature representation and the computational cost of the model, this invention preferably sets the feature dimension of the latent space. .
[0076] At this point, the deep feature extraction backbone network has completed the high-level semantic encoding of the source-level correction features, and outputs high-level feature representations of the source domain for both the source domain input sample set and the target domain input sample set. , ,in, The total number of source domain samples in the current training batch, and the high-level feature representation of the target domain. , ,in, Given the total number of target domain samples in the current training batch, these two sets of high-level feature representations will serve as the direct input base for the next stage of the conduction-level correction module.
[0077] S4: Construct a global-local collaborative alignment module to perform transmission-level distribution offset correction on the high-level feature representations of the source domain and the high-level features of the target domain, and calculate the global-local collaborative alignment loss; In a specific embodiment of the present invention, the source domain high-level feature representation output by the deep feature extraction backbone network in S3 is addressed. High-level feature representation of the target domain ,in, This represents the total number of samples in the source domain within the current training batch. The total number of samples in the target domain within the current training batch represents the total number of samples. Although source-level frequency band channel correction is completed by S2, the inherent physiological heterogeneity between individuals still causes severe transmission-level distribution shifts in this high-dimensional latent space due to the error amplification effect of deep neural networks when performing nonlinear high-dimensional mappings. To fundamentally eliminate this distribution drift, this step constructs a global-local collaborative alignment module, which performs deep distribution distortion correction by coordinating the first and second-order statistics of the feature space.
[0078] S41: Calculate the difference between the high-level feature representations of the source domain and the high-level feature representations of the target domain in the first-order statistics, and construct the global macro-distribution alignment loss; To achieve overall overlap across the subject feature space, it is first necessary to eliminate the spatial positional deviation between the source and target domains at a macro scale, and independently compute the source domain feature matrix along the sample batch dimension. With the target domain feature matrix The mathematical expectation of the source domain feature matrix The mathematical expectation function expression is shown in equation (15), and the target domain feature matrix is... The functional expression for the mathematical expectation is shown in equation (16): (15) in: Let be the center vector of the first-order statistics of the source domain. , Let the center vector of the first-order statistic of the target domain be . , This represents the total number of source domain samples in the current training batch. The first characteristic matrix of the source domain is the first characteristic matrix of the source domain. Row vectors.
[0079] (16) in: The first feature matrix of the target domain Row vectors.
[0080] Source domain feature matrix With the target domain feature matrix The mathematical expectation precisely characterizes the global macroscopic centroid position of each physiological feature sample group in the high-dimensional latent space in a geometric and topological sense.
[0081] After accurately obtaining the macroscopic centroid coordinates of the two domains, to eliminate the global distribution drift of the two in the latent space, the squared Euclidean distance is used to measure and penalize the spatial dispersion between the two mean vectors, thereby constructing a global macroscopic distribution alignment loss function. Furthermore, to prevent the accumulation of high-dimensional spatial distances during backpropagation from causing gradient magnitude explosion and thus affecting the convergence stability of the model optimization, this invention introduces a feature dimension. The distance scalar is averaged and penalized as a scaling factor, and the loss is aligned with the global macroscopic distribution. By minimizing the constraint, the network can forcefully drive the feature sample group of the target domain to undergo a global translation in the latent space during iterative optimization, making its macroscopic centroid infinitely approach the centroid coordinates of the source domain. This completely eliminates the basic distribution shift across subject features at the first-order statistical level, thus eliminating the global macroscopic distribution alignment loss. As shown in equation (17):
[0082] in: For global macroscopic distribution alignment loss, It is an L2 norm.
[0083] By adjusting the alignment loss of the global macroscopic distribution By minimizing constraints, the network can forcefully drive the feature sample group of the target domain to undergo an overall translation in the latent space during iterative optimization, making its macroscopic centroid infinitely approach the centroid coordinates of the source domain, thereby completely eliminating the basic distribution shift across subject features at the level of first-order statistics.
[0084] S42: Calculate the difference between the high-level feature representations of the source domain and the high-level feature representations of the target domain in the second-order statistic, and construct the local correlation geometric alignment loss; While global alignment at the first-order statistic level can eliminate translational bias in the overall distribution, it often overlooks distortions in the internal structure of the feature subspace caused by differences in individual cognitive patterns among subjects. These distortions manifest as a warped correlation between feature dimensions in the high-dimensional latent space. Therefore, to more accurately correct this misalignment in the internal refined spectral-spatial correlation, S42 calculates the second-order statistic of the feature distribution, namely the unbiased covariance matrix.
[0085] In a specific embodiment of the present invention, step S42 is as follows: S421: Utilizing the central vector of the first-order statistic in the source domain With the center vector of the first-order statistic of the target domain For the source domain characteristic matrix respectively and target domain feature matrix The centering process is performed to obtain the feature matrix of the source domain after centering and the feature matrix of the target domain after centering.
[0086] S422: Using the eigenvalue matrix of the centered source domain and the eigenvalue matrix of the centered target domain, calculate the unbiased covariance matrix of the source domain, respectively. , The unbiased covariance matrix of the features of the target domain , As shown in equations (18)-(19);
[0087] in: Let be the unbiased covariance matrix of the source domain. This represents the total number of source domain samples in the current training batch. The source domain characteristic matrix, For length is A column vector of all 1s It is the transpose of the center vector of the first-order statistics of the source domain.
[0088]
[0089] in: Let be the unbiased covariance matrix of the target domain features. This represents the total number of target domain samples in the current training batch. The feature matrix of the target domain, For length is A column vector of all 1s It is the transpose of the center vector of the first-order statistics of the target domain.
[0090] The covariance matrix algorithm described above can accurately capture the local geometric correlations between various feature dimensions in the latent space and the co-occurrence patterns of neural activation distributions.
[0091] S423: Calculate the structural divergence difference between the unbiased covariance matrix of the source domain and the unbiased covariance matrix of the target domain features using the squared Frobenius norm, and construct a local alignment loss. As shown in Equation (20), it is used to force the feature geometry of the target domain to be rotated and scaled to align with the geometry of the source domain.
[0092]
[0093] in: This is an empirical normalization scaling factor derived based on the norm bounds of high-dimensional matrices. For local alignment loss, It is the Frobenius norm.
[0094] In this invention, the empirical normalization scaling factor derived from the norm limit of high-dimensional matrices strictly constrains the numerical magnitude balance between the second-order covariance loss and the first-order centroid loss in backpropagation calculation, preventing gradient domination caused by excessively large covariance matrix element values.
[0095] S43: Directly superimpose the global macro-distribution alignment loss and the local alignment loss to construct a global-local cooperative joint loss; Existing traditional domain adaptation and alignment methods typically require the introduction of additional artificial hyperparameters to balance the alignment strength between macroscopic drift terms and local distortion terms. However, when dealing with highly dynamic cross-subject EEG signals with extremely low signal-to-noise ratios, artificially fixing hyperparameters often leads the model into local suboptimal solutions.
[0096] This invention proposes to directly incorporate the global macroscopic distribution alignment loss (global centroid drift penalty term) by strictly unifying physical dimensions and abandoning complex hyperparameter searches. Local alignment loss with (local covariance distortion penalty term) By performing unweighted direct superposition, a parameter-independent joint alignment loss is obtained. As shown in equation (21):
[0097] in: To coordinate and align joint losses.
[0098] Through a dual constraint and collaborative mechanism of global and local constraints, the network can not only forcefully narrow the physical distance between the feature distributions of different subjects at the macro level, but also strictly protect and maintain the fragile emotional discriminative feature signature contained in the feature space at the micro level. This collaborative strategy effectively breaks through the bottleneck of excessive alignment technology in traditional domain adaptation algorithms, which loses classification features in pursuit of alignment, from the mathematical foundation, and achieves perfect compatibility between distribution invariance and emotional discriminability.
[0099] S5: Construct a deep clustering module with learnable centers. Based on the source domain feature representation and the target domain feature representation, calculate the supervised contrast loss and the learnable center loss, and construct the joint loss for deep clustering based on the supervised contrast loss and the learnable center loss.
[0100] Although the global-local collaborative alignment module in S4 effectively eliminates the feature domain distribution offset across subjects, in complex EEG emotion recognition tasks, eliminating inter-domain differences often inevitably leads to category overlap. That is, in a unified high-dimensional latent space, neural representations belonging to different emotion categories may still have semantic interleaving, resulting in blurred boundaries of the final classification decision. To completely solve this terminal-level emotion semantic misalignment problem, this invention constructs a learning-center deep clustering module, which enforces dual geometric topological constraints of intra-class compactness and inter-class separation within the latent space.
[0101] S51: High-level feature representation matrix for the current batch Each sample feature vector in the sample is processed Norm normalization is performed, and the temperature-scaled cosine similarity between each pair of samples within a batch is calculated based on the normalized eigenvectors to obtain the calibrated spatial similarity. .
[0102] In a specific embodiment of the present invention, step S51 is as follows: S511: Extracting high-level feature representation matrix from the feature stream processed by the alignment network. , and its corresponding sentiment tag vector , ,in, This refers to the batch sample size. For the latent space feature dimension, For emotional labeling space.
[0103] S512: Utilizing Norm pairs of high-level feature representation matrices High-level feature vector of each independent sample in After normalization, the vector is projected onto a high-dimensional unit hypersphere to obtain a new unit eigenvector. This eliminates the difference in absolute norm of high-dimensional features caused by drastic energy fluctuations in different channels, making subsequent clustering metrics purely dependent on the angular direction of the feature vectors in space. The new unit feature vector... The calculation process is shown in equation (22):
[0104] in: For each independent sample, the high-level feature vector The new unit eigenvectors are obtained by normalization. For each independent sample, the high-level feature vector, for Norm.
[0105] S513: Based on the new unit feature vectors of each sample within a batch, calculate the cosine similarity between pairs of samples, and after adjustment by the temperature hyperparameter τ, obtain the similarity of any sample within the batch. i and j The calibrated spatial similarity between .
[0106] In a specific embodiment of the present invention, the calibrated spatial similarity As shown in equation (23):
[0107] in: For any two samples within a batch and The calibrated spatial similarity between them For each independent sample, the high-level feature vector The new eigenvectors obtained by normalization For temperature overparameter, For each independent sample, the high-level feature vector The new eigenvectors are obtained by normalization.
[0108] In a specific embodiment of the present invention, the calibrated spatial similarity It provides a high-resolution distance metric benchmark for subsequent contrastive learning and semantic clustering.
[0109] S52: Using the sample sentiment category label as the supervision signal, the samples within the batch are compared with the anchor samples. i Other samples with the same label are defined as the positive sample set. Samples with different labels are defined as the negative sample set, and a supervised contrastive loss function is constructed. .
[0110] Obtaining spatial similarity after calibration Subsequently, to effectively broaden the latent space decision boundaries for different emotion categories, this invention introduces a supervised contrastive learning paradigm. Unlike unsupervised contrastive learning, which only treats different augmented versions of the same sample as positive samples, this invention fully utilizes the emotion category labels of EEG samples as strong supervision signals for the current anchor sample within a batch. Define the set of all other samples within a batch that share the same sentiment label as the positive sample set, and denote its index set as . Samples with different sentiment labels naturally constitute the negative sample set. This is achieved by constructing a supervised contrastive loss function. During backpropagation, the algorithm is forced to maximize the spatial cohesion of the anchor sample and all positive samples on the hypersphere, while minimizing the similarity between the anchor sample and all negative samples (i.e., strengthening the inter-class repulsion).
[0111] In a specific embodiment of the present invention, step S52 is as follows: S521: Obtain the calibrated spatial similarity between all pairs of samples within a batch. And using the sample sentiment category label as a supervision signal, the samples within the batch are compared with the anchor samples. i Other samples with the same label are defined as the positive sample set. P ( i Samples with different labels are defined as the negative sample set; S522: Constructing the supervised contrastive loss function As shown in Equation (24), during backpropagation, the spatial cohesion between the anchor sample and all positive samples on the hypersphere is maximized, while the similarity between the anchor sample and all negative samples is minimized.
[0112] in: To supervise the contrastive loss function, This represents the total number of samples in the current training batch. This is the index of the current anchor sample in the batch. For anchor point samples A set of indices of other samples with the same sentiment label. Positive sample set The number of elements in the middle, for Index in For the sample With sample The calibrated spatial similarity between them To iterate over all that are not equal to Sample index, For the sample With sample The calibrated spatial similarity between them It is an exponential function. It is the natural logarithm function.
[0113] In a specific implementation of this application, the supervised contrastive loss function is minimized. While aligning the source and target domains, the network can open up a non-aliasing discrimination region with a large margin in the latent space for multi-classification tasks by means of an explicit semantic isolation mechanism.
[0114] S53: Initialize a learnable class center vector for each sentiment category in the latent space, and calculate the mean squared error between the high-level feature vector of each sample in the batch and its corresponding class center vector to construct a center loss function to achieve the ultimate compression of intra-class variance.
[0115] While relying solely on supervised contrastive loss can broaden inter-class boundaries, the lack of absolute anchor points in space leads to loose manifold distributions of similar features in the latent space, making the model extremely sensitive to noise in the target domain. To address this issue, S53 introduces a learnable class center constraint mechanism.
[0116] S531: In the hidden space, for Each sentiment category initializes a set of learnable center vector matrices. These center vector matrices serve as a set of independent parameters for the network and are iteratively updated synchronously with the network weights during backpropagation. S532: Calculate the high-level feature vector for each independent sample within a batch. Compared to its true emotional label The corresponding specific category center The mean squared error between them is used to construct the central loss function. As shown in equation (25), and a minimization constraint is applied;
[0117] in: For the center loss function, This represents the total number of samples in the current training batch. For the first The original high-level feature vector of each sample, To be consistent with the sample Authentic Labels The corresponding learnable category center vector.
[0118] In the specific implementation of this application, the center loss function is minimized, and the network generates a strong centripetal force, forcing all feature samples belonging to the same emotion category to collapse toward their corresponding dynamic center point, thereby achieving the ultimate compression of intra-class variance in multi-dimensional space and greatly enhancing the robustness of single emotion representation.
[0119] S54: The supervised contrast loss function and the center loss function are weighted and combined to construct a deep clustering loss with learnable centers.
[0120] In the specific implementation of this application, the monitored comparative loss will be... With center loss function To combine and utilize the balance coefficient By adjusting the weight ratio of the two, the deep clustering loss with learnable centers can be obtained. As shown in equation (26):
[0121] in: For the center-based deep clustering loss, in this application, This is an empirically optimal value (a real number), which can be determined by tuning on the validation set based on different datasets. The optimal value is... [0.1,1].
[0122] Finally, contrast loss expands inter-class margin as a repulsive force to improve classification discrimination, while center loss compresses intra-class divergence as an attractive force to improve perturbation resistance. Under these dual constraints, the problems of class overlap and sentiment semantic misalignment left over after cross-subject distribution alignment are completely corrected, and the topological structure of the feature space is completely reconstructed.
[0123] S6: Based on the loss function constructed in S4-S5, the cascaded correction network is trained. The cascaded correction network consists of a dual-path attention correction module, a deep feature extraction backbone network, and a global-local collaborative alignment module and a deep clustering module with learnable centers connected in parallel after the deep feature extraction backbone network. The trained network is then used to identify the emotional state of the EEG signals of the subjects in the target domain.
[0124] The three-stage cascaded correction network proposed in this invention is a completely end-to-end unified architecture. In order to completely block the cascade amplification effect of individual heterogeneity on the neural conduction chain, the relevant modules need to achieve strict closed-loop collaboration through multi-objective joint optimization.
[0125] S61: The high-level feature representations of the source domain and the target domain output from S3 are input into a linear emotion classifier to obtain the predicted probability distribution. The classification loss is then calculated based on the true labels of the source domain samples to construct the overall target loss function. In a specific embodiment of this invention, the high-level feature representations output from S3 are input into a linear emotion classifier consisting of a fully connected layer and a Soft max activation function to obtain the model's predicted probability distributions for various emotions. The classification loss is then calculated based on the true emotion labels of the source domain samples. Calculate classification and identification loss To systematically neutralize multiple heterogeneous biases along the transmission chain and establish a dynamic balance between feature discrimination and distribution invariance, this invention globally and physically integrates the various independent correction objectives to construct a total objective loss function. As shown in equation (27):
[0126] In the formula, This is the domain-adaptive alignment strength adjustment coefficient, used to penalize the convergence of feature distributions within the global-local collaborative alignment module. This is a semantic clustering strength adjustment coefficient, used to control the constraint strength of intra-class feature compression and inter-class feature repulsion in the learnable center deep clustering module. Let the total objective loss function be... To classify and identify losses, for Norm weight regularization term, For regularization weights.
[0127] In this invention, the domain-adaptive alignment strength adjustment coefficient Semantic clustering strength adjustment coefficient Regularization weights This is an empirically optimized value, which can be determined by fine-tuning on the validation set based on different datasets. The preferred value is... [0.01, 1.0] [0.1, 1.0] [1e -5 , 1e -3 ].
[0128] In a specific embodiment of the present invention, in constructing the total objective loss function At that time, considering the objective limitations of limited sample size and extremely high dimensionality of EEG signals, the overall objective loss function... Introduction Norm weight regularization term And paired with regularized weights The model dynamically decays and limits the magnitude of all learnable parameters of the network, thereby effectively suppressing the high-dimensional overfitting phenomenon that is prone to occur in deep and complex neural networks in cross-subject feature learning, and significantly improving the global stability of numerical calculation and model convergence.
[0129] S62: Backpropagate the network error based on the total objective loss function, and use the optimization algorithm to perform end-to-end synchronous iterative updates on the network weights until the total objective loss function converges, thus obtaining a cascaded correction network with fixed weights. In a specific embodiment of the present invention, during the offline joint training phase, based on the construction of the overall target loss function... Error backpropagation is performed on the overall computational graph of the network. An optimization algorithm with weight decay characteristics is employed to perform end-to-end synchronous iterative updates on the network weights of the front-end dual-path attention correction module, the deep feature extraction backbone network, the class center vector matrix, and the linear classifier. This closed-loop optimization mechanism mathematically establishes a powerful inter-module interdependence and dynamic reinforcement effect, improving the classification loss. Guide the backbone network to extract basic features with preliminary sentiment differences; cooperative alignment loss While driving the backbone network to eliminate inter-domain distribution drift, its computed gradient flow is further backpropagated to the front-end dual-path attention correction module, and the spatial calibration weights of frequency bands and channels are fine-tuned. This provides the system with a cleaner initial input with lower physiological noise during the next forward propagation, while the deep clustering loss... Then, by utilizing its updated learnable category centers, it provides category-aware semantic boundary feedback for the alignment process of the global-local collaborative alignment model. As a result, the source denoising at the front end significantly reduces the difficulty of the distributed alignment at the back end, while the semantic feedback at the back end further accurately corrects the physical correction parameters of the front end, enabling the entire three-stage cascaded architecture to achieve global optimal convergence through repeated joint backpropagation.
[0130] S63: Collect EEG signals from new subjects and extract differential entropy feature matrix. After processing by a cascaded correction network with fixed weights, output high-level feature representation, which is then input into a linear emotion classifier and passed through a Soft max layer to obtain the emotion category prediction result.
[0131] In a specific embodiment of the present invention, step S63 is as follows: S631: During the offline joint training phase, the cascaded correction network is trained using the training set until the total loss function is reached. Convergence completes the decoupling and learning of individual physiological heterogeneity characteristics, and solidifies the weight parameters of all network nodes within the model. When the model is deployed in a real brain-computer interface or emotion computing system for online evaluation, it enters a simplified forward inference mode. S632: Real-time acquisition of EEG signals from new subjects who have not participated in training, and frame segmentation, baseline correction, and frequency band separation according to a preset standard preprocessing procedure, and extraction of multi-channel and multi-frequency band differential entropy feature matrix; S633: The differential entropy feature matrix is input into the cascaded correction network with solidified weights. Adaptive source-specific denoising and frequency-spatial feature amplification are performed through the dual-path attention correction module. Then, advanced semantic reconstruction and high-dimensional latent space mapping are performed through the deep feature extraction backbone network to output advanced feature representation. The advanced feature representation is directly input into the linear classifier. The maximum probability score is calculated through the Softmax layer to output the discrete sentiment category prediction result.
[0132] It is worth noting that this invention exhibits high inference efficiency in its architectural design. During offline training, the cross-subject distribution barriers in the latent space are completely broken down by the global-local collaborative alignment module, and the discriminative semantic boundaries of various emotions are strictly defined by the learnable center deep clustering module. Therefore, during online evaluation, there is no need to compute complex collaborative distribution alignment and deep clustering. Instead, the high-level feature representation output from the backbone network is directly input into the linear classifier for extremely rapid forward propagation. Finally, based on the maximum probability score of the Softmax layer, the linear classifier outputs the accurate discrete emotion category prediction result for the new subject in real time.
[0133] The cascaded correction network based on dual-path attention and global-local collaborative alignment proposed in this invention was used in cross-subject emotion recognition experiments on the internationally recognized EEG emotion benchmark datasets SEED and SEED-IV, and compared with state-of-the-art methods in this field over the past five years. To rigorously verify the generalization performance of the algorithm, a leave-one-subject cross-validation scheme was used in the experiment. The relevant comparison results are shown in Tables 1 and 2.
[0134] Experimental results show that in complex cross-subject emotion recognition tasks, the proposed method achieves highly competitive performance on two mainstream datasets. On the SEED dataset shown in Table 1, the proposed method achieves a recognition accuracy of 74.70%, and on the SEED-IV dataset shown in Table 2, which has more categories and is more difficult, the recognition accuracy reaches 55.13%, comprehensively surpassing existing methods based on graph convolution and conventional domain adaptation. In particular, the proposed method demonstrates a significant technical advantage in terms of performance stability.
[0135] Experimental data clearly show that the performance standard deviation of CM-Net (Cascade Modulation Network) remains below 5% on both datasets. Compared with the current state-of-the-art benchmark methods, its performance fluctuation is reduced by about 56.8%. Analysis suggests that this extremely high stability is directly due to the systematic design of the three-stage cascaded correction in this invention.
[0136] Table 1. Comparison with state-of-the-art cross-subject methods on the SEED dataset.
[0137] Table 2 Comparison with state-of-the-art methods on the SEEDIV dataset
[0138] In summary, this invention proposes an end-to-end three-stage processing framework. It suppresses physiological noise at the source through front-end dual-path attention, alleviates distribution shift during transmission through mid-stage global-local collaborative alignment, and locks semantic boundaries at the terminal through back-end deep clustering. The experimental results not only verify the advanced nature of this invention in improving cross-subject recognition accuracy, but also fully demonstrate its excellent robustness in the face of heterogeneity among different subjects, providing reliable technical support for brain-computer interface technology to move from the laboratory to real-world complex scenarios.
[0139] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0140] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A cross-subject EEG emotion recognition method based on dual-path attention and global-local collaborative integration alignment, characterized in that, Includes the following steps: S1: Acquire EEG signals from the source and target domains, extract multi-band, multi-channel EEG differential entropy feature matrices as raw inputs, and standardize the raw inputs to obtain the standardized initial feature matrix; S2: Construct a dual-path attention correction module for frequency band feature branches and channel feature branches; S3: Construct a deep feature extraction backbone network consisting of a cascaded dynamic graph convolution module, a multi-head self-attention module, and a feature mapping module; S4: Construct a global-local cooperative alignment module to perform distribution offset correction on the high-level feature representations of the source domain and the target domain, and calculate the global-local cooperative joint loss; S5: Construct a deep clustering module with learnable centers. Based on the high-level feature representations of the source domain and the target domain, calculate the supervised contrastive loss and the learnable center loss, and construct the joint loss for deep clustering based on the supervised contrastive loss and the learnable center loss. S6: Based on the loss function constructed in S4-S5, the cascaded correction network is trained. The cascaded correction network consists of a dual-path attention correction module, a deep feature extraction backbone network, and a global-local collaborative alignment module and a deep clustering module with learnable centers connected in parallel after the deep feature extraction backbone network. The trained network is then used to identify the emotional state of the EEG signals of the target domain subjects.
2. The cross-subject EEG emotion recognition method based on dual-path attention and global-local collaborative integration alignment according to claim 1, characterized in that, The specific steps for S1 are as follows: S11: Using a multi-channel EEG acquisition device, raw EEG signals from the scalp of different subjects under a preset emotional evoked paradigm were obtained, and power frequency filtering, bandpass filtering, and independent component analysis were performed to remove artifacts and obtain pure EEG signals. S12: Framing, windowing, and baseline correction are performed on the pure EEG signal to obtain a baseline-corrected multi-channel EEG timing signal; S13: The baseline-corrected multichannel EEG time-series signal is decomposed into several frequency bands, and the differential entropy characteristics of each frequency band are calculated, as shown in Equation (1): (1) in: EEG signals in a specific frequency band The eigenvalues of the differential entropy, Let be the probability density function of a time series signal in a specific frequency band. A time series of EEG signals in a specific frequency band. For the instantaneous values of the EEG signal time series, The variance of the time-series signal in this specific frequency band. It is a natural constant. A time series of EEG signals in a specific frequency band. It is the natural logarithm. Pi; S14: Concatenate and stack the differential entropy eigenvalues obtained within a single time window to reconstruct a two-dimensional feature matrix, and then perform a standardization operation on the two-dimensional feature matrix to obtain a standardized initial feature matrix. .
3. The cross-subject EEG emotion recognition method based on dual-path attention and global-local collaborative integration alignment according to claim 1, characterized in that, The specific steps for S2 are as follows: S21: Construct a frequency band feature branch, transpose, pool, and perform nonlinear mapping on the standardized initial feature matrix to obtain the frequency band attention weight vector, construct a diagonal matrix using the frequency band attention weight vector, and perform right multiplication calibration on the standardized initial feature matrix to obtain the frequency band calibration feature matrix. S22: Construct channel feature branches, perform pooling and nonlinear mapping on the standardized initial feature matrix to obtain channel attention weight vectors, construct diagonal matrices using channel attention weight vectors, and perform left multiplication on the standardized initial feature matrix to obtain channel space calibration feature matrices; S23: The frequency band calibration feature matrix and the channel space calibration feature matrix are nonlinearly recombined using the Hadamard product to obtain the correction feature matrix.
4. The method for cross-subject EEG emotion recognition based on dual-path attention and global-local collaborative integration alignment according to claim 1, characterized in that, The specific steps for S3 are as follows: S31: The scalp EEG electrode channels are abstracted as vertices of a mathematical graph model, and a spatial adjacency matrix is constructed using electrode spatial coordinates. The Laplace matrix is calculated, and the graph convolution kernel is truncated based on Chebyshev polynomials. The correction feature matrix is then graph convolved to obtain the spatial node fusion feature matrix. S32: Using three sets of learnable weight matrices, the spatial node fusion feature matrix is projected into a query matrix, a key matrix, and a value matrix. The feature is divided into multiple parallel attention heads along the feature dimension, and the context-weighted features of a single attention head are calculated separately. The local weighted features of all attention heads are concatenated along the channel dimension and then fused to reduce the dimension, thus obtaining the attention feature matrix. S33: Map the attention feature matrix to a high-dimensional latent feature space to obtain the high-level feature representations of the source domain and the target domain, respectively.
5. The cross-subject EEG emotion recognition method based on dual-path attention and global-local collaborative integration alignment according to claim 1, characterized in that, The specific steps for S4 are as follows: S41: Calculate the difference between the high-level feature representation of the source domain and the high-level feature representation of the target domain in the first-order statistic, and construct the global macro-distribution alignment loss, as shown in Equation (2). in: For global macroscopic distribution alignment loss, It is the L2 norm. Let be the center vector of the first-order statistics of the source domain. Let the center vector of the first-order statistic of the target domain be . The dimension of the high-dimensional latent feature space; S42: Calculate the difference between the high-level feature representation of the source domain and the high-level feature representation of the target domain in the second-order statistic, and construct the local alignment loss as shown in Equation (3): in: This is an empirical normalization scaling factor derived based on the norm bounds of high-dimensional matrices. For local alignment loss, For the Frobenius norm, Let be the unbiased covariance matrix of the source domain. Let be the unbiased covariance matrix of the target domain features; S43: Superimpose the global macro-distribution alignment loss with the local alignment loss to construct a global-local collaborative joint loss.
6. The cross-subject EEG emotion recognition method based on dual-path attention and global-local collaborative integration alignment according to claim 1, characterized in that, The specific steps for S5 are as follows: S51: Perform L2 norm normalization on each sample feature vector in the high-level feature representation matrix of the current batch, and calculate the temperature scaling cosine similarity between each pair of samples in the batch based on the normalized feature vectors to obtain the calibrated spatial similarity, as shown in Equation (4): in: For any two samples within a batch and The calibrated spatial similarity between them For each independent sample, the high-level feature vector The new eigenvectors obtained by normalization For temperature overparameter, For each independent sample, the high-level feature vector The new feature vector obtained by normalization; S52: Using the sample sentiment category label as the supervision signal, other samples with the same label as the anchor sample within the batch are defined as the positive sample set, and samples with different labels are defined as the negative sample set. A supervised contrastive loss function is constructed based on the calibrated spatial similarity, as shown in Equation (5): in: To supervise the contrastive loss function, This represents the total number of samples in the current training batch. This is the index of the current anchor sample in the batch. For anchor point samples A set of indices of other samples with the same sentiment label. Positive sample set The number of elements in the middle, for Index in For the sample With sample The calibrated spatial similarity between them To iterate over all that are not equal to Sample index, For the sample With sample The calibrated spatial similarity between them It is an exponential function. It is the natural logarithm function; S53: Initialize a learnable class center vector for each sentiment category in the latent space, calculate the mean squared error between the high-level feature vector of each sample in the batch and its corresponding class center vector, and construct the center loss function based on the mean squared error, as shown in Equation (6): in: For the center loss function, For the first The original high-level feature vector of each sample, To be consistent with the sample Authentic Labels The corresponding learnable category center vector; S54: The supervised contrast loss function and the center loss function are weighted and combined to construct a deep clustering joint loss.
7. The cross-subject EEG emotion recognition method based on dual-path attention and global-local collaborative integration alignment according to claim 1, characterized in that, The specific steps for S6 are as follows: S61: Input the source domain high-level feature representation and the target domain high-level feature representation output from S3 into a linear sentiment classifier to obtain the predicted probability distribution and construct the total target loss function; S62: Backpropagate the network error based on the total objective loss function, and use the optimization algorithm to perform end-to-end synchronous iterative updates on the network weights until the total objective loss function converges, thus obtaining a cascaded correction network with fixed weights. S63: Collect EEG signals from new subjects and extract differential entropy feature matrix. After processing by a cascaded correction network with fixed weights, output high-level feature representation, which is then input into a linear emotion classifier and passed through a Soft max layer to obtain the emotion category prediction result.