Auditory attention decoding method and device based on EEG signal and medium
Through the multi-view spatial selection mechanism based on EEG signals and the adaptive time attention module, combined with the twin comparison learning strategy, the problem of failure to accurately distinguish auditory attention in the existing technology is solved, and more efficient auditory attention decoding is achieved.
Patent Information
- Application Number
- CN202510518134.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-08
AI Technical Summary
Existing auditory attention decoding techniques are difficult to accurately distinguish the sound sources of listeners’ attention in complex environments, especially in the case of multiple sound sources and noise mixing, and fail to fully consider the impact of early brain selection process on attention relocation.
A auditory attention decoding method based on EEG signals is designed, using a multi-view spatial selection mechanism and an adaptive time attention module, combined with a twin comparison learning strategy, to capture the spatiotemporal interaction characteristics and selection deviation characteristics of the EEG signal to form an auditory attention decoding model.
A more expressive auditory attention feature extraction is achieved, overcoming the impact of early brain selection on attention relocation, and improving the accuracy and robustness of auditory attention decoding.
Smart Images

Figure CN120448776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electroencephalogram (EEG) signal processing, and more particularly to an auditory attention decoding method, device, and medium based on EEG signals. Background Art
[0002] In a complex "cocktail party" scene, there is usually a mixture of voices from multiple speakers and background noise. However, when there are multiple sound sources, hearing-impaired people show insufficient ability to focus on the voice of the target speaker. Although modern hearing aids improve the hearing experience of hearing-impaired people in multi-speaker scenes by enhancing the clarity of the target sound source and suppressing background noise at the same time, they are usually still unable to accurately distinguish the specific sound source that the listener wants to focus on. Therefore, auditory attention decoding (AAD) research has become a potential solution to this problem.
[0003] Neuroscience research has found that an individual's auditory choices can be obtained by decoding brain activity signals. Existing tools for capturing brain activity include electrocorticography (ECoG), magnetoencephalography (MEG), and electroencephalography (EEG). In recent years, methods for decoding auditory attention have focused on two paradigms: one relies on stimulus reconstruction or speech envelope reconstruction, and the other relies on physiological signal processing, namely auditory spatial attention detection. However, real environments are usually mixed with multiple sound sources and noise, making it extremely challenging to obtain the required clean auditory stimuli. Relatively speaking, due to the natural advantages of EEG being non-invasive and low-cost, it has become the most popular AAD tool and provides the possibility of reflecting an individual's auditory focus in real time and accurately.
[0004] Advances in neuroscience research have revealed that the human brain exhibits complex physiological activities when processing auditory stimuli. These activities involve interactions between multiple brain regions, and these interactions are crucial for regulating auditory attention. Therefore, the paper "Cai, Siqi, Tanja Schultz, and Haizhou Li. "Brain topology modeling with EEG-graphs for auditory spatial attention detection." IEEE Transactions on Biomedical Engineering (2023)." uses graph networks to model the brain's functional connectivity and achieves advanced results in AAD decoding performance. The paper "Fan, Cunhang, et al. "DGSD: Dynamical graph self-distillation for EEG-based auditory spatial attention detection." Neural Networks 179 (2024): 106580." devises a method that combines graph convolution with a self-distillation algorithm. After each graph convolution layer, a deep model guides a shallow model, further enhancing its performance. Furthermore, psychoacoustics suggests that human attention is a dynamic and aggregated process. Based on this discovery, some studies have introduced attention mechanisms to pay long-term and short-term attention to different temporal features in the signal to improve the understanding of the characteristics of the entire sequence. However, extracting only single-dimensional features may be difficult to cope with complex signal patterns. Therefore, the document "Ni, Qinke, et al." Dbpnet: Dual-branch parallel network with temporal-frequency fusion for auditory attention detection." Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI 2024). 2024." proposed a dual-branch parallel time-frequency fusion network that fuses the features of time-varying and spectral space to achieve multi-dimensional capture of auditory attention patterns. This suggests that comprehensive consideration of the multi-channel feature correlation and temporal dynamic changes of EEG signals is crucial for the further development of AAD systems.
[0005] It's noteworthy that the human attention mechanism may engage in selective filtering in its early stages, suppressing the voice of unattended speakers and preventing them from participating in subsequent neural processing. Furthermore, during the process of switching attentional focus, frontal-parietal regions exhibit varying levels of change. Building on this, the paper "Sengupta, Ankita, et al. "The right posterior parietal cortex mediates spatial reorienting of attentional choice bias." Nature Communications 15.1(2024):6938. further found that the right posterior parietal cortex can influence the direction of attention by modulating attentional choice bias in a spatial attention reorientation task. Essentially, focusing attention on different speakers can be considered a form of attention reorientation. However, most existing methods fail to fully account for the brain's early selection process when focusing on different speakers, limiting their ability to capture the characteristics of selection bias during attention reorientation. Therefore, designing data-driven mechanisms guided by early attention selection has both research value and practical significance. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides an auditory attention decoding method, device and medium based on EEG signals, and designs an auditory attention learning system (SMCL-AAN) based on a twin architecture. Specifically, a multi-perspective spatial selection mechanism and an adaptive temporal attention module are carefully designed to exploit the spatiotemporal interaction characteristics of EEG signals to explore the dependency between local and global features, thereby generating more expressive auditory attention features. At the same time, in order to overcome the influence of early brain selection on attention relocation, an efficient twin-like contrastive learning strategy is designed to capture the intrinsic feature differences of the brain during attention relocation, thereby achieving robust extraction of selection bias features.
[0007] In a first aspect, the present invention provides an auditory attention decoding method based on EEG signals, the method comprising:
[0008] Acquire an EEG signal, and apply a sliding window to the EEG signal to obtain a sample pair;
[0009] Constructing a multi-scale temporal feature extraction module, inputting the sample pairs into the multi-scale temporal feature extraction module, and obtaining a time-varying feature map through multi-layer temporal convolution;
[0010] Constructing a spatiotemporal feature interactor, the spatiotemporal feature interactor including a multi-view spatial selection mechanism and an adaptive attention module, inputting the time-varying feature map into the spatiotemporal feature interactor, processing the time-varying feature map through the multi-view spatial selection mechanism and the adaptive attention module respectively to obtain a first feature map and a second feature map, and fusing the first feature map and the second feature map to obtain a spatiotemporal feature map;
[0011] The multi-scale temporal feature extraction module and the spatiotemporal feature interactor are combined to form an auditory attention decoding model. The auditory attention decoding model is trained based on a twin-class contrastive learning strategy to extract selection bias features using the trained auditory attention decoding model.
[0012] In a second aspect, the present invention provides an auditory attention decoding device based on EEG signals, the device comprising:
[0013] a sample generation module configured to acquire an EEG signal and apply a sliding window to the EEG signal to obtain a sample pair;
[0014] a preliminary characterization extraction module configured to construct a multi-scale temporal feature extraction module, input the sample pair into the multi-scale temporal feature extraction module, and obtain a time-varying feature map through multi-layer temporal convolution;
[0015] a spatiotemporal interaction feature extraction module configured to construct a spatiotemporal feature interactor, the spatiotemporal feature interactor including a multi-view spatial selection mechanism and an adaptive attention module, inputting the time-varying feature map into the spatiotemporal feature interactor, processing the time-varying feature map through the multi-view spatial selection mechanism and the adaptive attention module respectively to obtain a first feature map and a second feature map, and fusing the first feature map and the second feature map to obtain a spatiotemporal feature map;
[0016] The selection bias feature capture module is configured to combine the multi-scale temporal feature extraction module and the spatiotemporal feature interactor to form an auditory attention decoding model, train the auditory attention decoding model based on the twin-class contrastive learning strategy, and use the trained auditory attention decoding model to extract selection bias features.
[0017] In a third aspect, the present invention provides a readable storage medium, wherein the readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method as described above.
[0018] The present invention has at least the following beneficial effects:
[0019] 1. The present invention effectively utilizes the spatiotemporal interaction characteristics of EEG signals to generate more expressive auditory attention features.
[0020] 2. This paper effectively guides early attention selection as a data-driven mechanism and designs an efficient twin-class contrastive learning strategy, which realizes the robust extraction of selection bias features for the first time.
[0021] 3. This invention explores for the first time the intrinsic characteristic differences of the brain during attention reorientation, overcomes the influence of the brain's early selection on attention reorientation, and provides a solution for the further development of auditory attention decoding technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 FIG2 shows an architecture diagram of a twin architecture-based auditory attention learning system (SMCL-AAN) according to an embodiment of the present invention;
[0023] Figure 2 A flowchart of an auditory attention decoding method based on EEG signals according to an embodiment of the present invention is shown;
[0024] Figure 3 A schematic diagram showing the performance comparison of SMCL-AAN according to an embodiment of the present invention and other advanced technologies in a cross-subject experiment is shown;
[0025] Figure 4 Schematic diagrams of confusion matrices for two data sets according to an embodiment of the present invention are shown; (a) is the KUL data set; (b) is the DTU data set;
[0026] Figure 5 A schematic diagram showing the feature distribution of a KUL dataset according to an embodiment of the present invention is shown;
[0027] Figure 6 A schematic diagram of the characteristic distribution of a DTU dataset according to an embodiment of the present invention is shown;
[0028] Figure 7 A schematic diagram of visualizing the spatial weight of the KUL dataset according to an embodiment of the present invention is shown;
[0029] Figure 8 A schematic diagram of visualizing the spatial weights of a DTU dataset according to an embodiment of the present invention is shown;
[0030] Figure 9 The figure shows a structural diagram of an auditory attention decoding device based on EEG signals according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. The embodiments of the present invention are further described in detail below with reference to the accompanying drawings and specific embodiments, but are not intended to limit the present invention. For the various steps described herein, if there is no necessity for a contextual relationship between each other, the order in which they are described as examples herein should not be regarded as limiting, and those skilled in the art should know that they can be adjusted in order as long as the logic between them is not destroyed, resulting in the inability to implement the entire process.
[0032] Example 1: Auditory Attention Decoding Method Based on EEG Signals
[0033] In neurophysiology, auditory attention exhibits complex neural activity patterns. At the same time, in terms of data drive, effectively mining these patterns has always been a challenge faced by researchers. It is worth noting that electroencephalogram (EEG) signals are widely used due to their intuitiveness and low cost. However, previous methods often ignore the early selection process of the brain when decoding auditory attention patterns, which limits the learning ability of selection bias features in the process of attention relocation. In addition, there is still room for further exploration of the spatiotemporal interaction characteristics of EEG signals. In order to solve these problems, an embodiment of the present invention provides an auditory attention decoding method based on EEG signals, which develops an auditory attention learning system (SMCL-AAN) based on a twin architecture, such as Figure 1 Figure 2 shows the architecture of an auditory attention learning system based on a twin architecture. It embeds spatiotemporal representations of EEG signals and designs an efficient class-based learning strategy for accurate auditory attention decoding. This approach first introduces a multi-view convolutional block to capture EEG patterns at different levels, serving as a preliminary representation for subsequent research. Then, a multi-view spatial selection mechanism and an adaptive temporal attention module are carefully designed to exploit the spatiotemporal interactions of EEG signals, thereby generating more expressive auditory attention features. Furthermore, to further account for the influence of early brain selection on attention reorientation, the aforementioned feature extraction process is replicated and a twin-class contrastive learning strategy is designed to robustly extract selection bias features. The proposed SMCL-AAN learning system is validated on two public datasets, KUL and DTU, demonstrating significantly superior decoding performance. Furthermore, this approach has a relatively small number of trainable parameters, providing a lightweight, high-performance solution for auditory attention decoding.
[0034] This method can be based on Figure 1This is achieved using the auditory attention learning system shown in the figure. Its core functions include: extracting preliminary representations, exploiting spatiotemporal interaction characteristics, and capturing selection bias features. Overall, after sliding window processing of the original signal, two samples are fed into a shared weight model to extract preliminary features. These features are then fed into the multi-view spatial selection mechanism and the adaptive temporal attention module. Subsequently, these features are fused and feature similarity is calculated, and finally, the classification result is output based on the similarity.
[0035] Specifically, if Figure 2 FIG. 1 is a flow chart of an auditory attention decoding method based on EEG signals, wherein the auditory attention decoding method based on EEG signals comprises the following steps S100-S400.
[0036] S100 , acquiring an EEG signal, and applying a sliding window to the EEG signal to obtain a sample pair.
[0037] It should be noted that the EEG signal obtained in step S100 is a signal obtained by pre-processing the raw data collected by the EEG signal acquisition device. The co-spatial pattern (CSP) method is used to optimize spatial filtering to enhance signal separability, where C and T represent the number of electrode channels and time points, respectively, to obtain the EEG signal. EEG signals are typically composed of multi-channel time series, with each channel corresponding to a different electrode location on the cerebral cortex.
[0038] In some embodiments, this embodiment applies a sliding window to the processed EEG signal to obtain a series of short-duration sub-windows, called decision windows, which are expressed as E = [x1, ..., x s ……,x K ]∈R C×T'×K , where K represents the number of windows, T' represents the length of the time series under a window, and C represents the number of electrode channels. A decision window can be expressed as x i =[c1,……,c k ...c C ]∈R C×T' , where c i ∈R 1×T' is a time series from the ith channel containing T' samples.
[0039] In order to make full use of the characteristic information of the brain when paying attention to different spatial locations, SPC can be used to generate sample pairs to constrain the learning process of the model.
[0040] First, a decision window is randomly selected as the first input sample, that is, x i ∈R C×T'Then, we traverse the entire list of decision windows and randomly select another window x j Perform sample pair construction, that is The sample pairs with the same label are regarded as positive sample pairs, and the sample pairs with different labels are regarded as negative sample pairs. Finally, the constructed sample pairs are sent to the downstream task for preliminary representation extraction (i.e., step S200).
[0041] S200: Construct a multi-scale time feature extraction module, input the sample pair into the multi-scale time feature extraction module, and obtain a time-varying feature map through multi-layer time convolution.
[0042] The human brain's auditory system is highly sensitive to temporal patterns. By capturing temporal features from EEG signals, we can understand their patterns. However, previous studies have often focused solely on local temporal patterns in EEG signals, ignoring the dependencies between local and global temporal patterns. Therefore, to capture local and global temporal feature dependencies, this embodiment employs a multi-scale temporal feature extraction module, implemented through multi-layer temporal convolution.
[0043] In some embodiments, the specific implementation steps of step S200 include:
[0044] S201, use convolution kernels of different scales to compare sample x i Slide along the time dimension to capture the instantaneous changes of EEG signals at different scales. After each convolution operation, batch normalization is used to alleviate the gradient vanishing problem and speed up the model training process. In addition, in order to keep the nonlinear characteristics of the features, the activation function is used to generate nonlinear feature representation. It is defined as:
[0045]
[0046] in, and Respectively represent The weights and biases of the temporal convolution block of the layer, where M and N represent the convolution kernel size. BN(.) represents the batch normalization operation, σ(·) represents the GELU activation function, m and n represent the length and width of the convolution kernel, k represents the temporal convolution block number, x(o+m,p+n) represents the position value of the input data and the convolution kernel size, and o and p represent the starting position of the current convolution kernel.
[0047] S202, the generated feature maps containing time information of different scales are spliced according to the channel dimension, and the fine-grained features are extracted by point-by-point convolution and aggregated to construct F(X) containing time features of different scales. ω ):
[0048]
[0049] In the formula, Concat represents feature map concatenation, w (ω) and b (ω) Represents the weight and bias of the convolution, F (ω) Represents the concatenated feature tensor, F(X ω ) represents the aggregate feature map;
[0050] S203. In order to obtain more expressive features, we expand the number of feature maps to enrich its feature representation and generate the final feature tensor containing local-global time-varying features. Where C' represents the number of feature maps generated. Time-varying feature maps It can be described by the following equation:
[0051]
[0052] Among them, w (k) and b (k) represents the weight and bias term of the k-th feature map, Represents the RELU activation function.
[0053] S300. Construct a spatiotemporal feature interactor, which includes a multi-perspective spatial selection mechanism and an adaptive attention module. Input the time-varying feature map into the spatiotemporal feature interactor, and obtain a first feature map and a second feature map after being processed by the multi-perspective spatial selection mechanism and the adaptive attention module respectively. The first feature map and the second feature map are fused to obtain a spatiotemporal feature map.
[0054] EEG can record the collaborative functions between the cerebral cortex, and these functions show that human attention changes dynamically over time. Further research has shown that there is a close correlation between this temporal change pattern and the activity state of the cerebral cortex. Based on the above theoretical basis, this embodiment designs an attention module that can simultaneously capture the dynamic patterns of changes in brain regions over time and promote interaction between the two. Therefore, this embodiment designs a spatiotemporal feature interactor, which consists of a multi-perspective spatial selection mechanism and an adaptive temporal attention module.
[0055] In some embodiments, step S300 includes the following steps S301 - S304 .
[0056] S301, group the time-varying feature map by channel dimension to obtain group features Where G is the number of groups.
[0057] Specifically, this embodiment converts the feature tensor of the time-varying feature map into Group by channel dimension to obtain group features Where G is the number of groups to be grouped.
[0058] S302, perform branching operation on the grouping features to obtain in is the first input feature, is the second input feature.
[0059] In this embodiment, A branching operation is performed so that the feature maps within the group are input into the following two attention modules respectively to capture different feature patterns.
[0060] S303: Process the first input feature based on a multi-view space selection mechanism to obtain a first feature map.
[0061] In some embodiments, the first input feature is processed by the following steps to obtain a first feature map:
[0062] S3031: For the first input feature Will Channel compression is performed to suppress channels with smaller effects and reduce computational complexity. The output intermediate feature map It can be expressed as:
[0063]
[0064] Among them, w (s) and b (s) Represents the weights and biases of the convolution.
[0065] S3032, after n convolution layers with convolution kernels of different sizes, generate n feature maps with different scale information. Through the feature map splicing operation, a set of feature tensors containing local-global spatial features can be obtained. Represents the dimension of the feature tensor. It is defined as follows:
[0066]
[0067] Where, Feature maps representing information at different scales
[0068] S3033. For the feature tensor F (Ω) , the peak and average response of the spatial features of the EEG signal are extracted respectively through global maximum pooling and global average pooling, and then the two pooling results are spliced along the channel dimension and fused through a two-dimensional convolution layer to enhance the feature expression. Finally, the attention map is generated by the Sigmoid activation function, and the attention map is mapped back to the original channel dimension through a one-dimensional convolution to obtain the spatial weight map Spatial weight map The calculation details are as follows:
[0069]
[0070] Where, and represents the mapping weight and bias, f ψ Represents the feature representation after the concatenation of the peak feature and the average response feature, Indicates feature aggregation, e represents a natural constant, f ψ (o+m,p+n) represents the position value of the input data and the size of the convolution kernel. represents the aggregation bias, represents the aggregation weight;
[0071] S3034, input the original feature map After global average pooling and global maximum pooling to compress the time step dimension, and add the two features, the representative feature E of each channel is obtained. c Next, two fully connected layers are used to achieve dimensionality increase and dimensionality reduction, and an activation function is introduced between the two layers to achieve nonlinear transformation to learn the dependencies between channels. Finally, a weight map containing channel information is obtained through the Sigmoid activation function. The calculation details are as follows:
[0072]
[0073] Where, E c Represents the maximum and average aggregate features of time features, w (c2) Represents the first layer channel weight matrix, w (c1) Represents the second layer channel weight matrix;
[0074] S3035, spatial weight map With channel weight map Do matrix broadcast multiplication and apply it to the original feature map Get a new feature map that integrates spatial information and channel information The process is as follows:
[0075]
[0076] Therefore, the processing process of steps S3031-S3035 above is adopted to capture the dependency of local-global spatial features through multiple perspectives and utilize the spatial attention mechanism to emphasize task-related features. At the same time, a channel attention is used to dynamically learn the importance weight of each channel. By combining the two, irrelevant noise interference can be effectively suppressed, thereby generating more expressive features.
[0077] S304: Process the second input feature based on the adaptive attention module to obtain a second feature map.
[0078] In some embodiments, the second input feature is processed by the following steps to obtain a second feature map:
[0079] S3041. For the second input feature Will After transposition, a linear transformation is performed to generate the query Q and key value K:
[0080]
[0081] Among them, W (q) , W (k) ∈R T'×d is the projection weight matrix, b (q) ,b (k) represents the bias of the linear transformation, and d represents the output dimension.
[0082] S3042. In order to better capture the global relationship of time features, we set a learnable weight W (β) =[w1,……,w i ,……,w d ] T . Perform matrix multiplication of query Q and learnable weights to obtain the global time vector β∈R T'×1 ,Right now:
[0083] β=Q*W (β)
[0084] S3043, perform broadcast multiplication on the obtained global time vector β and the query Q, and sum them according to the channel dimension to obtain a query vector Q containing multi-channel global time information (β) ∈R T'×1 :
[0085]
[0086] S3044, in order to achieve the interaction between global time features, use the query vector Q (β) Perform broadcast multiplication with the key value K. Then use a linear transformation to enhance the feature expression ability. Finally, use a residual connection query vector Q to prevent the feature from losing its original time characteristics. Through this series of operations, a feature map containing global time information is obtained. The process is as follows:
[0087]
[0088] S3045. Use a linear transformation to map back to the original channel dimension to obtain the final feature map containing global important time information Right now:
[0089]
[0090] Through the above steps S3041-S3045, the focus on important time nodes is enhanced and the interference of noise or irrelevant information is reduced, so that the model can more accurately understand and analyze complex patterns and long-term dependencies in time features.
[0091] In some embodiments, the first feature map and the second feature map are fused to obtain a spatiotemporal feature map in the following manner:
[0092] After the above two attention processes, G groups of feature maps are obtained. Each group consists of a feature map containing spatial relationships and a feature map containing global temporal dependencies, which are represented here as In order to fully integrate the spatial and temporal features between different groups, channel-by-channel convolution is used to extract features from each feature map, and point-by-point convolution is used to fuse its spatiotemporal features. In addition, in order to retain representative features, average pooling is used to obtain the spatiotemporal feature map. It can be expressed as follows:
[0093]
[0094] Where, Represents the spatiotemporal feature map, Avgpool represents average pooling, Represents the RELU activation function, BN represents batch normalization, m and n represent the length and width of the convolution kernel, represents the G group feature map, and represents the feature maps of group 1 and group G, represents the first feature map, Represents the second feature map, o and p represent the starting position of the convolution kernel, Indicates that the convolution kernel is at position The weight value, w (g) represents the weight matrix.
[0095] So far, this embodiment has obtained a feature map containing global spatiotemporal features. Subsequently, an adaptive average pooling can be used to retain important spatiotemporal features and provide a priori representation for the learning model in the subsequent step S400.
[0096] S400. Combining the multi-scale temporal feature extraction module and the spatiotemporal feature interactor to form an auditory attention decoding model. Based on a twin-class contrastive learning strategy, the auditory attention decoding model is trained to extract selection bias features using the trained auditory attention decoding model.
[0097] In some embodiments, step S400 includes the following steps S401 and S402.
[0098] S401, Forward Propagation: To consider the impact of early brain selection on attention reorientation, we further designed a twin-class contrastive learning strategy based on the above module, aiming to achieve robust extraction of selection bias features. The twin architecture consists of two sub-networks with the same structure and shared weights. These two sub-networks extract features from two inputs respectively. To this end, we represent the feature vector of the sample pair obtained after the above feature extraction and feature interaction of the input sample pair as In order to measure the similarity between two samples, we first calculate the Euclidean distance between them, which is expressed as
[0099]
[0100] Get the Euclidean distance D between two sample vectors Euclidean After that, the similarity between sample vectors is calculated, which is expressed as
[0101]
[0102] in Indicates whether the sample pair is of the same type, and m represents a boundary value used to distinguish the distance between sample pairs so that they can be better applied to downstream tasks. In addition, in order to complete the target task, a classifier is required to classify each sample.
[0103] Subsequently, the cross entropy loss function is used to calculate its classification loss, which is expressed as
[0104]
[0105] in and Indicates the true label corresponding to the sample pair, and Indicates the predicted label corresponding to the sample pair. In this embodiment, the objective optimization function of SMCL-AAN is as follows:
[0106]
[0107] Here, α and β are learnable factors that balance the two losses, ranging from 0 to 1, with α + β = 1. γ represents the regularization parameter, Ψ(·) is a trainable model parameter, and Ω1 and Ω2 are trainable parameters for each sub-network in the twin architecture. This joint loss optimization enables the model to better capture spatial patterns and temporal dependencies in the signal, achieving robust extraction of selection bias features.
[0108] S402, Back Propagation: Use the backward gradient propagation optimization method. For the trainable parameters Ω1 and Ω2 of the two sub-networks, the update process can be expressed as:
[0109]
[0110] Where, represents the symbol of partial derivative;
[0111] Since the trainable parameters of the two sub-networks share weights, they can also update each other, which can be expressed as follows:
[0112]
[0113] Where λ represents the learning rate.
[0114] Example 2: Experimental Design
[0115] In this embodiment, the proposed SMCL-AAN learning system is verified on the following two publicly available datasets, as shown in Table 1. For detailed information on the KUL and DTU datasets, please refer to the literature "Das, Neetha, et al." The effect of head-related filtering and ear-specific decoding bias on auditory attention detection." Journal of neural engineering 13.5 (2016): 056014.", "Das, Neetha, Tom Francart, and Alexander Bertrand." Auditory attention detection dataset KULeuven." Zenodo (2020).", "Fuglsang, Asp, Torsten Dau, and Jens "Noise-robust cortical tracking of attended speech in real-worldacoustic scenes."NeuroImage 156(2017):435-444." and "Fuglsang, A.,DDWong,and Jens "EEG and audio dataset for auditory attention decoding."Zenodo(2018).》obtained.
[0116] Table 1 Dataset details
[0117]
[0118] The KUL dataset contains 64-channel EEG data from 16 subjects, consisting of 8 male and 8 female subjects. The auditory stimuli consisted of four Dutch stories told by different male speakers, and the subjects were instructed to focus on one of the two speakers. The two speakers were positioned 90 degrees to the left and right of the speaker as they told the stories. The EEG data was recorded using a BioSemi ActiveTwo device at a sampling rate of 8164 Hz. For each subject, eight trials were collected, each lasting 6 minutes, resulting in a total of 48 minutes of EEG data.
[0119] The DTU dataset contains 64-channel EEG data from 18 subjects. Auditory stimuli were presented to male and female speakers in Danish. During the experiment, subjects were required to focus on one of the two speakers and ignore the other. The speakers' voices were positioned at +60° and -60° from the subject's angle, respectively. The EEG data was recorded by BioSemiActive at a sampling rate of 512 Hz. Each subject completed 60 trials, each lasting 50 seconds, resulting in a total of 50 minutes of EEG data.
[0120] To ensure the fairness of the experiment, this embodiment performed specific preprocessing steps on each dataset. For the KUL dataset, this embodiment performed a bandpass filter from 0.1Hz to 50Hz and downsampled it to 128Hz. For the DTU dataset, this embodiment first filtered out 50Hz linear noise, then downsampled it to 64Hz, and finally used a joint decorrelation framework to remove eye artifacts. In this embodiment, the performance of the model under different decision windows, including windows of 0.1 seconds, 1 second, and 2 seconds. For a single subject of KUL, this embodiment obtained a total of 3104 windows under a decision of 1 second, and a total of 49664 windows for 16 subjects. For DTU, this embodiment obtained a total of 3000 windows under a decision of 1 second for a single subject, and a total of 54000 windows for 18 subjects.
[0121] The primary task of auditory attention detection is to discern the direction of a sound source, specifically distinguishing left from right. We evaluated the proposed model on the KUL and DTU datasets. Using the KUL dataset with a 1-second decision window as an example, we describe the implementation details, including training settings and network configuration.
[0122] Specifically, in this embodiment, after sliding window processing, the batch size is 64, the training rounds are 100, Adam is used as the optimizer, the learning rate is set to 0.001, and the two hyperparameters α and β are 0.6 and 0.4 respectively.
[0123] Before training, this embodiment uses CSP to improve the signal-to-noise ratio of the original signal so that the subsequent model can better learn useful feature representations and obtain feature representations. Obtain feature tensors of temporal features containing different information from different perspectives Then, in order to capture the temporal and spatial interaction patterns in the features, the feature tensors are fed into different attention modules at the group level. For multi-view spatial attention, three convolutional layers with different convolution kernel sizes are used to capture spatial information of different shapes, and spatial attention is used in conjunction with channel attention to enhance the feature representation of spatial and channel information, resulting in For adaptive temporal attention, the feature map is transformed and then subjected to a linear transformation to obtain the query Q∈R 128×d and key value K∈R 128×d , d is the output dimension. Using a learnable vector W (β) ∈R d ×1 Multiplying with the query Q yields the global time vector β∈R 128×1 Then, the global time vector β is broadcast multiplied with the query Q and summed according to the channel dimension to obtain the vector Q containing multi-channel global time information (β) ∈R 128×1Finally, a residual link is used to prevent the loss of the original features, and a linear transformation is used to map back to the original feature dimension to obtain the final feature map containing global important time information. After two attention modules, the important spatiotemporal features are fused and retained through depth-wise separable convolution and adaptive average pooling, and then the contrast loss is calculated.
[0124] The specific parameters of the model are shown in Table 2, and it is run under the PyTorch framework.
[0125] Table 2 Model specific parameters
[0126]
[0127]
[0128] Example 3: Single-subject experiment
[0129] Single-subject validation was performed on each subject in the two datasets. In this embodiment, each experimental data of each subject was randomly divided into a training set, a validation set, and a test set in a ratio of 8:1:1, and then training, validation, and testing were performed on a single subject. In order to more accurately evaluate the performance of the model of the present invention, this embodiment chose to compare it with other advanced AAD models on the KUL and DTU datasets, including STAnet, DGSD
[14] , AGSLNet, DBPNet, and EEG-Graph Net. According to the results shown in Table 3, our model showed superior performance under 0.1 second, 0.5 second, and 2 second decision windows.
[0130] Table 3 Performance comparison of the proposed SMCL-AAN and other advanced techniques in single-subject experiments
[0131]
[0132]
[0133] Specifically, on the KUL dataset, our model performed exceptionally well with decision windows of 0.1 seconds (average: 93.1%, standard deviation: 6.1%), 0.5 seconds (average: 96.8%, standard deviation: 4.4%), and 2 seconds (average: 97.4%, standard deviation: 3.7%). Notably, for STAnet and AGSLNet, two methods that also capture spatiotemporal features, our method outperformed STAnet by 12.8%, 6.7%, and 6% in accuracy at 0.1 seconds, 1 second, and 2 seconds on the KUL dataset, respectively, and outperformed the state-of-the-art AGSLNet by 5%, 3.2%, and 4.1%. On the DTU dataset, the two methods also outperformed by 16.3%, 16.7%, and 17.2%, and by 5.9%, 7.8%, and 9.6%, respectively. This is because our method considers spatial features from multiple perspectives of the EEG signal, thereby capturing more discriminative features. Furthermore, our method outperformed other state-of-the-art methods. Because other methods fail to fully account for feature biases in the brain's attention selection process, they fail to capture this potentially differentiating feature. The twin architecture proposed in this paper can better capture this biased feature in the brain's attention selection process, thereby learning a richer feature representation. The inventors further found that as the EEG signal window length increased, the AAD detection accuracy of most subjects significantly improved. Although there were a few exceptions, where detection accuracy decreased slightly for individual subjects as the window length increased, this phenomenon did not change the overall trend. Therefore, it can be concluded that longer signal windows help improve the accuracy of AAD detection. Table 3 further observes that in the DTU dataset, the average accuracy at 0.1s, 1s, and 2s was 82.0% (SD: 4.9%), 88.6% (SD: 4.1%), and 90.9% (SD: 4.8%), respectively. This trend in detection accuracy is consistent with previous findings in the KUL dataset.
[0134] In addition, Tables 4 and 5 also show the changing trends of the detection accuracy of a single subject in the two datasets under three different time windows. It is worth noting that the average detection accuracy on the KUL dataset is significantly higher than that on the DTU dataset. We believe that the reasons for this difference may be the following factors: (1) Gender differences of speakers. Since the KUL dataset only contains speech stimuli from male speakers, the DTU dataset contains female pronunciations in addition to male pronunciations. (2) Different directions of sound stimulation. The speech stimuli of the KUL dataset come from the 90° directions on the left and right, while the speech stimuli of the DTU dataset come from the 60° directions on the left and right. (3) Different environmental reverberations. The different environmental reverberations collected by the two datasets may reduce the accuracy of cortical speech tracking.
[0135] Table 4 Changes in detection accuracy of a single subject in different time windows of the KUL dataset
[0136]
[0137] Table 5. Changes in detection accuracy of the DTU dataset
[0138]
[0139]
[0140] Example 4: Cross-subject experiment
[0141] In practical applications, cross-subject experiments are of great reference significance for testing the generalization ability of the method. In order to evaluate the generalization performance of the proposed SMCL-ANN between different subjects, this embodiment adopts a method called "leave-one-subject-out". Specifically, the data of one subject is selected from the data set as the test set, and the data of all other subjects are used to train the model. In this process, the data of each subject will be used as a test set once in turn, so as to obtain multiple detection accuracy results, and finally calculate the average detection accuracy. This method can not only comprehensively evaluate the generalization ability of the model between different individuals, but also effectively prevent overfitting problems and ensure the robustness and reliability of the model.
[0142] Tables 6 and 7 show the cross-subject performance of individual subjects in both datasets. We found that some subjects had lower detection accuracy. This is due to individual differences in EEG activity between subjects, which results in poor generalization of feature patterns obtained from other subjects to the individual subject. When analyzing the DTU and KUL datasets, we observed some significant differences. In the DTU dataset, overall performance across all participants was relatively low. The best-performing participant, S08, achieved a detection accuracy of 66.77%, while the lowest-performing participant, S14, achieved only 50.40%, a gap of 16.37 percentage points. In contrast, in the KUL dataset, the range of detection accuracy was much wider: Subject S16 achieved the highest accuracy of 98.11%, while the lowest-performing participant, S03, achieved only 52.12%, a difference of approximately 46 percentage points.
[0143] Based on these findings, we can speculate that the following possible reasons may have contributed to the above differences: First, there may be large differences in EEG features between different participants in the KUL dataset, which may have affected the model's performance across subjects, while in the DTU dataset, such differences between individual subjects may be relatively small; Second, the model of the present invention may have failed to effectively extract features from the data of some participants that can be well generalized to other individuals, thereby limiting its performance on specific participants.
[0144] Performance of cross-subject experiments compared with other methods Figure 2 As shown, the results are compared with AGSLNet, GCN, MBSSFCC, and DGSD methods. It can be observed that compared with the within-subject experiment, the cross-subject performance is more challenging and the detection accuracy is lower, but the detection accuracy is still better than the results of other models. In particular, the performance on the KUL dataset is that the average accuracy of the proposed SMCL-AAN reaches 73.6%, which is 9.7% higher than that of DGSD, and 14% higher than that of the GCN model with the highest detection accuracy. The results on DTU are also higher than those of other methods. These results show that the model proposed in this invention has better generalization and robustness.
[0145] Table 6 Auditory decoding effects of the proposed SMCL-AAN on different subjects in the KUL dataset
[0146]
[0147]
[0148] Table 7 Auditory decoding effects of the proposed SMCL-AAN on different subjects in the DTU dataset
[0149]
[0150] like Figure 3 Figure 2 shows a performance comparison diagram of the proposed SMCL-AAN and other advanced technologies in a cross-subject experiment. The data are all from published papers (the DTU detection data of AGSLNet comes from the reproduced model).
[0151] Example 5: Ablation Experiment
[0152] To evaluate the contribution of each model component, this example conducted an ablation study. This experiment primarily eliminated specific modules to create network variants and evaluated their performance on the KUL and DTU datasets. As shown in Table 8, this example performed the analysis by removing two key components: the MSAM and LATM modules.
[0153] Table 8 Ablation study results
[0154]
[0155] When both modules were removed simultaneously, the model's performance dropped significantly across three different time windows (0.1, 1, and 2 seconds). Taking the KUL dataset as an example, when using only the base model (MSC), the average accuracy for the 0.1, 1, and 2 second windows was 83.2% (SD: 9.27%), 88.9% (SD: 8.53%), and 90.2% (SD: 6.64%), respectively. When using only the MSAM module, the average accuracy for these three time windows was 89.6% (SD: 7.05%), 94.6% (SD: 5.17%), and 95.3% (SD: 4.02%), respectively, representing improvements of 6.4%, 5.7%, and 5.1%, respectively, compared to the MSC. When using only the LATM module, the average accuracy for the three time windows improved by 5%, 4.9%, and 4.6%, respectively.
[0156] These results demonstrate that both the proposed MSAM and LATM modules significantly improve model performance. By integrating these two modules, the model achieved accuracies of 93.1% (SD: 6.17%), 96.8% (SD: 4.43%), and 97.4% (SD: 3.78%) across three time windows, respectively. These performance improvements are superior to using either module alone or without either module.
[0157] Similar results were obtained from ablation experiments on the DTU dataset. When both components were used simultaneously, the model performance was consistently better than when only one component was used, or when no component was used at all. This finding demonstrates that the proposed model can effectively extract discriminative features from EEG signals for AAD-related tasks.
[0158] Example 6: Comparison of model operation efficiency
[0159] This example compares the trainable parameter count of the model with that of SSF-CNN, DBPNet, DGSD, and AGSLNet, and the results are summarized in Table 9. Specifically, the proposed SMCL-AAN has only 0.06M trainable parameters, which is a 98.6-fold reduction compared to AGSLNet, a 15.1-fold reduction compared to DBPNet, and a 2.5-fold reduction compared to DGSD.
[0160] Table 9 Comparison of the running efficiency of the proposed SMCL-AAN and other methods
[0161]
[0162] In addition, this embodiment also compared the test time and found that under the same configuration environment, the time required for SSF-CNN to train a single subject for one epoch was 3.42 seconds, while the model of the present invention only needed 1.44 seconds. This means that compared with SSF-CNN, the model of the present invention can reduce the time consumption by approximately 57.9% on the same task. Although DGSD has advanced detection accuracy and a low number of parameters, the model of the present invention also performs well in training time. This further proves that even in a resource-constrained environment, the model of the present invention can not only maintain excellent detection performance, but also achieve high detection accuracy in a shorter training time.
[0163] Example 7: Visualization and Spatial Attention Analysis
[0164] Based on Examples 1 to 6, it can be seen that the SMCL-AAN proposed in this invention can effectively capture the spatiotemporal features of EEG signals. In particular, the design of the class contrast learning strategy can optimize the spatial embedding of features to better adapt to downstream classification tasks. In this example, the classification accuracy of SMCL-AAN on each category is presented, as shown in Figure 2. Figure 4 In addition, this example also studies the feature distribution of different subjects learned by SMCL-AAN, as shown in Figure 5 and Figure 6 In order to evaluate whether the model of the present invention is inspired by biological neurons, the weight of spatial attention is visualized on the brain topology map to better explain the relationship between auditory attention and brain spatial patterns.
[0165] Visualization:
[0166] In this embodiment, the confusion matrix is used to visualize the classification performance of the model for different auditory attention directions on the KUL dataset and the DTU dataset, as shown in Figure 4 As shown in Figure 3, the proposed SMCL-AAN is more easily able to identify right-ear auditory attention, achieving 97% and 89% accuracy on the KUL and DTU datasets, respectively, both 1% higher than the left-ear. Furthermore, SMCL-AAN demonstrates higher recognition accuracy for both left and right-ear auditory attention on the KUL dataset, while its classification performance on the DTU dataset is relatively inferior. This conclusion is consistent with the results of Example 3. This further demonstrates that differences in the subject, external environment, and equipment significantly impact EEG signals.
[0167] In order to more intuitively demonstrate the effectiveness of the proposed SMCL-AAN system, this example visualizes the feature distribution after learning different subjects. Feature dimension reduction is achieved by t-distributed stochastic neighbor embedding (t-SNE) technology. Figure 5 and Figure 6In the example, the left and right auditory attention are represented by red and blue respectively. This example selects 6 representative test subject feature distributions from two data sets to show the distribution of the test subject features. Figure 5 and Figure 6 It can be clearly seen that the proposed method can successfully distinguish the two target tasks, especially achieving impressive results in distinguishing the left and right ear features of subject 16 in the KUL dataset. Notably, SMCL-AAN achieves significantly better feature discrimination on the KUL dataset than the DTU dataset, which corresponds to the previous conclusion that SMCL-AAN demonstrates higher recognition accuracy for left and right ear auditory attention in KUL.
[0168] Spatial Attention Analysis:
[0169] In terms of auditory attention, the channel signals recorded by different electrodes on the pars plana cortex have different degrees of contribution. The method we proposed can be used to explain the relationship between the spatial weight differences in EEG signals and the improvement of performance. We visualized the electrode channel weights in two datasets and selected 8 representative subjects for each to illustrate, as shown in the figure below. Figure 7 and Figure 8 As shown. It can be found that the electrodes placed in the frontal lobe and bilateral temporal lobe areas have higher weights than other areas, which is consistent with previous studies. This is because the temporal lobe area is responsible for processing auditory information collected from the outside world, and then transmitting it to the frontal lobe area for cognitive decision-making processing and analysis. We also found that the cortex in the right posterior parietal-occipital area was more active, because the right posterior parietal lobe mediates the spatial relocation of attention. This shows that considering the early selection process of the brain is necessary for the model to capture more discriminative auditory attention features. And the parietal lobe is responsible for the brain's processing of perceptual functions, and when the brain receives auditory information from the outside world, it is inspired to respond perceptually. In addition, the EEG signals between different subjects are specific. In addition to the above-mentioned commonalities, the spatial weights also have corresponding characteristics. For example Figure 8 The frontal regions S2 and S14 in the MRI scans did not show significant activation. Furthermore, some subjects showed significant hemispheric lateralization, along with hemispheric symmetry. Therefore, the SMCL-AAN learning system proposed in this paper can dynamically learn individual signal differences and assign different weights based on these specific differences, potentially promoting the development of AAD.
[0170] In summary, the present invention develops a SMCL-AAN learning system that achieves accurate decoding of auditory attention by embedding the spatiotemporal representation of EEG signals and designing an efficient class learning strategy. In addition, SMCL-AAN takes into account the impact of the brain's early selection on attention relocation, making the captured features more robust. The results on the KUL and DTU datasets show that the proposed SMCL-AAN learning system not only outperforms existing advanced models in decoding accuracy, but also exhibits an impressive low parameter count, demonstrating its high practical application value. At the same time, the ablation study highlights the contribution of each designed module. In addition, the visualization of the learned feature distribution and spatial weight distribution further demonstrates the rationality and generalization of the learning system. The present invention provides a solution for achieving lightweight, high-performance auditory attention decoding.
[0171] Example 8: Auditory Attention Decoding Device Based on EEG Signals
[0172] The embodiment of the present invention also provides an auditory attention decoding device based on EEG signals, such as Figure 9 As shown, the device includes:
[0173] The sample generation module 901 is configured to acquire an EEG signal and apply a sliding window to the EEG signal to obtain a sample pair;
[0174] The preliminary characterization extraction module 902 is configured to construct a multi-scale temporal feature extraction module, input the sample pair into the multi-scale temporal feature extraction module, and obtain a time-varying feature map through multi-layer temporal convolution;
[0175] The spatiotemporal interaction feature extraction module 903 is configured to construct a spatiotemporal feature interactor, which includes a multi-view spatial selection mechanism and an adaptive attention module. The time-varying feature map is input into the spatiotemporal feature interactor, and processed by the multi-view spatial selection mechanism and the adaptive attention module to obtain a first feature map and a second feature map. The first feature map and the second feature map are fused to obtain a spatiotemporal feature map.
[0176] The selection bias feature capture module 904 is configured to combine the multi-scale temporal feature extraction module and the spatiotemporal feature interactor to form an auditory attention decoding model, train the auditory attention decoding model based on the twin-class contrastive learning strategy, and use the trained auditory attention decoding model to extract selection bias features.
[0177] It should be noted that the structures of the various EEG signal-based auditory attention decoding devices described in this embodiment belong to the same technical concept as the previously described EEG signal-based auditory attention decoding method, and achieve the same beneficial effects through the same principles, which will not be repeated here.
[0178] An embodiment of the present invention further provides a readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the method described in any of the above embodiments.
[0179] The above description is intended to be illustrative rather than restrictive. For example, the above examples (or one or more of their solutions) can be used in combination with each other. For example, those of ordinary skill in the art may use other embodiments when reading the above description. In addition, in the above-mentioned specific embodiments, various features can be grouped together to simplify the present invention. This should not be interpreted as an intention that a feature of an invention that is not claimed for protection is necessary for any claim. On the contrary, the subject matter of the present invention may be less than all the features of the embodiments of a particular invention. Thus, the following claims are incorporated into the specific embodiments as examples or embodiments, wherein each claim is independently a separate embodiment, and it is considered that these embodiments can be combined with each other in various combinations or arrangements. The scope of the present invention should be determined with reference to the appended claims and the full scope of equivalents to which these claims are entitled.
Claims
1. A method for auditory attention decoding based on EEG signals, characterized in that: The method comprises: Acquire an EEG signal, and apply a sliding window to the EEG signal to obtain a sample pair; Constructing a multi-scale temporal feature extraction module, inputting the sample pairs into the multi-scale temporal feature extraction module, and obtaining a time-varying feature map through multi-layer temporal convolution; Constructing a spatiotemporal feature interactor, the spatiotemporal feature interactor including a multi-view spatial selection mechanism and an adaptive attention module, inputting the time-varying feature map into the spatiotemporal feature interactor, processing the time-varying feature map through the multi-view spatial selection mechanism and the adaptive attention module respectively to obtain a first feature map and a second feature map, and fusing the first feature map and the second feature map to obtain a spatiotemporal feature map; The multi-scale temporal feature extraction module and the spatiotemporal feature interactor are combined to form an auditory attention decoding model. The auditory attention decoding model is trained based on a twin-class contrastive learning strategy to extract selection bias features using the trained auditory attention decoding model.
2. The auditory attention decoding method based on EEG signals according to claim 1, characterized in that An EEG signal is acquired and a sliding window is applied to the EEG signal to obtain sample pairs, including: Applying a sliding window to the EEG signal, a series of short-duration sub-windows are obtained as decision windows, which are expressed as E = [x1, ..., x s ……,x K ]∈R C×T'×K , where K represents the number of windows, T' represents the length of the time series under a window, C represents the number of electrode channels, R represents the real space, x1, x s and x K represent the 1st, sth and Kth decision windows respectively; Randomly select a decision window as the first input sample, denoted as x i ∈R C×T' ; Traverse the entire decision window list and randomly select another decision window to construct a sample pair to obtain the sample pair where x j For another input sample; The sample pairs with the same label are regarded as positive sample pairs, and the sample pairs with different labels are regarded as negative sample pairs.
3. The auditory attention decoding method based on EEG signals according to claim 2, characterized in that The sample pairs are input into the multi-scale temporal feature extraction module, and a time-varying feature map is obtained through multi-layer temporal convolution, including: Use convolution kernels of different scales to compare sample x i Slide along the time dimension to capture the instantaneous changes of EEG signals at different scales. After each convolution operation, batch normalization is used to alleviate the gradient vanishing problem and speed up the model training process. In order to keep the nonlinear characteristics of the features, the activation function is used for nonlinearization to generate feature maps containing time information at different scales. The calculation process is expressed as: Where, and Represent the weight and bias of the s-th layer of temporal convolution block, M and N represent the convolution kernel size, BN(·) represents the batch normalization operation, σ(·) represents the GELU activation function, m and n represent the length and width of the convolution kernel, k represents the temporal convolution block number, x(o+m,p+n) represents the position value of the input data and the convolution kernel size, o and p represent the starting position of the current convolution kernel; The generated feature maps containing time information of different scales are spliced according to the channel dimension, and their fine-grained features are extracted by point-by-point convolution and aggregated to construct an aggregated feature map containing time features of different scales. The calculation process is expressed as: In the formula, Concat represents feature map concatenation, w (ω) and b (ω) Represents the weight and bias of the convolution, F (ω) Represents the concatenated feature tensor, F(X ω ) represents the aggregate feature map; Expand the number of feature maps for the aggregated feature map to generate a time-varying feature map containing local-global time-varying features The calculation process is expressed as: Where w (k) and b (k) represents the weight and bias term of the k-th feature map, represents the RELU activation function, and C' represents the number of feature maps generated.
4. The auditory attention decoding method based on EEG signals according to claim 3, characterized in that The time-varying feature map is input into the spatiotemporal feature interactor, and processed by the multi-view spatial selection mechanism and the adaptive attention module respectively to obtain a first feature map and a second feature map, including: Group the time-varying feature maps by channel dimension to obtain grouped features Where G is the number of groups; Perform branching operations on the grouping features to obtain in is the first input feature, is the second input feature; Processing the first input feature based on a multi-view spatial selection mechanism to obtain a first feature map; Based on the adaptive attention module, the second input feature is processed to obtain a second feature map.
5. The auditory attention decoding method based on EEG signals according to claim 4, characterized in that Based on the multi-view space selection mechanism, the first input feature is processed to obtain a first feature map, including: A channel compression operation is performed on the first input feature to obtain an intermediate feature map, which is represented as: Where, represents the intermediate feature map, Represents the dimension of the feature tensor, w (s) and b (s) Represents the weights and biases of the convolution; After passing through n convolution layers with convolution kernels of different sizes, n feature maps with different scale information are generated. A set of feature tensors containing local-global spatial features are obtained through feature map splicing operations. Expressed as: Where, One of the feature maps representing information at different scales; For the feature tensor F (Ω) The peak and average response of the spatial features of the EEG signal are extracted by global maximum pooling and global average pooling respectively. Then the two pooling results are spliced along the channel dimension and fused through a two-dimensional convolution layer to enhance the feature expression. The attention map is generated by the Sigmoid activation function, and the attention map is mapped back to the original channel dimension through one-dimensional convolution to obtain the spatial weight map. Spatial weight map The calculation process is expressed as: Where, and represents the mapping weight and bias, f ψ Represents the feature representation after the concatenation of the peak feature and the average response feature, Indicates feature aggregation, e represents a natural constant, f ψ (o+m,p+n) represents the position value of the input data and the size of the convolution kernel. represents the aggregation bias, represents the aggregation weight; At the same time, the first input feature map After global average pooling and global maximum pooling to compress the time step dimension, and add the two features, the representative feature E of each channel is obtained. c , through two fully connected layers to achieve dimensionality increase and dimensionality reduction, and introduce activation function between the two layers to achieve nonlinear transformation, in order to learn the dependency between channels, and finally obtain a channel weight map containing channel information through Sigmoid activation function The calculation process is expressed as: Where, E c Represents the maximum and average aggregate features of time features, w (c2) Represents the first layer channel weight matrix, w (c1) Represents the second layer channel weight matrix; The spatial weight map With channel weight map Do matrix broadcast multiplication and apply it to the first input feature map Get the first feature map of comprehensive spatial information and channel information 6. The auditory attention decoding method based on EEG signals according to claim 4, characterized in that The second input feature is processed based on the adaptive attention module to obtain a second feature map, including: The second input feature After transposition, a linear transformation is performed to generate query Q and key value K; Set a learnable weight and perform matrix multiplication on the query Q and the learnable weight to obtain the global time vector β∈R T'×1 ; Perform broadcast multiplication on the obtained global time vector β and the query Q, and sum them according to the channel dimension to obtain a query vector Q containing multi-channel global time information (β) ∈R T'×1 ; Using query vector Q (β) Perform broadcast multiplication with the key value K and use a linear transformation to enhance the feature expression capability. Finally, use a residual connection query vector Q to obtain a feature map containing global time information Use a linear transformation to transform the feature map Mapping back to the original channel dimension, we get the final second feature map containing global important time information 7. The auditory attention decoding method based on EEG signals according to claim 4, characterized in that The spatiotemporal feature map is obtained by fusing the first feature map and the second feature map using the following formula: Where, Represents the spatiotemporal feature map, Avgpool represents average pooling, Represents the RELU activation function, BN represents batch normalization, m and n represent the length and width of the convolution kernel, represents the G group feature map, and represents the feature maps of group 1 and group G, represents the first feature map, Represents the second feature map, o and p represent the starting position of the convolution kernel, Indicates that the convolution kernel is at position The weight value, w (g) represents the weight matrix.
8. The auditory attention decoding method based on EEG signals according to claim 1, characterized in that The auditory attention decoding model is trained based on a twin-class contrastive learning strategy, including: Construct a twin architecture, which includes two sub-networks with the same structure and shared weights. The two sub-networks extract features from the two samples of the sample pair respectively, and the feature vector of the sample pair is expressed as The Euclidean distance of the feature vectors of the sample pairs is calculated using the following formula: Where, represents the Euclidean distance of the feature vectors of the sample pair, Represents sample x respectively i and sample x j The spatiotemporal feature map of Calculate the similarity between sample vectors, expressed as In the formula Indicates whether the sample pair is of the same type, and m represents a boundary value used to distinguish the distance between sample pairs; A classifier is used to classify each sample; then, the classification loss is calculated using the cross entropy loss function, which is expressed as In the formula and Indicates the true label corresponding to the sample pair, and Represents the predicted label corresponding to the sample pair; Based on the similarity between sample vectors and the cross entropy loss function, the target optimization function is determined as follows: Where α and β are learnable factors that balance the two losses, ranging between 0 and 1, and α + β = 1, γ represents the regularization parameter, Ψ(·) is the model trainable parameter, and Ω1, Ω2 are the trainable parameters of each sub-network in the twin architecture; For the trainable parameters Ω1 and Ω2 of the two sub-networks, the update process is expressed as: Where, represents the symbol of partial derivative; The trainable parameters of the two sub-networks are updated with each other through the following formula: Where λ represents the learning rate.
9. An auditory attention decoding device based on EEG signals, characterized in that: The device comprises: a sample generation module configured to acquire an EEG signal and apply a sliding window to the EEG signal to obtain a sample pair; a preliminary characterization extraction module configured to construct a multi-scale temporal feature extraction module, input the sample pair into the multi-scale temporal feature extraction module, and obtain a time-varying feature map through multi-layer temporal convolution; a spatiotemporal interaction feature extraction module configured to construct a spatiotemporal feature interactor, the spatiotemporal feature interactor including a multi-view spatial selection mechanism and an adaptive attention module, inputting the time-varying feature map into the spatiotemporal feature interactor, processing the time-varying feature map through the multi-view spatial selection mechanism and the adaptive attention module respectively to obtain a first feature map and a second feature map, and fusing the first feature map and the second feature map to obtain a spatiotemporal feature map; The selection bias feature capture module is configured to combine the multi-scale temporal feature extraction module and the spatiotemporal feature interactor to form an auditory attention decoding model, train the auditory attention decoding model based on the twin-class contrastive learning strategy, and use the trained auditory attention decoding model to extract selection bias features. 10 . A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, perform the method according to claim 1 .