Auditory attention decoding method and system based on multi-scale frequency-space attention network
By introducing multi-scale frequency-space attention networks into auditory attention decoding technology, the problem of insufficient decoding accuracy in the existing technology in complex acoustic environments and short decision windows is solved, efficient feature extraction and information capture are achieved, and decoding performance is significantly improved.
Patent Information
- Application Number
- CN202510696114.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The existing auditory attention decoding technology has bottlenecks in feature extraction, model training and information capture, and it is difficult to achieve high-accuracy decoding in complex acoustic environments and short decision windows.
A auditory attention decoding method based on multi-scale frequency-space attention network (MSSANet) is proposed. Through the multi-scale time domain convolution module, frequency-space attention module and full-connection layer classification module, the time domain and frequency-space characteristics of the EEG signal are extracted and integrated to realize adaptive frequency domain information acquisition and global dependency capture.
It significantly improves the accuracy and computing efficiency of auditory attention decoding, especially the decoding performance under a short time window, enhances the real-time and adaptability of the system, and provides more accurate and efficient technical support for neural-driven auditory assistive devices.
Smart Images

Figure CN120216936A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of auditory brain-computer interfaces, and particularly relates to an auditory attention decoding method and system based on a multi-scale spatio-temporal attention network, which can be applied to the development of neurally-driven auditory assistive devices (cochlear implants or hearing aids), etc. Background Art
[0002] In a complex acoustic environment, humans can exhibit the "cocktail party effect", that is, focus on the target speech of interest in an environment with multiple mixed sounds while ignoring the interference of other sounds. However, for people with hearing loss, the masking effect of background noise on the target speech is significantly enhanced, resulting in impaired auditory selective attention ability and difficulty in effectively focusing on the target speech. Neuroscientific research shows that compared with non-target speech, the neural activity in the cerebral cortex shows a stronger similarity to the amplitude envelope of the target speech. Therefore, decoding the direction of auditory attention (i.e., auditory attention decoding) based on electroencephalogram (EEG) signals can provide key technical support for neurally-driven auditory assistive devices.
[0003] At present, existing research has confirmed the feasibility of decoding auditory attention from EEG, and EEG-based auditory attention decoding algorithms are mainly divided into two types: stimulus reconstruction and direct classification. The stimulus reconstruction algorithm faces great challenges in practical applications because it needs to separate pure speech from mixed speech, which is extremely difficult in real-world scenarios. Although the direct classification method has more potential in practical applications, traditional linear decoders have obvious defects. Due to the non-linear characteristics of the brain's auditory system, traditional linear decoders are difficult to capture the non-linear mapping relationship in EEG signals, resulting in a longer decision time window. Moreover, as the decoding window length shortens, the decoding accuracy will rapidly decline. In recent years, deep learning techniques have been widely used in auditory attention decoding research, but existing methods still have many problems. 1) Complex and suboptimal frequency-domain feature extraction: When extracting EEG frequency-domain features, existing methods usually need to pre-filter the EEG signals and then manually extract differential entropy features in fixed frequency bands. This operation not only increases the complexity of data preprocessing, but also makes it difficult to obtain the optimal frequency-domain decoding range that best matches auditory attention decoding due to the fixed frequency band division method, resulting in the inability to fully exploit the frequency-domain information in EEG signals. 2) Poor adaptability of convolutional kernels: Most methods based on convolutional neural networks use convolutional kernels of fixed size to learn local features. However, in actual situations, the optimal convolutional kernel size varies among different subjects and at different time points. The method with a fixed convolutional kernel size cannot adapt to this variation, limiting the effective extraction of features from different EEG data. 3) Limitations of 3D convolution: Some studies have tried to map two-dimensional EEG into three-dimensional data and use 3D convolution to process spatio-temporal or frequency-spatial features to utilize the spatial distribution features of EEG signals. However, 3D convolution faces many difficulties during training, with high computational complexity and difficult model optimization. At the same time, it is also difficult to capture the long-range dependence relationships and dynamic change information between different brain regions in EEG signals, and cannot comprehensively reflect the activity characteristics of the brain during the auditory attention process. 4) Lack of an effective attention mechanism: Currently, in the field of auditory attention decoding, there is insufficient research on attention mechanisms based on spatial and frequency-domain features. Due to the lack of a mechanism that can comprehensively integrate and analyze spatial and frequency-domain information in EEG data, existing methods cannot fully capture the key information related to auditory attention, resulting in limitations in the performance of the model.
[0004] In summary, existing auditory attention decoding technologies still have bottlenecks in feature extraction, model training, and information capture. There is an urgent need for an innovative method to optimize the decoding strategy, improve the decoding accuracy in a short time window, and enhance the real-time performance and adaptability of the system, so as to provide more accurate and efficient technical support for neuro-driven auditory assistive devices. Summary of the Invention
[0005] In view of the deficiencies of the existing auditory attention decoding technology, the present invention proposes an auditory attention decoding method and system based on a multi-scale spatio-temporal attention network, aiming to improve the decoding accuracy and computational efficiency, especially the decoding performance in complex acoustic environments and short decision windows.
[0006] The technical solution adopted to achieve the purpose of the present invention is as follows: An auditory attention decoding method based on a multi-scale spatio-temporal attention network, comprising the following steps: Step 1, obtaining electroencephalogram (EEG) data when attentively listening to voices in different directions, and using a sliding window to divide the EEG data to generate a series of decision windows, each decision window containing a segment of EEG signal; Step 2, inputting the EEG signal of the decision window into a multi-scale spatio-temporal attention network (MSSANet model), the multi-scale spatio-temporal attention network including a multi-scale time-domain convolution module, a spatio-temporal attention module, and a fully connected layer classification module: Step 2.1, the multi-scale time-domain convolution module processes the input EEG signal to extract time-domain features in different frequency ranges: the multi-scale time-domain convolution module includes a multi-scale residual convolution unit and a time-domain logarithmic variance calculation unit. In the multi-scale residual convolution unit, N convolution kernels of size 1×1 are used to perform a dimensionality increase operation on a single input sample R and then the output Y after dimensionality increase is divided into K groups according to the channel dimension. For each group of output Y b depth convolution is performed using different convolution kernels to obtain and the results after grouped depth convolution are concatenated according to the output channel dimension to obtain a multi-scale convolution output The multi-scale convolution output is obtained through time-domain logarithmic variance calculation to get ; Meanwhile, for each group of output Y b a residual convolution output res is obtained through processing by a fixed convolution operation Conv The residual convolution output is obtained through time-domain logarithmic variance calculation to get ; and are added together to obtain the output of the multi-scale time-domain convolution module, that is, the time-domain features in different frequency ranges ; Step 2.2, the spatio-temporal attention module subjects the time-domain features Convert it into a spectro-spatial feature map, and further capture the global dependencies between different brain regions through the self-attention mechanism and learnable position encoding to extract spectro-spatial information related to auditory attention; Step 2.3, the fully connected layer classification module outputs the probabilities of predicting the speech direction as the left or right direction based on the spectro-spatial information.
[0007] In the above technical solution, in step 1, before dividing the EEG data, there is also a data preprocessing step, and the data preprocessing step includes downsampling, filtering, artifact removal, and / or channel normalization.
[0008] If the EEG data comes from the KUL dataset, first downsample the EEG data to 128 Hz, then perform band-pass filtering with an 8th-order Butterworth filter in the range of 0.1 - 50 Hz, and finally perform channel normalization; If the EEG data comes from the DTU dataset, first filter out 50 Hz linear noise and artifacts, remove eye artifacts through joint decorrelation analysis, perform whole-brain average re-reference, then downsample to 128 Hz and perform channel normalization.
[0009] In the above technical solution, , where, is a convolution with a kernel of 1, R is the EEG signal within each decision window, W1 is the weight matrix of the 1×1 convolution kernel, b1 is the bias vector, Y ∈ R N×C×T is the output after dimensionality increase, N is the number of convolutions, C is the number of channels of the EEG signal, and T is the number of samples within each decision window.
[0010] In the above technical solution, Y = [Y1, Y2,..., Y K , where Y b ∈ R N / K×C×T , b = 1, 2,..., K, the b-th group uses a convolution kernel of size (1, k b ) for depth convolution, , where, is the depth convolution, k b is the different convolution kernel sizes, W 2b is the weight matrix of the b -th group of convolution kernels, b 2b is the bias vector, Z b is the EEG data after depth convolution, Z b ∈ R N / K×C×T , .
[0011] In the above technical solution, , where, is a residual convolution, W3 is the weight matrix of the residual convolution kernel, and b3 is the bias vector.
[0012] In the above technical solution, the calculation formula of the time-domain logarithmic variance is , where ∈ represents the i th sample point of the lead, is the stride, represents the variance of the sample points, and the sample points are or .
[0013] In the above technical solution, in step 2.2, the output obtained by the multi-scale time-domain convolution module is converted into a frequency-time spatial feature map F ∈ R N×M with a size of N × M, where M = C × D, D represents the number of divisions of the EEG time-domain length according to the T' step size. The frequency-time spatial feature map F is passed through a learnable position encoding to retain the spatial position information of the EEG signal and output the feature P , and the Transformer Encoder is used to perform cross-frequency domain processing on the feature P to obtain , , is the spectral spatial information related to auditory attention.
[0014] In the above technical solution, in step 2.3, the frequency-time spatial attention feature is first flattened, and then passed through two fully connected layers to predict the probability of the auditory attention decoding direction to obtain , , where is the predicted probability output by the model, W4 is the weight matrix of the conversion, and b4 is the bias vector.
[0015] In the above technical solution, in step 2, the multi-scale frequency-time spatial attention network is evaluated using the cross-entropy loss function , , where represents the number of samples, is the number of classifications, is the true value, is the predicted value, i is the i rd sample, c is the number of categories, corresponding to left or right.
[0016] On the other hand, the present invention further includes a system capable of implementing the auditory attention decoding method based on the multi-scale spatio-frequency attention network, including a data import module, a data preprocessing module, the multi-scale spatio-frequency attention network, a model training module, and a result visualization module; The data import module is used to select different types of data sets and import data. The data preprocessing module preprocesses the imported data, and the preprocessing includes downsampling, high-pass filtering, low-pass filtering, and / or normalization; The model training module is used to set the length of the time window and the data set division ratio to optimize the training effect of the multi-scale spatio-frequency attention network to meet the requirements of different auditory attention decoding tasks; The result visualization module visually displays the accuracy rate of the model prediction results after the multi-scale spatio-frequency attention network is trained. When the probability of the predicted speech direction being high on the left side in step 2.3, the model prediction result is "left side". When the probability of the predicted speech direction being high on the right side in step 2.3, the model prediction result is "right side". The accuracy rate is the percentage of the correctly predicted ones among all samples in the total samples.
[0017] Compared with the prior art, the beneficial effects of the present invention are: (1) Efficient feature extraction and adaptive frequency domain information acquisition: The MSSANet model proposed by the present invention extracts the EEG local features in different frequency domain ranges through multi-scale time domain convolution, which can simulate the filtering process, avoid the complex preprocessing of manually extracting frequency domain features, and can adaptively obtain the frequency domain information related to auditory attention decoding. When the time domain logarithmic variance calculation unit processes the EEG signal, on the one hand, it can efficiently extract the time domain information, and on the other hand, it cleverly retains the spatial information in the signal. Through experiments, compared with the traditional convolution and pooling operations, it significantly improves the auditory attention decoding accuracy rate; (2) Spatio-frequency attention module for enhancing model performance: This module effectively captures the long-range dependencies and global dependencies between different brain regions with the help of the self-attention mechanism and the learnable position encoding, comprehensively obtains the frequency domain and spatial information related to auditory attention, and improves the model performance. The self-attention mechanism can suppress the interference of irrelevant frequency bands / brain regions by enhancing the interaction weights of the brain region pairs related to auditory attention. This module realizes the explicit modeling of the cross-band spatial dependence relationship in the EEG signal and provides a more global perspective feature representation for auditory attention decoding; (3)Excellent experimental performance and high practical value: In the experiments on the KUL and DTU public datasets, MSSANet demonstrated its superior strength, achieving the highest classification accuracy under extremely short decision windows of 0.1 second, 0.5 second, and the conventional 1 second. This excellent performance under short decoding time windows lays a solid foundation for real-time auditory attention decoding, greatly enhancing the response speed and accuracy. With this advantage, it has great practical value in practical application scenarios such as neurally-driven auditory assistive devices (cochlear implants or hearing aids), can effectively meet real-world needs, and injects strong impetus into the development of related fields; (4)Constructing an efficient automated decoding system: The present invention develops an auditory attention decoding system based on the multi-scale spatio-temporal attention network framework, integrating modules such as data import, preprocessing, model training, and result visualization. It supports the import of multiple datasets and formats. After preprocessing such as downsampling, the model is optimized through training parameters such as adjustable time windows, and the results are presented visually in terms of accuracy, realizing full-process automated analysis. It has the ability to flexibly adjust parameters and efficiently process data, reducing human errors, and providing a powerful tool for research and application related to auditory attention. Description of the Drawings
[0018] Figure 1 It is a flowchart of the auditory attention decoding method based on the multi-scale spatio-temporal attention network.
[0019] Figure 2 It is a schematic diagram of the multi-scale spatio-temporal attention network.
[0020] Figure 3 It is the decoding effect of the multi-scale spatio-temporal attention network in two public datasets.
[0021] Figure 4 It is the decoding effect of the ablation experiment of the multi-scale spatio-temporal attention network.
[0022] Figure 5 It is the visualization result of the ablation experiment features.
[0023] Figure 6 It is the auditory attention decoding system based on the multi-scale spatio-temporal attention network. Detailed Implementation Modes
[0024] The following further elaborates on the present invention in conjunction with specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0025] Embodiment 1
[0026] An auditory attention decoding method based on the multi-scale spatio-temporal attention network includes the following steps: Step 1: Obtain EEG data, preprocess the EEG data, and then divide it to generate a series of decision windows: Step 1.1: EEG data preprocessing steps. The data preprocessing steps include downsampling, filtering, artifact removal, and normalization. If the EEG data comes from the KUL dataset, first downsample the EEG data to 128 Hz, then perform band-pass filtering from 0.1 - 50 Hz using an 8th-order Butterworth filter, and finally perform channel normalization. If the EEG data comes from the DTU dataset, first filter to remove 50 Hz linear noise and artifacts, remove eye artifacts through joint decorrelation analysis, perform whole-brain average rereferencing, then downsample to 128 Hz and perform channel normalization. Step 1.2: Use a sliding window to divide the preprocessed EEG data to generate a series of decision windows, where each decision window contains a segment of EEG signal (such as 0.1 second, 0.5 second, and 1 second). Step 2: Input the EEG signal of the decision window into a multi-scale frequency-space attention network (MSSANet model). The multi-scale frequency-space attention network includes a multi-scale time-domain convolution module, a frequency-space attention module, and a fully connected layer classification module. Step 2.1: The multi-scale time-domain convolution module processes the input EEG signal to extract time-domain features in different frequency ranges. In the multi-scale residual convolution unit, use N convolution kernels of size 1×1 for dimensionality increase operation, specifically expressed as: (1); Among them, R is the EEG signal within each decision window, is the convolution with a kernel of 1, W1 is the weight matrix of the 1×1 convolution kernel, b1 is the bias vector, Y ∈ R N×C×T is the output after dimensionality increase, N is the number of convolutions, C is the number of channels of the EEG signal, T is the number of samples within each decision window. After the above steps, the dimension of the input data R can be changed to N×C×T; Then divide the output Y after dimensionality increase into K groups along the channel dimension, that is, Y = [Y1, Y2,..., Y K , where Y b ∈ R N / K×C×T , b = 1, 2,..., K. The b-th group uses a convolution kernel of size (1, k b ) ( k b are different convolution kernel sizes) for depth convolution, specifically expressed as: (2); Among them, is the depth convolution, and W 2b is the b weight matrix of the 2b th group of convolution kernels, and b is the bias vector, and b ∈R N / K×C×T . The results after grouped depth convolution are concatenated according to the output channel dimension to obtain: (3); is the multi-scale convolution output.
[0027] Since the time-domain convolution has a filtering-like effect, the model can dynamically extract multi-scale features in the EEG signal during operation. In addition, a residual path with fixed convolution kernels is introduced, and the residual connection mechanism is used to enhance the feature extraction ability and improve the network's ability to capture long-term dependence relationships. Assume that through a fixed convolution operation Conv res processes the input to obtain the residual convolution output : (4); where is the residual convolution, W3 is the weight matrix of the residual convolution kernel, and b3 is the bias vector.
[0028] In the time-domain log variance calculation unit: To efficiently extract time-domain features and retain the spatial information in the EEG signal, the log variance is used to calculate the variance of the time series of each lead of the EEG and take the logarithm. The calculation formula for the log variance operation is as follows: (5); where ∈ represents the i th sample point of the th lead, is the stride, represents the variance of the sample point, and the sample point is or . Based on formula (5), the time-domain log variance calculation unit is applied respectively after the multi-scale convolution ( ) and the residual convolution ( ) to perform the log variance operation with a stride of .
[0029] Thus, the multi-scale convolution output ( ) obtains ∈R N×C×D after the log variance operation, and the residual convolution output (O) obtains ∈RN×C×D , where , D represents the number of divisions of the EEG time domain length according to T' the step size, is the number of samples in each decision window. After passing through the logarithmic variance operation module, the results of the two paths are finally added element-wise: (6); is the output of the multi-scale time domain convolution module; Step 2.2, the frequency-space attention module processes the frequency-space feature map, captures the global dependencies between different brain regions, and extracts the spectral-spatial information related to auditory attention; The processing process of the frequency-space attention module includes the following steps: The output obtained after passing through the multi-size time domain convolution module N ×M is converted into a frequency-space feature map F ∈ R of size N×M P , where M = C×D. Further, the frequency-space feature map F is passed through a learnable position encoding to retain the spatial position information of the EEG signal and output the feature P , and a Transformer Encoder is used to perform cross-frequency domain processing on the feature to enhance the model's ability to extract frequency-space distribution information. Finally, the output of the model can be obtained through the above steps : is the spectral-spatial information related to auditory attention; Step 2.3, the fully connected layer classification module outputs the probability of predicting the left or right direction based on the spectral-spatial information, and then defines the one with the higher probability as the result predicted by the model. Further, the percentage of correctly predicted samples in all samples to the total number of samples is the final accuracy rate.
[0030] The fully connected layer classification module uses two fully connected layers as classifiers. First, the spectral-spatial information related to auditory attention is flattened in the frequency domain, and then the probability of the auditory attention decoding direction is predicted through two fully connected layers: (8); where is the predicted probability output by the model, W4 is the transformation weight matrix, and b4 is the bias vector; Furthermore, a cross-entropy loss function is adopted to measure the difference between the prediction result and the true result of the MSSANet model, so as to guide the training of the MSSANet model: (9); wherein, represents the number of samples, is the number of classifications (in the present invention, it is 2: left or right), is the true value, is the predicted value, i is the i th sample, c is the number of categories, corresponding to left or right.
[0031] Example 2 This example verifies a method for auditory attention decoding based on a multi-scale spatio-temporal attention network in Example 1.
[0032] This example conducts experimental verification based on two publicly available datasets (KUL and DTU) widely used in the field of auditory attention research.
[0033] The KUL dataset contains EEG data of 16 normal-hearing subjects. The 64-channel EEG data is recorded using the BioSemi ActivateTwo System at a sampling rate of 8196 Hz. The experiment is conducted in a shielded room. The speech materials are read by 3 male native Dutch speakers, filtered at 4 kHz and set to an intensity of 60 dB, and then played through in-ear headphones. The experiment includes two auditory conditions. Each subject listens to 8 trials, and each trial lasts for 6 minutes.
[0034] The DTU dataset contains data from 18 normal-hearing subjects, which are collected through a 64-channel BioSemi ActiveTwodevice at a sampling rate of 512 Hz. The subjects need to listen to one of the two simultaneously presented speakers. The speaker's voice is played at an angle of 60° relative to each other. The auditory materials are played through ER-2 in-ear headphones at an intensity of 60 dB and include Danish audiobooks read by 3 male and 3 female readers. Each subject conducts 60 trials of the experiment, and each trial lasts for 50 seconds.
[0035] Data preprocessing: To ensure the fairness and reliability of experimental results, specific data preprocessing strategies are implemented for different datasets. For the KUL dataset, first, the EEG data is downsampled to 128 Hz to retain key information while reducing the data dimension. Then, an 8th-order Butterworth filter is used for band-pass filtering in the range of 0.1 - 50 Hz to effectively remove noise while maximizing the retention of EEG information. Finally, channel normalization is performed to make the mean of each lead sample point 0 and the standard deviation 1, ensuring that the data of each lead is on the same scale for subsequent analysis. For the DTU dataset, first, the data is filtered to remove 50 Hz linear noise and other artifacts to ensure data quality. Then, eye artifacts are removed through joint decorrelation analysis to reduce the interference of eye movements on the EEG signal. Whole-brain average re-referencing is performed to optimize the signal reference standard. After that, the data is downsampled to 128 Hz and channel normalization is completed to make the data meet the requirements of subsequent experiments.
[0036] Dataset division: To ensure the effectiveness and generalization ability of model training, the dataset is divided as follows. The first 90% of the data is intercepted from each trial as the training set for model training and parameter learning; the last 10% of the data is used as the test set to evaluate the generalization ability of the model on unknown data. Using the sliding window technique, samples are extracted from the training set and the test set with an overlap rate of 50% to increase data utilization and diversity. The training set samples are further randomly divided into a training set (90% proportion) and a validation set (10% proportion). Taking the KUL dataset as an example, after the above operations, each subject can obtain 4658 training samples, 518 validation samples, and 568 test samples. The present invention strictly follows the dataset division principle to ensure data independence during the experiment, avoid data leakage, and improve the reliability of the results.
[0037] To improve the training effect of the model, this embodiment adopts an optimization strategy. Specifically, the cross-entropy loss function is selected as the index to measure the difference between the model prediction result and the true label, guiding the training direction of the model. The batch size is set to 32 to balance memory usage and training efficiency; the epochs of the model are set to 100 to ensure that the model has enough training times to learn data features. An early stopping strategy is adopted. When the loss function value on the validation set does not decrease for 10 consecutive epochs, the training is stopped to prevent the model from overfitting. The Adam optimizer is used to train the model, and the learning rate is set to 0.0005 and the weight decay is set to 0.0003 to adjust the model parameters and make the model converge towards the optimal solution. The MSSANet model of the present invention is implemented in the Python 3.10 environment with the help of the PyTorch framework and accelerated by the GeForce RTX 3090 GPU to improve the training efficiency. In addition, the present invention uses the grid search method for hyperparameter tuning to ensure that the model reaches the best configuration. The parameter settings of the MSSANet model are shown in Table 1 in detail; ;
[0038] Model performance comparison and analysis: The present invention conducts comprehensive experiments under decision windows of 0.1 s, 0.5 s, 1 s, etc. that have important practical significance, and focuses on analyzing the decoding accuracy of the MSSANet model (presented in the form of mean ± standard deviation). The results are shown in Figure 3. The accuracies of the MSSANet model of the present invention on the KUL dataset are 92.4% ± 6.90 (0.1 s time window), 95.0% ± 5.09 (0.5 s time window), and 95.3% ± 4.94 (1 s time window) respectively; the accuracies on the DTU dataset are 78.5% ± 6.51 (0.1 s time window), 84.5% ± 5.75 (0.5 s time window), and 83.4% ± 6.80 (0.1 s time window) respectively. Further, to verify the superiority of the MSSANet model of the present invention, it is compared with a variety of advanced models, including STAnet, XAnet, MBSSFCC, BSAnet, DenseNet-3D, DBPNet, DARnet, etc. The results are shown in Table 2 in detail. The model of the present invention performs excellently under each decision window, significantly improving the decoding performance; ;
[0039] Ablation experiment verification: To deeply explore the specific contributions of each module of the model to the overall performance, ablation experiments are carried out to comprehensively analyze the impact of different modules on the model performance. The specific experimental operations and analysis processes are as follows: 1) Key module adjustment experiment: By replacing specific modules in the model or modifying their parameters, here we focus on adjusting the multi-scale time-domain convolution module and the size of the convolution kernel. The results are as described in Table 3; 2) Verification of the effectiveness of the frequency-space attention module: To prove the effectiveness of the frequency-space attention module in the model, we removed this module from the model while keeping other experimental settings the same as the original MSSANet model, and analyzed the changes in the auditory attention decoding performance of this module. The experiment is as Figure 4 shown. It can be seen from the figure that after removing the frequency-space attention module, the average accuracy rates in the decision-making windows of 0.1 s, 0.5 s, and 1 s decreased by 0.25%, 0.51%, and 0.74% respectively; 3) Feature visualization analysis: To visually compare the capabilities of different modules in feature extraction and information representation, the t-SNE technique is used to visualize the features before the fully connected layer classification module (the original data representation refers to the data without being processed by the model; convolution: refers to using a one-dimensional convolution kernel to replace the time-domain logarithmic variance calculation unit; average: refers to using the mean to replace the time-domain logarithmic variance calculation unit; kernel is 17: refers to using a one-dimensional convolution (kernel of size 1*17) to replace the multi-scale time-domain convolution module; without the frequency-space attention module means removing the frequency-space attention module from the model), as Figure 5 shown. It can be more clearly understood the differences in feature representation of different modules, and the model of the present invention has stronger feature separability (the red and blue feature points are more separated). The results of this experiment show that the frequency-space attention module and the time-domain logarithmic variance calculation unit play a key role in improving the decoding performance, further verifying the effectiveness of the present invention; ;
[0040] Example 3 A system for an auditory attention decoding method based on a multi-scale frequency-space attention network according to Example 1, as Figure 6 shown. This system can realize the full-process operation from the processing of auditory attention-related data to model training and result presentation, effectively improving the auditory attention decoding efficiency. The specific component modules are as follows: Data import module: Used to select different types of data sets and import data. It supports data sets such as the KUI data set and the DTU data set, and can also import newly collected data stored in formats such as.mff,.cnt, and curry. Through this module, users can select appropriate data set sources according to actual needs, providing a data basis for subsequent analysis. This system is suitable for the diverse requirements of data sources for auditory attention decoding in different research or application scenarios, such as using specific publicly available data sets in scientific research experiments or importing newly collected on-site data in practical applications.
[0041] Data preprocessing module: Perform a series of preprocessing operations on the imported data. It includes steps such as downsampling (which can be set to 128), high-pass filtering (with the parameter set to 48 Hz), low-pass filtering (with the parameter of 0.5 Hz), and normalization range setting (set to 0 - 1). Through these preprocessing operations, the module removes interference factors, ensures the accuracy and consistency of the data, and improves the reliability of subsequent analysis to meet the requirements of subsequent model training.
[0042] Model training module: Users can select training parameters for model training. The time window (such as 0.1 second, 0.5 second, 1 second, etc.) and the dataset division ratio (such as training set 0.9, training set 0.8, or other ratios) can be set. By adjusting these parameters, the training effect of the model can be optimized to meet the requirements of different auditory attention decoding tasks. This module can flexibly adjust the training parameters according to specific tasks, improving the adaptability and accuracy of the model for different auditory attention decoding tasks.
[0043] Result visualization module: Visualize the results after model training in the form of accuracy. For example, the accuracy corresponding to a 0.1-second time window is 92% (green light), the accuracy corresponding to a 0.5-second time window is 94% (green light), and the accuracy corresponding to a 1-second time window is being calculated (red light), which is convenient for users to intuitively understand the performance of the model under different parameter settings. Researchers or users can quickly evaluate the training effect of the model through this module, compare the accuracies of the model under different parameter settings such as different time windows, and thus select the optimal parameter configuration to guide subsequent auditory attention decoding applications.
[0044] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An auditory attention decoding method based on a multi-scale spatio-temporal attention network, characterized in that It includes the following steps: Step 1: Obtain the electroencephalogram (EEG) data when attentively listening to voices in different directions, and use a sliding window to divide the EEG data to generate a series of decision windows, where each decision window contains a segment of EEG signal; Step 2: Input the EEG signal of the decision window into a multi-scale spatio-temporal attention network, which includes a multi-scale time-domain convolution module, a spatio-temporal attention module, and a fully-connected layer classification module: Step 2.1, the multi-scale time-domain convolution module processes the input EEG signals to extract time-domain features in different frequency ranges: The multi-scale time-domain convolution module includes a multi-scale residual convolution unit and a time-domain logarithmic variance calculation unit. In the multi-scale residual convolution unit, N convolution kernels of size 1×1 are used to perform a dimensionality increase operation on a single input sample R , and then the output Y after dimensionality increase is divided into K groups according to the channel dimension. For each group of output Y b , depth convolution is further performed using different convolution kernels to obtain . The results after grouped depth convolution are concatenated according to the output channel dimension to obtain the multi-scale convolution output . The multi-scale convolution output is calculated through the time-domain logarithmic variance to obtain ; Meanwhile, for each group of outputs Y b Through the fixed convolution operation Conv res The residual convolution output is obtained through processing , the residual convolution output Is obtained through time-domain logarithmic variance calculation ; and are added together to obtain the output of the multi-scale time-domain convolution module, that is, the time-domain features in different frequency ranges ; Step 2.2, the frequency-space attention module converts the time-domain features into a frequency-space feature map, and further captures the global dependencies between different brain regions through the self-attention mechanism and learnable position encoding to extract the spectral-space information related to auditory attention; Step 2.3: The fully-connected layer classification module outputs the probabilities of predicting that the voice direction is the left or the right direction based on the spectral-spatial information.
2. The auditory attention decoding method based on the multi-scale spatio-temporal attention network according to claim 1, characterized in that, In Step 1, before dividing the EEG data, it further includes a data preprocessing step, and the data preprocessing step includes downsampling, filtering, artifact removal, and / or channel normalization.
3. The auditory attention decoding method based on the multi-scale spatio-temporal attention network according to claim 1, characterized in that , where is a convolution with a kernel of 1, R is the EEG signal within each decision window, W1 is the weight matrix of the 1×1 convolution kernel, and b1 is the bias vector. is the output after dimensionality increase, N is the number of convolutions, C is the number of channels of the EEG signal, and T is the number of sample points representing each decision window.
4. The auditory attention decoding method based on the multi-scale spatio-temporal attention network according to claim 3, wherein , where Y b ∈R N / K×C×T , b = 1, 2, …, K, the b-th group performs depth convolution using a convolution kernel of size (1, k b ), , where, is the depth convolution, k b are different convolution kernel sizes, W 2b is the b -th group's convolution kernel weight matrix, b 2b is the bias vector, Z b is the EEG data after depth convolution, Z b ∈R N / K×C×T , .
5. The auditory attention decoding method based on the multi-scale spatio-temporal attention network according to claim 1, characterized in that, , where, is a residual convolution, W3 is the weight matrix of the residual convolution kernel, and b3 is the bias vector.
6. The auditory attention decoding method based on the multi-scale spatio-temporal attention network according to claim 1, characterized in that The calculation formula for the time-domain logarithmic variance is , where represents the i th sample point of the lead, is the stride, represents the variance of the sample points, and the sample points are or .
7. The auditory attention decoding method based on a multi-scale spatio-temporal attention network according to claim 1, wherein In step 2.2, the output obtained from the multi-scale time-domain convolution module is converted into a spatio-frequency feature map F ∈ R of size N×M N ×M , where M = C×D, D represents the number of divisions of the EEG time-domain length according to T' the step size. The spatio-frequency feature map F is passed through learnable positional encoding to retain the spatial position information of the EEG signal and output the feature P , and the Transformer Encoder is used to perform cross-frequency domain processing on the feature P to obtain , , which is the spectral-spatial information related to auditory attention.
8. The auditory attention decoding method based on a multi-scale spatio-temporal attention network according to claim 1, wherein In step 2.3, first flatten the frequency-space attention feature and then pass it through two fully connected layers to predict the probability of the auditory attention decoding direction to obtain , , where is the predicted probability output by the model, W4 is the transformed weight matrix, and b4 is the bias vector.
9. The auditory attention decoding method based on a multi-scale spatio-temporal attention network according to claim 1, characterized in that In step 2, the cross-entropy loss function is used to evaluate the multi-scale spatio-temporal attention network, , where represents the number of samples, is the number of classifications, is the true value, is the predicted value, i is the i th sample, c is the number of categories, corresponding to left or right.
10. A system based on the auditory attention decoding method of the multi-scale spatio-temporal attention network as described in claim 1, characterized in that, It includes a data import module, a data preprocessing module, the multi-scale spatio-temporal attention network, a model training module, and a result visualization module; The data import module is used to select different types of data sets and import data, and the data preprocessing module preprocesses the imported data, and the preprocessing includes downsampling, high-pass filtering, low-pass filtering, and / or normalization; The model training module is used to set the length of the time window and the data set division ratio to optimize the training effect of the multi-scale spatio-temporal attention network to meet the requirements of different auditory attention decoding tasks; The result visualization module visually displays the accuracy rate of the model prediction results after the multi-scale spatio-temporal attention network is trained. When the probability of predicting the voice direction as the left is high in Step 2.3, the model prediction result is "left", and when the probability of predicting the voice direction as the right is high in Step 2.3, the model prediction result is "right", and the accuracy rate is the percentage of the correctly predicted ones in all samples to the total samples.
Citation Information
Patent Citations
Auditory attention detection method and system based on time-frequency domain fusion
CN118121192A
Auditory orientation attention decoding method and device
CN118885872A
Auditory attention decoding method and device based on multi-sound-source scene, equipment and medium
CN119138899A
EEG auditory attention classification method and system based on time-frequency attention mechanism
CN119498864A
Cross-session brainprint recognition method based on tensorized spatial-frequency attention network (TSFAN) with domain adaptation
US20240282439A1
Cited By
Local-global time sequence interaction network for unilateral directional motor imagery electroencephalogram decoding
CN120849917A
Noise robustness auditory decoding method and system based on adversarial domain adaptation network
CN121561652A