A method for detecting auditory attention based on pulse inspiration

CN119046753BActive Publication Date: 2026-09-29DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411084249.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-09-29
Estimated Expiration
2044-08-08

AI Technical Summary

Technical Problem

自1953年提出以来,鸡尾酒会问题一直是研究的热点领域,尽管正常听力的人可以通过集中注意力利用空间定位轻松分析复杂的语音场景,但一些听力受损的人(例如感觉神经性听力损失)即便佩戴有助听器或人工耳蜗,在存在多个说话者的声学环境中进行语音交流是极具挑战性的

Benefits of technology

[0041]本发明的有益效果:本发明将脉冲神经网络引入听觉注意力检测任务,模拟人脑信息编码机制,解决当前听觉注意力检测模型存在的可解释性弱、实时性差的问题,实现低延迟、短决策窗口下的高准确率,与其他方法相比,本发明提出的方法具有更好的泛化性,在不同受试者上均有稳定的检测结果。并且脉冲神经网络在配合专门的神经硬件进行使用时能够大幅降低模型的能耗,有助于硬件开发,为开发脑控助听器提供了有效方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119046753B_ABST
    Figure CN119046753B_ABST
Patent Text Reader

Abstract

The present application relates to the field of brain-computer interface in the computer and neuroscience cross discipline, and proposes an auditory attention detection method based on pulse inspiration. A threshold calculation module in the method provides a threshold for a pulse layer, and an electroencephalogram branch and an audio branch are used to process time-synchronized electroencephalogram signals and speech envelopes respectively. The electroencephalogram branch uses the pulse layer to extract time features and encode, and a channel attention module for the electroencephalogram is arranged therein; the audio branch uses the pulse layer to extract features and encode for two input audios respectively. Then a feature fusion module is used to extract fusion features of the electroencephalogram and the two audios. Finally, a classifier is used to obtain the result of auditory attention detection. Experiments show that the method realizes high accuracy under low delay and short decision window, has better generalization, and has stable detection results on different subjects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of brain-computer interfaces in computer science, and more particularly to a method for auditory attention detection based on electroencephalogram (EEG) data and audio. Background Technology

[0002] The cocktail party problem is known as how people can isolate and understand the voice of the speaker of interest from background noise when multiple speakers are speaking simultaneously in a noisy environment. Since its inception in 1953, the cocktail party problem has been a hot research area. While people with normal hearing can easily analyze complex speech scenarios by focusing their attention and using spatial localization, it is extremely challenging for some people with hearing loss (such as sensorineural hearing loss) to communicate in acoustic environments with multiple speakers, even with hearing aids or cochlear implants. According to the latest Global Burden of Disease study, the burden caused by hearing loss is increasing over time due to population aging, thus increasing the global demand for hearing aids. Over the past decade, hearing aids have made substantial progress in suppressing background noise that interferes with speech, reducing simple environmental noise and amplifying sound sources in front of the user. High-performance automatic speaker separation and speech enhancement algorithms have also emerged. However, they cannot enhance the target speaker without knowing which speaker the listener is talking to. Therefore, detecting the voice of interest to the user in noisy environments is crucial.

[0003] Recent studies on the characteristics of speech representation in the human auditory cortex have revealed that speakers who receive attention have stronger representations compared to unattended sound sources. These findings have sparked the prospect of brain-controlled assisted hearing devices that continuously monitor the listener's brainwaves and compare them to sounds in the environment to determine the speaker the subject is most likely to be paying attention to. The device can then amplify the voice of the speaker receiving attention relative to others, making the speaker more easily heard in a crowd. This process is known as Auditory Attention Decoding (AAD), a field of research that has seen considerable progress in recent years.

[0004] Auditory attention detection is a typical application of brain-computer interface technology and is of great significance in the study of brain function and neuroscience. In addition, researchers have made significant progress in fields such as affective computing, speech recognition, and natural language processing through further exploration of brain function and neuroscience. Research on auditory attention detection methods can not only solve the cocktail party problem and develop brain-controlled hearing aids, but also provide a foundation for in-depth research in brain science.

[0005] Currently, research on AAD based on EEG signals, both domestically and internationally, can be mainly divided into two categories: linear methods and nonlinear methods.

[0006] Linear methods seek the connection between EEG and speech through various modeling techniques, broadly categorized into forward modeling (modeling speech to EEG) and backward modeling (modeling EEG to speech). Stimulus reconstruction is a classic algorithm for backward modeling. This method applies a neural decoder to the EEG channel, reconstructs the involved speech envelope or spectrogram, and then compares it with the speech representations of all speakers. The speaker with the highest Pearson correlation coefficient is identified as the present speaker. The AAD method based on the time-domain response function belongs to forward modeling. It estimates the EEG signal from the speech, then performs correlation analysis between the estimated EEG signal and the real signal, and finally classifies based on the correlation. There are also methods based on canonical correlation analysis, which perform both forward and backward modeling. All of the above linear methods suffer from the problem of low correlation between the reconstructed data and the original data. They assume a linear relationship between EEG and speech signals, failing to adequately fit the nonlinear characteristics of the auditory system.

[0007] Nonlinear methods have been proposed in recent years, typically based on deep learning frameworks for auditory attention detection. Examples include convolutional neural networks, long short-term memory networks, and attention networks. These networks can learn the nonlinear relationship between EEG signals and audio, better simulating the human brain's data processing and addressing the issues of poor real-time performance and low accuracy associated with linear models. However, research using neural networks for auditory attention detection fails to consider the characteristics of EEG signals and neurocognitive mechanisms, leading to a "black box" problem and poor model interpretability.

[0008] In contrast, this invention introduces a bio-inspired spiking neural network that mimics the workings of a biological nervous system, making it more interpretable, capable of handling time-series data well, and with lower energy consumption. Summary of the Invention

[0009] This invention addresses the shortcomings of existing methods by proposing a pulse-inspired auditory attention detection method, which can effectively detect auditory attention in short sequences. This method is based on deep learning and utilizes synchronously acquired EEG signals and audio information for short-sequence auditory attention detection. It encodes the EEG signals and audio information by simulating the information processing mechanism of the human auditory system during perception using a spiking neural network, and then inputs this information into a subsequent feature extraction network for classification.

[0010] The technical solution of this invention is as follows:

[0011] A pulse-inspired auditory attention detection method is proposed. Simultaneously acquired EEG signals and two audio signals are preprocessed. First, a spiking neural network threshold is calculated from the EEG signals, and then input into both the EEG and audio branches. The EEG branch first extracts and encodes temporal features through a spiking layer, and then applies a channel attention module to capture the inter-channel relationships of the EEG signals. For the two simultaneous audio signals, separate audio branches are created, each extracting and encoding sequential temporal features through a spiking layer. The outputs of the two audio branches are concatenated with the output of the EEG branch along the channel dimension. The concatenated results are then input into a feature fusion module to obtain the fused features between the EEG and the two audio signals. A fully connected layer maps the fused features to an output dimension equal to the number of categories. Finally, a softmax function is used to obtain the classification result.

[0012] The specific steps are as follows:

[0013] Step 1: Preprocessing of EEG signals and audio

[0014] The EEG signal is first high-pass filtered at a cutoff frequency of 0.5 Hz to eliminate DC components and electrode drift; then the EEG signal is rereferenced to the EEG Cz channel, downsampled to 128 Hz and band-pass filtered.

[0015] The audio is processed using an auditory filter with power-law compression to obtain the speech envelope, and finally downsampled to 128Hz and bandpass filtered.

[0016] To make the data more suitable for subsequent networks, the EEG signals are standardized and normalized, and the speech envelope is normalized. To effectively verify the detection accuracy of the invention in a short time, a sliding window is used to divide the EEG signals and speech envelope into continuous decision windows, and the EEG signals and audio are always kept synchronized.

[0017] Step 2: Calculate the threshold of the spiking neural network based on EEG signals

[0018] Threshold calculation is used to provide thresholds that are more suitable for different subjects in subsequent pulse layers, thereby improving the generalization and robustness of the network.

[0019] The threshold is calculated based on given EEG data. After a period of training, the EEG signal is input into the threshold calculation module, and then the threshold is fixed after a period of training to train other network layers.

[0020] The threshold calculation module first calculates the mean and standard deviation of each channel of the input EEG signal E, then combines the mean and standard deviation to obtain a new tensor u. This tensor is then input into a convolutional layer, where batch normalization is performed. The Sigmoid function is applied to map the values ​​to the (0,1) range, and the average value is calculated to obtain the spiking neural network threshold g, as shown in the following expression:

[0021] u=Concat(Mean(E),Std(E)) (1)

[0022] g=Avg(σ(BN(Conv(u)))) (2)

[0023] Where Mean represents calculating the mean, Std represents calculating the standard deviation, Concat represents connecting along the channel dimension, Conv represents a one-dimensional convolutional layer, BN represents a batch normalization layer, σ represents a sigmoid activation layer, Avg represents calculating the mean, and g represents the threshold of a spiking neural network.

[0024] Step 3: Input the preprocessed EEG signals and audio signals into the EEG branch and audio branch respectively.

[0025] The EEG branch is based on the Spike-CNN architecture, comprising a spiking layer and a channel attention module. To better utilize the temporal feature extraction capabilities of spiking neurons, the input EEG signal is divided into K consecutive segments, which are then sequentially input into the spiking layer for temporal feature extraction and encoding. The output of the spiking layer is then concatenated along the length dimension of the EEG signal. The output of the spiking layer is then fed into the channel attention module to capture the inter-channel relationships of the EEG signal.

[0026] Both the EEG and audio branches utilize LIF neurons in their pulse layers, with the membrane potential V of the l-th layer neurons at timett. t,l This can be expressed as equations (3) to (6):

[0027] V t,l =H t,l +Z t,l-1 (3)

[0028] z t,l =f(V t-1,l -g) (4)

[0029] H t,l =(αV t-1,l (1-Z) t,l-1 (5)

[0030] Z t,0 =E t (6)

[0031] Where f(·) is the step function; g is the membrane potential threshold; α is the leakage factor of the LIF neuron; E t This refers to the t-th segment of the EEG decision window or the speech envelope decision window; Z t,l H represents the output pulse of the l-th layer at timestamp t; t,l This indicates that the membrane potential of the l-th layer at timestamp t is updated based on the leakage factor and the output pulse.

[0032] The channel attention module in EEG branches adaptively assigns different weights to channels to capture the relationships between them. The expression is:

[0033] x′=x·σ(W2·ReLU(W1·AvgPoo12d(x))) (7)

[0034] x represents the input of the channel attention module, x′ represents the output of the channel attention module, σ represents the Sigmoid activation layer, W1 and W2 are the weight matrices of the fully connected layer, ReLU(·) represents the ReLU activation function, and AvgPool2d(·) represents global average pooling.

[0035] The audio branch contains a pulse layer that divides the input speech envelope into K consecutive segments aligned with the EEG signal. These segments are then sequentially input into the pulse layer for sequential time feature extraction and encoding. The outputs of the pulse layer are then concatenated along the length dimension of the speech envelope.

[0036] Step 4: Classify based on classifier

[0037] The outputs of the two audio branches are concatenated with the output of the EEG branch along the channel dimension to obtain features M1 and M2. M1 and M2 are then flattened into two-dimensional tensors and input into the feature fusion module to obtain fused audio and EEG features T1 and T2. These fused features T1 and T2 are concatenated along the channel dimension and then passed through a fully connected layer and a softmax activation function to obtain the classification result. The expression for the feature fusion module is as follows:

[0038] Y=MaxPool(ReLU(Conu2d(X))) (8)

[0039] Conv2d represents the convolution operation, where X is the input data, Y is the output data, ReLU represents the activation function applied to the result of the convolution operation, and MaxPool represents the adaptive max pooling operation, which adjusts the activated feature map to the specified output size.

[0040] This invention uses a threshold calculation module to provide a threshold for the pulse layer, and uses EEG and audio branches to process time-synchronized EEG signals and audio, respectively. The EEG branch uses the pulse layer to extract and encode temporal features, including a channel attention module for EEG; the audio branch extracts and encodes features from the two input audio signals separately using the pulse layer. Then, a feature fusion module is used to extract the fused features of the EEG and the two audio signals. Finally, a fully connected layer and softmax are used to complete the classification. It is important to note that the EEG and audio branches form a loop structure; the input EEG and audio signals need to be segmented into K consecutive segments, and then sequentially entered into the EEG and audio branches to perform the above steps.

[0041] The beneficial effects of this invention are as follows: This invention introduces spiking neural networks into the auditory attention detection task, simulating the information encoding mechanism of the human brain. This solves the problems of weak interpretability and poor real-time performance in current auditory attention detection models, achieving high accuracy with low latency and a short decision window. Compared with other methods, the method proposed in this invention has better generalization ability and stable detection results on different subjects. Furthermore, when used in conjunction with specialized neural hardware, spiking neural networks can significantly reduce the model's energy consumption, which is beneficial for hardware development and provides an effective method for developing brain-controlled hearing aids. Attached Figure Description

[0042] Figure 1 This is the network structure diagram of this method. Detailed Implementation

[0043] The present invention will be further described in detail below with reference to specific embodiments, but the present invention is not limited to the specific embodiments.

[0044] An auditory attention detection method based on spiking neural networks is proposed, which includes data preprocessing, network model training, and testing.

[0045] Experiments were conducted using the KUL dataset, which is widely used in this task. This dataset contains synchronized EEG signals and audio, and experiments were performed on multiple subjects, meeting the requirements of this invention. Preprocessing of both EEG and audio data was performed in Matlab: EEG data was high-pass filtered at 0.5Hz and downsampled to 128Hz. Further steps involved re-referencing the EEG data based on the Cz channel, followed by bandpass filtering. Audio was downsampled to 12kHz, then filtered using a Gammatone filter bank. The filtered signal underwent full-wave rectification and power-law compression, and the average of all filters was downsampled to 128Hz. A second-order Butterworth filter was then used for bandpass filtering to obtain the speech envelope. Before feeding the EEG signals and speech envelope into the network, the mean and variance of the EEG data for each subject were calculated, and the EEG data was standardized. To make the data more suitable for input to the pulse layer, both the EEG data and the speech envelope were normalized. To verify the detection results of this invention in a shorter time window, the EEG signal and speech envelope were segmented into continuous decision windows. The decision windows can take values ​​within 2 seconds, with no data overlap, and always maintain synchronization between the data.

[0046] The proposed network model was implemented in PyTorch, using cross-entropy as the loss function. The threshold calculation module was disabled during the initial N epochs of training, with the initial threshold set to 'a' as a guide. After N epochs, the threshold calculation module was enabled for M epochs. During the last b epochs of training the threshold-calculated model, the thresholds calculated by the module were recorded, and their mean was calculated. This mean was used as the fixed threshold for subsequent model training, where 0 ≤ N ≤ 100 and 0 ≤ M ≤ 100. Experiments showed that this strategy of initially using 'a' as the threshold guide and then using the mean threshold obtained after a certain training period as the fixed threshold for subsequent training exhibited good accuracy. The model was trained using a stochastic gradient descent (SGD) optimizer with a momentum of 0.9. The batch size was 16, the initial learning rate was 0.001, and the leakage factor for the spurious layer was set to 0.2.

[0047] The network model parameters are further explained below: The EEG signals in the dataset are 64 channels. During preprocessing, the EEG signals were resampled to 128Hz and truncated into 2-second decision windows for subsequent analysis, meaning each window contains 128 × 2 = 256 sampling points. The EEG signals are processed by a threshold calculation module to obtain the spiking neural network threshold. The expression for the threshold calculation module is as follows:

[0048] u=Concat(Mean(E),Std(E)) (1)

[0049] g=Avg(σ(BN(Con(u)))) (2)

[0050] Mean represents calculating the mean, Std represents calculating the standard deviation, Concat represents connecting along the channel dimension, Conv represents a one-dimensional convolutional layer, BN represents a batch normalization layer, σ represents a sigmoid activation layer, Avg represents calculating the mean, and g represents the threshold of the spiking neural network. Specifically, the Conv convolutional layer has a kernel size of 2 and 1 input / output channel; the BN layer is a one-dimensional batch normalization layer with 1 channel.

[0051] The preprocessed EEG signal and audio are then input into the EEG branch and audio branch, respectively. The spiking neural network threshold g calculated in the previous step is used in both branches. The EEG branch contains one spiking layer and one channel attention module. The input EEG signal is divided into K consecutive segments, which are then sequentially input into the spiking layer. The outputs of the spiking layer are concatenated along the length dimension of the EEG signal. The audio branch has only one spiking layer. Both the EEG and audio branches use LIF neurons in their spiking layers, with the membrane potential V of the l-th layer neuron at timett being... t,l This can be expressed as equations (3) to (6):

[0052] V t,l =H t,l +Z t,l-1 (3)

[0053] Z t,l =f(V t-1,l -g) (4)

[0054] H t,l =(αV t-1,l (1-Z) t,l-1 (5)

[0055] Z t,0 =E t (6)

[0056] Where f(·) is the step function; g is the membrane potential threshold; α is the leakage factor of the LIF neuron; E t This refers to the t-th segment of the EEG decision window or the speech envelope decision window; H t,l Z t,l These represent the membrane potential and output pulse of the l-th layer at timestamp t, respectively. In this embodiment, the input EEG signal or audio is divided into segments with four sampling points each.

[0057] The expression for the channel attention module in the EEG branch is:

[0058] x′=x·σ(W2·ReLU(W1·AvgPool2d(x))) (7)

[0059] x represents the input to the channel attention module, x′ represents the output of the channel attention module, σ represents the Sigmoid activation layer, W1 and W2 are the weight matrices of the fully connected layers, ReLU(·) represents the ReLU activation function, and AvgPool2d(·) represents global average pooling. AvgPool2d(·) is an adaptive two-dimensional average pooling operation that adjusts the height and width of the tensor to 1×1. The first linear layer has a 64-dimensional input and a 4-dimensional output, and the second linear layer has a 4-dimensional input and a 64-dimensional output.

[0060] Finally, classification is performed based on the classifier. The outputs of the two audio branches are concatenated with the output of the EEG branch along the channel dimension to obtain features M1 and M2. Then, M1 and M2 are input into the feature fusion module to obtain the fused audio and EEG features T1 and T2. The expression of the feature fusion module is as follows:

[0061] Y=MaxPool(ReLU(Conv2d(X))) (8)

[0062] Conv2d represents a convolution operation, where X is the input data, Y is the output data, and ReLU represents the activation function applied to the result of the convolution operation. MaxPool represents an adaptive max pooling operation, which adjusts the activated feature map to a specified output size. Specifically, Conv2d has a 65×9 kernel size, 1 input channel, and 16 output channels; MaxPool is an adaptive 2D max pooling operation that adjusts the height and width of the tensor to 1×2, where 2 represents a decision window of 2 seconds.

[0063] The fused features T1 and T2 are flattened into two-dimensional tensors, then concatenated along the channel dimension, and finally passed through a fully connected layer and a softmax activation function to obtain the classification result. The expression is as follows:

[0064] T x_reshaped =reshape(T) x , shape=(batch_size,-1)) (9)

[0065] R = Softmax(W3·T) reshaped (10)

[0066] R represents the final classification result, T x Indicates T1 or T2, T x_reshaped T represents the fused features after adjustment to a two-dimensional tensor. reshaped Represents the fusion feature T 1_reshape T 2_reshapedThe data is concatenated along the channel dimension. Softmax represents the softmax activation function, and W3 represents the weight matrix of the fully connected layer. The fully connected layer has a 48-dimensional input and a 2-dimensional output, and the batch_size is the batch size of 16 mentioned earlier.

Claims

1. A pulse-inspired auditory attention detection method, characterized in that: The synchronously acquired EEG signals and two audio signals, after preprocessing, first calculate the threshold of the spiking neural network through the EEG signals, and then input them into the EEG branch and the audio branch respectively; The EEG branch first extracts and encodes temporal features through the pulse layer, and then applies the channel attention module to capture the inter-channel relationships of the EEG signals; there is also an audio branch with two audio signals processed separately, where the audio signals extract and encode sequential temporal features through the pulse layer. The output of the EEG branch is concatenated with the outputs of the two audio branches along the channel dimension. The concatenated results are then fed into two feature fusion modules to obtain the fused features between the EEG and the two audio samples. A fully connected layer maps the fused features to an output dimension equal to the number of categories. Finally, a softmax function is used to obtain the classification result. The specific steps are as follows: Step 1: Preprocessing of EEG signals and audio The EEG signal was high-pass filtered at a cutoff frequency of 0.5 Hz, and the EEG signal was rereferenced to the EEG Cz channel, downsampled to 128 Hz and band-pass filtered. The audio is processed using an auditory filter with power-law compression to obtain the speech envelope, and finally downsampled to 128Hz and bandpass filtered. The EEG signals were standardized and normalized, and the speech envelope was normalized. A sliding window is used to segment the EEG signal and speech envelope into continuous segments, while maintaining synchronization between the EEG signal and the audio at all times; Step 2: Calculate the threshold of the spiking neural network based on EEG signals The threshold calculation module first calculates the mean and standard deviation of each channel of the input EEG signal E, and then combines the mean and standard deviation to obtain a new tensor. This tensor is then input into a convolutional layer, followed by batch normalization. The values ​​are then mapped to the (0,1) range using the Sigmoid function. Finally, the average value is calculated to obtain the spiking neural network threshold g, as shown in the following expression: Where Mean represents the calculation of the mean, Std represents the calculation of the standard deviation, Concat represents the concatenation along the channel dimension, Conv represents a one-dimensional convolutional layer, and BN represents a batch normalization layer. denoted as Sigmoid activation layer, Avg represents mean calculation, and g represents threshold of spiking neural network; Step 3: Input the preprocessed EEG signals and audio signals into the EEG branch and audio branch respectively. The EEG branch is based on the Spike-CNN architecture, which includes a spike layer and a channel attention module. The input EEG signal is divided into N consecutive segments, which are then sequentially fed into the spike layer for temporal feature extraction and encoding. The output of the spike layer is concatenated along the length dimension of the EEG signal. The output of the spike layer is then fed into the channel attention module to capture the inter-channel relationships of the EEG signal. Both the EEG and audio branches utilize LIF neurons in their pulse layers at time stamp t. l Membrane potential of layer neurons This can be expressed as equations (3) to (6): in, (·) represents the step function; g represents the membrane potential threshold; α represents the leakage factor of the LIF neuron; For the t-th segment of the EEG decision window or the speech envelope decision window; This indicates the time stamp t. l The membrane potential is updated based on the attenuation factor and the firing pulse; This indicates the time stamp t. The layer's output pulse; The expression for the channel attention module in the EEG branch is: This represents the input to the channel attention module. This represents the output of the channel attention module. Indicates the Sigmoid activation layer. and These are the weight matrices of the fully connected layer. This represents the ReLU activation function. (·) indicates global average pooling; The audio branch contains a pulse layer; the input speech envelope is divided into N consecutive segments aligned with the EEG signal, and then sequentially input into the pulse layer for sequential time feature extraction and encoding. The output of the pulse layer is then spliced ​​together along the length dimension of the speech envelope. Step 4: Classify based on classifier The outputs of the two audio branches are concatenated with the outputs of the EEG branch along the channel dimension to obtain features. and characteristics ,Then and Flattening the data into a two-dimensional tensor and inputting it into the feature fusion module yields the fused features of audio and EEG. , The classification result is then obtained through a fully connected layer and a softmax activation function; the mathematical expression for the feature fusion module is as follows: Conv2d represents the convolution operation, where X is the input data. This refers to the output data; ReLU represents the activation function applied to the result of the convolution operation; MaxPool represents the adaptive max pooling operation, which adjusts the activated feature map to the specified output size; The threshold calculation module is not enabled during the initial N epochs of training. The threshold is set to 'a' as the initial threshold to guide the training. After training for N epochs, the threshold calculation module is enabled to continue training for M epochs. During the last b epochs of training the threshold calculation model, the thresholds calculated by the threshold calculation module are recorded and their mean is calculated. The mean is used as the fixed threshold for subsequent model training, where 0≤N≤100 and 0≤M≤100.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method and system for accompanying robot

    CN113947127A

  • Selection and configuration of an automated robotic process

    US20210356941A1