Auditory attention detection method, device, computer equipment and readable storage medium

By converting EEG signals without speech stimulation into weighted pulsed EEG features, and using long-term memory network models for feature extraction and classification, the problems of feature mismatch and high computational cost in the prior art are solved, and efficient and accurate auditory attention detection is achieved.

CN119564208BActive Publication Date: 2025-05-13THE CHINESE UNIV OF HONG KONG (SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510136253.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-13
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

When detecting EEG signals without speech stimulation, the prior art has the problem of feature mismatch, which affects the accuracy of auditory attention detection and is highly computationally cost-effective.

Method used

An auditory attention detection method is used to obtain the EEG signals without speech stimulation, convert them into weighted pulsed EEG features in the attention model, and use the long-term memory network model to extract these features to obtain the hidden state of the EEG and classify it according to the hidden state to obtain auditory attention information.

Benefits of technology

This method can reduce calculation costs and avoid sequential processing of EEG signals on the basis of ensuring accuracy, thereby improving detection efficiency. Through the combination of attention model and long-term memory network model, the limitations of long-distance dependence are alleviated and the accuracy of signal recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119564208B_ABST
    Figure CN119564208B_ABST
Patent Text Reader

Abstract

The present application relates to an auditory attention detection method, device, computer equipment, computer-readable storage medium and computer program product. The method comprises: obtaining an EEG signal without speech stimulation; converting the EEG signal into a weighted pulse EEG feature in an attention model; extracting features of a sequence composed of the weighted pulse EEG features through a long short-term memory network model to obtain an EEG hidden state; classifying according to the EEG hidden state to obtain auditory attention information; the auditory attention information indicates the location of the sound source of the speech felt by the object to which the EEG signal without speech stimulation belongs. The use of this method can reduce the computational cost while ensuring accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of auditory detection technology, and in particular to an auditory attention detection method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] With the continuous development of EEG signal detection technology, EEG signals can be detected through neural network models to obtain corresponding auditory attention, thus realizing auditory attention detection under competing sound sources. In traditional technology, clean speech stimulation is used for model training.

[0003] However, such clean speech stimuli may not always be available at inference time, and although one can use speaker separation techniques to obtain single speech stimuli, there is a feature mismatch between training and inference, which may affect the accuracy of auditory attention detection. Summary of the invention

[0004] Based on this, it is necessary to provide an auditory attention detection method, device, computer equipment, computer readable storage medium and computer program product to address the above-mentioned technical problems, which can reduce computing costs while ensuring accuracy.

[0005] In a first aspect, the present application provides an auditory attention detection method, comprising:

[0006] Obtain EEG signals without speech stimulation;

[0007] Converting the EEG signal into a weighted pulse EEG feature in an attention model;

[0008] By using a long short-term memory network model, feature extraction is performed on the sequence composed of the weighted pulse EEG features to obtain the EEG hidden state;

[0009] Classification is performed according to the EEG hidden state to obtain auditory attention information; the auditory attention information represents the location of the sound source of the speech felt by the subject to whom the EEG signal without speech stimulation belongs.

[0010] In one embodiment, converting the EEG signal into a weighted pulse EEG feature in an attention model includes:

[0011] Determining a pulse event based on the current of the EEG signal, and obtaining an EEG pulse signal according to the pulse event;

[0012] Converting the EEG pulse signal to obtain a pulse EEG attention feature in an attention model;

[0013] The weighted pulse EEG features in the attention model are determined based on the pulse EEG attention features.

[0014] In one embodiment, the step of determining a pulse event based on the current of the EEG signal comprises:

[0015] According to the connection weight between the i-th neuron of the previous layer of neurons and the j-th neuron of the current layer of neurons, the EEG signal pulse of the i-th neuron is adjusted to obtain the weighted EEG signal pulse;

[0016] According to the constant current value injected by the current layer neurons into the jth neuron, the weighted EEG signal pulse is adjusted to obtain the input current of the jth neuron;

[0017] The membrane potential of the jth neuron at the previous time step is adjusted according to the leakage factor to obtain the adjusted membrane potential of the previous time step; the pulse value of the jth neuron at the previous time step is adjusted based on the pulse emission threshold to obtain the adjusted pulse value of the previous time step; the membrane potential of the jth neuron at the current time step is determined according to the difference value of the sum of the adjusted membrane potential of the previous time step and the input current of the jth neuron relative to the adjusted pulse value of the previous time step;

[0018] When the membrane potential of the j-th neuron at the current time step is greater than the emission threshold, a pulse event is obtained; i and j are both positive integers.

[0019] In one embodiment, the step of determining a pulse event based on the current of the EEG signal comprises:

[0020] The pulse linear projection is performed on the current of the EEG signal in each time step to obtain the pulse event.

[0021] In one embodiment, determining the weighted pulse EEG feature in the attention model based on the pulse EEG attention feature includes:

[0022] Based on the attention model without a normalization function and a scaling factor, determining a combination result of a query feature and a key feature contained in the pulse EEG attention feature;

[0023] Determining an EEG attention weight according to the combination result;

[0024] According to the EEG attention weight, the value feature contained in the pulse EEG attention feature is adjusted to obtain a weighted pulse EEG feature.

[0025] In one embodiment, the extracting the features of the sequence of weighted pulse EEG features by using a long short-term memory network model to obtain the EEG hidden state includes:

[0026] By using the long short-term memory network model, the weighted pulse EEG features of each time step are sequentially extracted in the order in which the weighted pulse EEG features are generated, so as to obtain the EEG hidden state of each time step;

[0027] The step of classifying according to the EEG hidden state to obtain auditory attention information includes:

[0028] Classification and identification are performed based on the number of pulses generated by the EEG hidden state to obtain auditory attention information.

[0029] In a second aspect, the present application also provides an auditory attention detection device, comprising:

[0030] An acquisition module, used to obtain EEG signals without speech stimulation;

[0031] An attention module, used for converting the EEG signal into a weighted pulse EEG feature in an attention model;

[0032] An extraction module, used for extracting features from the sequence of weighted pulse EEG features through a long short-term memory network model to obtain EEG hidden states;

[0033] A classification module is used to classify according to the EEG hidden state to obtain auditory attention information; the auditory attention information represents the sound source position of the speech felt by the object to which the EEG signal without speech stimulation belongs.

[0034] In a third aspect, the present application further provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of auditory attention detection in any of the above embodiments are implemented.

[0035] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of auditory attention detection in any of the above embodiments are implemented.

[0036] In a fifth aspect, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of auditory attention detection in any of the above embodiments are implemented.

[0037] The above-mentioned auditory attention detection method, device, computer equipment, computer-readable storage medium and computer program product obtain an EEG signal without speech stimulation, and on this basis, convert the EEG signal into a weighted pulse EEG feature in the attention model. Since the attention model can dynamically assign weights to different features of the input stimulation when running, and allocate appropriate computing resources to the EEG signal with the most information, it can not only reduce the computing cost, but also avoid the EEG signal being processed in sequence, thereby ensuring accuracy. Moreover, the attention model itself can capture the context distance of the EEG signal for a long time, which helps to alleviate the limitations of long-distance dependence to ensure the accuracy of signal recognition. At the same time, through the long short-term memory network model, the sequence composed of the weighted pulse EEG features is feature extracted to obtain the EEG hidden state. Since the long short-term memory network model controls the flow of information based on the cell state and the corresponding gate control mechanism, it can maintain information for a long time, which can help to alleviate the limitations of long-distance dependence to ensure the accuracy of signal recognition; finally, the EEG hidden state is classified to obtain the auditory attention information; the auditory attention information indicates the sound source position of the voice felt by the object of the EEG signal without speech stimulation. Therefore, for EEG signals without speech stimulation, the use of the attention model and the long short-term memory network model can accurately and efficiently detect auditory attention information. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0039] Figure 1 is an application environment diagram of an auditory attention detection method in one embodiment;

[0040] Figure 2 is a flow chart of an auditory attention detection method in one embodiment;

[0041] Figure 3 A schematic diagram of a process for obtaining weighted pulse EEG features in one embodiment;

[0042] Figure 4 A schematic diagram of the structure of an ST-LSTM model in one embodiment;

[0043] Figure 5 The figure is a schematic diagram of the accuracy effect of LSTM and T-LSTM models in one embodiment.

[0044] Figure 6 Schematic diagram of a mask form of attention weight in one embodiment;

[0045] Figure 7 is a structural block diagram of an auditory attention detection device in one embodiment;

[0046] Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0048] The auditory attention detection method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 through a network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers.

[0049] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car devices, projection devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. This application can be executed by the server 104 alone, or by the terminal 102 alone, or it can be implemented through the interaction process between the terminal 102 and the server 104.

[0050] In an exemplary embodiment, Figure 2 As shown, a method for detecting auditory attention is provided, and the method is applied to Figure 1 The server 104 in the example is used as an example to illustrate, including the following steps 202 to 208. Among them:

[0051] Step 202, obtaining an EEG signal without speech stimulation.

[0052] No speech stimulation refers to an environment in which the subject of the EEG signal is not subjected to speech stimulation. No speech stimulation means that no speech stimulation is used as a reference. In this case, the EEG signal can also be used to analyze brain activity patterns to identify the neural correlates of auditory attention in different brain regions, thereby reflecting the spatial location of the sound source that a person is paying attention to.

[0053] In this case, the temporal dynamic information of EEG signals can be captured by convolutional neural networks (CNN) and recurrent neural networks (RNN), and the traditional recurrent neural network (RNN) is expected to solve the limitations of convolutional neural networks (CNN) in capturing long-distance dependencies of EEG signals. However, it is worth noting that due to the inherent sequential nature of traditional recurrent neural networks (RNN), it involves a large number of parameters and high computational costs, which is not suitable for implementation on neural guided hearing aids. Therefore, it is necessary to use the attention model and the long short-term memory network model to cooperate with each other to balance accuracy and cost.

[0054] Step 204: convert the EEG signal into a weighted pulse EEG feature in the attention model.

[0055] The weighted pulse EEG feature is obtained by weighting the value features contained in the attention model based on the EEG attention weight.

[0056] The EEG attention weight is the weight corresponding to the value feature in the attention model. The attention model can be a neural network model using the self-attention mechanism, and the framework of the attention model can be a Transformer model. The EEG attention weight can be obtained by directly encoding the EEG signal into a vector and converting it according to the weight calculation method of the attention system.

[0057] Exemplarily, at least three weight matrices are used to perform a linear transformation on each eigenvector of the EEG signal to obtain a corresponding attention feature; the attention feature includes three features, namely, a query feature (Q), a key feature (K) and a value feature (V), each of which is obtained by converting the input brain signal feature using at least one corresponding weight matrix; the query feature and the key feature are dot-multiplied to obtain an inner product value between the query feature and the key feature; the inner product value is then scaled according to a scaling factor of the normalized inner product value in the attention model to obtain a scaled inner product value; the scaling factor of the normalized inner product value can be the number of dimensions of the key feature; the scaled inner product value is normalized to obtain an EEG attention weight; the value feature is adjusted by the EEG attention weight to obtain a weighted pulse EEG feature.

[0058] Optionally, when the EEG signal is converted into a discrete value, the weight calculation method in the attention model can be adjusted to convert the EEG signal in the adjusted manner to obtain the EEG attention weight. For example, when the EEG signal is converted into a discrete EEG pulse signal, the inner product value between the query feature and the key feature of the EEG pulse signal can be used as the EEG attention weight in the attention model to facilitate the corresponding calculation, thereby ensuring processing efficiency.

[0059] Because the attention model uses the inner product value between the query feature and the key feature to determine the EEG attention weight when it is running, the value feature can be adaptively adjusted based on the query feature and the key feature, so that weights can be dynamically assigned to different features of the input stimulus, so that each position can directly access all the information of the entire sequence during calculation, thereby capturing long-distance dependencies without relying on the order or structure of the input, and will not be limited by the time step in the traditional network, so that appropriate computing resources are allocated to the most informative EEG signals, thereby reducing computing costs and avoiding sequential processing of EEG signals to ensure accuracy. In addition, the attention model itself can capture the long context distance of EEG signals, which helps to alleviate the limitations of long-distance dependencies to ensure the accuracy of signal recognition.

[0060] Step 206, using the long short-term memory network model, feature extraction is performed on the sequence consisting of the EEG attention weights to obtain the EEG hidden state.

[0061] The Long Short-Term Memory (LSTM) network model is a recurrent neural network that uses long short-term memory units to extract dependencies between adjacent features in a sequence, which can alleviate the gradient vanishing or exploding problem of traditional recurrent neural networks. In this embodiment, the result output by the long short-term memory network model is the EEG hidden state. The EEG hidden state is the hidden state output by the long short-term memory network model, which is obtained by extracting weighted pulse EEG features by the long short-term memory network model. The EEG hidden state is the hidden state of the recurrent layer.

[0062] When processing sequence data, the long short-term memory network model also processes in the order of time steps. At each time step, based on the gating mechanism of the long short-term memory network model, as well as the current input and the EEG hidden state of the previous time step, the cell state is updated and the new EEG hidden state is output; the gating mechanism includes functions for implementing input gates, forget gates, and output gates. For example, based on the function of the forget gate, the information of the current cell state is deleted, and based on the function of the input gate, information is selected from the information input at the current time step in the sequence and stored in the current cell state, so as to obtain the updated current cell state through the functions of the forget gate and the input gate. Finally, based on the function of the output gate, the information in the current cell state is selected to obtain the EEG hidden state used to affect the current time step.

[0063] Using long short-term memory units to extract dependencies between adjacent features in a sequence means that the long short-term memory network model will sequentially generate EEG hidden states at each time step according to the sequence of weighted pulse EEG features. For example, the long short-term memory network model relies on the previous EEG hidden state generated previously and the information stored in the cell state at the previous time step to determine the next EEG hidden state. This dependency reflects the sequential nature of the long short-term memory network model.

[0064] Optionally, the number of cells in the long short-term memory network model may be equal to the number of time slices in the EEG sequence, and the number of cells in the long short-term memory network model may also be greater than the number of time slices in the EEG sequence.

[0065] Step 208, classify according to the EEG hidden state to obtain auditory attention information; the auditory attention information represents the location of the sound source of the speech felt by the subject of the EEG signal without speech stimulation.

[0066] Auditory attention detection information is the sound source location information reflected through the auditory dimension. The sound source location information is the location information of the sound source relative to the object to which the EEG signal belongs. Auditory attention detection information is a type of auditory attention detection information (Auditory Attention Detection, AAD). Auditory attention information is used to determine where a person focuses his or her auditory attention through EEG signals without speech stimulation.

[0067] The auditory attention detection information at least relates to the direction in the binary classification task, especially in the scenario where the subject of the EEG signal focuses on the sound on the left or right. Optionally, the auditory attention detection information can be represented in a probabilistic form; when the auditory attention detection information indicates the left side with a greater probability, the sound source position of the speech felt by the subject of the EEG signal is the left side of the subject; when the auditory attention detection information indicates the right side with a greater probability, the sound source position of the speech felt by the subject of the EEG signal is the right side of the subject.

[0068] In an optional embodiment, auditory attention information is obtained by classification based on EEG hidden states, including: inputting the EEG hidden states after global mean pooling processing into a fully connected layer with a normalization function to obtain a probability vector; the probability vector is used to indicate the left or right side of the object to which the EEG signal belongs.

[0069] In the above auditory attention detection method, an EEG signal without speech stimulation is obtained, and on this basis, the EEG signal is converted into a weighted pulse EEG feature in the attention model. Since the attention model can dynamically assign weights to different features of the input stimulation when running, and allocate appropriate computing resources to the EEG signal with the most information, it can not only reduce the computing cost, but also avoid the sequential processing of EEG signals, thereby ensuring accuracy. Moreover, the attention model itself can capture the long context distance of the EEG signal, which helps to alleviate the limitations of long-distance dependence to ensure the accuracy of signal recognition. At the same time, through the long short-term memory network model, the sequence composed of weighted pulse EEG features is feature extracted to obtain the EEG hidden state. Since the long short-term memory network model controls the flow of information based on the cell state and the corresponding gate control mechanism, it can maintain information for a long time, which can help to alleviate the limitations of long-distance dependence to ensure the accuracy of signal recognition; finally, the auditory attention information is obtained by classification according to the EEG hidden state; the auditory attention information represents the sound source position of the speech felt by the object of the EEG signal without speech stimulation. Therefore, for EEG signals without speech stimulation, the use of the attention model and the long short-term memory network model can accurately and efficiently detect auditory attention information.

[0070] In an exemplary embodiment, Figure 3 As shown, converting the EEG signal into the weighted pulse EEG feature in the attention model includes steps 302 to 306. Among them:

[0071] Step 302, determining a pulse event based on the current of the EEG signal, and obtaining an EEG pulse signal according to the pulse event.

[0072] The spiking neuron network is a neural network model that simulates the transmission of information by biological neurons. Simulating the transmission of information by biological neurons refers to the way of converting data such as the current value of the input current to identify the pulse event of the input current, and then encoding it through the corresponding pulse event to obtain the EEG pulse signal. Specifically, the neurons of the spiking neuron network can use neurons with real-valued activation functions, allowing the transmission of real-valued signals and pulse events calculated using real numbers. Multiple spiking neurons convert real-valued EEG signals into binary pulse sequences. The functions of these spiking neurons are similar to those of biological neurons.

[0073] A spike event is an event triggered by a time step when the input current meets certain conditions. For example, a spike event can be a time step when the input current has a brief and significant voltage change (EEG spikes), which can be a time step where a peak or a valley occurs.

[0074] A binary pulse sequence can be formed according to whether there is a pulse event at each time step; the binary pulse sequence is an EEG pulse signal. The EEG pulse signal is a binary pulse sequence. Since the EEG pulse signal is a low-frequency pulse time sequence, the EEG pulse signal can accurately retain the information conveyed by the EEG signal when it is input, and the EEG pulse signal is in the form of discrete values, which helps to reduce the calculation cost, so that the solution of this embodiment can be applied to mobile terminals such as hearing aids.

[0075] Step 304, converting the EEG pulse signal to obtain the pulse EEG attention feature in the attention model.

[0076] The pulse EEG attention feature is the attention feature obtained by converting the EEG pulse signal according to the attention model.

[0077] The pulse EEG attention feature can be a feature that directly encodes the pulse EEG signal into a vector, matrix, etc. For example, each feature vector of the pulse EEG signal is linearly transformed using at least three weight matrices to obtain the corresponding pulse EEG attention feature; the pulse EEG attention feature includes three features, namely, query feature (Q), key feature (K) and value feature (V).

[0078] Step 306, determining the weighted pulse EEG features in the attention model based on the pulse EEG attention features.

[0079] The weighted pulse EEG feature is a weighted feature of attention for discrete values. When an EEG pulse signal is obtained according to a pulse event, the EEG signal has been converted into a discrete value. At this time, the EEG attention weight in the attention model can be determined based on the inner product value between the query feature and the key feature contained in the pulse EEG attention feature, so that the corresponding EEG attention weight can be used to adjust the value feature contained in the pulse EEG attention feature, thereby ensuring processing efficiency.

[0080] It can be understood that when the weighted pulse EEG feature is obtained by weighting the pulse EEG attention feature, the above-mentioned attention model can be called a pulse temporal attention block, the above-mentioned long short-term memory network model can be called a pulse long short-term memory network model, and the above-mentioned model for classification can be called a pulse classifier block.

[0081] In this embodiment, pulse events are determined based on the input current of the EEG signal in the pulse neuron network, and the EEG pulse signal is obtained according to the pulse event. This can convert the continuous input current into an event-driven form, and convert the continuous input current into a discrete-value EEG pulse signal, thereby significantly reducing the required computing resources.

[0082] In an exemplary embodiment, a pulse event is determined based on the current of an EEG signal, including: adjusting the EEG signal pulse of the ith neuron according to the connection weight between the ith neuron of the previous layer of neurons and the jth neuron of the current layer of neurons to obtain a weighted EEG signal pulse; adjusting the weighted EEG signal pulse according to the constant current value injected into the jth neuron by the current layer of neurons to obtain the input current of the jth neuron; adjusting the membrane potential of the jth neuron at the previous time step according to the leakage factor to obtain the adjusted membrane potential of the previous time step; adjusting the pulse value of the jth neuron at the previous time step based on the pulse emission threshold to obtain the adjusted pulse value of the previous time step; determining the membrane potential of the jth neuron at the current time step according to the difference value between the sum of the adjusted membrane potential of the previous time step and the input current of the jth neuron and the adjusted pulse value of the previous time step; determining a pulse event when the membrane potential of the jth neuron at the current time step is greater than the emission threshold; i and j are both positive integers.

[0083] The connection weight is the weight parameter between neurons in adjacent layers in a spiking neuron network, which can be a preset weight or a dynamically changing one. The input current is the current input to the jth neuron, which is used to represent the signal to be processed at the current input of the jth neuron. The weighted EEG signal pulse is obtained by adjusting the EEG signal pulse of the neurons in the current layer based on the connection weight. j represents any neuron in the previous layer, and j represents any neuron in the current layer.

[0084] The leakage factor represents the decrease of membrane potential over time, and is used to simulate the characteristics of ion channels in neuronal cell membranes that produce leaking ions. By adjusting the membrane potential of the jth neuron at the previous time step according to the leakage factor, the membrane potential can be kept at a relatively stable level when the input current is small, thus achieving the effect of noise reduction.

[0085] The sum of the adjusted membrane potential at the previous time step and the input current of the jth neuron is used to reflect that the current neuron accumulates the input signals it receives over time and adjusts its membrane potential according to these inputs until the threshold is reached and a pulse is emitted.

[0086] The emission threshold is the threshold when a pulse event exists, and the emission threshold is used to adjust the pulse value of the jth neuron in the previous time step to obtain the adjusted membrane potential of the previous time step. If the jth neuron did not emit a pulse in the previous time step, the emission threshold has no effect on the membrane potential; if the jth neuron emitted a pulse in the previous time step, the membrane potential will reduce the emission threshold to achieve a reset of the membrane potential, ensuring that the neuron will only emit pulses again when it receives enough input current to accumulate the membrane potential above the threshold again. At the same time, since the pulse is in binary form, the pulse emission threshold is used to adjust the pulse value of the previous time step of the current layer of neurons. If there is a pulse event in the previous time step, the pulse value of the previous time step can be lowered according to the pulse emission threshold; if there is no pulse event in the previous time step, the pulse value of the previous time step can be retained, thereby simulating the refractory period of biological neurons after the action potential is emitted, so that the pulse value of each time step can be more accurate.

[0087] The difference value relative to the adjusted pulse value of the previous time step can be in the form of a difference value or adjusted based on the difference value, based on the sum of the adjusted membrane potential of the previous time step and the input current of the jth neuron. The difference value is used to simulate the reset behavior of the neuron after emitting a pulse to prevent continuous pulse events.

[0088] When the membrane potential of the jth neuron at the current time step is greater than the emission threshold, it is determined that the pulse event exists, and in this case, the current pulse value of the jth neuron is reset to the resting potential. Correspondingly, when the membrane potential of the jth neuron at the current time step is less than the emission threshold, it is determined that the pulse event does not exist.

[0089] Exemplarily, the above-mentioned step of determining the pulse event can be implemented by a neuron model, and the neuron model can be a leaky integrate-and-fire (LIF); in this case, the connection weight, leakage factor, and emission threshold are all parameters in the leaky integrate-and-fire neural network model.

[0090] In this embodiment, there are corresponding connection weights between the neurons in the previous layer and the neurons in the current layer, so that the input current of each layer of neurons is obtained by adaptively adjusting the neurons layer by layer, which helps to more accurately determine the weighted EEG signal pulses; at the same time, a constant current value is used as an additional input, so that each layer of neurons does not completely rely on the output results of the neurons in the previous layer, so that the growth rate of the membrane potential can reach the threshold of the trigger pulse. At the same time, based on the leakage factor, the noise reduction effect is achieved, and the pulse emission threshold is used to adjust the pulse value of the previous time step of the neurons in the current layer to simulate the refractory period of biological neurons after the action potential is released. On this basis, the membrane potential after leakage processing is added with the input current of the current time step, and then the adjusted pulse value of the previous time step is subtracted to simulate the reset behavior of the neuron after the pulse is released, so that the pulse value of each time step can be more accurate.

[0091] In an exemplary embodiment, determining a pulse event based on the current of the EEG signal includes: performing pulse linear projection on the current of the EEG signal in each time step to obtain a pulse event.

[0092] The pulse linear projection can be realized by the above-mentioned leaky integral and emission model, or it can be obtained by comparing the current of the EEG signal in each time step with the corresponding threshold. For example, when the input current of the current time step is greater than the current threshold, it is determined that the pulse event has occurred; when the input current of the current time step is less than the current threshold, it is determined that the pulse event has not occurred.

[0093] In this embodiment, pulse linear projection is performed on the current of the EEG signal in each time step to adaptively obtain pulse events, thereby using less computing resources.

[0094] In an exemplary embodiment, a weighted pulse EEG feature in an attention model is determined based on a pulse EEG attention feature, including: determining a combination result of a query feature and a key feature contained in the pulse EEG attention feature based on an attention model that does not contain a normalization function and a scaling factor; determining an EEG attention weight based on the combination result; and adjusting a value feature contained in the pulse EEG attention feature based on the EEG attention weight to obtain a weighted pulse EEG feature.

[0095] The query feature and the key feature are two features in the pulse EEG attention feature. The combination result of the two features is obtained by combining the values ​​of the two features, and is used to characterize the combination relationship of the two features. Optionally, the combination result can be the dot product result of the query feature and the key feature contained in the pulse EEG attention feature, or it can be the feature processing result in other attention models. Optionally, the combination result can be used as the EEG attention weight; or an operation without normalization processing can be performed on the combination result to obtain the EEG attention weight.

[0096] Exemplarily, the inner product value between the query feature and the key feature contained in the pulse EEG attention feature can be determined, and the inner product value is the dot product result of the query feature and the key feature contained in the pulse EEG attention feature; the inner product value is used as the EEG attention weight in the attention model.

[0097] In this embodiment, when the pulse attention feature is a discrete value, the normalization function and the scaling factor are omitted to facilitate the calculation process of the discrete value, and the normalization process and the scaling process used in the corresponding calculation process can be omitted to more accurately determine the combination result between the query feature and the key feature, thereby determining the EEG attention weight with fewer computing resources.

[0098] In an exemplary embodiment, a sequence of weighted pulse EEG features is subjected to feature extraction through a long short-term memory network model to obtain an EEG hidden state, including: extracting features of the weighted pulse EEG features of each time step in sequence according to the generation order of the weighted pulse EEG features through a long short-term memory network model to obtain an EEG hidden state of each time step;

[0099] Correspondingly, the auditory attention information is obtained by classifying according to the EEG hidden state, including: classifying and identifying the number of pulses generated based on the EEG hidden state to obtain the auditory attention information.

[0100] Time steps are ordered time periods. Based on whether there is a pulse at each time step, feature extraction can be performed in sequence to obtain the EEG hidden state of each time step in sequence.

[0101] The attention-weighted pulse EEG feature is the result of adjusting the value feature contained in the pulse EEG attention feature; wherein, the attention-weighted pulse EEG feature can still be a form of electroencephalogram. The number of pulses is the number of peaks of the EEG hidden state, and the number of peaks is the number of local maximum values ​​of the EEG hidden state.

[0102] In this embodiment, the original features of the attention model are replaced by the EEG hidden state at each time step, and the EEG hidden state at each time step is used to ensure the accuracy of classification and recognition under the premise of low computing cost; at the same time, in view of the fact that the EEG signal is an EEG pulse signal, the classification method of the original EEG hidden state is changed, and the number of pulses generated by the EEG hidden state at each time step is counted to classify and identify the corresponding pulses, so as to accurately obtain auditory attention information with less computing resources.

[0103] In one embodiment, the above steps 202 to 208 constitute a framework of a temporal long short-term memory network model (T-LSTM), which combines temporal attention (TA) with long short-term memory (LSTM) to effectively simulate the dynamic information of EEG signals. The temporal dynamics of EEG signals (EEG) within a time window are modeled to achieve accurate AAD decision-making.

[0104] T-LSTM consists of three components: temporal attention module, LSTM module and classifier. The input EEG signal is defined as E∈ , T represents the number of samples within the window, and N is the number of channels. By utilizing this EEG as input current, T-LSTM is able to predict the spatial location of the sound source of interest using a binary decision (i.e., left or right).

[0105] In the temporal attention module, attention regulation plays a key role in various human cognitive processes. In selective listening, auditory attention effectively filters out irrelevant speech stimuli, allowing individuals to pay special attention to the speech of interest. Inspired by this idea, the attention mechanism has been implemented in deep learning models, which dynamically assigns weights to different samples of input stimuli at runtime. The mechanism has demonstrated its effectiveness in allocating appropriate computing resources to the most informative signal samples. Given that the brain's response to speech stimuli evolves over time, the temporal attention mechanism is a natural choice for decoding auditory attention in EEG signals. In this embodiment, a three-step process is used to implement the temporal attention mechanism to implement step 204.

[0106] In the first step, the EEG or other EEG signals are input into the attention model through linear transformation, that is, E is projected onto the query ( ),key( ) and value ( ), its expression is as follows:

[0107]

[0108] in, ; Represents the ReLU activation function.

[0109] Step 2: Query ( ) and key( ) is calculated by dot product and softmax function as follows:

[0110]

[0111] in, is the temporal attention mask, i.e., the attention weight in step 204 above; is the scaling factor of the normalized inner product value, ie, the scaling factor in the above step 204 .

[0112] Step 3: Temporal Attention Mask is used to dynamically assign values ​​to the EEG signal ( ) to obtain the weighted sum of attention , i.e., the weighted pulse EEG feature in the above step 204, the expression of the weighted pulse EEG feature is as follows:

[0113]

[0114] For the above step 206, a long short-term memory module (LSTM module) is used to learn the temporal dynamics of Ea. The LSTM architecture contains long short-term memory units, which is intended to alleviate the problem of gradient disappearance or explosion in traditional RNN. In this embodiment, the same number of LSTM layer cells as the number of time slices in the EEG sequence is used. The output of LSTM is the EEG hidden state of the recurrent layer, and the expression of LSTM is as follows;

[0115]

[0116]

[0117] in, is the number of hidden nodes.

[0118] Finally, the above step 208 is implemented based on the AAD classifier. The T-LSTM model is an end-to-end solution that uses the AAD classifier as the backend to identify auditory attention. Specifically, the output generated by the LSTM is received by the classifier and then assigned to the appropriate output category. In this embodiment, the binary cross entropy loss is used as the loss function for training until the model converges.

[0119] A global average pooling layer is applied on the time axis of the EEG features, followed by an fc layer with a sigmoid activation function to average the EEG features. Convert to probability vector , which is expressed as follows:

[0120]

[0121] It is to define the weights and biases of the fc layer;

[0122] is the Sigmoid activation function;

[0123]

[0124]

[0125] in, is the label of the mth decision window, and M is the batch size.

[0126] Further, after the above steps 302-306 and the corresponding embodiments are applied to the above step 204, a spike temporal LSTM (ST-LSTM) is formed. In this way, the three components in T-LSTM are redefined in ST-LSTM as a spike temporal attention block, a spike LSTM module, and a spike classifier block. The processing process is as follows: Figure 4 As shown in the figure, ST-LSTM, including a spike time attention mechanism, a spike LSTM module and a classifier. Taking EEG data as input, ST-LSTM decodes auditory attention through a binary decision process. ST-LSTM is a spike implementation of T-LSTM designed specifically for AAD tasks. The model encodes brain signals as spikes and processes them in an event-driven manner, which significantly reduces computational overhead.

[0127] On the basis of T-LSTM, neurons with real-valued activation functions are used to allow the transmission of real-valued signals and to calculate the spike events of T-LSTM using real numbers, namely ST-LSTM, and multiple spike neurons are used to convert real-valued EEG signals into binary spike sequences. The function of these spike neurons is similar to that of biological neurons. Considering that EEG signals are low-frequency spike time sequences, this conversion is expected to preserve the input EEG signals well while generating discrete-valued signals, namely EEG spike signals, for low-cost calculations. The ST-LSTM model learns temporal information and extracts discriminative features in a biologically reasonable way. Then, the ST-LSTM model uses a spike time attention mechanism to assign different weights to EEG spikes, followed by a spike LSTM module that captures the temporal features of the spike EEG data, and finally a spike AAD classifier that detects auditory attention by making binary decisions.

[0128] The above-mentioned spiking neuron model may be a leaky integrate-and-fire (LIF) neuron model, which can balance biological accuracy and computational efficiency.

[0129] At time step t, the membrane potential of LIF neuron j in layer l is It can be expressed as follows:

[0130]

[0131] Where λ is the leakage factor, represents the emission threshold, The current that corresponds to neuron j received from its predecessor neuron. is the connection weight between the pre-spiking neuron i and the post-spiking neuron j, represents the constant current injected into neuron j in layer l.

[0132] For the case of a spike event generated by a LIF neuron or not, the expression is as follows:

[0133]

[0134] In LIF neurons, when the membrane potential of neuron j in layer l Reaching the emission threshold When a spike event occurs at time t, the neuron's membrane potential is reset to its resting potential. , and the neuron enters a refractory period of a specific duration. In this embodiment, real-valued inputs are used as time-varying input currents, which are directly applied to the following expressions at the first time step:

[0135] .

[0136] Pulse counting plays a bridge role between ANN and SNN. Specifically, the number of pulses counted in neuron i of layer l is For example, the expression used by LIF neurons to determine the number of spikes is as follows:

[0137]

[0138] in, Represents the total number of time steps.

[0139] Furthermore, the aforementioned spike temporal attention module is an innovative implementation of the temporal attention mechanism in SNNs, which has great potential in processing sequence data. By leveraging the inherent ability of SNNs to encode temporal information through spike timing and extract key temporal features through the attention mechanism, it is able to capture complex temporal patterns from EEG.

[0140] Query ,key Sum The input currents are determined by performing a linear projection of the pulse at each time step.

[0141]

[0142] where (·) represents a spike linear function that encodes the EEG input into a spike train.

[0143] Then, the vector It can be defined as follows:

[0144]

[0145] Next, we calculate the query vector and key vector The relationship between. Used to get The expression of input current is as follows:

[0146]

[0147] It is worth noting that the softmax function and scaling factors are removed to facilitate spike calculation.

[0148] Then, the EEG spikes are masked by the spike timing attention mask at each time step Dynamic weighting, thus producing attention-weighted pulse EEG, is expressed as , the attention weighted pulse EEG is a weighted pulse EEG feature in step 304, and the expression of the attention weighted pulse EEG is as follows:

[0149]

[0150] Finally, using the expression used by the LIF neuron to determine the number of spike counts, we obtain the following expression for the spike count of the EEG:

[0151]

[0152] The spiking LSTM module is designed to process sequential data by combining the temporal properties of spiking neurons with the memory retention capabilities of traditional LSTM. In practice, the input current of the spiking LSTM module will be defined by the attention-weighted spiking EEG at each time step:

[0153]

[0154] in, is the attention-weighted pulse EEG After the attention-weighted pulse EEG is input into the pulse long short-term memory module, the expression of LSTM is executed to output the EEG hidden state.

[0155] like Figure 4 As shown, the spike AAD classifier follows the ST-LSTM model and is trained in an end-to-end manner. The auditory attention decision is output using the SNN classifier. Specifically, the spike count C generated by the spike LSTM is received by the classifier and then assigned to one of the output categories. In this embodiment, the binary cross entropy loss is used as the learning target.

[0156] A global average pooling layer is applied on the time axis of the EEG pulse features, followed by an fc layer with a sigmoid activation function to average the EEG pulse features. Convert to probability vector .

[0157]

[0158] are the weights and biases that define the fc layer, and σ(·) is the sigmoid activation function.

[0159] The loss function is the same as the loss function in the above step 208, but the parameters need to be changed into probability vectors for adaptive adjustment.

[0160]

[0161] Here, M represents the batch size and C' is the spike count output by the SNN classifier.

[0162]

[0163]

[0164] in is the label of the mth decision window, M is the batch size, and C' is the spike count output by the SNN classifier.

[0165] In one embodiment, the effectiveness of the method is verified by experimental data. In order to verify the effectiveness of the temporal feature representation, a T-LSTM model is first trained on the publicly available AAD dataset. In order to compare the T-LSTM model with other previous CNN-based AAD methods, a ST-LSTM model is further trained and compared with T-LSTM in terms of accuracy and computational cost.

[0166] During the use process, there is a selection process for a specific dataset: In order to facilitate comparative analysis, experiments were conducted on the KUL dataset. This dataset has been widely used in previous studies to explore attention mechanisms from EEG signals. Making it a reliable benchmark for evaluating new AAD methods. Specifically, the dataset contains 64-channel EEG signals from 16 individuals who claim to have normal hearing. In the experiment, participants were instructed to pay attention to one of two competing speakers. The stimuli included multiple Dutch stories told by different male speakers. Each experiment involved playing two auditory streams, with the position 90 degrees to the left and right of the listener. The position of the target speaker changed randomly throughout the experiment. In the original experiment on the KUL dataset, a total of 72 minutes of EEG data were collected for each subject. However, the last 24 minutes of data consisted of repeated recordings of the same story, which were excluded from the analysis. Therefore, approximately 48 minutes of EEG data were used for each subject in this embodiment.

[0167] During the experiment, there was a specific dataset preparation process: the EEG signal was first bandpass filtered from 1 to 32 Hz and then further downsampled to 128 Hz to match the frequency range used in previous AAD studies. To ensure consistency, the EEG signal in each trial was normalized. Subsequently, the EEG data was segmented into shorter durations, called decision windows, using a sliding window method, and the overlapping parts (i.e., the tail of the training window may end in the test set) were called repeated segments. In order to maintain the uniqueness of the training, validation, and test sets, all repeated segments were excluded.

[0168] Given that humans can shift their attention from one speaker to another in less than 2 seconds, developing low-latency AAD solutions is critical for real-world applications. Therefore, this example pays special attention to low-latency settings and uses decision windows of different lengths, including 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, and 2 seconds. After preprocessing, a total of 5752 decision windows were obtained for each subject under the 1-second condition, with a total of 92,032 decision windows.

[0169] During the experiment, the training and evaluation process includes:

[0170] For both T-LSTM and ST-LSTM, the AAD performance was evaluated in a subject-dependent manner using a 5-fold cross validation (CV) technique. The accuracy of AAD was calculated as the percentage of correct decisions in all decision windows. For the SNN simulations, all features were encoded using short time windows of 10 time steps. The discrete and non-differentiable nature of SNNs poses challenges when directly applying the error back-propagation method during training. To overcome this issue, a technique called tandem learning (TL) is used to train the ST-LSTM model. The tandem learning technique combines an SNN with a coupled artificial neural network (ANN) for parameter optimization. The coupled ANN acts as an auxiliary structure that allows error back-propagation at the spike train level and facilitates the training of the SNN.

[0171] For the coupled neural network, taking a 1-second decision window as an example, the proposed ST-LSTM transforms EEG data (i.e. 128 samples, 64 channels) as input. The temporal attention module consists of three linear layers and =4, and get the output In the LSTM module, the hidden size = 4. Therefore, the output of the LSTM layer is . Then, a global average pooling layer is applied along the time dimension. The data is flattened into a one-dimensional vector as input to the fc layer (input: 4, output: 2) to decode auditory attention. The hyperparameters are selected by grid search on the validation set. During training, the network is updated using the Adam optimization technique with a learning rate of To prevent overfitting and improve generalization, dropout and batch normalization techniques are used. It is worth noting that ST-LSTM has the same hyperparameters as the coupled ANN. Emission threshold It was empirically adjusted to 0.7.

[0172] After the experimental process, the results include multiple aspects:

[0173] First, the proposed LSTM and previous AAD methods based on CNN models are compared. Then, ablation analysis is performed using a decision window of 1 second to verify the effectiveness of the temporal attention mechanism. The AAD performance of two models (LSTM and LSTM with temporal attention, referred to as T-LSTM hereafter) is compared. Finally, the performance of spiking implementation, especially the ST-LSTM model, is evaluated under low latency settings (from 0.1 seconds to 2 seconds).

[0174] A. LSTM VS CNN Decoder

[0175] For a fair comparison, the experimental setup of the CNN-based AAD model in

[24] was replicated. The CNN and LSTM models were fed with the same EEG windows as input. In brief, the CNN model architecture consisted of a convolutional layer with a 64x17 kernel, followed by average pooling and two (fc) layers (input: 5, hidden layers: 5, output: 2). The hyperparameters of the CNN model were tuned following the same procedure as the LSTM model in

[24] .

[0176] The LSTM model showed superior performance to the CNN model, with improvements of 1.4% and 2.5% in decision windows of 1 second and 2 seconds, respectively. Notably, the LSTM model consistently outperformed the CNN model within a decision window of 1 second. Furthermore, a paired t-test was performed on the performance of both models at all decision window lengths, yielding statistically significant results (p=0.0017, paired t-test). In terms of model complexity, the CNN contains approximately 5,500 parameters, as reported by Van-DeCappelle et al. While the LSTM model has a more efficient design with only 1,200 parameters. These findings highlight the potential of LSTM as a powerful tool for analyzing and interpreting time-related patterns on EEG data, thereby contributing to its outstanding performance in the AAD task, as shown in Table 1.

[0177] Table 1

[0178]

[0179] B. Temporal LSTM

[0180] like Figure 5 Figure 2. AAD performance of LSTM and T-LSTM models across 16 subjects with a decision window of 1 second. Subjects are ranked by T-LSTM performance. Statistically significant differences: ***p<0.001. Thus, the AAD performance of the LSTM model and the LSTM with temporal attention is reported with a decision window of 1 second across all subjects. The LSTM model achieved an average AAD accuracy of 85.5% with a standard deviation of 5.87%. The T-LSTM model improved on average by 5.4% over the LSTM model with an average accuracy of 90.9% (standard deviation of 4.76%). These consistent improvements were observed across all subjects. Furthermore, there was a significant difference in AAD accuracy between the LSTM model and the T-LSTM (p<0.001, paired t-test). These results highlight the effectiveness of temporal attention in assigning dynamic weights to EEG samples and promote the exploration of the temporal dynamics of attention in EEG signals.

[0181] This ablation analysis demonstrates the advantage of integrating temporal attention with the LSTM framework and verifies the effectiveness of the proposed T-LSTM approach in modeling temporal dynamics in EEG-based auditory attention detection.

[0182] B. Spike Timing LSTM

[0183] To evaluate the impact of pulse implementation, we turn to the ST-LSTM model. The ST-LSTM model was trained with decision windows of different lengths, namely 0.2, 0.5, and 1. The average AAD accuracy and subject-based AAD accuracy of the ST-LSTM model at different decision window sizes are presented. When using a decision window of 1 second, the ST-LSTM model achieved an accuracy of 86.2% (SD: 4.41%), while a decision window of 2 seconds achieved an accuracy of 88.1% (SD: 422%). This is consistent with previous findings that the AAD accuracy generally decreases with decreasing decision window length. Specifically, for a decision window of 0.5 seconds, the ST-LSTM model achieved an accuracy of 82.9% (SD: 4.78%). However, it is worth noting that promising results were obtained with a decision window of 0.2 seconds, achieving an accuracy of 79.6% (SD: 5.13%).

[0184] To further evaluate the adaptability and robustness of ST-LSTM, it was tested under subject-independent conditions. Different from previous studies that mainly focused on subject-dependent scenarios, subject-independent AAD provides promising implications for BCI applications. Specifically, the model was trained using data from all participants of the KUL dataset, excluding one subject in each iteration. Taking a 1-second decision window as an example, the average accuracy of ST-LSTM was 73.2% (standard deviation: 10.3%). These results demonstrate the competitive performance of ST-LSTM even in subject-independent environments, indicating its ability to dynamically learn temporal information from EEG signals.

[0185] In summary, this example highlights the effectiveness of the ST-LSTM model in accurately decoding auditory attention. In addition, incorporating spiking neurons into the FER model has potential advantages in terms of resource allocation and energy efficiency.

[0186] The accuracy (%) of different models under different decision window lengths is shown in Table 2. Table 2 is as follows:

[0187] Table 2

[0188]

[0189] It is observed that the T-LSTM achieves 85.2% (standard deviation: 6.26%) accuracy in a 0.2 second decision window, 87.7% (standard deviation: 5.83%) accuracy in a 0.5 second decision window, 90.9% (standard deviation: 4.76%) accuracy in a 1 second decision window, and 92.3% (standard deviation: 4.51%) accuracy in a 2 second decision window. Notably, the T-LSTM model achieves competitive decoding accuracy (>85%) with an EEG signal duration of only 200 milliseconds (0.2 seconds). In contrast to two recent studies, the T-LSTM model achieves 85.2% (standard deviation: 6.26%) accuracy in a 0.2 second decision window.

[0190] Comparing 2 recently proposed AAD models, namely CSP-based and RGC-based models, the T-SIM model shows significant advantages, with an average accuracy improvement of more than 11% across all decision window sizes. In addition, T-LSTM achieves state-of-the-art (SOTA) results across all decision windows, surpassing other competitive baselines such as STAnet (paired t-test, p<0.05). It is worth noting that the T-LSTM and STAnet models are comparable in size, both containing about 5,000 parameters. These findings suggest that the T-LSTM model has great potential in fast and accurate decoding of auditory (spatial) attention. Regarding the peak implementation of T-LSTM, namely ST-LSTM, the results in Table II show that ST-LSTM outperforms CSP and RGC models. On average, the ST-LSTM model improves AAD accuracy by 6.2% and 6.3% across four different decision windows. In addition, when compared with other spiking models (such as NI-AAD model, spiking CNN), the ST-LSTM model improves AAD accuracy by 6.3%. and SGCN (Pulse GCN) over various decision window sizes, ST-LSTM consistently shows superior performance. Although ST-LSTM is not as competitive as STAnet and T-LSTM models in terms of AAD accuracy, it benefits from event-driven information processing. This role feature enhances its scalability and suitability for deployment on resource-constrained platforms. Specifically, the ST-LSTM model shows significant advantages in computational efficiency, and the computational cost (PJ) comparison of ST-LSTM and T-LSTM models is shown in Table 3, which is as follows:

[0191] Table 3

[0192]

[0193] B. Comparison of Computational Costs

[0194] We now further compare the computational cost between the proposed spiking implementation, i.e., the ST-LSTM model, and traditional neural networks, especially the T-LSTM model. The overall computational cost of a model is usually measured by the number of floating-point operations per second (FLOPs) required.

[0195] The ST-LSTM model has a significant low power consumption characteristic, which is achieved by activating neurons when a certain number of input pulses exceed a preset threshold. This approach can put inactive neurons into low power mode, thereby achieving significant energy savings. The calculations within the ST-LSTM model run in an event-driven manner, using binary pulse processing, where pulses are represented as 1 or 0. This approach significantly reduces the computational requirements to only floating point additions. In contrast, the T-LSTM model is computationally inefficient due to the need to perform floating point additions and multiplications in each multiply-accumulate (MAC) operation. As shown in Table I, the pulse implementation of has significantly reduced computational cost compared to the LSTM implementation in the KUL database. It is worth noting that the ST-LSTM model has an average computational cost reduction of 88.7% compared to the T-LSTM model.

[0196] In summary, traditional EEG-based AAD architectures face the challenge of high computational cost, which makes them impractical in resource-limited devices such as neural-guided hearing aids. In contrast, our novel spiking AAD model provides a solution that delivers superior energy efficiency while requiring less computational resources and processing power. This increase in efficiency allows the model to scale effectively and run on resource-constrained platforms such as neuromorphic hardware or edge devices. Therefore, the ST-LSTM model creates new possibilities for practical applications that require low-power and efficient learning algorithms in the real world.

[0197] C. Bio-Visualization

[0198] To better understand the proposed AAD model, visualizations of temporal attention masks generated by T-LSTM and ST-LSTM models using 1-second decision windows are shown. Analysis of these attention masks provides valuable insights into how the temporal attention mechanism modulates EEG signal dynamics to identify key samples that contribute significantly to the AAD task.

[0199] Figure 6 Attention maps of five randomly selected subjects are shown. Figure 6In the figure, (a)-(e) are the attention maps generated by T-LSTM. (f)-(j) are the attention maps generated by ST-LSTM. The attention maps of five randomly selected subjects are shown here. The color of the cell indicates the weight, and lighter colors indicate higher weights. Note that the attention maps of the T-LSTM and ST-LSTM models belong to the same EEG decision window. It is worth noting that the attention maps of the T-LSTM and ST-LSTM models correspond to the same EEG decision window. Since each 1-second decision window contains 128 EEG samples, the dimension of the attention mask is 128x128. As shown Figure 6 As shown in (a)-(e) in Figure 3, different EEG samples are assigned different attention weights, which highlights the effectiveness of temporal attention in dynamically regulating EEG input. In contrast, the attention mask generated by the ST-LSTM model, such as Figure 6 The contained (f)-(j) are different. Compared with the T-LSTM model, Figure 6 The included (f)-(j) show a significant reduction in salient EEG samples. This observation suggests that the ST-LSTM model provides a more energy-efficient solution, requiring less computational resources and processing power than the T-LSTM model.

[0200] This application provides valuable insights into the application of the ST-LSTM model in EEG-based AAD. However, further research is needed to improve the robustness and generalizability of the model before it can be integrated into hearing aids. Future investigations should include experiments using various EEG datasets and evaluate the performance of AAD under different noise levels to better assess the adaptability of the model in real-world settings. Although the proposed ST-LSTM model exhibits energy efficiency compared to traditional ANNs, its performance may face challenges when scaling to larger datasets or handling more complex tasks. To address these limitations, hybrid approaches that combine ST-LSTM with other advanced architectures should be explored to combine their advantages.

[0201] This example introduces the ST-LSTM model, a spike temporal attention model for capturing temporal information in EEG signals in AAD tasks. Results show that the proposed ST-LSTM model performs well in low-latency scenarios and is an advanced approach that surpasses traditional techniques. In addition, the ST-LSTM model has significant advantages in low power consumption, making it a practical choice for integration into neural-guided hearing aids and other BCI systems. Furthermore, the visualization of the ST-LSTM provides new insights into the temporal dynamics of neural activity for selective hearing. Overall, the ST-LSTM model provides superior performance, energy efficiency, and a deeper understanding of neural processes, indicating its great potential in smart healthcare applications.

[0202] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time step, but can be executed at different time steps, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0203] Based on the same inventive concept, the embodiment of the present application also provides an auditory attention detection device for implementing the auditory attention detection method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more auditory attention detection device embodiments provided below can refer to the limitations of the auditory attention detection method above, and will not be repeated here.

[0204] In an exemplary embodiment, Figure 7 As shown, a device for detecting auditory attention is provided, comprising:

[0205] The acquisition module 702 is used to acquire the EEG signal without speech stimulation;

[0206] An attention module 704, used to convert the EEG signal into a weighted pulse EEG feature in an attention model;

[0207] An extraction module 706 is used to extract features from the sequence of weighted pulse EEG features through a long short-term memory network model to obtain an EEG hidden state;

[0208] The classification module 708 is used to classify according to the EEG hidden state to obtain auditory attention information; the auditory attention information represents the sound source position of the speech felt by the subject of the EEG signal without speech stimulation.

[0209] In one embodiment, the attention module 704 is used to:

[0210] Determining a pulse event based on the current of the EEG signal, and obtaining an EEG pulse signal according to the pulse event;

[0211] Converting the EEG pulse signal to obtain a pulse EEG attention feature in an attention model;

[0212] The weighted pulse EEG features in the attention model are determined based on the pulse EEG attention features.

[0213] In one embodiment, the attention module 704 is used to:

[0214] According to the connection weight between the i-th neuron of the previous layer of neurons and the j-th neuron of the current layer of neurons, the EEG signal pulse of the i-th neuron is adjusted to obtain the weighted EEG signal pulse;

[0215] According to the constant current value injected by the current layer neurons into the jth neuron, the weighted EEG signal pulse is adjusted to obtain the input current of the jth neuron;

[0216] The membrane potential of the jth neuron at the previous time step is adjusted according to the leakage factor to obtain the adjusted membrane potential of the previous time step; the pulse value of the jth neuron at the previous time step is adjusted based on the pulse emission threshold to obtain the adjusted pulse value of the previous time step; the membrane potential of the jth neuron at the current time step is determined according to the difference value of the sum of the adjusted membrane potential of the previous time step and the input current of the jth neuron relative to the adjusted pulse value of the previous time step;

[0217] When the membrane potential of the j-th neuron at the current time step is greater than the emission threshold, a pulse event is obtained; i and j are both positive integers.

[0218] In one embodiment, the attention module 704 is used to:

[0219] The pulse linear projection is performed on the current of the EEG signal in each time step to obtain the pulse event.

[0220] In one embodiment, the attention module 704 is used to:

[0221] Based on the attention model without a normalization function and a scaling factor, determining a combination result of a query feature and a key feature contained in the pulse EEG attention feature;

[0222] Determining an EEG attention weight according to the combination result;

[0223] According to the EEG attention weight, the value feature contained in the pulse EEG attention feature is adjusted to obtain a weighted pulse EEG feature.

[0224] In one embodiment, the extraction module 706 is used to: extract the weighted pulse EEG features of each time step in sequence according to the generation order of the weighted pulse EEG features through the long short-term memory network model to obtain the EEG hidden state of each time step;

[0225] Correspondingly, the classification module 708 is used to: perform classification and identification based on the number of pulses generated by the EEG hidden state to obtain auditory attention information.

[0226] Each module in the above-mentioned auditory attention detection device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0227] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an auditory attention detection method is implemented.

[0228] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0229] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0230] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0231] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0232] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0233] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0234] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0235] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for detecting auditory attention, characterized in that: The method comprises: Acquiring an EEG signal without speech stimulation; wherein the absence of speech stimulation means that speech stimulation is not used as a reference; Determine a pulse event based on the current of the EEG signal, and obtain an EEG pulse signal according to the pulse event; convert the EEG pulse signal to obtain a pulse EEG attention feature in an attention model; determine a weighted pulse EEG feature in the attention model based on the pulse EEG attention feature; By using the long short-term memory network model, the weighted pulse EEG features of each time step are sequentially extracted in the order in which the weighted pulse EEG features are generated, so as to obtain the EEG hidden state of each time step; Classification and identification are performed based on the number of pulses generated by the EEG hidden state to obtain auditory attention information; the auditory attention information represents the location of the sound source of the speech felt by the object to which the EEG signal without speech stimulation belongs.

2. The method according to claim 1, characterized in that The current-based pulse event determination of the EEG signal comprises: According to the connection weight between the i-th neuron of the previous layer and the j-th neuron of the current layer, the EEG signal pulse of the i-th neuron is adjusted to obtain a weighted EEG signal pulse; According to the constant current value injected by the current layer neurons into the jth neuron, the weighted EEG signal pulse is adjusted to obtain the input current of the jth neuron; The membrane potential of the jth neuron at the previous time step is adjusted according to the leakage factor to obtain the adjusted membrane potential of the previous time step; the pulse value of the jth neuron at the previous time step is adjusted based on the pulse emission threshold to obtain the adjusted pulse value of the previous time step; the membrane potential of the jth neuron at the current time step is determined according to the difference value of the sum of the adjusted membrane potential of the previous time step and the input current of the jth neuron relative to the adjusted pulse value of the previous time step; When the membrane potential of the j-th neuron at the current time step is greater than the emission threshold, a pulse event is obtained; i and j are both positive integers.

3. The method according to claim 1, characterized in that The current-based pulse event determination of the EEG signal comprises: The pulse linear projection is performed on the current of the EEG signal in each time step to obtain the pulse event.

4. The method according to claim 1, characterized in that: The step of determining the weighted pulse EEG feature in the attention model based on the pulse EEG attention feature comprises: Based on the attention model without a normalization function and a scaling factor, determining a combination result of a query feature and a key feature contained in the pulse EEG attention feature; Determining an EEG attention weight according to the combination result; According to the EEG attention weight, the value feature contained in the pulse EEG attention feature is adjusted to obtain a weighted pulse EEG feature.

5. An auditory attention detection device, characterized in that: The device comprises: An acquisition module, used for acquiring EEG signals without speech stimulation; the absence of speech stimulation means that speech stimulation is not used as a reference; An attention module is used to determine a pulse event based on the current of the EEG signal, and obtain an EEG pulse signal according to the pulse event; convert the EEG pulse signal to obtain a pulse EEG attention feature in the attention model; and determine a weighted pulse EEG feature in the attention model based on the pulse EEG attention feature; An extraction module is used to extract the weighted pulse EEG features of each time step in sequence according to the generation order of the weighted pulse EEG features through a long short-term memory network model to obtain the EEG hidden state of each time step; A classification module is used to perform classification and identification based on the number of pulses generated by the hidden state of the electroencephalogram to obtain auditory attention information; the auditory attention information represents the location of the sound source of the speech felt by the object to which the electroencephalogram signal without speech stimulation belongs.

6. The device according to claim 5, characterized in that The attention module is used to: According to the connection weight between the i-th neuron of the previous layer and the j-th neuron of the current layer, the EEG signal pulse of the i-th neuron is adjusted to obtain a weighted EEG signal pulse; According to the constant current value injected by the current layer neurons into the jth neuron, the weighted EEG signal pulse is adjusted to obtain the input current of the jth neuron; The membrane potential of the jth neuron at the previous time step is adjusted according to the leakage factor to obtain the adjusted membrane potential of the previous time step; the pulse value of the jth neuron at the previous time step is adjusted based on the pulse emission threshold to obtain the adjusted pulse value of the previous time step; the membrane potential of the jth neuron at the current time step is determined according to the difference value of the sum of the adjusted membrane potential of the previous time step and the input current of the jth neuron relative to the adjusted pulse value of the previous time step; When the membrane potential of the j-th neuron at the current time step is greater than the emission threshold, a pulse event is obtained; i and j are both positive integers.

7. The device according to claim 5, characterized in that The attention module is used to: The pulse linear projection is performed on the current of the EEG signal in each time step to obtain the pulse event.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Auditory attention detection method based on pulse inspiration

    CN119046753A

  • Hybrid CNN-SNN architecture for EEG-based auditory attention detection

    WO2024107619A1