EEG auditory attention detection method, system, terminal and medium based on contrastive learning and multi-task learning
By combining contrastive learning and multi-task learning, a cross-subject auditory attention detection model was constructed, which solved the problem of insufficient generalization ability caused by differences in EEG signal data across subjects, and achieved more efficient auditory attention detection and intelligent adaptability of hearing aid equipment.
Patent Information
- Application Number
- CN202510586850.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The large differences in EEG signal data across subjects lead to insufficient model generalization ability, affecting the decoding performance of cross-subject tasks.
A cross-subject auditory attention detection model was constructed using a method based on contrastive learning and multi-task learning. Data was collected through a multi-channel EEG device. After preprocessing, common features were extracted using an EEG temporal feature encoder, a contrastive feature encoder, and a classifier. The cross-entropy loss and contrastive loss were optimized to improve the model's adaptability and decoding performance.
The model's adaptability and decoding accuracy among different subjects are improved, the hearing assistance effect of hearing aid devices in noisy environments is enhanced, the needs of individual differences are adapted, and the universality and robustness of the equipment are improved.
Smart Images

Figure CN120105168B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of physiological signal processing, and in particular to an EEG auditory attention detection method, system, terminal and medium based on contrastive learning and multi-task learning. Background Art
[0002] The "cocktail party effect" refers to the ability of humans to selectively focus on the voice of a particular speaker when there are multiple speakers in a crowd, effectively ignoring other background noise. This ability enables humans to quickly identify important information in a noisy environment without being distracted by other sounds in the environment. The latest neuroscience research shows that the brain exhibits characteristics of lateral auditory spatial attention when processing target sounds and competing sounds from different spatial locations, which can enhance the perception of target sounds and suppress the interference of background noise. Therefore, auditory attention is a neural activity that can be decoded by EEG signals, which is called auditory attention detection. As a non-invasive, non-invasive and efficient technology, EEG signals have been widely used in auditory attention detection, promoting the development of many innovative methods and systems.
[0003] Auditory attention detection tasks based on EEG signals can generally be divided into two types: within-subject tasks and cross-subject tasks. Within-subject tasks refer to training a specific model for each individual subject. This model can focus on the unique EEG characteristics of a specific subject during training, and thus can usually achieve a higher decoding accuracy. However, the application scope of within-subject tasks is relatively limited and cannot adapt to the individual differences and needs of different users. In contrast, cross-subject tasks refer to training and testing data from different subjects, focusing on evaluating the generalization ability of the auditory attention detection system and its ability to adapt to different individual differences. The importance of cross-subject tasks lies in that in actual applications, personalized devices such as hearing aids have the ability to adapt to different users, and there is no need to adjust the device separately for each user. Therefore, the study of cross-subject tasks is of great significance to improving the universality and convenience of equipment, and is one of the key technologies to improve the effectiveness of hearing assistive devices.
[0004] Although decoding accuracy has achieved good results in the within-subject task, decoding performance in the cross-subject task still faces a significant decline. This is mainly due to the large differences in EEG signal data between different subjects, resulting in insufficient generalization of the model, which in turn affects decoding performance.
[0005] Therefore, relevant technologies still need to be improved and developed. Summary of the Invention
[0006] The main purpose of this application is to provide an EEG auditory attention detection method, system, terminal and medium based on contrastive learning and multi-task learning, aiming to solve the problem in related technologies that the EEG signal data between different subjects are quite different, resulting in insufficient generalization ability of the model, thereby affecting the decoding performance across test tasks.
[0007] A first aspect of an embodiment of the present application provides an EEG auditory attention detection method based on contrastive learning and multi-task learning, and the EEG auditory attention detection method based on contrastive learning and multi-task learning includes the following steps: obtaining EEG data of multiple subjects; constructing a cross-subject auditory attention detection model based on contrastive learning and multi-task learning, and training the cross-subject auditory attention detection model based on the EEG data to obtain a trained cross-subject auditory attention detection model; obtaining the user's actual EEG data, inputting the actual EEG data into the trained cross-subject auditory attention detection model to obtain the user's target auditory attention direction, and controlling the hearing aid device to enhance the sound source in the target auditory attention direction according to the target auditory attention direction.
[0008] Optionally, in one embodiment of the present application, the EEG data is a target EEG signal; and obtaining the EEG data of multiple subjects specifically includes: obtaining original EEG signals of multiple subjects collected by a multi-channel EEG device; and preprocessing the original EEG signals to obtain target EEG signals.
[0009] Optionally, in one embodiment of the present application, the preprocessing of the EEG signal to obtain a target EEG signal specifically includes: re-referencing the original EEG signal to the average response of all electrodes to obtain a first EEG signal; band-pass filtering the first EEG signal within a preset Hz range to obtain a second EEG signal; down-sampling the second EEG signal to a target Hz to obtain a third EEG signal; and performing Z-value normalization on the third EEG signal to obtain a target EEG signal.
[0010] Optionally, in one embodiment of the present application, the cross-subject auditory attention detection model includes an EEG time feature encoder, an EEG contrast feature encoder, an EEG feature refinement module and a classifier; the training of the cross-subject auditory attention detection model based on the EEG data to obtain a trained cross-subject auditory attention detection model specifically includes: inputting the target EEG signal into the cross-subject auditory attention detection model; the EEG time feature encoder extracts time features from the target EEG signal; the EEG contrast feature encoder extracts contrast features from the target EEG signal and obtains contrast loss based on the contrast features; the EEG feature refinement module interacts the time features and the contrast features to obtain refined features; the classifier obtains a binary cross entropy loss based on the refined features; and the binary cross entropy loss and the contrast loss are optimized to obtain a trained cross-subject auditory attention detection model.
[0011] Optionally, in one embodiment of the present application, the EEG time feature encoder includes a two-dimensional convolutional layer, a position encoding layer and a multi-head self-attention layer; the EEG time feature encoder extracts time features from the target EEG signal, specifically including: the two-dimensional convolutional layer captures the local time features corresponding to the target EEG signal; the position encoding layer obtains the position information corresponding to the local time features; the multi-head self-attention layer obtains the time features based on the local time features and the position information.
[0012] Optionally, in one embodiment of the present application, the EEG feature refinement module interacts the time feature and the contrast feature to obtain a refined feature, specifically including: the EEG feature refinement module obtains a cross-subject shared feature based on the contrast feature, and performs dynamic feature fusion on the time feature based on the cross-subject shared feature to obtain a refined transition feature; the EEG feature refinement module uses a residual connection to add the time feature and the refined transition feature to obtain a refined feature.
[0013] Optionally, in one embodiment of the present application, the classifier obtains a binary cross entropy loss based on the refined features, specifically including: the classifier classifies the refined features to obtain a predicted label; and uses a binary cross entropy loss function to compare the predicted label with the true label to obtain a binary cross entropy loss.
[0014] A second aspect of the embodiments of the present application further provides an EEG auditory attention detection system based on contrastive learning and multi-task learning, wherein the EEG auditory attention detection system based on contrastive learning and multi-task learning includes:
[0015] An EEG data acquisition module, used to acquire EEG data of the subject;
[0016] A detection model construction and training module, configured to construct a cross-subject auditory attention detection model based on contrastive learning and multi-task learning, and train the cross-subject auditory attention detection model according to the EEG data to obtain a trained cross-subject auditory attention detection model;
[0017] The auditory attention direction decoding module is used to obtain the user's actual EEG data and input the actual EEG data into the trained cross-subject auditory attention detection model to obtain the user's target auditory attention direction, so that the hearing aid device can enhance the sound source in the target auditory attention direction.
[0018] The third aspect of an embodiment of the present application also provides a terminal, wherein the terminal includes: a memory, a processor, and an EEG auditory attention detection program based on contrastive learning and multi-task learning stored in the memory and runnable on the processor, wherein the EEG auditory attention detection program based on contrastive learning and multi-task learning, when executed by the processor, implements the steps of the EEG auditory attention detection method based on contrastive learning and multi-task learning as described above.
[0019] The fourth aspect of an embodiment of the present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an EEG auditory attention detection program based on contrastive learning and multi-task learning, and when the EEG auditory attention detection program based on contrastive learning and multi-task learning is executed by a processor, the steps of the EEG auditory attention detection method based on contrastive learning and multi-task learning as described above are implemented.
[0020] Beneficial effects: The present application provides an EEG auditory attention detection method, system, terminal and medium based on contrastive learning and multi-task learning. Through data collection, model establishment and training, and model application, the present application extracts more common features among subjects in cross-subject tasks, thereby improving the adaptability and decoding performance of the model, thereby achieving the effect of improving the decoding ability across cross-subject tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 This is a flowchart of a preferred embodiment of the EEG auditory attention detection method based on contrastive learning and multi-task learning of the present application;
[0023] Figure 2This is a flowchart of preprocessing of EEG signals in a preferred embodiment of the EEG auditory attention detection method based on contrastive learning and multi-task learning of the present application;
[0024] Figure 3 This is a diagram of the EEG auditory attention detection model framework in a preferred embodiment of the EEG auditory attention detection method based on contrastive learning and multi-task learning of the present application;
[0025] Figure 4 This is a diagram of the EEG temporal feature encoder framework of the auditory attention detection model in a preferred embodiment of the EEG auditory attention detection method based on contrastive learning and multi-task learning of the present application;
[0026] Figure 5 This is a diagram of the EEG contrast feature encoder framework of the auditory attention detection model in a preferred embodiment of the EEG auditory attention detection method based on contrastive learning and multi-task learning of the present application;
[0027] Figure 6 This is a structural diagram of a preferred embodiment of the EEG auditory attention detection system based on contrastive learning and multi-task learning of the present application;
[0028] Figure 7 This is a structural diagram of a preferred embodiment of the terminal of this application.
[0029] Description of reference numerals:
[0030] 100. Electrical data acquisition module; 200. Detection model construction and training module; 300. Auditory attention direction decoding module. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solutions and effects of this application clearer and more specific, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. The described embodiments are only possible technical implementations of this application and are not all possible implementations. Based on the embodiments in this application, those skilled in the art can fully combine the embodiments of this application to obtain other embodiments without creative work, and these embodiments are also within the scope of protection of this application.
[0032] The following describes the EEG auditory attention detection method, system, terminal and medium based on contrastive learning and multi-task learning of the embodiment of the present application with reference to the accompanying drawings. In view of the problem that the EEG signal data between different subjects in the above-mentioned related technology are quite different, resulting in insufficient generalization ability of the model, thereby affecting the decoding performance across test tasks, the present application provides an EEG auditory attention detection method based on contrastive learning and multi-task learning. In this method, through data collection, model establishment and training, and model application, more common features between subjects are extracted in cross-subject tasks, thereby improving the adaptability and decoding performance of the model, and achieving the effect of improving the decoding ability across test tasks. Thus, it solves the technical problem that the EEG signal data between different subjects in the related technology are quite different, resulting in insufficient generalization ability of the model, thereby affecting the decoding performance across test tasks.
[0033] This application can solve the problem of insufficient decoding accuracy in existing cross-subject tasks. This application introduces contrastive learning to enhance feature consistency between different subjects, and combines multi-task learning to improve the generalization ability of the system. It can extract features that are more common between subjects and improve decoding performance in cross-subject tasks. The application of this application will provide more efficient hearing assistance services for hearing-impaired individuals, promote the application and development of auditory attention detection technology in intelligent hearing aids, and provide strong support for solving the universality problem of personalized hearing devices.
[0034] This application effectively combines contrastive learning and multi-task learning techniques, significantly improving the performance and universality of auditory attention detection. Specifically, the EEG signals collected by a multi-channel EEG device are first preprocessed to remove noise and perform signal enhancement, and then the processed EEG signals are input into a cross-subject auditory attention detection model based on contrastive learning and multi-task learning. This model can efficiently extract cross-subject shared features with strong generalization capabilities from EEG signals, and use these features to achieve high-precision auditory attention decoding on new subjects. The decoded auditory attention direction is used as input to control the audio processing function of the hearing aid device, enhance the target sound source, and thus improve the user's listening experience in a noisy environment. The cross-subject auditory attention detection method of this application breaks through the limitations of related technologies, has better universality and robustness, can adapt to the differentiated needs of different individuals, and is widely used in various personalized devices, improving the intelligence and adaptability of hearing aids and other hearing-assistive devices.
[0035] The following specific embodiments are used to describe the technical solution of the present application in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0036] The EEG auditory attention detection method based on contrastive learning and multi-task learning described in the preferred embodiment of this application is as follows: Figure 1 As shown, the EEG auditory attention detection method based on contrastive learning and multi-task learning includes the following steps:
[0037] In step S101 , EEG data of multiple subjects are acquired.
[0038] It should be noted that this application combines contrastive learning and multi-task learning techniques to assist training. The temporal characteristics of EEG signals can reflect the activity changes of brain regions and the interactions between regions. By extracting universal and effective temporal features, it is possible to identify local patterns, timing changes, and synchronous activities between brain regions in EEG signals, thereby more accurately decoding the direction of auditory attention across subjects and meeting the application scenarios and needs of different subjects.
[0039] In one possible implementation, the EEG data is a target EEG signal. Raw EEG signals from multiple subjects are acquired using a multi-channel EEG device. The raw EEG signals are preprocessed to obtain the target EEG signal. The EEG signals from multiple subjects are collected using a multi-channel EEG device and then preprocessed.
[0040] Specifically, EEG data from multiple subjects is collected and preprocessed using a multi-channel EEG device. During the collection process, the data is ensured to reflect distinct differences in EEG activity under different auditory tasks. The preprocessing step ensures data quality and consistency, providing high-quality input for subsequent model training.
[0041] It is worth noting that this application collects EEG data from multiple subjects through multi-channel EEG equipment. The collection of these data is to reflect the differences in brain activity under different auditory tasks, so that they can be used to train models to identify the direction of auditory attention; the use of multi-channel EEG equipment can capture the electrical activity of multiple areas of the brain and provide more comprehensive information; the data comes from multiple subjects, which is to build a cross-subject model, that is, the model can be generalized to new, untrained subjects.
[0042] In one possible implementation, the original EEG signal is re-referenced to the average response of all electrodes to obtain a first EEG signal; the first EEG signal is band-pass filtered within a preset Hz range to obtain a second EEG signal; the second EEG signal is downsampled to a target Hz (for example, 128 Hz or other values) to obtain a third EEG signal; and the third EEG signal is Z-normalized to obtain a target EEG signal.
[0043] Specifically, such as Figure 2As shown in the figure, in the preprocessing step of the EEG signal, the original multi-channel EEG signal is re-referenced to the average response of all electrodes to improve the signal quality and reduce the deviation of the reference electrode, making the signal more stable and reliable; the EEG signal is band-pass filtered at 1-50 Hz to remove low-frequency drift and high-frequency noise to ensure that the signal retains valid information related to auditory attention, because the components related to auditory attention in the EEG signal are mainly concentrated in this frequency range; the filtered EEG signal is downsampled to 128 Hz to reduce the amount of data, thereby speeding up subsequent processing while retaining sufficient signal details. Downsampling to a suitable frequency can reduce computational complexity while maintaining information integrity; the downsampled signal is Z-score normalized to ensure that the data mean is 0 and the standard deviation is 1 to achieve consistency and comparability of data from different subjects. Normalization helps to eliminate signal amplitude differences due to individual differences between different subjects, so that the model can more easily learn common features. It is understood that the Z-score, also known as the standard score or Z value, is a standardization method widely used in statistics, which expresses the distance of a value from the mean in units of standard deviation.
[0044] In step S102, a cross-subject auditory attention detection model based on contrastive learning and multi-task learning is constructed, and the cross-subject auditory attention detection model is trained according to the EEG data to obtain a trained cross-subject auditory attention detection model.
[0045] In one possible implementation, the cross-subject auditory attention detection model includes an EEG temporal feature encoder, an EEG contrast feature encoder, an EEG feature refinement module, and a classifier. The EEG temporal feature encoder extracts temporal features from the target EEG signal; the EEG contrast feature encoder extracts contrast features from the target EEG signal and obtains contrast loss based on the contrast features; the EEG feature refinement module interacts the temporal features with the contrast features to obtain refined features; the classifier obtains binary cross entropy loss based on the refined features; the binary cross entropy loss and the contrast loss are optimized to obtain a trained cross-subject auditory attention detection model. The cross-subject auditory attention detection model based on contrastive learning and multi-task learning is trained to effectively extract cross-subject shared features for decoding the auditory attention direction of new subjects.
[0046] Specifically, such as Figure 3As shown in the figure, a cross-subject auditory attention detection model is established: Leveraging convolutional neural networks, self-attention mechanisms, cross-attention mechanisms, contrastive learning, and multi-task learning, a cross-subject auditory attention detection model based on EEG signals is constructed. Auditory attention directions and target EEG signals are used as training data for the detection model, which efficiently identifies and extracts common EEG features across subjects from EEG signals.
[0047] It should be noted that a two-dimensional convolutional layer and a self-attention mechanism are used to extract EEG temporal features; a contrastive loss function is used to maximize the similarity of features within the same label and minimize the differences between EEG features of different subjects; a cross-head attention mechanism is used to refine features, combining temporal and contrastive features; a classifier is used to classify the optimized features and decode auditory attention direction; and cross-entropy loss and contrastive loss are optimized simultaneously to improve model performance. This application can efficiently extract cross-subject shared features with strong generalization capabilities from the target EEG signal and use them to decode auditory attention direction.
[0048] In one possible implementation, the EEG temporal feature encoder includes a two-dimensional convolutional layer, a position encoding layer, and a multi-head self-attention layer. The two-dimensional convolutional layer captures the local temporal features corresponding to the target EEG signal; the position encoding layer obtains the position information corresponding to the local temporal features; and the multi-head self-attention layer obtains the temporal features based on the local temporal features and the position information.
[0049] Specifically, such as Figure 4 As shown, the input preprocessed multi-channel EEG signal E , using a two-dimensional temporal convolutional layer to capture local temporal features E t ;Introduce the absolute position encoding layer P , providing location information for each time step; applying the multi-head self-attention mechanism to capture the interaction between time steps and assigning relevant weights to each time step to obtain the time feature E a The two-dimensional temporal convolution layer is the basis of feature extraction, providing local temporal features; the position encoding layer enhances the model's ability to capture temporal dependencies; and the multi-head self-attention mechanism further captures the global relationship between time steps.
[0050] In one possible implementation, the EEG feature refinement module obtains cross-subject shared features based on the contrast features, and performs dynamic feature fusion on the time features based on the cross-subject shared features to obtain refined transition features; the EEG feature refinement module uses residual connection to add the time features and the refined transition features to obtain refined features.
[0051] Specifically, such as Figure 5 As shown, the input preprocessed multi-channel EEG signal E ;Use 2D temporal convolution layer to extract temporal features E c Through contrastive learning, the similarity of features within the same label is maximized while minimizing the differences between EEG features across subjects. The InfoNCE contrastive loss function is used for optimization. Two-dimensional temporal convolutional layers are again used for feature extraction, but here the focus is on contrastive learning. Contrastive learning aims to learn universal EEG features and enhance the model's cross-subject capabilities. The InfoNCE loss function is used to guide the contrastive learning process.
[0052] Specifically, the input is the temporal feature from the EEG temporal feature encoder E a and features from EEG contrast feature encoder E c ; Use the cross-head attention mechanism to make E a and E c interact with each other; E c The cross-subject shared features captured in further refine the temporal features E a ; Use residual connections to avoid key information loss and obtain the final refined features F The cross-multi-head attention mechanism realizes the interaction and refinement of the two parts of features; the residual connection ensures the integrity of information during the feature extraction process.
[0053] Furthermore, the temporal features output by the EEG temporal feature encoder are E a As a query, the contrast feature output by the EEG contrast feature encoder E c It serves as both key and value, enabling the model to actively use time features to "query" relevant information in contrast features; it uses a 4-head attention mechanism for parallel computing, with each attention head independently learning different feature interaction patterns. E c The shared features across subjects are captured by contrastive learning. These features contain common patterns that are discriminative between different subjects. E a When querying these shared features, it is equivalent to calibrating the temporal features in the universal feature space to make it contain more subject-invariant information; the features output by each attention head are concatenated and linearly transformed to obtain a refined feature representation (dynamic feature fusion); in order to avoid feature loss, the original temporal features are converted into E aAdded to the refined features, this residual structure ensures that the model can learn new interactive features during the feature refinement process while retaining important information in the original temporal features.
[0054] In a possible implementation, the classifier classifies the refined features to obtain a predicted label; and a binary cross entropy loss function is used to compare the predicted label with the true label to obtain a binary cross entropy loss.
[0055] Specifically, input refined features F The prediction results are processed through two fully connected layers and optimized using a binary cross-entropy loss function. Global average pooling reduces feature dimensionality and avoids overfitting. The fully connected layers perform the classification task. The binary cross-entropy loss function guides the classifier training process.
[0056] It is worth noting that the EEG temporal feature encoder and the EEG contrast feature encoder work in parallel, extracting temporal features and contrast features from the original EEG signal, respectively. The EEG feature refinement module interacts and refines the output features of the two encoders to obtain a higher-level feature representation. The classifier classifies the refined features to achieve decoding of auditory attention direction. The entire model adopts a multi-task learning strategy to improve the model's decoding performance on known and unknown subject data by simultaneously optimizing cross-entropy loss and contrast loss. By combining multiple steps such as temporal feature extraction, contrastive learning, and feature refinement, the model aims to effectively extract features related to auditory attention from preprocessed multi-channel EEG signals. Through the multi-task learning strategy, the model can not only achieve good decoding results on known subject data, but also enhance its generalization ability on unknown subject data. This design makes the model more suitable for cross-subject tasks and improves the accuracy and robustness of auditory attention detection.
[0057] It should be noted that refined features are comprehensive feature representations formed after multi-stage information fusion and interaction, which can simultaneously retain temporal dynamic characteristics, cross-individual commonalities, and enhance nonlinear correlations between features. F It is through the multi-head attention mechanism to analyze the temporal features E a and contrasting features E c Refined features generated by deep interaction F The time feature E a and contrasting features E cCross-attention interaction is carried out to achieve joint time-space optimization, injecting common spatial patterns (contrast features) into temporal features to enhance the spatial interpretability of temporal features; and individual-group information fusion is achieved, introducing group commonality constraints while retaining individual-specific temporal features. Refined features achieve multimodal feature fusion through the attention mechanism, breaking through the limitations of focusing only on a single dimension of time or space, and resolving the contradiction between individual specificity and group commonality in cross-subject tasks. It can provide hearing aids with a solution for mapping neural activity patterns to acoustic parameters, enabling more accurate analysis of the neural dynamics of attention focus shifts when processing complex auditory scenarios (such as multi-person conversation environments), thereby achieving real-time intelligent control of hearing aid gain.
[0058] Furthermore, in the construction and training process of the cross-subject auditory attention detection model: the EEG temporal feature encoder is used to extract EEG temporal features from the preprocessed multi-channel EEG signals, which includes a two-dimensional convolution layer, a position encoding layer and a multi-head self-attention layer. First, a two-dimensional temporal convolution layer is used to capture local temporal features. E t :
[0059] ;
[0060] in, E t Represents the local time features after two-dimensional time convolution and activation function processing, E ∈R N×T Represents the preprocessed multi-channel EEG signal, N is the number of channels (i.e., the number of electrodes), T is the time step of the EEG signal (i.e., the number of sampling points on each channel), R is a set of real numbers, and k is the number of convolution kernels (filters) (which determines the number of features extracted). Conv2d(⋅) indicates a two-dimensional convolution operation using k filters (convolution kernel size = N×5, meaning a temporal convolution with a width of 5 is applied simultaneously to N channels), ELU(⋅) indicates the Exponential Linear Unit activation function, and Reshape(⋅) indicates that R k×1×T Reshape into R T×k Reshape operation.
[0061] Next, an absolute position encoding layer is introduced to provide position information for each time step, thereby enhancing the model's ability to capture temporal dependencies. Subsequently, a multi-head self-attention mechanism is applied to effectively capture the interactions between time steps and assign relevant weights to each time step. The formula is as follows:
[0062] ;
[0063] in, E aRepresents the time features after processing by the multi-head self-attention mechanism, MultiHeadAttention() represents the 4-head self-attention mechanism; P ∈R T×k Represents the absolute position encoding layer, which is a E t A matrix of the same dimensions that provides position information for the time steps.
[0064] The EEG contrast feature encoder introduces contrastive learning, which aims to maximize the similarity of features within the same label while minimizing the differences between EEG features of different subjects. Especially in cross-subject tasks, through contrastive learning, the model can learn common EEG features, thereby avoiding feature drift caused by individual differences, making the model more capable of decoding new subjects. First, a two-dimensional temporal convolutional layer is used to extract the temporal features of the EEG. E c , the formula is as follows:
[0065] ;
[0066] in, E c Represents the EEG temporal features extracted by another two-dimensional temporal convolution layer for comparative learning, and E t have the same dimensions;
[0067] Next, contrastive learning will E c Compare the results with multiple positive and negative samples from the training set. Positive samples consist of EEG segments with the same label from different subjects, while negative samples come from EEG segments with different labels. The InfoNCE contrastive loss function (which maximizes the similarity between positive samples and minimizes the similarity between negative samples) is used, with a temperature scaling factor of 0.1 to control the sharpness of the feature vector similarity values.
[0068] The EEG feature refinement module refines features through a multi-head attention mechanism. Specifically, E a As a query, the EEG contrast feature encoder E c As keys and values. The cross-attention mechanism enables these two parts of features to interact with each other. E c The cross-subject shared features captured in further refine the temporal features E a In order to avoid losing key information during feature extraction, this module uses residual connections to obtain the final refined features.F :
[0069] ;
[0070] in, F It represents the refined features after the cross-multi-head attention mechanism is refined, and CrossMultiHeadAttention(,) represents the cross-multi-head attention mechanism.
[0071] The classifier is used for classification tasks, and the optimized features F A global average pooling operation is applied, and then processed through two fully connected layers (with 64 and 2 units respectively). The binary cross entropy loss function is used for optimization, and the formula is as follows:
[0072] ;
[0073] in, L BCE Represents binary cross entropy loss, which is used for optimization of classification tasks; y It's an EEG signal E The true label of (i.e., the true value of the auditory attention direction); y’ is the probability of the auditory attention direction predicted by the model.
[0074] The cross-subject auditory attention detection model adopts a multi-task learning strategy to improve model performance by simultaneously optimizing two tasks: first, decoding the auditory attention direction through cross-entropy loss, and second, learning more common features across different subjects through contrastive loss. In this way, the model can not only achieve good decoding results on known subject data, but also enhance its generalization ability on unknown subject data, improving decoding performance across test tasks. The total loss is calculated as the weighted sum of binary cross-entropy loss and contrastive loss, as follows:
[0075] ;
[0076] in, L Total represents the total loss, which is the weighted sum of binary cross entropy loss and contrast loss; γ is a temperature scaling factor; L InfoNCE represents the InfoNCE contrast loss.
[0077] The Adam optimizer (with adaptive learning rate adjustment feature for model training) is used for model training, and the learning rate is set to 1×10 -3The total training time is 50 epochs (indicating the number of full dataset iterations during training, i.e. the total number of training rounds), and a learning rate scheduler is introduced to reduce the learning rate to 10% of the original value when the validation loss does not improve within 20 epochs.
[0078] In step S103, the user's measured EEG data is obtained and input into the trained cross-subject auditory attention detection model to determine the user's target auditory attention direction. Based on the target auditory attention direction, the hearing aid device is controlled to enhance the sound source in the auditory attention direction. The auditory attention method decoded by the cross-subject model is output and combined with the hearing aid device to enhance the sound source in the specific direction.
[0079] Specifically, after model training is complete, a complete auditory attention detection system will be constructed using the established cross-subject detection model. This system will adjust the audio processing functions of hearing aid devices based on the auditory attention direction output by the model. Specifically, by decoding the user's auditory attention direction in real time and feeding the decoding results back to the hearing aid device, it can enhance specific sound sources, helping hearing-impaired individuals focus on target sounds and ignore background noise in noisy environments.
[0080] In this application, contrastive learning and multi-task learning are combined: by combining contrastive learning and multi-task learning, the advantages of both are fully utilized. Contrastive learning learns common features across subjects by maximizing the similarity between similar features while minimizing the feature differences between different subjects, thereby improving the model's generalization ability and decoding performance across test tasks. Multi-task learning improves the model's ability to transfer knowledge between different tasks by simultaneously optimizing auditory attention direction decoding and feature learning across test tasks, ensuring the model's performance on new subject data.
[0081] In this application, the self-attention and cross-attention mechanisms are combined: self-attention is used to capture the temporal dependencies and local features in EEG signals, while the cross-attention mechanism is used to further refine the extraction of cross-subject features. This combination of mechanisms helps the model better handle the relationship between different time steps, while enhancing feature refinement and improving decoding accuracy.
[0082] In this application, decoding performance across multiple subjects is improved: Through comparative learning, the system can extract universal features from data from multiple subjects, thereby avoiding the impact of individual differences between subjects on decoding performance. Traditional methods often experience a significant drop in decoding performance across multiple subjects, but this application effectively improves decoding capabilities across subjects by learning universal features.
[0083] Next, an EEG auditory attention detection system based on contrastive learning and multi-task learning proposed in accordance with an embodiment of the present application will be described with reference to the accompanying drawings.
[0084] Figure 6 4 is a structural diagram of an EEG auditory attention detection system based on contrastive learning and multi-task learning in an embodiment of the present application.
[0085] like Figure 6 As shown, the EEG auditory attention detection system based on contrastive learning and multi-task learning includes: an EEG data acquisition module 100, a detection model construction and training module 200 and an auditory attention direction decoding module 300.
[0086] Specifically, the EEG data acquisition module 100 is used to acquire EEG data of a subject;
[0087] A detection model construction and training module 200 is used to construct a cross-subject auditory attention detection model based on contrastive learning and multi-task learning, and train the cross-subject auditory attention detection model according to the EEG data to obtain a trained cross-subject auditory attention detection model;
[0088] The auditory attention direction decoding module 300 is used to obtain the user's actual EEG data and input the actual EEG data into the trained cross-subject auditory attention detection model to obtain the user's target auditory attention direction, so that the hearing aid device can enhance the sound source in the target auditory attention direction.
[0089] Figure 7 This is a diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may include:
[0090] Memory 501 , processor 502 , and computer programs stored in the memory 501 and executable on the processor 502 .
[0091] When the processor 502 executes the program, the EEG auditory attention detection method based on contrastive learning and multi-task learning provided in the above embodiment is implemented.
[0092] Furthermore, the terminal further includes:
[0093] The communication interface 503 is used for communication between the memory 501 and the processor 502 .
[0094] The memory 501 is used to store computer programs that can be run on the processor 502 .
[0095] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0096] If the memory 501, processor 502, and communication interface 503 are implemented independently, the communication interface 503, memory 501, and processor 502 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EIS) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0097] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.
[0098] The processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0099] This embodiment also provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the above-mentioned EEG auditory attention detection method based on contrastive learning and multi-task learning is implemented.
[0100] One embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the Figure 1 Any of the corresponding embodiments provides an EEG auditory attention detection method based on contrastive learning and multi-task learning.
[0101] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to user analysis data, user storage data, user display data, etc.) and signals involved in the present invention are all information, data and signals authorized by the user or fully authorized by all parties; and the collection, use and processing of relevant information, data and signals comply with the laws, regulations and standards of relevant countries and regions.
[0102] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0103] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0104] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0105] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (not exhaustive) of computer-readable storage media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable storage medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0106] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0107] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0108] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0109] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
[0110] It should be understood that the application of this application is not limited to the above examples. For ordinary technicians in this field, they can make improvements or changes based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to this application.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for EEG auditory attention detection based on contrastive learning and multi-task learning, characterized in that: The EEG auditory attention detection method based on contrastive learning and multi-task learning includes: Acquire EEG data from multiple subjects; Constructing a cross-subject auditory attention detection model based on contrastive learning and multi-task learning, and training the cross-subject auditory attention detection model according to the EEG data to obtain a trained cross-subject auditory attention detection model; Obtaining measured EEG data of the user, inputting the measured EEG data into the trained cross-subject auditory attention detection model to obtain the user's target auditory attention direction, and controlling the hearing aid device to enhance the sound source in the target auditory attention direction according to the target auditory attention direction; The cross-subject auditory attention detection model includes an EEG temporal feature encoder, an EEG contrast feature encoder, an EEG feature refinement module, and a classifier; The step of training the cross-subject auditory attention detection model according to the EEG data to obtain a trained cross-subject auditory attention detection model specifically includes: inputting the target EEG signal into the cross-subject auditory attention detection model; The EEG temporal feature encoder extracts temporal features from the target EEG signal; The EEG contrast feature encoder extracts contrast features from the target EEG signal and obtains contrast loss based on the contrast features; The EEG feature refinement module interacts the time feature and the contrast feature to obtain a refined feature; The classifier obtains a binary cross entropy loss based on the refined features; The binary cross entropy loss and the contrastive loss are optimized to obtain a trained cross-subject auditory attention detection model.
2. The EEG auditory attention detection method based on contrastive learning and multi-task learning according to claim 1 is characterized in that: The EEG data is a target EEG signal; The obtaining of EEG data of multiple subjects specifically includes: Obtaining original EEG signals of multiple subjects collected by multi-channel EEG equipment; The original EEG signal is preprocessed to obtain a target EEG signal.
3. The EEG auditory attention detection method based on contrastive learning and multi-task learning according to claim 2 is characterized in that: The preprocessing of the original EEG signal to obtain the target EEG signal specifically includes: Re-referencing the original EEG signal to the average response of all electrodes to obtain a first EEG signal; performing a bandpass filter within a preset hertz range on the first EEG signal to obtain a second EEG signal; downsampling the second EEG signal to a target Hz to obtain a third EEG signal; Perform Z-value normalization on the third EEG signal to obtain a target EEG signal.
4. The EEG auditory attention detection method based on contrastive learning and multi-task learning according to claim 1 is characterized in that: The EEG temporal feature encoder includes a two-dimensional convolutional layer, a position encoding layer, and a multi-head self-attention layer; The EEG time feature encoder extracts time features from the target EEG signal, specifically comprising: The two-dimensional convolutional layer captures the local temporal features corresponding to the target EEG signal; The position encoding layer obtains position information corresponding to the local time feature; The multi-head self-attention layer obtains a temporal feature based on the local temporal feature and the position information.
5. The EEG auditory attention detection method based on contrastive learning and multi-task learning according to claim 1 is characterized in that: The EEG feature refinement module interacts the time feature and the contrast feature to obtain a refined feature, specifically including: The EEG feature refinement module obtains cross-subject shared features based on the contrast features, and performs dynamic feature fusion on the time features based on the cross-subject shared features to obtain refined transition features; The EEG feature refinement module uses a residual connection to add the time feature and the refined transition feature to obtain a refined feature.
6. The EEG auditory attention detection method based on contrastive learning and multi-task learning according to claim 1 is characterized in that: The classifier obtains a binary cross entropy loss based on the refined features, specifically including: The classifier classifies the refined features to obtain a predicted label; The predicted label and the true label are compared using a binary cross entropy loss function to obtain a binary cross entropy loss.
7. An EEG auditory attention detection system based on contrastive learning and multi-task learning, characterized in that: The EEG auditory attention detection system based on contrastive learning and multi-task learning is applied to the EEG auditory attention detection method based on contrastive learning and multi-task learning according to any one of claims 1 to 6; The EEG auditory attention detection system based on contrastive learning and multi-task learning includes: An EEG data acquisition module, used to acquire EEG data of the subject; A detection model construction and training module, configured to construct a cross-subject auditory attention detection model based on contrastive learning and multi-task learning, and train the cross-subject auditory attention detection model according to the EEG data to obtain a trained cross-subject auditory attention detection model; The auditory attention direction decoding module is used to obtain the user's actual EEG data and input the actual EEG data into the trained cross-subject auditory attention detection model to obtain the user's target auditory attention direction, so that the hearing aid device can enhance the sound source in the target auditory attention direction.
8. A terminal, characterized in that: The terminal includes: a memory, a processor, and an EEG auditory attention detection program based on contrastive learning and multi-task learning stored in the memory and runnable on the processor. When the EEG auditory attention detection program based on contrastive learning and multi-task learning is executed by the processor, the steps of the EEG auditory attention detection method based on contrastive learning and multi-task learning as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an EEG auditory attention detection program based on contrastive learning and multi-task learning. When the EEG auditory attention detection program based on contrastive learning and multi-task learning is executed by a processor, the steps of the EEG auditory attention detection method based on contrastive learning and multi-task learning as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Auditory attention decoding method and device in audio-visual scene and hearing aid system
CN117992909A
Multi-task auditory attention detection method and device based on electroencephalogram signals
CN118709119A