Attention state recognition method based on neurovascular bimodal collaborative learning

Through the attention state recognition method based on neurovascular bimodal collaborative learning, combined with EEG-fNIRS collaborative adversarial training and fNIRS enhancement training, the problems of data acquisition interference and limited fNIRS data in attention state recognition in the prior art are solved, achieving higher accuracy and robustness.

CN120148811APending Publication Date: 2025-06-13JIANGSU SUCAI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510034345.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When using EEG-fNIRS combined for attention state recognition, the prior art faces problems such as interference factors during data acquisition, complex multimodal data fusion analysis, expensive equipment and susceptible to perturbation. In particular, fNIRS data is limited and overfitting problems exist, and it has failed to fully exert its application value.

Method used

The attention state recognition method based on neurovascular dual-modal collaborative learning is adopted, and the accuracy and robustness of fNIRS attention state recognition is improved through the three-stage framework of EEG-fNIRS collaborative adversarial training, fNIRS enhancement training and fNIRS single peak testing. This framework includes aligning the EEG-fNIRS feature distribution in the collaborative adversarial training stage, designing dynamic attention networks for efficient classification in the fNIRS enhancement training stage, and only fNIRS single-modal inference is required for inference in the test stage.

Benefits of technology

Through collaborative learning and dynamic attention network, the accuracy and robustness of fNIRS in attention state recognition are improved, the operation process is simplified, the requirements for equipment are reduced, and interference problems in multimodal data fusion analysis are effectively avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148811A_ABST
    Figure CN120148811A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of cognitive states of nervous systems, and particularly relates to an attention state recognition method based on nerve and blood vessel bimodal collaborative learning, which comprises the following steps: step S1, synchronously acquiring electroencephalogram and functional near infrared spectrum data under an attention experiment normal form, and preprocessing the data to obtain EEG-fNIRS data; s2, taking EEG-fNIRS data corresponding to a task stage and enhancing the EEG-fNIRS data by using sliding window data to obtain EEG-fNIRS input data, firstly enabling an EEG-fNIRS primary feature extractor to align with EEG-fNIRS feature distribution through EEG-fNIRS collaborative adversarial training, and then realizing fNIRS efficient classification through fNIRS enhanced training, and S3, in an fNIRS single-peak test stage, taking the fNIRS data, and obtaining the EEG-fNIRS input data. And an identification result is obtained through the trained fNIRS primary extractor and the trained fNIRS dynamic attention network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of the cognitive state of the nervous system, and particularly relates to a method for recognizing the attention state based on the collaborative learning of neurovascular bimodality. Background Art

[0002] Attention is a core concept in cognitive psychology, which refers to an individual's ability to selectively focus on specific information. In today's fast-paced information age, understanding and quantifying the attention state is of great significance in many fields such as education, medical care, and human-computer interaction. Traditional assessment methods rely on subjective reports and behavioral tests, but these methods often lack real-time and objectivity. In recent years, with the development of brain imaging technologies, especially the combined application of electroencephalogram (EEG) and functional near-infrared spectroscopy (fNIRS), new ways have been provided for achieving more accurate and real-time monitoring of the attention state;

[0003] Using EEG or fNIRS alone has limitations: EEG mainly reflects the electrical activity of neurons, while fNIRS focuses on blood flow and oxygenation; EEG is sensitive to noise and has poor spatial localization, while the temporal resolution of fNIRS is not as good as that of EEG. Combining the two can complement their respective advantages and provide a more comprehensive view of brain activity. For example, in the recognition of the attention state, EEG can detect instantaneous attention fluctuations, while fNIRS can reveal the impact of continuous attention load on brain metabolism;

[0004] Although EEG-fNIRS shows great potential in the recognition of the attention state, it still faces some challenges. The data acquisition process may be interfered by various factors, such as motion artifacts, external light, etc.; more advanced algorithms are also required for multimodal data fusion analysis. In addition, the EEG-fNIRS joint acquisition device is expensive, vulnerable to disturbance during the acquisition process, and multimodal joint acquisition also increases the risk of data loss. While fNIRS is convenient to acquire, robust, and has better application value. However, the fNIRS data is limited and there is a hemodynamic delay, resulting in a serious overfitting problem of fNIRS. Therefore, it is not appropriate to construct an effective network architecture by increasing the number of convolutional kernels and the network depth. The current research focuses on EEG-fNIRS multimodal fusion or is limited to single modality, and fails to fully utilize the application value of fNIRS;

[0005] Collaborative learning is a method that enables multiple models or modalities to jointly complete specific tasks or improve their respective performances through mutual cooperation. This method draws on the concept of cooperation in human society and aims to enhance the overall learning efficiency and effectiveness through knowledge sharing and interaction among models. Collaborative learning can promote EEG-fNIRS information interaction, improve the generalization performance and accuracy of fNIRS, achieve EEG-fNIRS multimodal training and fNIRS unimodal testing, and can achieve higher accuracy and robustness while retaining the computational complexity of the unimodal modality, thus better empowering attention-based brain-computer interface applications. Summary of the Invention

[0006] The objective of the present invention is to provide an attention state recognition method based on neurovascular bimodal collaborative learning, which is convenient for fNIRS signal acquisition and has strong robustness, and has better application value in attention state recognition. The present invention aims to provide an fNIRS enhancement framework based on neurovascular bimodal collaborative learning, which can improve the accuracy and robustness of fNIRS attention state recognition. This framework consists of three key stages: EEG-fNIRS collaborative adversarial training, fNIRS enhancement training, and fNIRS unimodal testing. In the collaborative adversarial stage, the EEG-fNIRS feature distributions are aligned through domain adversarial methods, and multimodal optimization is performed simultaneously; in the fNIRS enhancement training stage, an fNIRS dynamic attention network is designed to efficiently classify attention states; in the testing stage, only fNIRS unimodal is required for inference.

[0007] To achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:

[0008] An attention state recognition method based on neurovascular bimodal collaborative learning, comprising the following steps:

[0009] Step S1, synchronously collect electroencephalogram and functional near-infrared spectroscopy data under an attention experiment paradigm, and preprocess the data to obtain EEG-fNIRS data;

[0010] Step S2, take the EEG-fNIRS data corresponding to the task stage and use sliding window data augmentation to obtain EEG-fNIRS input data. First, through EEG-fNIRS collaborative adversarial training, align the EEG-fNIRS feature distributions of the EEG-fNIRS primary feature extractors, and then through fNIRS enhancement training, achieve efficient fNIRS classification.

[0011] Step S3, in the fNIRS unimodal testing stage, take fNIRS data and obtain recognition results through the trained fNIRS primary extractor and fNIRS dynamic attention network respectively.

[0012] In step S1, a word generation experimental paradigm is adopted to synchronously collect EEG and fNIRS signals during the WG task. The subjects perform word generation according to the visual cues and brief beeping signals displayed on the screen. Each subject participates in three sessions, and each session contains 10 WG and resting trials. Each task includes a 2s instruction period, a 10s task period, a 1s stop period, and a 13 - 15s rest period. Finally, the synchronous EEG - fNIRS signals are pre - processed respectively.

[0013] The pre - processing of step S1 first downsamples the signals to 200Hz, corrects the mean baseline using the 2s of the instruction stage, and finally performs a 1Hz - 40Hz band - pass filter. During the pre - processing of the fNIRS signals, the signals are first downsampled to 10Hz, the fNIRS is converted to oxy - and deoxy - hemoglobin concentrations through the modified Beer - Lambert law, corrected the mean baseline using the 2s of the instruction stage and then performs a 0.01Hz - 0.1Hz band - pass filter, and finally splices HbO and Hbr in the spatial dimension.

[0014] The specific implementation process of step S2 is as follows:

[0015] Step S21, take the pre - processed data representative of the task stage in step S1 and perform data augmentation using a sliding window.

[0016] Step S22, take the EEG - fNIRS data processed in S21 for collaborative adversarial training, align the EEG - fNIRS feature distributions, and then perform enhanced training on fNIRS using a dynamic attention network.

[0017] The specific implementation process of S22 is as follows:

[0018] Step S221, the EEG - fNIRS data obtains primary features E b and F b , the EEG feature extractor consists of strip convolutions with convolution kernels of (1, 25) and (30, 1) and a pooling layer of (1, 75) in series, and the fNIRS feature extractor consists of strip convolutions of (1, 30) and (14, 1) in series. E b and F b After passing through the gradient reversal layer, they are input into a domain discriminator with shared weights to obtain the domain category d e and d f and are optimized through cross - entropy loss. E b and F b are input into a classifier with shared weights to obtain the classification result y e and y fIt is optimized by the cross - entropy loss with label smoothing. Both the domain discriminator and the classifier are fully - connected layers. The trained EEG and fNIRS primary feature extractors can map EEG - fNIRS to the common feature distribution space;

[0019] Step S222: The fNIRS is input into the trained fNIRS primary feature extractor to obtain the fNIRS primary feature F b , and then it is input into the dynamic attention network. The dynamic attention network consists of a self - attention module and a convolutional layer. The self - attention module is a multi - head attention network with head = 6. First, F b extracts features through the self - attention module, and then n weights w are obtained through a fully - connected layer and Sigmoid. At the same time, F b passes through n different convolutions in parallel to obtain n attention features. The weighted sum of w and the attention map is used to obtain the high - order feature. Finally, the classification result yff is obtained through a fully - connected layer. The n parallel convolutions consist of convolution kernels of different sizes and are used to extract attention features of different fine - grained levels. The calculation process of the dynamic attention network is as follows:

[0020] W = Sigmoid(Linear(MHA(F b )))

[0021]

[0022] where MHA is the multi - head attention mechanism, Linear is the linear layer, Sigmoid is the activation function, FC is the fully - connected layer, and yff is optimized by the cross - entropy loss with label smoothing. During the training process, the update of the primary feature extractor needs to be frozen, and finally, the trained model weight parameters are saved.

[0023] The specific implementation process of step S3 is as follows: Take the weight parameters saved in step S2, import the fNIRS primary feature extractor and the dynamic attention network to form the fNIRS recognition network. In practical applications, only the new fNIRS data needs to be input into the fNIRS recognition network to identify the attention state of the subject.

[0024] The present invention proposes an fNIRS attention state recognition framework based on neurovascular bimodal collaborative learning. This framework makes full use of the complementarity of EEG - fNIRS, improves the robustness of fNIRS recognition during the EEG - fNIRS collaborative learning process, introduces the domain adversarial mechanism to align the EEG - fNIRS feature distribution space; at the same time, a dynamic attention network is proposed to achieve the efficient recognition of the fNIRS attention state. In practical applications, it can improve the accuracy of fNIRS recognition while maintaining the computational complexity of a single modality.

[0025] The present invention has the following technical advantages compared with the prior art:

[0026] 1. Through EEG-fNIRS collaborative adversarial training, by leveraging the complementarity of EEG and fNIRS data, the accuracy of fNIRS in attention state recognition is enhanced.

[0027] 2. Adversarial training helps reduce the differences between data modalities, improve the generalization ability of the model under different data conditions, and thus enhance the robustness.

[0028] 3. When traditional multi-modal fusion methods process the jointly collected EEG and fNIRS data, they often face problems such as optoelectronic interference. The present invention effectively avoids these problems through collaborative learning without sacrificing data quality.

[0029] 4. The fNIRS single-modal enhancement training further ensures that in practical applications, reliable results can be obtained even when only fNIRS data is available.

[0030] 5. By introducing a dynamic attention network, the classification efficiency of fNIRS data is improved, making the recognition of attention states faster and more accurate.

[0031] 6. In practical applications, only fNIRS data input is required to recognize the attention state, which simplifies the operation process and reduces the requirements for equipment.

[0032] 7. Detailed preprocessing steps ensure data quality, including baseline correction, band-pass filtering, etc., which helps extract purer and more useful signal features. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The present invention can be further illustrated by the non-limiting embodiments given in the drawings.

[0034] Figure 1 is the flowchart of the preprocessing of neurovascular modal signals;

[0035] Figure 2 is the attention state recognition framework based on neurovascular bimodal collaborative learning proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] In order to enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be further described below in conjunction with the drawings and embodiments.

[0037] As Figure 1-2As shown in the figure, the attention state recognition method based on neurovascular bimodal collaborative learning of the present invention has complementary EEG and fNIRS, and there will be no optoelectronic interference problem during the joint acquisition process. Most current research focuses on multimodal fusion, but problems such as perturbations are likely to occur during the joint acquisition process. The fNIRS unimodal has better application value. The present invention aims to improve the accuracy and robustness of fNIRS recognition through the process of EEG-fNIRS collaborative learning, and achieve efficient classification of fNIRS through a dynamic attention network to promote the reliable application of fNIRS attention state recognition. The implementation of the present invention mainly includes three steps: (1) Neurovascular data acquisition and preprocessing; (2) EEG-fNIRS collaborative adversarial training and fNIRS enhanced training; (3) fNIRS unimodal testing.

[0038] The following will explain each step in detail.

[0039] Step S1: Neurovascular data acquisition and preprocessing

[0040] According to the WG experimental paradigm, experiments are carried out, and EEG and fNIRS signals during the WG task are synchronously collected. The subjects perform word generation (thinking of words starting with a certain letter) according to the visual cues and short beeping signals displayed on the screen. Each subject participates in three sessions, and each session contains 10 WG and rest (BL) trials. Each task includes a 2s instruction period, a 10s task period (WG or BL), a 1s stop period, and a 13 - 15s rest period. Finally, the synchronous EEG-fNIRS signals are preprocessed respectively. EEG is recorded at 1000Hz through 30 electrodes, and fNIRS records data of 36 channels at 10.4Hz.

[0041] Figure 1 For the data preprocessing process of EEG and fNIRS, during EEG preprocessing, first downsample the signal to 200Hz, perform mean baseline correction with the 2s of the instruction stage, and finally perform 1Hz - 40Hz band-pass filtering. During the preprocessing process of fNIRS signals, first downsample the signal to 10Hz, convert fNIRS to oxyhemoglobin and deoxyhemoglobin concentrations (Hbo, Hbr) through the modified Beer-Lambert law, perform 0.01Hz - 0.1Hz band-pass filtering after mean baseline correction with the 2s of the instruction stage, and finally splice Hbo and Hbr in the spatial dimension.

[0042] Step S2: EEG-fNIRS collaborative adversarial training and fNIRS enhanced training

[0043] The fNIRS attention state recognition framework based on multimodal collaborative learning involved in the present invention is as Figure 2 shown, and the specific implementation steps are as follows:

[0044] Step S21: Extract 10s data of the task phase in S1 for experiments. The dimension of the EEG dataset is (N, 60, 30, 2000), representing the number of subjects, the number of experiments, the number of EEG channels, and the time series in sequence. The dimension of the fNIRS dataset is (N, 60, 72, 100), representing the number of subjects, the number of experiments, the number of fNIRS channels, and the time series in sequence. Among them, the number of channels is composed of the superposition of HBO and Hbr (36×2). Since the EEG-fNIRS data is limited, data augmentation is performed through a sliding window with a size of 3s and a step size of 1s to obtain the EEG-fNIRS input data. The sizes of the EEG-fNIRS segments are (30, 600) and (72, 30) respectively. The first dimension is the number of channels, and the second dimension is the 3s time series.

[0045] Step S22: Input the EEG-fNIRS data into the EEG and fNIRS primary feature extractors respectively to obtain the EEG and fNIRS primary features E b and F b , and then input E b and F b into the domain discriminator and classifier respectively to obtain the domain category d e and d f and the classification result y e and yf. Among them, there is a gradient reversal layer in front of the domain discriminator. The domain category and the classification result are optimized through the cross-entropy loss and the cross-entropy loss with label smoothing respectively. The trained primary feature extractor can align the EEG and fNIRS feature distributions. During the fNIRS enhanced training process, F b is classified through a dynamic attention network, where the dynamic attention network is composed of a self-attention module and a convolutional layer, and the network is optimized through the cross-entropy loss with label smoothing optimization.

[0046] Step S221: The EEG-fNIRS data passes through the feature extractor to obtain the primary features E b and F b . The EEG feature extractor is serially composed of strip convolutions with convolution kernels of (1, 25) and (30, 1) and a pooling layer of (1, 75). The fNIRS feature extractor is serially composed of strip convolutions of (1, 30) and (14, 1). E b and F b After passing through the gradient reversal layer, they are input into the domain discriminator with shared weights to obtain the domain category d e and d f and are optimized through the cross-entropy loss. E b and F b are input into the classifier with shared weights to obtain the classification result y eIt is optimized with the cross - entropy loss smoothed by the label and yf. After the iteration is completed, the EEG - fNIRS primary feature extractor can map EEG - fNIRS to the common feature distribution space. The model training adopts a five - fold cross - validation strategy, with a total of 120 epochs trained, and the initial learning rate is 10 -3 , and a dynamic learning rate strategy is adopted. The learning rate is multiplied by a decay factor of 0.1 at epochs 60 and 90 respectively to reduce the learning rate, and the model weight parameters of the last round of training are saved.

[0047] Step S222: Import the model weight parameters into the fNIRS primary feature extractor, and extract fNIRS features to obtain fNIRS primary features F b , and input it into the dynamic attention network. The dynamic attention network consists of a self - attention module and a convolutional layer. The self - attention module is a multi - head attention network with 6 heads. First, F b extracts features through the self - attention module, then obtains 4 weights w through a linear layer and Sigmoid. At the same time, F b passes through 4 different convolutions in parallel to obtain 4 attention features. The weighted sum of w and the attention map is used to obtain high - order features, and finally the classification result yff is obtained using a fully - connected layer. The sizes of the 4 parallel convolution kernels are (1,1), (3,3), (5,5), and (7,7) respectively, which are used to extract attention features of different fine - grained levels. The calculation process of the dynamic attention network is as follows:

[0048] W = Sigmoid(Linear(MHA(F b )))

[0049]

[0050] where MHA is the multi - head attention mechanism, Linear is the linear layer, Sigmoid is the activation function, and FC is the fully - connected layer. During the model training process, y ff is optimized with the cross - entropy loss smoothed by the label. It is trained for 120 epochs by the method of five - fold cross - validation. During the training process, the fNIRS primary feature extractor needs to be frozen (that is, its weight parameters are fixed during the training process), and a dynamic learning rate strategy is adopted. The learning rate is multiplied by a decay factor of 0.1 at epochs 60 and 90 respectively to reduce the learning rate, and the model weight parameters of the last round of training are saved.

[0051] Step S3: fNIRS unimodal test

[0052] Take the weight parameters saved in step S2 and import them into the fNIRS primary feature extractor and the dynamic attention network to form the fNIRS recognition network. In practical applications, only by inputting new fNIRS data into the fNIRS recognition network can the attention state of the subject be recognized.

[0053] The above embodiments are only used to exemplarily illustrate the principles and effects of the present invention, rather than to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A method for identifying attention states based on neurovascular dual-modal collaborative learning, characterized by: The following steps are involved: Step S1, synchronously collect EEG and functional near-infrared spectroscopy data under the attention experimental paradigm, and preprocess the data to obtain EEG-fNIRS data; Step S2, taking the EEG-fNIRS data corresponding to the task stage and using the sliding window data enhancement to obtain the EEG-fNIRS input data, firstly, using EEG-fNIRS collaborative adversarial training to align the EEG-fNIRS primary feature extractor with the EEG-fNIRS feature distribution, and then using fNIRS enhancement training to achieve fNIRS efficient classification; Step S3, fNIRS unimodal testing phase, takes fNIRS data and obtains recognition results through the trained fNIRS primary extractor and fNIRS dynamic attention network.

2. The method for attention state recognition based on neurovascular dual-modal collaborative learning according to claim 1, characterized in that: In the step S1, a word generation experimental paradigm is adopted to synchronously collect EEG and fNIRS signals during the WG task. The subjects generate words according to the visual cues displayed on the screen and the short beep signal instructions. Each subject participates in three sessions, each session contains 10 WG and rest trials, and each task includes a 2s instruction period, a 10s task period, a 1s stop period and a 13-15s rest period. Finally, the synchronous EEG-fNIRS signals are preprocessed respectively.

3. The method for attention state recognition based on neurovascular dual-modal collaborative learning according to claim 1, characterized in that: The preprocessing of step S1 first downsamples the signal to 200 Hz, performs mean baseline correction using the 2 s of the instruction phase, and finally performs 1 Hz-40 Hz bandpass filtering. In the preprocessing of the fNIRS signal, the signal is first downsampled to 10 Hz, and the fNIRS is converted into oxygenated and deoxygenated hemoglobin concentrations by modifying the Beer-Lambert law, and then performs 0.01 Hz-0.1 Hz bandpass filtering after 2 s mean baseline correction using the instruction phase, and finally Hbo and Hbr are spliced ​​in the spatial dimension.

4. The method for identifying attention states based on neurovascular dual-modal collaborative learning according to claim 1, characterized in that: The specific implementation process of step S2 is as follows: Step S21, take the preprocessed data represented by the task stage of step S1, and perform data enhancement using a sliding window. In step S22, the EEG-fNIRS data processed in S21 is taken for collaborative adversarial training, the EEG-fNIRS feature distribution is aligned, and then the fNIRS is enhanced with a dynamic attention network.

5. The method for attention state recognition based on neurovascular dual-modal collaborative learning according to claim 4, characterized in that: The specific implementation process of S22 is as follows: Step S221, the EEG-fNIRS data is extracted by a feature extractor to obtain primary features E b and F b The EEG feature extractor consists of a series of strip convolutions with kernels (1,25) and (30,1) and a series of pooling layers with kernels (1,75). The fNIRS feature extractor consists of a series of strip convolutions with kernels (1,30) and (14,1). b and F b After the gradient reversal layer, the domain discriminator with weight sharing is input to obtain the domain category d e With d f And optimized by cross entropy loss, E b and F b Input the weight-sharing classifier to get the classification result y e With y f The training is optimized by cross entropy loss with label smoothing, where both the domain discriminator and classifier are fully connected layers. The trained EEG and fNIRS primary feature extractors are able to map EEG-fNIRS to a common feature distribution space. Step S222, fNIRS inputs the trained fNIRS primary feature extractor to obtain the fNIRS primary feature F b , and then input the dynamic attention network, which consists of a self-attention module and a convolutional layer. The self-attention module is a multi-head attention network with 6 heads. First, F b The features are extracted through the self-attention module, and then n weights w are obtained through the fully connected layer and Sigmoid. b Through n different convolutions in parallel, n attention features are obtained. The high-order features are obtained by weighted summing w and the attention map. Finally, the classification result yff is obtained by using a fully connected layer. The n parallel convolutions are composed of convolution kernels of different sizes, which are used to extract attention features of different fine-grained levels. The calculation process of the dynamic attention network is as follows: W=Sigmoid(Linear(MHA(F b ))) Among them, MHA is the multi-head attention mechanism, Linear is the linear layer, Sigmoid is the activation function, FC is the fully connected layer, yff is optimized by the label-smoothed cross entropy loss, and the update of the primary feature extractor needs to be frozen during the training process, and finally the weight parameters of the trained model are saved.

6. The method for attention state recognition based on neurovascular dual-modal collaborative learning according to claim 1, characterized in that: The specific implementation process of step S3 is: taking the weight parameters saved in step S2, importing the fNIRS primary feature extractor and the dynamic attention network to form an fNIRS recognition network. In practical applications, only the new fNIRS data needs to be input into the fNIRS recognition network to identify the subject's attention state.