Cognitive task classification method based on double-signal alternating guidance network
By adopting a dual-signal alternating guidance network method in the cognitive task classification method, fusing the spatiotemporal characteristics of EEG and fNIRS signals, and using attention mechanism for late fusion, the problems of data redundancy and information interaction in multimodal fusion are solved, and a more efficient cognitive task classification is achieved.
Patent Information
- Application Number
- CN202510249213.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The existing cognitive task classification methods fail to fully utilize the spatial and temporal complementarity advantages of EEG and fNIRS in the early fusion of multimodal, and there are problems of data alignment and data redundancy; in the late fusion of multimodal, it is impossible to effectively promote information interaction between different modal features, and there is a potential homogeneity loss problem.
Using a method based on a dual-signal alternating guidance network, the channel activation matrix of the fNIRS signal is generated by fusion of the time-varying power information of the EEG signal and the hemodynamic response of the fNIRS signal, the task is selected, and the modal weight is dynamically allocated in the late fusion stage to achieve deep fusion of EEG and fNIRS features.
It effectively solves the data redundancy problem in early multimodal fusion, and promotes information interaction between modal features in the late fusion stage, improving the performance and efficiency of cognitive task classification.
Smart Images

Figure CN120052915A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cognitive task classification, and particularly relates to a cognitive task classification method based on a dual-signal alternating guidance network. Background Art
[0002] Electroencephalogram (EEG) and functional near-infrared spectroscopy (fNIRS) technologies are two commonly used non-invasive technologies for studying brain activities, and are widely applied to brain cognitive detection and recognition, such as multiple fields including motor imagery (MI), emotion recognition, mild cognitive impairment (MCI) assessment, and workload monitoring. EEG records weak electrical signals generated by brain neuron activities and monitors the changes in brain electrical activities in real time, with high temporal resolution. fNIRS uses near-infrared light to measure the concentration changes of oxyhemoglobin (HbO) and deoxyhemoglobin (HbR) in the brain, with relatively high spatial resolution. The EEG and fNIRS imaging technologies each have their own advantages and disadvantages. Due to the inevitable volume conduction effect, EEG has a relatively low spatial resolution. fNIRS has disadvantages such as low temporal resolution, inherent delay caused by hemodynamic response (HRF), and low frequency. Therefore, researchers have proposed a research method for fusing dual-modal signals of EEG and fNIRS to overcome the limitations of single-modal information, make full use of the spatio-temporal complementary characteristics of the signals, and understand the functional activities of the brain more comprehensively, so as to obtain richer information in neuroscience research.
[0003] According to the occurrence stage of the multimodal fusion strategy, it can be divided into early, middle, and late fusion. Early fusion completes the fusion of modalities before feature extraction, emphasizing the interaction relationship between EEG and fnirs signals, which helps to extract more valuable feature information. Some models have designed an fNIRS-guided attention layer in the early fusion stage, which can not only capture the spatial correlation joint information of the two signals but also exclude the EEG channels interfered by noise and motion artifacts. There are also studies that integrate the spectral features of EEG into the fNIRS data processing, achieving the extraction of low-dimensional features of fNIRS signals without additional feature selection. However, there are significant differences in the sampling frequencies of EEG and fNIRS data, and issues such as data alignment and data redundancy need to be carefully considered during early fusion. In contrast, late fusion first independently extracts modal features and then fuses the features or the decision results of local modalities in a concatenated manner, which is more suitable for handling asynchronous modal information. For example, late fusion processing of the concatenated features of EEG and fNIRS signals is applicable to MI and mental arithmetic (MA) tasks. Or a reliability-based decision fusion mechanism is used to fuse the prediction results of EEG signals and fNIRS signals to obtain a comprehensive final prediction result. However, this fusion strategy is prone to potential loss of homogeneity during the feature extraction process, which may have a negative impact on the final fusion effect. In addition, during the model training process, the requirements of multimodal fusion for computing resources are also a challenge that needs to be overcome. The optimization of the model needs to consider how to remove redundant data, reduce processing time, and improve the classification performance of the network. The current cognitive task classification methods have the following problems:
[0004] (1) In multimodal early fusion, the spatio-temporal complementary advantages of EEG and fNIRS have not been fully utilized, and the problems of data alignment and data redundancy remain to be solved. Existing studies often manually select task-related signal channels based on prior knowledge of neuroscience to remove redundant data and reduce processing time. Although this method is simple and easy to implement, it ignores the synergistic effect during multimodal data fusion, which may lead to the result of channel selection being unable to fully utilize the complementary information between multimodal data, thus affecting the final fusion effect and performance.
[0005] (2) In the multimodal late fusion stage, feature concatenation or decision fusion is generally used, but this method cannot effectively promote the information interaction between different modal features. Concatenation simply juxtaposes the features of each branch, lacking the modeling of the mutual relationship between features. Therefore, it may limit the model's learning ability of bimodal interaction features and there is also a potential problem of loss of homogeneity. Summary of the Invention
[0006] In view of the above deficiencies in the prior art, a cognitive task classification method based on a dual-signal alternating guidance network provided by the present invention solves the problems of data redundancy in multi-modal early fusion and the inability to effectively promote information interaction between different modal features in the multi-modal late fusion stage in the current cognitive task classification methods.
[0007] In order to achieve the above invention purpose, the technical solution adopted by the present invention is: a cognitive task classification method based on a dual-signal alternating guidance network, including the following steps:
[0008] S1. Obtain the original EEG data and the original fNIRS data, preprocess them to obtain EEG signals and fNIRS signals;
[0009] S2. Fuse the time-varying power information of the EEG signal and the hemodynamic response of the fNIRS signal to generate a channel activation matrix of the fNIRS signal;
[0010] S3. Calculate the channel maximum activation value vector according to the channel activation matrix of the fNIRS signal, and select the task-active brain regions;
[0011] S4. Extract features from the EEG signal and the fNIRS signal of the corresponding channels in the task-active brain regions to obtain the temporal features of the EEG and the spatial features of the fNIRS;
[0012] S5. Perform late fusion on the temporal features of the EEG and the spatial features of the fNIRS to generate the brain cognitive task classification result.
[0013] Further: In the above S1, the method for preprocessing the original EEG data includes filtering, first segmentation, baseline correction, and second segmentation operations;
[0014] The method for preprocessing the original fNIRS data includes the following sub-steps:
[0015] A1. Collect the original fNIRS data, which is specifically the change amount of the near-infrared light intensity absorbed or scattered by the brain tissue;
[0016] A2. Use the modified Beer-Lambert law to convert the change amount of the near-infrared light intensity into the relative change amount of the oxyhemoglobin concentration and the relative change amount of the deoxyhemoglobin concentration;
[0017] A3. Perform filtering, first segmentation, baseline correction, and second segmentation operations on the relative change amount of the oxyhemoglobin concentration and the relative change amount of the deoxyhemoglobin concentration in sequence to obtain the concentration change of the oxyhemoglobin and the concentration change of the deoxyhemoglobin, and use them as the fNIRS signal;
[0018] In the above A2, the expressions for the relative change amount ΔHbO of the oxyhemoglobin concentration and the relative change amount ΔHbR of the deoxyhemoglobin concentration are specifically as follows:
[0019]
[0020] In the formula, d is the differential path length factor, λ 1 is the first illumination wavelength, λ 2 is the second illumination wavelength, l is the distance between the light source and the detector, ε HbO is the extinction coefficient of the chromophore HbO, ε HbR is the extinction coefficient of the chromophore HbR, is the change in the optical density of λ 1 and is the change in the optical density of λ 2 .
[0021] Furthermore: The above S2 includes the following sub-steps:
[0022] S21. Obtain the average channel based on the average value of all channels of the EEG signal, and obtain the time-varying power of the time series of the average channel through short-time Fourier transform;
[0023] S22. Perform time-varying power spectrum analysis based on the time-varying power of the time series, determine the time interval from the start of the task to the appearance of the peak of the time-varying power and establish a Boxcar function;
[0024] S23. Establish an HRF function based on the fNIRS signal, convolve the HRF function with the Boxcar function, and construct a response curve;
[0025] S24. Fit the response curve and the fNIRS signal through GLM to obtain the activation level of the fNIRs channel under a specific task, and further obtain the channel activation matrix of the fNIRS signal.
[0026] Furthermore: In the above S21, the expression for the time-varying power P(τ, ω) of the time series u(t) of the average channel is specifically as follows:
[0027]
[0028] In the formula, τ is the time index of the PSD, ω is the frequency index of the PSD, W(t - τ) is the Hanning window function, e -jωt is the complex exponential term for STFT, and t is the time scale;
[0029] In the above S22, the expression for the Boxcar function Boxcar(z) is specifically as follows:
[0030] Boxcar(z) = A(H(z - a) - H(z - b))
[0031] Wherein, A is a constant, a is the lower limit of the interval [a, b], b is the upper limit of the interval [a, b], and H(·) is the Heaviside function, and its specific expression is:
[0032]
[0033] Wherein, is the starting point of the Boxcar function, is the total time elapsed by the actual task;
[0034] In the above S23, the specific expression of the HRF function is:
[0035]
[0036] Wherein, HRF is the HRF function, is the time point of the fNIRS signal, is the dispersion time constant of the peak, is the dispersion time constant of the undershoot period, is the peak time, is the undershoot time, represents the ratio of the amplitude of the peak to the undershoot period, and Γ(·) is the gamma function;
[0037] In the above S24, the specific expression of fitting the response curve and the fNIRS signal by GLM is:
[0038] Y = Gβ + ε
[0039] Wherein, Y is the fNIRS signal, β is the regression coefficient between the expected and actual hemodynamic responses, which reflects the activation level of the fNIRS channel under a specific task, and ε is the error term between the expected and actual hemodynamic responses;
[0040] The method for obtaining the channel activation matrix of the fNIRS signal is specifically:
[0041] Based on the data obtained by fitting the response curve and the fNIRS signal, use the sliding window technique with a size of 2 and a step size of 1 for segmentation to generate a time series containing several regression coefficients. According to the time series, construct the channel activation matrices of oxyhemoglobin and deoxyhemoglobin, and merge the channel activation matrices of oxyhemoglobin and deoxyhemoglobin to obtain the channel activation matrix of the NIRS signal.
[0042] Furthermore: The above S3 includes the following sub-steps:
[0043] S31. Use the equidistant azimuthal projection technique to map the three-dimensional electrode positions in the EEG signal and fNIRS signal onto a two-dimensional matrix, and divide the regional subsets of the prefrontal lobe, left parietal lobe, right parietal lobe, and occipital lobe of the brain according to the electrode layouts of the EEG signal and fNIRS signal respectively;
[0044] S32. Calculate the channel maximum activation value vector through the channel activation matrix of the fNIRS signal according to the divided regional subsets;
[0045] S33. Determine the task-active brain regions according to the channel maximum activation value vector.
[0046] Furthermore: In the above S32, the expression of the channel maximum activation value vector AM max is specifically:
[0047] AM max ={max(AM 1 ), max(AM 2 ),... max(AM C )}
[0048] In the formula, max(AM 1 ) is the maximum activation information of the channel activation matrix of the first channel, max(AM 2 ) is the maximum activation information of the channel activation matrix of the second channel, max(AM c ) is the maximum activation information of the channel activation matrix of the c-th channel, and c is the number of channels;
[0049] The above S33 is specifically:
[0050] Normalize the maximum activation value vector, obtain the fNIRs channels whose maximum activation information in the maximum activation value vector exceeds 0.6, and take the brain regions where the obtained channels are located as the task-active brain regions.
[0051] Furthermore: In the above S4, the method for obtaining the time features of EEG is specifically:
[0052] B1. Input the EEG signal into the first two-dimensional convolutional layer, the first batch normalization layer, the first exponential linear unit, and the average pooling layer in sequence to generate the first feature;
[0053] B2. Input the first feature into the multi-scale convolutional network in sequence to obtain the second feature;
[0054] Among them, the MSTCN network includes the first residual block, the second residual block, and the third residual block connected in sequence;
[0055] The structures of the first residual block, the second residual block, and the third residual block are the same, and each includes a first dilated causal convolution sub-module and a second dilated causal convolution sub-module that are connected to each other. Among them, the input of the first dilated causal convolution sub-module serves as the input of this structure, and the sum of the input of the first dilated causal convolution sub-module and the output of the second dilated causal convolution sub-module serves as the output of this structure;
[0056] The structures of the first dilated causal convolution sub-module and the second dilated causal convolution sub-module are the same, and each includes a dilated causal convolution sub-layer, a batch normalization sub-layer, an ELU activation function sub-layer, and a Dropout sub-layer that are connected in sequence;
[0057] The dilated causal convolution sub-layer includes a first dilated causal convolution unit, a second dilated causal convolution unit, and a third dilated causal convolution unit that are connected in sequence;
[0058] B3. Input the second feature into the temporal attention module to obtain the temporal feature of the EEG.
[0059] Further: In the above B3, the temporal feature F of the EEG is obtained. e_out The specific expression is:
[0060]
[0061] In the formula, d eeg is the dimension size of K eeg , T is the transpose symbol, softmax(·) is the softmax activation function, Q eeg is the first query matrix, K eeg is the first key matrix, and V eeg is the first value matrix;
[0062] Q eeg = F e_CCV W e_Q
[0063] K eeg = F e_CCV W e_K
[0064] V eeg = F e_CCV W e_V
[0065] In the formula, F e_CCV is the second feature, W e_Q is the first weight matrix, W e_K is the second weight matrix, and W e_V is the third weight matrix, which is used to perform a linear transformation on the second feature.
[0066] Further: In the step S4, the method for obtaining the spatial features of fNIRS is specifically as follows:
[0067] C1. Input the fNIRS signal into the spatial attention module to obtain the spatial attention feature F f_sa ;
[0068]
[0069] In the formula, softmax(·) is the softmax activation function, Q fnirs is the second query matrix, K fnirs is the second key matrix, d fnirs is the dimension size of K fnirs and V fnirs is the second value matrix;
[0070] Q fnirs = X fnirs W f_Q
[0071] K fnirs = X fnirs W f_K
[0072] V fnirs = X fnirs W f_V
[0073] In the formula, X fnirs is the fNIRS signal, W f_Q is the fourth weight matrix, W f_K is the fifth weight matrix, W f_V is the sixth weight matrix, which is used to perform the linear transformation of the fNIRS signal;
[0074] C2. Input the spatial attention feature into the spatial convolution module to obtain the spatial features of fNIRS;
[0075] Among them, the spatial convolution module includes a second two-dimensional convolutional layer, a second batch normalization layer, and a thermal exponential linear unit connected in sequence. The number of filters of the second two-dimensional convolution is 32, and the convolution kernel size is (C 2 , 1), where C 2 is the number of channels of fNIRS in the task-active brain region.
[0076] Further: The step S5 includes the following sub-steps:
[0077] S51. Perform a linear transformation on the spatial features of fNIRS to obtain the third query matrix Q fusion and the third key matrix K fusion , and perform a linear transformation on the temporal features of EEG to obtain the third value matrix Vfusion ;
[0078] Q fusion = W T_q F f_out
[0079] K fusion = W T_k F f_out
[0080] V fusion = W T_v F e_out
[0081] Wherein, W T_q is the fourth weight matrix, W T_k is the fifth weight matrix, W T_v is the sixth weight matrix;
[0082] S52. Calculate the autocorrelation matrix according to the third query matrix and the third key matrix, and calculate the bimodal fusion feature F using the autocorrelation matrix and the third value matrix st ;
[0083]
[0084] Wherein, d fusion is the dimension size of K fusion , W fusion is the autocorrelation matrix, and its expression is specifically:
[0085]
[0086] S53. Convert the bimodal fusion feature into a one-dimensional vector, input it into the classification layer, calculate the classification probability of the output classification result of the classification layer through the softmax function, and use the classification result with the highest probability as the brain cognitive task classification result; among them, the expression of the loss function L of the classification layer is specifically:
[0087]
[0088] Wherein, N is the number of samples, C is the number of categories, ω(·) is the exponential function, y n is the predicted label of the nth sample, l c is the true label of the cth category, p c is the conditional probability of the output cth category.
[0089] The beneficial effects of the present invention are:
[0090] (1) The present invention provides a cognitive task classification method based on a dual-signal alternating guidance network. Based on the neurophysiological association between EEG signals and fNIRS signals, by utilizing the time-varying power characteristics of EEG signals and the hemodynamic response function of fNIRS, a task-related response curve is generated. The GLM is used to fit the response curve and the fNIRS signal to obtain the fNIRS channel activation matrix, realizing the early fusion of bimodal information.
[0091] (2) The present invention utilizes the high spatial resolution advantage of fNIRS signals. By means of the activation matrix and threshold setting, task-related brain regions and signal channels are selected, reducing data redundancy, enhancing the efficiency of data processing, improving the multi-domain feature extraction of bimodal signals, and realizing multi-cognitive task classification.
[0092] (3) In the multi-modal late fusion stage, the present invention is based on the attention mechanism. The autocorrelation matrix is calculated using the spatial features of fNIRS and used as the spatial attention weight to guide the temporal features of EEG for fusion, dynamically allocating the weights of the modalities, highlighting important modal features, improving the data fusion efficiency, and realizing the late fusion of the temporal features of EEG and the spatial features of NIRS driven by data.
[0093] (4) The method of the present invention utilizes the EEG-fNIRS alternating guidance network. By gradually fusing modal features at different stages, it fully utilizes the spatio-temporal complementary advantages of bimodal signals and captures the complex interaction relationships from low-order to high-order between modalities. Brief Description of the Drawings
[0094] Figure 1 It is a flowchart of a cognitive task classification method based on a dual-signal alternating guidance network of the present invention.
[0095] Figure 2 It is a schematic diagram of the MSTCN network of the present invention.
[0096] Figure 3 It is a schematic diagram of the residual block in the MSTCN network of the present invention.
[0097] Figure 4 It is a schematic diagram of the SA-SCV of the present invention.
[0098] Figure 5 It is a flowchart of the method of the present invention using the EEG-fNIRS alternating guidance network. Detailed Embodiments
[0099] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0100] As Figure 1 shown, in an embodiment of the present invention, a cognitive task classification method based on a dual-signal alternating guidance network includes the following steps:
[0101] S1. Obtain the EEG raw data and fNIRS raw data, preprocess them, and obtain the EEG signal and fNIRS signal;
[0102] S2. Fuse the time-varying power information of the EEG signal with the hemodynamic response of the fNIRS signal to generate a channel activation matrix of the fNIRS signal;
[0103] S3. Calculate the channel maximum activation value vector according to the channel activation matrix of the fNIRS signal, and select the task-active brain regions;
[0104] S4. Extract features from the EEG signal and fNIRS signal corresponding to the channels in the task-active brain regions to obtain the temporal features of the EEG and the spatial features of the fNIRS;
[0105] S5. Perform late fusion on the temporal features of the EEG and the spatial features of the fNIRS to generate the brain cognitive task classification result.
[0106] In this embodiment, the present invention uses EEG and fNIRS acquisition devices to synchronously monitor the voltage change and blood oxygen change of brain activities to obtain the EEG raw data and fNIRS raw data.
[0107] In the step S1, the method for preprocessing the EEG raw data includes filtering, first segmentation, baseline correction, and second segmentation operations;
[0108] The method for preprocessing the fNIRS raw data includes the following sub-steps:
[0109] A1. Collect the fNIRS raw data, which is specifically the change amount of the near-infrared light intensity absorbed or scattered by brain tissue;
[0110] A2. Convert the change amount of the near-infrared light intensity into the relative change amount of oxyhemoglobin concentration and the relative change amount of deoxyhemoglobin concentration by using the modified Beer-Lambert law;
[0111] A3. Perform filtering, first segmentation, baseline correction, and second segmentation operations on the relative change in oxyhemoglobin concentration and the relative change in deoxyhemoglobin concentration in sequence to obtain the concentration change of oxyhemoglobin and the concentration change of deoxyhemoglobin, and use them as fNIRS signals;
[0112] In A2, the expressions for the relative change in oxyhemoglobin concentration ΔHbO and the relative change in deoxyhemoglobin concentration ΔHbR are specifically:
[0113]
[0114] In the formula, d is the differential path length factor, λ 1 is the first illumination wavelength, λ 2 is the second illumination wavelength, l is the distance between the light source and the detector, ε HbO is the extinction coefficient of the chromophore HbO, ε HbR is the extinction coefficient of the chromophore HbR, is the change in the optical density of λ 1 and is the change in the optical density of λ 2 .
[0115] S2 includes the following sub - steps:
[0116] S21. Obtain the average channel according to the average value of all channels of the EEG signal, and obtain the time - varying power of the time series of the average channel through short - time Fourier transform;
[0117] S22. Perform time - varying power spectrum analysis according to the time - varying power of the time series, determine the time interval from the start of the task to the peak of the time - varying power and establish a Boxcar function;
[0118] S23. Establish an HRF function according to the fNIRS signal, convolve the HRF function with the Boxcar function to construct a response curve;
[0119] S24. Fit the response curve and the fNIRS signal through GLM to obtain the activation level of the fNIRs channel under a specific task, and then obtain the channel activation matrix of the fNIRS signal.
[0120] In S21, the expression for the time - varying power P(τ, ω) of the time series u(t) of the average channel is specifically:
[0121]
[0122] In the formula, τ is the time index of the PSD, ω is the frequency index of the PSD, W(t - τ) is the Hanning window function, e -jωt is the complex exponential term for STFT, and t is the time scale;
[0123] In S22, its latency can be determined according to time-varying power spectrum analysis. The latency is the time span from the start of the task to the appearance of the power peak, and the Boxcar function is defined based on this. The specific expression of the Boxcar function Boxcar(z) is as follows:
[0124] Boxcar(z) = A(H(z - a) - H(z - b))
[0125] In the formula, A is a constant, a is the lower limit of the interval [a, b], and b is the upper limit of the interval [a, b]. In this embodiment, the fixed parameters of the Boxcar function are set as: A = 1, a = 0, b = 140. H(·) is the Heaviside function, and its specific expression is:
[0126]
[0127] In the formula, is the starting point of the Boxcar function, is the total time elapsed for the actual task;
[0128] In S23, the specific expression of the HRF function is:
[0129]
[0130] In the formula, HRF is the HRF function, is the time point of the fNIRS signal, is the dispersion time constant of the peak, is the dispersion time constant of the undershoot period, is the peak time, is the undershoot time, represents the ratio of the amplitude of the peak to the undershoot period, Γ(·) is the gamma function. In this embodiment, c = 6;
[0131] In S24, the expression for fitting the response curve and the fNIRS signal through GLM is:
[0132]
[0133] In the formula, Y is the fNIRS signal, β is the regression coefficient between the expected and actual hemodynamic responses, which reflects the activation level of the fNIRS channel under a specific task, and ε is the error term between the expected and actual hemodynamic responses;
[0134] The method for obtaining the channel activation matrix of the fNIRS signal is specifically:
[0135] Based on the data obtained by fitting the response curve and the fNIRS signal, a sliding window technique with a size of 2 and a step size of 1 is used for segmentation to generate a time series containing a number of regression coefficients. According to the time series, a channel activation matrix of oxyhemoglobin and deoxyhemoglobin is constructed, and the channel activation matrices of oxyhemoglobin and deoxyhemoglobin are combined to obtain the channel activation matrix of the NIRS signal.
[0136] In this embodiment, segmentation is performed using a sliding window to generate a time series {β 1 , β 2 ,..., β m} containing a number of regression coefficients, where m is the number of sliding windows, thereby constructing a channel activation matrix of the NIRS signal for all channels.
[0137] S3 includes the following sub-steps:
[0138] S31. Using an equidistant azimuthal projection technique, map the three-dimensional electrode positions in the EEG signal and the fNIRS signal onto a two-dimensional matrix, and divide the regional subsets of the prefrontal lobe, left parietal lobe, right parietal lobe, and occipital lobe of the brain according to the electrode layouts of the EEG signal and the fNIRS signal respectively;
[0139] S32. Calculate the channel maximum activation value vector based on the channel activation matrix of the fNIRS signal according to the divided regional subsets;
[0140] S33. Determine the task-active brain regions based on the channel maximum activation value vector.
[0141] In S32, the expression of the channel maximum activation value vector AM max is specifically:
[0142] AM max = {max(AM 1 ), max(AM 2 ),... max(AM C )}
[0143] In the formula, max(AM 1 ) is the maximum activation information of the channel activation matrix of the first channel, max(AM 2 ) is the maximum activation information of the channel activation matrix of the second channel, max(AM c ) is the maximum activation information of the channel activation matrix of the c-th channel, and c is the number of channels;
[0144] S33 is specifically:
[0145] Normalize the maximum activation value vector, obtain the fNIRs channels in the maximum activation value vector where the maximum activation information exceeds 0.6, and use the brain regions where the obtained channels are located as the task-active brain regions.
[0146] In S4, the method for obtaining the time features of EEG is specifically as follows:
[0147] B1. Input the EEG signal into the first two-dimensional convolutional layer, the first batch normalization layer, the first exponential linear unit, and the average pooling layer in sequence to generate the first feature;
[0148] In this embodiment, the present invention uses the first two-dimensional convolutional layer to extract the global spatial information of the EEG signal. The number of its filters is set to 32, and the kernel size is set according to the number of electrode channels of the EEG signal to Then, introduce the batch normalization layer (BN) and the exponential linear unit (ELU). After the ELU processing, an average pooling layer with a size of (1, 20) is applied to the data, realizing 20-fold compression of the time data. In this way, the data dimension can be effectively reduced while retaining important time series information.
[0149] B2. Input the first feature into the multi-scale convolutional network in sequence to obtain the second feature;
[0150] As Figure 2 shown, the MSTCN network includes a first residual block, a second residual block, and a third residual block connected in sequence;
[0151] As Figure 3 shown, the structures of the first residual block, the second residual block, and the third residual block are the same, and each includes a first dilated causal convolutional sub-module and a second dilated causal convolutional sub-module connected to each other. Among them, the input of the first dilated causal convolutional sub-module is used as the input of this structure, and the sum of the input of the first dilated causal convolutional sub-module and the output of the second dilated causal convolutional sub-module is used as the output of this structure;
[0152] The structures of the first dilated causal convolutional sub-module and the second dilated causal convolutional sub-module are the same, and each includes a dilated causal convolutional sub-layer, a batch normalization sub-layer, an ELU activation function sub-layer, and a Dropout sub-layer connected in sequence;
[0153] The dilated causal convolutional sub-layer includes a first dilated causal convolutional unit, a second dilated causal convolutional unit, and a third dilated causal convolutional unit connected in sequence;
[0154] In this embodiment, the first dilated causal convolutional unit, the second dilated causal convolutional unit, and the third dilated causal convolutional unit all contain the same kernel size (k l = 4) but different dilation coefficients (dl The dilated causal convolution with dilation rate d (d = 1, 2, 4). After performing the dilated causal convolution in each dilated causal convolution sub-layer, a BN layer, an ELU activation function, and a Dropout layer with a value of 0.3 are immediately introduced.
[0155] B3. Input the second feature into the temporal attention module to obtain the temporal feature of the EEG.
[0156] In the above B3, the temporal feature F of the EEG is obtained. e_out The specific expression of is as follows:
[0157]
[0158] In the formula, d eeg is the dimension size of K eeg , T is the transpose symbol, softmax(·) is the softmax activation function, Q eeg is the first query matrix, K eeg is the first key matrix, and V eeg is the first value matrix;
[0159] Q eeg = F e_CCV W e_Q
[0160] K eeg = F _CCV W e_K
[0161] V eeg = F _CCV W e_V
[0162] In the formula, F e_CCV is the second feature, W e_Q is the first weight matrix, W e_K is the second weight matrix, and W e_V is the third weight matrix, which is used to perform the linear transformation of the second feature.
[0163] In this embodiment, the dot product of the first query matrix and the transpose of the first key matrix is calculated, and the result of the dot product is normalized by the softmax function. The normalized result is used as a weight matrix to perform weighted summation with the first value matrix, and the weighted result is combined with the first value matrix to obtain the temporal feature F of the EEG e_out ;
[0164] In the above S4, the method for obtaining the spatial feature of fNIRS is specifically as follows:
[0165] C1. Input the fNIRS signal into the spatial attention module to obtain the spatial attention feature F f_sa ;
[0166]
[0167] Wherein, softmax(·) is the softmax activation function, and Q fnirs is the second query matrix, and K fnirs is the second key matrix, d fnirs is the dimension size of K fnirs , and V fnirs is the second value matrix;
[0168] Q fnirs = X fnirs W f_Q
[0169] K fnirs = X fnirs W f_K
[0170] V fnirs = X fnirs W f_V
[0171] Wherein, X fnirs is the fNIRS signal, W f_Q is the fourth weight matrix, W f_K is the fifth weight matrix, W f_V is the sixth weight matrix, which is used to perform a linear transformation on the fNIRS signal;
[0172] In this embodiment, the present invention multiplies the query matrix by the transposed key matrix to calculate the attention weight matrix, so as to simulate the interaction between fNIRS features at different positions. Then, the attention weight matrix is normalized through the softmax activation function, and these weights are converted into a probability distribution, so as to realize the weighted aggregation of features at different positions. Finally, the weighted aggregated value matrix is used as the spatial attention feature output by the spatial attention module.
[0173] C2. Input the spatial attention feature into the spatial convolution module to obtain the spatial feature of fNIRS;
[0174] Wherein, the spatial convolution module includes a second two-dimensional convolutional layer, a second batch normalization layer, and a thermal exponential linear unit connected in sequence. The number of filters of the second two-dimensional convolution is 32, and the convolution kernel size is (C 2 , 1), where C 2 is the number of channels of fNIRS in the task active brain region.
[0175] In this embodiment, the interconnected spatial attention module and spatial convolution module (SA-SCV) are as Figure 4As shown. In the spatial convolution module, the number of filters for two-dimensional convolution is set to 32, and the convolution kernel size is set to (C 2 , 1), where C 2 is the number of channels of local fNIRS. The purpose of this process is to extract the spatial features related to the channels corresponding to each region of the brain from the weighted dual-signal channel activation matrix. Then, the BN layer and ELU are applied to achieve non-linear activation.
[0176] The S5 includes the following sub-steps:
[0177] S51. Linearly transform the spatial features of fNIRS to obtain the third query matrix Q fusion and the third key matrix K fusion , and linearly transform the temporal features of EEG to obtain the third value matrix V fusion ;
[0178] Q fusion = WT _q F f_out
[0179] K fusion = W T_k F f_out
[0180] v fusion = W T_v F e_out
[0181] In the formula, W T_q is the fourth weight matrix, W T_k is the fifth weight matrix, W T_v is the sixth weight matrix;
[0182] S52. Calculate the autocorrelation matrix according to the third query matrix and the third key matrix, and calculate the bimodal fusion feature F st ;
[0183]
[0184] In the formula, d fusion is the dimension size of K fusion , W fusion is the autocorrelation matrix, and its specific expression is:
[0185]
[0186] In this embodiment, by calculating the dot product of Q fusion and the transpose of K fusion , a spatial attention weight matrix W fusion, this matrix reflects the internal relationship between the spatial features of fNIRS. W fusion needs to be multiplied by the reciprocal of the dimension d of K fusion to ensure that the variance of the dot product is not affected by the change in the vector length. Then, the softmax function is used for normalization. Then, a weighted sum is performed on V fusion to generate the bimodal fusion feature F fusion . st .
[0187] S53. Convert the bimodal fusion feature into a one-dimensional vector, input it into the classification layer, calculate the classification probability of the output of the classification layer through the softmax function, and take the classification result with the highest probability as the classification result of the brain cognitive task; among them, the expression of the loss function L of the classification layer is specifically:
[0188]
[0189] In the formula, N is the number of samples, C is the number of categories, ω(·) is the exponential function, y n is the predicted label of the nth sample, l c is the true label of the cth category, and p c is the conditional probability of the output cth category.
[0190] The implementation process of the method of the present invention is as Figure 5 shown. First, obtain EEG signals and fNIRS signals, and based on the task-related EEG time-frequency information, construct a channel activation matrix of fNIRS signals to highlight the activity intensity of each channel in different time periods; then, use the maximum activation information in the channel activation matrix of fNIRS signals to select specific brain regions of EEG and fNIRS as task-active brain regions, and perform feature extraction for these regions respectively; secondly, for fNIRS signals, use multi-brain region collaborative analysis to capture their spatial features, and for EEG signals, focus on the in-depth extraction of time features; finally, in the late feature fusion stage, based on the attention mechanism, integrate the spatial features of fNIRS into the time features of EEG, realizing the in-depth fusion and complementary advantages of the two signal features.
[0191] In summary, based on the original EEG and fNIRs data, this solution first performs preprocessing, then uses the time dynamics of EEG signals and the spatial activity patterns of fNIRS to fuse the spatio-temporal features of bimodal signals, capture task-related active brain regions, and then based on the attention mechanism, realize the late fusion of EEG and fNIRS features. This solution captures the complex interaction relationship from low-order to high-order between modalities by gradually fusing modal features at different stages according to the contribution of each modality at different stages, which helps to improve the interpretability of the model.
[0192] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the number of technical features. Therefore, the features defined by "first", "second", and "third" may explicitly or implicitly include one or more of such features.
Claims
1. A cognitive task classification method based on a dual-signal alternating guidance network, characterized in that: The following steps are involved: S1, obtaining EEG raw data and fNIRS raw data, and preprocessing them to obtain EEG signals and fNIRS signals; S2, fusing the time-varying power information of the EEG signal with the hemodynamic response of the fNIRS signal to generate the channel activation matrix of the fNIRS signal; S3, calculate the channel maximum activation value vector based on the channel activation matrix of the fNIRS signal and select the task-active brain area; S4, extracting features from the EEG signals and fNIRS signals of the corresponding channels in the task-active brain areas to obtain the temporal features of the EEG and the spatial features of the fNIRS; S5. Late-fusion of EEG temporal features and fNIRS spatial features to generate brain cognitive task classification results.
2. The cognitive task classification method based on the dual-signal alternating guidance network according to claim 1 is characterized in that: In said S1, the method for preprocessing the EEG raw data includes filtering, first segmentation, baseline correction and second segmentation operations; The method for preprocessing fNIRS raw data includes the following steps: A1. Collect fNIRS raw data, which is specifically the change in the intensity of near-infrared light absorbed or scattered by brain tissue; A2. Using the modified Beer-Lambert law, the change in near-infrared light intensity is converted into the relative change in oxygenated hemoglobin concentration and the relative change in deoxygenated hemoglobin concentration; A3, filtering, first segmenting, baseline correction and second segmenting are performed on the relative changes of the oxygenated hemoglobin concentration and the relative changes of the deoxygenated hemoglobin concentration in sequence to obtain the concentration changes of the oxygenated hemoglobin and the deoxygenated hemoglobin, which are used as fNIRS signals; In A2, the relative change amount ΔHbO of the oxygenated hemoglobin concentration and the relative change amount ΔHbR of the deoxygenated hemoglobin concentration are specifically expressed as: Where d is the differential path length factor, λ1 is the first illumination wavelength, λ2 is the second illumination wavelength, l is the distance between the light source and the detector, and ε HbO is the extinction coefficient of the chromophore HbO, ε HbR is the extinction coefficient of the chromophore HbR, is the change in optical density of λ1, is the change in optical density of λ2.
3. The cognitive task classification method based on the dual-signal alternating guidance network according to claim 1 is characterized in that: The S2 comprises the following sub-steps: S21, obtaining an average channel according to the average value of all channels of the EEG signal, and obtaining the time-varying power of the time series of the average channel by short-time Fourier transform; S22, perform time-varying power spectrum analysis based on the time-varying power of the time series, determine the time interval from the start of the task to the peak of the time-varying power and establish a Boxcar function; S23, establishing the HRF function according to the fNIRS signal, convolving the HRF function with the Boxcar function, and constructing the response curve; S24. By fitting the response curve and fNIRS signal through GLM, the activation level of fNIRs channel under specific tasks is obtained, and then the channel activation matrix of fNIRS signal is obtained.
4. The cognitive task classification method based on the dual-signal alternating guidance network according to claim 3 is characterized in that: In S21, the expression of the time-varying power P(τ, ω) of the time series u(t) of the average channel is specifically: Where τ is the time index of PSD, ω is the frequency index of PSD, W(t-τ) is the Hanning window function, and e -jωt is the complex exponential term used for STFT, t is the time scale; In S22, the expression of the Boxcar function Boxcar(z) is specifically: Boxcar(z)=A(H(za)-H(zb)) Where A is a constant, a is the lower limit of the interval [a, b], b is the upper limit of the interval [a, b], and H(·) is the Heaviside function, which is expressed as follows: In the formula, is the starting point of the Boxcar function, is the total time of the actual task; In S23, the expression of the HRF function is specifically: Where HRF is the HRF function, is the time point of the fNIRS signal, is the peak dispersion time constant, is the dispersion time constant of the undershoot period, is the peak time, is the undershoot time, represents the ratio of the peak amplitude to the undershoot period, Γ(·) is the gamma function; In S24, the expression of the response curve and the fNIRS signal fitted by GLM is specifically: Y=Gβ+ε Where Y is the fNIRS signal, β is the regression coefficient between the expected and actual hemodynamic response, which reflects the activation level of the fNIRS channel under a specific task, and ε is the error term between the expected and actual hemodynamic response; The method for obtaining the channel activation matrix of fNIRS signals is as follows: Based on the data after fitting the response curve with the fNIRS signal, the sliding window technique with a size of 2 and a step size of 1 was used for segmentation to generate a time series containing several regression coefficients. The channel activation matrices of oxyhemoglobin and deoxyhemoglobin were constructed according to the time series, and the channel activation matrices of oxyhemoglobin and deoxyhemoglobin were merged to obtain the channel activation matrix of the NIRS signal.
5. The cognitive task classification method based on dual signal alternating guidance network according to claim 1 is characterized in that: The S3 comprises the following sub-steps: S31, using equidistant azimuth projection technology, the three-dimensional electrode positions in the EEG signal and fNIRS signal were mapped onto a two-dimensional matrix, and the regional subsets of the prefrontal lobe, left parietal lobe, right parietal lobe and occipital lobe of the brain were divided according to the electrode layout of the EEG signal and fNIRS signal; S32, calculating the channel maximum activation value vector through the channel activation matrix of the fNIRS signal according to the divided regional subsets; S33. Determine the task-active brain area based on the channel maximum activation value vector.
6. The cognitive task classification method based on the dual-signal alternating guidance network according to claim 5 is characterized in that: In S32, the channel maximum activation value vector AM max The specific expression is: AM max ={max(AM1),max(AM2),...max(AM C )} Wherein, max(AM1) is the maximum activation information of the channel activation matrix of the first channel, max(AM2) is the maximum activation information of the channel activation matrix of the second channel, and max(AM c ) is the maximum activation information of the channel activation matrix of the c-th channel, where c is the number of channels; The S33 is specifically: The maximum activation value vector was normalized to obtain the fNIRs channel whose maximum activation information in the maximum activation value vector exceeded 0.6, and the brain area where the obtained channel was located was used as the task-active brain area.
7. The cognitive task classification method based on dual signal alternating guidance network according to claim 1 is characterized in that: In S4, the method for obtaining the time characteristics of EEG is specifically as follows: B1, input the EEG signal into the first two-dimensional convolution layer, the first batch normalization layer, the first exponential linear unit and the average pooling layer in sequence to generate the first feature; B2. Input the first feature into the multi-scale convolutional network in sequence to obtain the second feature; The MSTCN network includes a first residual block, a second residual block and a third residual block connected in sequence; The first residual block, the second residual block and the third residual block have the same structure, and all include a first dilated causal convolution submodule and a second dilated causal convolution submodule connected to each other, wherein the input of the first dilated causal convolution submodule is used as the input of the structure, and the input of the first dilated causal convolution submodule is added to the output of the second dilated causal convolution submodule as the output of the structure; The first dilated causal convolution submodule and the second dilated causal convolution submodule have the same structure, both including a dilated causal convolution sublayer, a batch normalization sublayer, an ELU activation function sublayer, and a Dropout sublayer connected in sequence; The dilated causal convolution sublayer includes a first dilated causal convolution unit, a second dilated causal convolution unit, and a third dilated causal convolution unit connected in sequence; B3. Input the second feature into the temporal attention module to obtain the temporal feature of EEG.
8. The cognitive task classification method based on the dual-signal alternating guidance network according to claim 7 is characterized in that: In B3, the time feature F of EEG is obtained. e_out The specific expression is: Where, d eeg K eeg , T is the transposed symbol, softmax(·) is the softmax activation function, Q eeg is the first query matrix, K eeg is the first bond matrix, V eeg is the first value matrix; Q eeg =F e_CCV W e_Q K eeg =F e_CCV W e_K V eeg =F e_CCV W e_V In the formula, F e_CCV is the second feature, W e_Q is the first weight matrix, W e_K is the second weight matrix, W e_V is the third weight matrix, used to perform a linear transformation of the second feature.
9. The cognitive task classification method based on the dual-signal alternating guidance network according to claim 8 is characterized in that: In S4, the method for obtaining the spatial characteristics of fNIRS is specifically as follows: C1. Input the fNIRS signal into the spatial attention module to obtain the spatial attention feature F f_sa ; In the formula, softmax(·) is the softmax activation function, Q fnirs is the second query matrix, K fnirs is the second bond matrix, d fnirs K fnirs The dimension size, V fnirs is the second value matrix; Q fnirs =X fnirs W f_Q K fnirs =X fnirs W f_K V fnirs =X fnirs W f_V Where, X fnirs is the fNIRS signal, W f_Q is the fourth weight matrix, W f_K is the fifth weight matrix, W f_V is a sixth weight matrix, used to perform a linear transformation of the fNIRS signal; C2, input the spatial attention features into the spatial convolution module to obtain the spatial features of fNIRS; Among them, the spatial convolution module includes a second two-dimensional convolution layer, a second batch normalization layer and a first thermal exponential linear unit connected in sequence. The number of filters of the second two-dimensional convolution is 32, and the convolution kernel size is (C2,1), where C2 is the number of fNIRS channels in the task-active brain area.
10. The cognitive task classification method based on the dual-signal alternating guidance network according to claim 9 is characterized in that: The S5 comprises the following sub-steps: S51. Perform linear transformation on the spatial features of fNIRS to obtain the third query matrix Q fusion and the third bond matrix K fusion , linearly transform the time characteristics of EEG to obtain the third value matrix V fusion ; Q fusion =W T_q F f_out K fusion =W T_k F f_out V fusion =W T_v F e_out Where W T_q is the fourth weight matrix, W T_k is the fifth weight matrix, W T_v is the sixth weight matrix; S52, calculate the autocorrelation matrix according to the third query matrix and the third key matrix, and use the autocorrelation matrix and the third value matrix to calculate the bimodal fusion feature F st ; Where, d fusion K fusion The dimension size, W fusion is the autocorrelation matrix, and its specific expression is: S53, converting the bimodal fusion feature into a one-dimensional vector, inputting it into the classification layer, calculating the classification probability of the classification result output by the classification layer through the softmax function, and taking the classification result with the highest probability as the classification result of the brain cognitive task; wherein, the expression of the loss function L of the classification layer is specifically: Where N is the number of samples, C is the number of categories, ω(·) is the exponential function, and y n is the predicted label of the nth sample, l c is the true label of the cth category, p c is the conditional probability of the cth category of the output.
Citation Information
Patent Citations
Brain dysfunction auxiliary evaluation method based on multi-modal data fusion
CN115553752A
Bimodal signal fusion method based on adaptive space-time convolution attention network
CN118626940A
EEG-fNIRS coupling brain network construction method and system
CN119112215A
Cognitive state classification method based on EEG-fNIRS deep fusion of feature decoupling
CN119397339A
Neurovascular coupling analytical method based on electroencephalogram and functional near-infrared spectroscopy
US20210282694A1