A cognitive task classification method based on double-signal alternating guidance network
By generating a channel activation matrix for fNIRS signals and using a dual-signal alternation guidance network, the problems of data redundancy and information interaction in multimodal fusion are solved, achieving efficient feature extraction and cognitive task classification of EEG and fNIRS signals.
Patent Information
- Application Number
- CN202510249213.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-03-04
AI Technical Summary
In existing cognitive task classification methods, early multimodal fusion does not fully utilize the spatiotemporal complementary advantages of EEG and fNIRS, resulting in data alignment and redundancy issues. Late fusion cannot effectively promote information interaction between features of different modalities and has high computational resource requirements.
By generating a channel activation matrix of fNIRS signals, selecting task-active brain regions, and combining the temporal features of EEG and the spatial features of fNIRS, the network is used to perform early and late fusion using dual-signal alternation, dynamically allocating modality weights to achieve information interaction and feature extraction.
It effectively reduces data redundancy, improves data processing efficiency, enhances multimodal signal feature extraction, improves the accuracy and efficiency of cognitive task classification, and fully utilizes the spatiotemporal complementary advantages of dual-modal signals.
Smart Images

Figure CN120052915B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of cognitive task classification, and particularly relates to a cognitive task classification method based on a double-signal alternating guide network. BACKGROUND
[0002] Electroencephalogram (EEG) and functional near-infrared spectroscopy (fNIRS) are two commonly used non-invasive techniques for studying brain activity, and are widely used in brain cognitive detection and recognition, such as motor imagery (MI), emotion recognition, mild cognitive impairment (MCI) assessment, and workload monitoring. EEG records the weak electrical signals generated by the activity of brain neurons, and monitors the changes in brain electrical activity in real time, with high temporal resolution. fNIRS uses near-infrared light to measure the concentration changes of oxygenated hemoglobin (HbO) and deoxygenated hemoglobin (HbR) in the brain, with high spatial resolution. EEG and fNIRS imaging techniques each have advantages and disadvantages. Due to the inevitable volume conduction effect, EEG has a low spatial resolution. fNIRS has the disadvantages of low temporal resolution, inherent delay caused by hemodynamic response (HRF), and low frequency. Therefore, researchers have proposed a research method that fuses EEG and fNIRS dual-mode signals, overcomes the limitations of single-mode information, fully utilizes the temporal and spatial complementary characteristics of the signals, and more comprehensively understands the functional activity of the brain, thereby obtaining more abundant information in neuroscience research.
[0003] According to the occurrence stage of the multi-modal fusion strategy, it can be divided into early, middle and late fusion. Early fusion completes the fusion of modalities before feature extraction, emphasizing the interaction between EEG and fnirs signals, which helps to extract more valuable feature information. There are models in the early fusion stage that design attention layers guided by fNIRS, which not only capture the joint information of the spatial correlation of the two signals, but also exclude the EEG channels disturbed by noise and motion artifacts. There are also studies that integrate the spectral features of EEG into fNIRS data processing, achieving the extraction of low-dimensional features of fNIRS signals without additional feature selection. However, there is a significant difference in the sampling frequency of EEG and fNIRS data, and careful consideration of data alignment and data redundancy is needed when early fusion. In contrast, late fusion first extracts modal features independently, and then fuses features or local modal decision results in a splicing manner, which is easier to handle asynchronous modal information. For example, late fusion of spliced features of EEG and fNIRS signals is suitable for MI and mental arithmetic (MA) tasks. Or based on the reliability of the decision fusion mechanism, the prediction results of EEG signals and fNIRS signals are fused to obtain a comprehensive final prediction result. However, this fusion strategy may lose potential homogeneity in the feature extraction process, which may negatively affect the final fusion effect. In addition, during model training, the requirement of multi-modal fusion for computing resources is also a challenge that needs to be overcome, and the optimization of the model needs to consider how to remove redundant data, reduce processing time, and improve the classification performance of the network. The current cognitive task classification methods have the following problems:
[0004] (1) In the early fusion of multi-modal, the spatiotemporal complementary advantages of EEG and fNIRS have not been fully utilized, and the problems of data alignment and data redundancy remain to be solved. Existing researches often manually select task-related signal channels based on prior knowledge of neuroscience, remove redundant data, and reduce processing time. Although this method is simple to operate and easy to implement, it ignores the synergistic effect of multi-modal data fusion, which may result in the inability to fully utilize the complementary information between multi-modal data, thereby affecting the final fusion effect and performance.
[0005] (2) In the late fusion stage of multi-modal, feature splicing or decision fusion is commonly used, but this method cannot effectively promote the information interaction between different modal features. Splicing simply arranges the features of each branch side by side, lacks modeling of the relationship between features, and therefore may limit the model's ability to learn interactive features of dual modalities and exist potential homogeneity loss problems. SUMMARY
[0006] In view of the above problems in the prior art, the cognitive task classification method based on the dual-signal alternating guiding network solves the problems of data redundancy in the multi-modal early fusion and the problem that the information interaction between different modal features cannot be effectively promoted in the multi-modal late fusion stage in the current cognitive task classification method.
[0007] In order to achieve the above-mentioned purposes, the technical scheme adopted by the present application is as follows: a cognitive task classification method based on a dual-signal alternating guiding network, comprising the following steps:
[0008] S1, acquiring EEG original data and fNIRS original data, and preprocessing the same to obtain EEG signals and fNIRS signals;
[0009] S2, fusing time-varying power information of the EEG signals with hemodynamic responses of the fNIRS signals to generate a channel activation matrix of the fNIRS signals;
[0010] S3, calculating a channel maximum activation value vector according to the channel activation matrix of the fNIRS signals, and selecting a task active brain region;
[0011] S4, extracting features of the EEG signals and the fNIRS signals of the corresponding channels in the task active brain region to obtain time features of the EEG signals and spatial features of the fNIRS signals;
[0012] S5, late fusing the time features of the EEG signals and the spatial features of the fNIRS signals to generate a brain cognitive task classification result.
[0013] Further, in the S1, the method for preprocessing the EEG original data comprises filtering, first segmentation, baseline correction and second segmentation operations;
[0014] The method for preprocessing the fNIRS original data comprises the following steps:
[0015] A1, collecting fNIRS original data, which is specifically the change amount of near-infrared light intensity absorbed or scattered by brain tissue;
[0016] A2, converting the change amount of the near-infrared light intensity into a relative change amount of oxyhemoglobin concentration and a relative change amount of deoxyhemoglobin concentration by using a modified Beer-Lambert law;
[0017] A3, sequentially performing filtering, first segmentation, baseline correction and second segmentation operations on the relative change amount of the oxyhemoglobin concentration and the relative change amount of the deoxyhemoglobin concentration to obtain changes in the concentrations of oxyhemoglobin and deoxyhemoglobin, and taking the same as the fNIRS signals;
[0018] In the A2, the expression of the relative change amount of the oxyhemoglobin concentration ΔHbO and the relative change amount of the deoxyhemoglobin concentration ΔHbR is specifically:
[0019]
[0020] In the formula, d is a differential path length factor, λ1 is a first light wavelength, λ2 is a second light wavelength, l is a distance between a light source and a detector, ε HbO is an extinction coefficient of a chromophore HbO, ε HbR is an extinction coefficient of a chromophore HbR, is a change of optical density of λ1, is a change of optical density of λ2.
[0021] Further, the S2 includes the following steps:
[0022] S21, obtaining an average channel according to an average value of all channels of an EEG signal, and obtaining a time-varying power of a time sequence of the average channel through a short-time Fourier transform;
[0023] S22, performing time-varying power spectrum analysis according to the time-varying power of the time sequence, determining a time interval of the time-varying power from a task start to a peak value, and establishing a Boxcar function;
[0024] S23, establishing an HRF function according to an fNIRS signal, convolving the HRF function with the Boxcar function, and constructing a response curve;
[0025] S24, fitting the response curve with the fNIRS signal through a GLM, obtaining an activation level of the fNIRS channel under a specific task, and further obtaining a channel activation matrix of the fNIRS signal.
[0026] Further, in the S21, the expression of the time-varying power P(τ, ω) of the time sequence u(t) of the average channel is specifically:
[0027]
[0028] In the formula, τ is a time index of a PSD, ω is a frequency index of the PSD, W(t-τ) is a Hanning window function, e -jωt is a complex exponential term for STFT, and t is a time scale;
[0029] In the S22, the expression of the Boxcar function Boxcar(z) is specifically:
[0030] Boxcar(z)=A(H(z-a)-H(z-b))
[0031] wherein A is a constant, a is the lower limit of the interval [a, b], b is the upper limit of the interval [a, b], and H(·) is the Heaviside function, whose expression is specifically as follows:
[0032]
[0033] wherein, is the starting point of the Boxcar function, is the total time actually spent on the task;
[0034] In the S23, the expression of the HRF function is specifically as follows:
[0035]
[0036] wherein HRF is the HRF function, is the time point of the fNIRS signal, is the dispersion time constant of the peak, is the dispersion time constant of the undershoot period, is the peak time, is the undershoot time, is the ratio of the amplitude of the peak to the undershoot period, and Γ(·) is the Gamma function;
[0037] In the S24, the expression of the fitting of the response curve and the fNIRS signal by the GLM is specifically as follows:
[0038] Y=Gβ+ε
[0039] wherein Y is the fNIRS signal, β is the regression coefficient between the expected and actual hemodynamic responses, which reflects the activation level of the fNIRS channel under a specific task, and ε is the error term between the expected and actual hemodynamic responses;
[0040] The method for obtaining the channel activation matrix of the fNIRS signal is specifically as follows:
[0041] Based on the fitted data of the response curve and the fNIRS signal, a sliding window technique with a size of 2 and a step of 1 is used for segmentation to generate a time sequence containing a plurality of regression coefficients, and the channel activation matrix of oxyhemoglobin and deoxyhemoglobin is constructed according to the time sequence, and the channel activation matrix of oxyhemoglobin and deoxyhemoglobin is combined to obtain the channel activation matrix of the NIRS signal.
[0042] Further, the S3 comprises the following sub-steps:
[0043] S31, mapping the three-dimensional electrode positions in the EEG signal and the fNIRS signal to a two-dimensional matrix by using the equidistant azimuthal projection technology, and dividing subsets of regions of the frontal lobe, the left parietal lobe, the right parietal lobe and the occipital lobe according to the electrode layout of the EEG signal and the fNIRS signal;
[0044] S32, calculating a channel maximum activation value vector according to the channel activation matrix of the fNIRS signal according to the divided subsets of regions;
[0045] S33, determining a task active brain region according to the channel maximum activation value vector.
[0046] Further, in the S32, the channel maximum activation value vector AM max is specifically expressed as:
[0047] AM max ={max(AM1),max(AM2),...max(AM C )}
[0048] In the formula, max(AM1) is the maximum value activation information of the channel activation matrix of the first channel, max(AM2) is the maximum value activation information of the channel activation matrix of the second channel, and max(AM c ) is the maximum value activation information of the channel activation matrix of the cth channel, and c is the number of channels.
[0049] The S33 is specifically:
[0050] The maximum activation value vector is normalized to obtain the fNIRS channel whose maximum value activation information in the maximum activation value vector exceeds 0.6, and the brain region where the obtained channel is located is taken as the task active brain region.
[0051] Further, in the S4, the method for obtaining the time feature of the EEG is specifically:
[0052] B1, sequentially inputting the EEG signal into a first two-dimensional convolution layer, a first batch normalization layer, a first exponential linear unit and an average pooling layer to generate a first feature;
[0053] B2, sequentially inputting the first feature into a multi-scale convolution network to obtain a second feature;
[0054] The MSTCN network includes a first residual block, a second residual block and a third residual block connected in sequence.
[0055] The first residual block, the second residual block and the third residual block are the same in structure and each comprises a first dilated causal convolution submodule and a second dilated causal convolution submodule connected with each other, wherein the input of the first dilated causal convolution submodule is taken as the input of the structure, and the input of the first dilated causal convolution submodule is added with the output of the second dilated causal convolution submodule to be taken as the output of the structure;
[0056] The first dilated causal convolution submodule and the second dilated causal convolution submodule are the same in structure and each comprises a dilated causal convolution sublayer, a batch normalization sublayer, an ELU activation function sublayer and a Dropout sublayer connected in sequence;
[0057] The dilated causal convolution sublayer comprises a first dilated causal convolution unit, a second dilated causal convolution unit and a third dilated causal convolution unit connected in sequence;
[0058] B3, inputting the second feature into a time attention module to obtain a time feature of the EEG.
[0059] Further, in the B3, the time feature F of the EEG is obtained e_out The expression of the time feature F of the EEG is specifically as follows:
[0060]
[0061] In the expression, d eeg is the dimension size of K eeg , T is a transpose symbol, softmax(·) is a softmax activation function, Q eeg is a first query matrix, K eeg is a first key matrix, and V eeg is a first value matrix.
[0062] Q eeg =F e_CCV W e_Q
[0063] K eeg =F e_CCV W e_K
[0064] V eeg =F e_CCV W e_V
[0065] In the expression, F e_CCV is a second feature, W e_Q is a first weight matrix, W e_K is a second weight matrix, W e_V is a third weight matrix, and used for performing linear transformation on the second feature.
[0066] Further, in the S4, the method for obtaining the spatial feature of the fNIRS is specifically:
[0067] C1, input the fNIRS signal into a spatial attention module to obtain a spatial attention feature F f_sa ;
[0068]
[0069] In the formula, softmax(·) is a softmax activation function, Q fnirs is a second query matrix, K fnirs is a second key matrix, d fnirs is the dimension size of K fnirs , V fnirs is a second value matrix.
[0070] Q fnirs =X fnirs W f_Q
[0071] K fnirs =X fnirs W f_K
[0072] V fnirs =X fnirs W f_V
[0073] In the formula, X fnirs is the fNIRS signal, W f_Q is a fourth weight matrix, W f_K is a fifth weight matrix, and W f_V is a sixth weight matrix, which is used to perform linear transformation of the fNIRS signal.
[0074] C2, input the spatial attention feature into a spatial convolution module to obtain the spatial feature of the fNIRS.
[0075] The spatial convolution module includes a second two-dimensional convolution layer, a second batch normalization layer and a first exponential linear unit connected in sequence, the filter number of the second two-dimensional convolution is 32, and the convolution kernel size is (C2, 1), wherein C2 is the channel number of the fNIRS in the task active brain area.
[0076] Further, the S5 includes the following steps:
[0077] S51, linearly transform the spatial feature of the fNIRS to obtain a third query matrix Q fusion and a third key matrix K fusion , and linearly transform the time feature of the EEG to obtain a third value matrix V fusion ;
[0078] Q fusion = W T_q F f_out
[0079] K fusion = W T_k F f_out
[0080] V fusion = W T_v F e_out
[0081] In the formula, W T_q is a fourth weight matrix, W T_k is a fifth weight matrix, and W T_v is a sixth weight matrix.
[0082] S52, a self-correlation matrix is calculated according to the third query matrix and the third key matrix, and a bimodal fusion feature F st is calculated by using the self-correlation matrix and the third value matrix.
[0083]
[0084] In the formula, d fusion is the dimension size of K fusion , W fusion is a self-correlation matrix, and the expression thereof is specifically as follows:
[0085]
[0086] S53, the bimodal fusion feature is converted into a one-dimensional vector, which is input into a classification layer, a classification probability of a classification result output by the classification layer is calculated by using a softmax function, and the classification result with the highest probability is taken as a brain cognitive task classification result; wherein the expression of a loss function L of the classification layer is specifically as follows:
[0087]
[0088] In the formula, N is the sample number, C is the class number, ω(·) is an exponential function, y n is a predicted label of the nth sample, l c is a true label of the cth class, and p c is a conditional probability of the cth class output.
[0089] The present application has the following beneficial effects:
[0090] (1) The application provides a cognitive task classification method based on a dual-signal alternating guidance network, based on the neurophysiological correlation of EEG signals and fNIRS signals, using the time-varying power characteristics of the EEG signal and the hemodynamic response function of the fNIRS, generating a task-related response curve, using GLM to fit the response curve and the fNIRS signal, obtaining the fNIRS channel activation matrix, and realizing early fusion of dual-mode information.
[0091] (2) The application utilizes the high spatial resolution advantage of the fNIRS signal, selects the task-related brain region and signal channel through the activation matrix and threshold setting, reduces data redundancy, enhances the efficiency of data processing, improves the multi-domain feature extraction of dual-mode signals, and realizes multi-cognitive task classification.
[0092] (3) In the late fusion stage of multi-modal, the application is based on the attention mechanism, calculates the autocorrelation matrix based on the spatial features of fNIRS, and uses it as the spatial attention weight to guide the fusion of the time features of EEG, dynamically allocates the weight of the mode, highlights the important modal features, improves the data fusion efficiency, realizes the late fusion of the data-driven time features of EEG and the spatial features of NIRS.
[0093] (4) The method of the application utilizes the EEG-fNIRS alternating guidance network, gradually fuses the modal features at different stages, fully utilizes the time-space complementary advantage of dual-mode signals, and captures the complex interaction between modes from low-order to high-order. BRIEF DESCRIPTION OF DRAWINGS
[0094] Figure 1 A cognitive task classification method based on a dual-signal alternating guidance network of the application.
[0095] Figure 2 A schematic diagram of the MSTCN network of the application.
[0096] Figure 3 A schematic diagram of the residual block in the MSTCN network of the application.
[0097] Figure 4 A schematic diagram of the SA-SCV of the application.
[0098] Figure 5 A flowchart of the application method utilizing the EEG-fNIRS alternating guidance network. DETAILED DESCRIPTION
[0099] The specific embodiments of the present application are described below to facilitate the understanding of the present application for those skilled in the art, but it should be clear that the present application is not limited to the scope of the specific embodiments, and for those skilled in the art, it is obvious that various changes are within the spirit and scope of the present application defined and determined by the appended claims, and all the inventions utilizing the concept of the present application are within the scope of protection.
[0100] As shown in the figure, in one embodiment of the present application, a cognitive task classification method based on double-signal alternating guiding network comprises the following steps: Figure 1
[0101] S1, obtaining EEG raw data and fNIRS raw data, and preprocessing them to obtain EEG signals and fNIRS signals;
[0102] S2, fusing the time-varying power information of the EEG signal with the hemodynamic response of the fNIRS signal to generate a channel activation matrix of the fNIRS signal;
[0103] S3, calculating a channel maximum activation value vector according to the channel activation matrix of the fNIRS signal, and selecting a task active brain region;
[0104] S4, extracting features of the EEG signal and the fNIRS signal of the corresponding channel in the task active brain region to obtain time features of the EEG and spatial features of the fNIRS;
[0105] S5, late fusion of the time features of the EEG and the spatial features of the fNIRS to generate a brain cognitive task classification result.
[0106] In the present embodiment, the present application uses EEG and fNIRS acquisition devices to synchronously monitor the voltage change and blood oxygen change of brain activity to obtain EEG raw data and fNIRS raw data.
[0107] In S1, the method for preprocessing the EEG raw data includes filtering, first segmentation, baseline correction and second segmentation operations;
[0108] The method for preprocessing the fNIRS raw data includes the following steps:
[0109] A1, collecting fNIRS raw data, which is specifically the change amount of near-infrared light intensity absorbed or scattered by brain tissue;
[0110] A2, converting the change amount of near-infrared light intensity into the relative change amount of oxyhemoglobin concentration and the relative change amount of deoxyhemoglobin concentration by using the modified Beer-Lambert law;
[0111] A3, the relative change amount of the concentration of oxyhemoglobin and the relative change amount of the concentration of deoxyhemoglobin are sequentially subjected to filtering, first segmentation, baseline correction and second segmentation, to obtain the concentration change of oxyhemoglobin and the concentration change of deoxyhemoglobin, which are taken as the fNIRS signal;
[0112] In the A2, the expression of the relative change amount of the concentration of oxyhemoglobin ΔHbO and the relative change amount of the concentration of deoxyhemoglobin ΔHbR is specifically as follows:
[0113]
[0114] In the formula, d is a differential path length factor, λ1 is a first light wavelength, λ2 is a second light wavelength, l is the distance between the light source and the detector, ε HbO is the extinction coefficient of the chromophore HbO, ε HbR is the extinction coefficient of the chromophore HbR, is the change of optical density of λ1, is the change of optical density of λ2.
[0115] The S2 includes the following steps:
[0116] S21, obtaining an average channel according to the average value of all channels of the EEG signal, and obtaining the time-varying power of the time sequence of the average channel through short-time Fourier transform;
[0117] S22, performing time-varying power spectrum analysis according to the time-varying power of the time sequence, determining the time interval of the time-varying power from the start of the task to the peak value, and establishing a Boxcar function;
[0118] S23, establishing an HRF function according to the fNIRS signal, convolving the HRF function with the Boxcar function, and constructing a response curve;
[0119] S24, fitting the response curve with the fNIRS signal through GLM, obtaining the activation level of the fNIRs channel under a specific task, and further obtaining the channel activation matrix of the fNIRS signal.
[0120] In the S21, the expression of the time-varying power P(τ,ω) of the time sequence u(t) of the average channel is specifically as follows:
[0121]
[0122] In the formula, τ is the time index of the PSD, ω is the frequency index of the PSD, W(t-τ) is the Hanning window function, e -jωt is a complex exponential term for STFT, and t is a time scale;
[0123] In the S22, the latency, i.e. the time span from the beginning of the task to the appearance of the power peak, can be determined according to the time-varying power spectrum analysis, and the Boxcar function is defined according to the latency. The expression of the Boxcar function Boxcar(z) is specifically as follows:
[0124] Boxcar(z) = A(H(z-a)-H(z-b))
[0125] In the formula, A is a constant, a is the lower limit of the interval [a, b], and b is the upper limit of the interval [a, b]. In the embodiment, the fixed parameters of the Boxcar function are set as A = 1, a = 0, and b = 140. H(·) is the Heaviside function, and the expression of the Heaviside function is specifically as follows:
[0126]
[0127] In the formula, t is the starting point of the Boxcar function, is the total time actually spent by the task,
[0128] In the S23, the expression of the HRF function is specifically as follows:
[0129]
[0130] In the formula, HRF is the HRF function, is the time point of the fNIRS signal, is the dispersion time constant of the peak, is the dispersion time constant of the undershoot period, is the peak time, is the undershoot time, is the ratio of the amplitude of the peak to the undershoot period, and Γ(·) is the Gamma function. In the embodiment, Γ(·) is specifically as follows: c = 6;
[0131] In the S24, the expression of the GLM fitting response curve and the fNIRS signal is specifically as follows:
[0132]
[0133] In the formula, Y is the fNIRS signal, β is the regression coefficient between the expected and actual hemodynamic responses, which reflects the activation level of the fNIRS channel under a specific task, and ε is the error term between the expected and actual hemodynamic responses;
[0134] The method for obtaining the channel activation matrix of the fNIRS signal is specifically as follows:
[0135] Based on the data fitted by the response curve and the fNIRS signal, a sliding window technology with a size of 2 and a step of 1 is used for segmentation, a time sequence containing several regression coefficients is generated, an oxygenated hemoglobin and deoxygenated hemoglobin channel activation matrix is constructed according to the time sequence, and a channel activation matrix of the NIRS signal is obtained by merging the oxygenated hemoglobin and deoxygenated hemoglobin channel activation matrix.
[0136] In the embodiment, the sliding window is used for segmentation, a time sequence containing several regression coefficients {β1, β2,..., β m} is generated, and a channel activation matrix of the full-channel NIRS signal is constructed.
[0137] The S3 comprises the following steps:
[0138] S31, the equidistant azimuthal projection technology is used to map the three-dimensional electrode positions in the EEG signal and the fNIRS signal to a two-dimensional matrix, and the subsets of the frontal lobe, the left parietal lobe, the right parietal lobe and the occipital lobe of the brain are divided according to the electrode layout of the EEG signal and the fNIRS signal;
[0139] S32, the channel maximum activation value vector is calculated according to the divided subset through the channel activation matrix of the fNIRS signal;
[0140] S33, the task active brain area is determined according to the channel maximum activation value vector.
[0141] In the S32, the expression of the channel maximum activation value vector AM max is specifically as follows:
[0142] AM max ={max(AM1), max(AM2),... max(AM C )}
[0143] In the formula, max(AM1) is the maximum value activation information of the channel activation matrix of the first channel, max(AM2) is the maximum value activation information of the channel activation matrix of the second channel, max(AM c ) is the maximum value activation information of the channel activation matrix of the cth channel, and c is the channel number;
[0144] The S33 is specifically as follows:
[0145] The maximum activation value vector is normalized to obtain the fNIRs channel whose maximum value activation information exceeds 0.6 in the maximum activation value vector, and the brain area where the obtained channel is located is taken as the task active brain area.
[0146] In the S4, the method for obtaining the time characteristics of the EEG is specifically as follows:
[0147] B1, inputting the EEG signal into a first two-dimensional convolution layer, a first batch normalization layer, a first exponential linear unit and an average pooling layer in sequence to generate a first feature;
[0148] In the embodiment, the first two-dimensional convolution layer is used to extract the global spatial information of the EEG signal, the number of filters is set to 32, and the kernel size is set according to the number of electrode channels of the EEG signal The number of filters is set to 32, and the kernel size is set according to the number of electrode channels of the EEG signal Then, the batch normalization layer (BN) and the exponential linear unit (ELU) are introduced. After the ELU processing, an average pooling layer with a size of (1, 20) is applied to the data, realizing 20 times compression of the time data. In this way, the data dimension can be effectively reduced while the important time sequence information is retained.
[0149] B2, inputting the first feature into a multi-scale convolution network in sequence to obtain a second feature;
[0150] As shown in Figure 2 , the MSTCN network includes a first residual block, a second residual block and a third residual block connected in sequence;
[0151] As shown in Figure 3 , the first residual block, the second residual block and the third residual block have the same structure, and each includes a first dilated causal convolution submodule and a second dilated causal convolution submodule connected with each other, wherein the input of the first dilated causal convolution submodule is taken as the input of the structure, and the input of the first dilated causal convolution submodule is added with the output of the second dilated causal convolution submodule to be taken as the output of the structure;
[0152] The first dilated causal convolution submodule and the second dilated causal convolution submodule have the same structure, and each includes a dilated causal convolution submodule, a batch normalization submodule, an ELU activation function submodule and a Dropout submodule connected in sequence;
[0153] The dilated causal convolution submodule includes a first dilated causal convolution unit, a second dilated causal convolution unit and a third dilated causal convolution unit connected in sequence;
[0154] In the embodiment, the first dilated causal convolution unit, the second dilated causal convolution unit and the third dilated causal convolution unit each contain dilated causal convolutions with the same kernel size (k l =4) but different dilation coefficients (d l =1, 2, 4), and after the dilated causal convolution is performed at each dilated causal convolution submodule, a BN layer, an ELU activation function and a Dropout layer with a value of 0.3 are introduced in sequence.
[0155] B3, input the second feature into the time attention module to obtain a time feature of the EEG.
[0156] In the B3, the time feature F of the EEG is obtained e_out The expression of the time feature F of the EEG is specifically:
[0157]
[0158] In the expression, d eeg is the dimension size of K eeg , T is a transpose symbol, softmax(·) is a softmax activation function, Q eeg is a first query matrix, K eeg is a first key matrix, and V eeg is a first value matrix.
[0159] Q eeg =F e_CCV W e_Q
[0160] K eeg =F _CCV W e_K
[0161] V eeg =F _CCV W e_V
[0162] In the expression, F e_CCV is the second feature, W e_Q is a first weight matrix, W e_K is a second weight matrix, and W e_V is a third weight matrix, which is used to perform linear transformation on the second feature.
[0163] In the embodiment, the transpose of the first query matrix and the first key matrix is dot multiplied, the result after the dot multiplication is normalized by the softmax function, the normalized result is weighted and summed with the first value matrix as a weight matrix, and the weighted result is combined with the first value matrix to obtain the time feature F e_out of the EEG.
[0164] In the S4, the method for obtaining the spatial feature of the fNIRS is specifically:
[0165] C1, input the fNIRS signal into the spatial attention module to obtain a spatial attention feature F f_sa .
[0166]
[0167] In the expression, softmax(·) is a softmax activation function, Q fnirsis a second query matrix, K fnirs is a second key matrix, d fnirs is K fnirs a dimension size, V fnirs is a second value matrix;
[0168] Q fnirs = X fnirs W f_Q
[0169] K fnirs = X fnirs W f_K
[0170] V fnirs = X fnirs W f_V
[0171] In the formula, X fnirs is an fNIRS signal, W f_Q is a fourth weight matrix, W f_K is a fifth weight matrix, W f_V is a sixth weight matrix, used to perform linear transformation of the fNIRS signal;
[0172] In this embodiment, the present application multiplies the query matrix with the transposed key matrix to calculate the attention weight matrix, thereby simulating the interaction between different position fNIRS features. Then, the attention weight matrix is normalized by a softmax activation function to convert these weights into a probability distribution, thereby realizing weighted aggregation of different position features. Finally, the value matrix after weighted aggregation is taken as the spatial attention feature output by the spatial attention module.
[0173] C2, the spatial attention feature is input into a spatial convolution module to obtain the spatial feature of the fNIRS;
[0174] The spatial convolution module includes a second two-dimensional convolution layer, a second batch normalization layer and a second exponential linear unit connected in sequence, the filter number of the second two-dimensional convolution is 32, and the convolution kernel size is (C2, 1), wherein C2 is the channel number of the fNIRS in the task active brain area.
[0175] In this embodiment, the spatial attention module and the spatial convolution module (SA-SCV) connected with each other are as shown in Figure 4 In the spatial convolution module, the filter number of the two-dimensional convolution is set to 32, and the convolution kernel size is set to (C2, 1), wherein C2 is the channel number of the local fNIRS. The purpose of this process is to extract the spatial features related to the channels corresponding to each region of the brain from the weighted double signal channel activation matrix. Then, the BN layer and the ELU are applied to realize nonlinear activation.
[0176] S5 comprises the following sub-steps:
[0177] S51, linearly transforming the spatial features of fNIRS to obtain a third query matrix Q fusion and a third key matrix K fusion , linearly transforming the temporal features of EEG to obtain a third value matrix V fusion ;
[0178] Q fusion = WT _q F f_out
[0179] K fusion = W T_k F f_out
[0180] v fusion = W T_v F e_out
[0181] wherein W T_q is a fourth weight matrix, W T_k is a fifth weight matrix, and W T_v is a sixth weight matrix;
[0182] S52, calculating an autocorrelation matrix according to the third query matrix and the third key matrix, and calculating a bimodal fusion feature F st using the autocorrelation matrix and the third value matrix;
[0183]
[0184] wherein d fusion is the dimension size of K fusion , and W fusion is the autocorrelation matrix, and the expression thereof is specifically:
[0185]
[0186] In this embodiment, by calculating the dot product of Q fusion and the transpose of K fusion , a spatial attention weight matrix W fusion is constructed, which reflects the internal relationship between the spatial features of fNIRS. W fusion needs to be multiplied by the reciprocal of the dimension d fusion of K fusion to ensure that the variance of the dot product is not affected by the change of the vector length. Then, the softmax function is used for normalization processing. Then, V fusion is weighted and summed to generate the bimodal fusion feature F st .
[0187] S53, convert the bimodal fusion feature into a one-dimensional vector, input it into a classification layer, calculate the classification probability of the classification result of the classification layer output by a softmax function, and take the classification result with the highest probability as the brain cognitive task classification result; wherein the expression of the loss function L of the classification layer is specifically:
[0188]
[0189] In the formula, N is the number of samples, C is the number of categories, ω(·) is an exponential function, y n is the predicted label of the nth sample, l c is the true label of the cth category, p c is the conditional probability of the cth category of the output.
[0190] The implementation process of the method is shown in Figure 5 As shown in the figure, first, the EEG signal and the fNIRS signal are obtained, the channel activation matrix of the fNIRS signal is constructed based on the task-related EEG time-frequency information, so as to highlight the activity intensity of each channel in different time periods; subsequently, the maximum activation information in the channel activation matrix of the fNIRS signal is used to select specific brain regions of the EEG and the fNIRS as task active brain regions, and feature extraction is performed on these regions respectively; secondly, for the fNIRS signal, the spatial features are captured by multi-brain region collaborative analysis, and for the EEG signal, the depth extraction of the time features is focused; finally, in the feature late fusion stage, the spatial features of the fNIRS are integrated into the time features of the EEG based on the attention mechanism, realizing the deep fusion and complementary advantages of the features of the two signals.
[0191] In summary, based on the EEG and fNIRs original data, first, after preprocessing, the time dynamic change of the EEG signal and the spatial activity mode of the fNIRS are used to fuse the space-time features of the bimodal signals, capture the active brain regions related to the task, and then based on the attention mechanism, realize the late fusion of the EEG and fNIRS features. According to the contribution of each modality in different stages, the modality features are gradually fused in different stages, the complex interaction relationship between the modalities from low order to high order is captured, which is helpful to improve the explainability of the model.
[0192] In the description of the application, it needs to be understood that the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the application. In addition, the terms "first", "second", "third" are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implying the number of technical features indicated. Therefore, the features defined by "first", "second", "third" can explicitly or implicitly include one or more of the features.
Claims
1. A cognitive task classification method based on a dual-signal alternating guidance network, characterized in that, Includes the following steps: S1. Obtain raw EEG data and raw fNIRS data, preprocess them, and obtain EEG signal and fNIRS signal; S2. The time-varying power information of the EEG signal is fused with the hemodynamic response of the fNIRS signal to generate the channel activation matrix of the fNIRS signal. S3. Calculate the maximum activation value vector of the channels based on the channel activation matrix of the fNIRS signal, and select the brain regions active for the task. S4. Extract features from the EEG and fNIRS signals of the corresponding channels in the brain regions where the task is active to obtain the temporal features of EEG and the spatial features of fNIRS. S5. Perform late-stage fusion of the temporal features of EEG and the spatial features of fNIRS to generate brain cognitive task classification results.
2. The cognitive task classification method based on a dual-signal alternating guidance network according to claim 1, characterized in that, In S1, the method for preprocessing the raw EEG data includes filtering, first segmentation, baseline correction, and second segmentation. The method for preprocessing raw fNIRS data includes the following steps: A1. Collect raw fNIRS data, specifically the change in the intensity of near-infrared light absorbed or scattered by brain tissue. A2. The modified Beer-Lambert law is used to convert the change in near-infrared light intensity into the relative change in oxyhemoglobin concentration and the relative change in deoxyhemoglobin concentration. A3. The relative changes in oxyhemoglobin concentration and deoxyhemoglobin concentration are sequentially filtered, segmented for the first time, corrected for baseline, and segmented for the second time to obtain the concentration changes of oxyhemoglobin and deoxyhemoglobin, which are then used as fNIRS signals. In A2, the expressions for the relative changes in oxyhemoglobin concentration ΔHbO and deoxyhemoglobin concentration ΔHbR are as follows: In the formula, d is the differential path length factor, λ1 is the first illumination wavelength, λ2 is the second illumination wavelength, l is the distance between the light source and the detector, and ε HbO ε is the extinction coefficient of the chromophore HbO. HbR The extinction coefficient of the chromophore HbR is given. For the change in optical density of λ1, The change in optical density is λ2.
3. The cognitive task classification method based on a dual-signal alternating guidance network according to claim 1, characterized in that, S2 includes the following steps: S21. Obtain the average channel based on the average value of the entire EEG signal, and obtain the time-varying power of the average channel's time series through short-time Fourier transform. S22. Perform time-varying power spectrum analysis based on the time-varying power of the time series, determine the time interval from the start of the task to the occurrence of the peak power, and establish the Boxcar function. S23. Based on the fNIRS signal, establish an HRF function, convolve the HRF function with the Boxcar function, and construct the response curve; S24. By fitting the response curve and the fNIRS signal with GLM, the activation level of the fNIRs channel under a specific task is obtained, and then the channel activation matrix of the fNIRS signal is obtained.
4. The cognitive task classification method based on a dual-signal alternating guidance network according to claim 3, characterized in that, In S21, the expression for the time-varying power P(τ, ω) of the average channel's time series u(t) is specifically as follows: In the formula, τ is the time index of the PSD, ω is the frequency index of the PSD, W(t-τ) is the Hanning window function, and e -jωt Here is the complex exponential term used for STFT, and t is the time scale; In S22, the expression for the Boxcar function Boxcar(z) is as follows: Boxcar(z) = A(H(za) - H(zb)) In the formula, A is a constant, a is the lower limit of the interval [a,b], b is the upper limit of the interval [a,b], and H(·) is the Herveside function, whose expression is as follows: In the formula, This is the starting point of the Boxcar function. This represents the total time elapsed for the actual task. In S23, the expression for the HRF function is specifically as follows: In the formula, HRF is the HRF function. The time point of the fNIRS signal, The dispersion time constant of the peak value is... Let be the dispersion time constant of the downstroke period. Peak time, For the downstroke time, It represents the ratio of the peak amplitude to the undershoot period, and Γ(·) is the gamma function; In step S24, the expression for fitting the response curve and the fNIRS signal using GLM is specifically as follows: Y=Gβ+ε In the formula, Y is the fNIRS signal, β is the regression coefficient between the expected and actual hemodynamic response, which reflects the activation level of the fNIRS channel under a specific task, and ε is the error term between the expected and actual hemodynamic response. The specific method for obtaining the channel activation matrix of the fNIRS signal is as follows: Based on the data fitted to the response curve and the fNIRS signal, a sliding window technique with a size of 2 and a step size of 1 is used for segmentation to generate a time series containing several regression coefficients. Channel activation matrices for oxyhemoglobin and deoxyhemoglobin are constructed based on the time series. The channel activation matrices for oxyhemoglobin and deoxyhemoglobin are then merged to obtain the channel activation matrix of the NIRS signal.
5. The cognitive task classification method based on a dual-signal alternating guidance network according to claim 1, characterized in that, S3 includes the following steps: S31. Using equidistant azimuth projection technology, the three-dimensional electrode positions in EEG signals and fNIRS signals are mapped onto a two-dimensional matrix, and the prefrontal lobe, left parietal lobe, right parietal lobe and occipital lobe regions of the brain are divided according to the electrode layout of EEG signals and fNIRS signals. S32. Calculate the maximum activation value vector of the channels using the channel activation matrix of the fNIRS signal based on the divided region subsets; S33. Determine the brain regions active for the task based on the maximum activation value vector of the channel.
6. The cognitive task classification method based on a dual-signal alternating guidance network according to claim 5, characterized in that, In S32, the channel maximum activation value vector AM max The specific expression is: AM max ={max(AM1),max(AM2),...max(AM C )} In the formula, max(AM1) represents the maximum activation value of the channel activation matrix for the first channel, and max(AM2) represents the maximum activation value of the channel activation matrix for the second channel. c ) represents the maximum activation information of the channel activation matrix of the c-th channel, where c is the number of channels; Specifically, S33 is: The maximum activation value vector is normalized to obtain the fNIRs channels whose maximum activation information exceeds 0.6, and the brain regions where the obtained channels are located are taken as the task-active brain regions.
7. The cognitive task classification method based on a dual-signal alternating guidance network according to claim 1, characterized in that, In S4, the method for obtaining the temporal characteristics of EEG is as follows: B1. Input the EEG signal sequentially into the first two-dimensional convolutional layer, the first batch normalization layer, the first exponential linear unit and the average pooling layer to generate the first feature; B2. Input the first feature into a multi-scale convolutional network sequentially to obtain the second feature; The MSTCN network includes a first residual block, a second residual block, and a third residual block connected in sequence. The first residual block, the second residual block, and the third residual block have the same structure, each including a first dilated causal convolution submodule and a second dilated causal convolution submodule that are connected to each other. The input of the first dilated causal convolution submodule is used as the input of the structure, and the input of the first dilated causal convolution submodule is added to the output of the second dilated causal convolution submodule to get the output of the structure. The first and second dilated causal convolution submodules have the same structure, both including a dilated causal convolution sublayer, a batch normalization sublayer, an ELU activation function sublayer, and a Dropout sublayer connected in sequence. The dilated causal convolutional sublayer includes a first dilated causal convolutional unit, a second dilated causal convolutional unit, and a third dilated causal convolutional unit connected in sequence. B3. Input the second feature into the temporal attention module to obtain the temporal features of the EEG.
8. The cognitive task classification method based on a dual-signal alternating guidance network according to claim 7, characterized in that, In B3, the time feature F of EEG is obtained. e_out The specific expression is: In the formula, d eeg For K eeg The dimension size, T is the transpose sign, softmax(·) is the softmax activation function, Q eeg Let K be the first query matrix. eeg Let V be the first bond matrix. eeg This is the first-value matrix; Q eeg =F e_CCV W e_Q K eeg =F e_CCV W e_K V eeg =F e_CCV W e_V In the formula, F e_CCV As the second feature, W e_Q W is the first weight matrix. e_K W is the second weight matrix. e_V This is the third weight matrix, used to perform the linear transformation of the second feature.
9. The cognitive task classification method based on a dual-signal alternating guidance network according to claim 8, characterized in that, In S4, the method for obtaining the spatial features of fNIRS is as follows: C1. Input the fNIRS signal into the spatial attention module to obtain the spatial attention feature F. f_sa ; In the formula, softmax(·) is the softmax activation function, and Q... fnirs For the second query matrix, K fnirs Let d be the second bond matrix. fnirs For K fnirs The size of the dimension, V fnirs This is the second-value matrix; Q fnirs =X fnirs W f_Q K fnirs =X fnirs W f_K V fnirs =X fnirs W f_V In the formula, X fnirs For fNIRS signal, W f_Q W is the fourth weight matrix. f_K W is the fifth weight matrix. f_V This is the sixth weight matrix, used to perform the linear transformation of the fNIRS signal; C2. Input the spatial attention features into the spatial convolution module to obtain the spatial features of fNIRS; The spatial convolution module includes a second two-dimensional convolutional layer, a second batch normalization layer, and a thermal exponential linear unit connected in sequence. The second two-dimensional convolution has 32 filters and a kernel size of (C2,1), where C2 is the number of channels of fNIRS in the task-active brain region.
10. The cognitive task classification method based on a dual-signal alternating guidance network according to claim 9, characterized in that, S5 includes the following steps: S51. Perform a linear transformation on the spatial features of fNIRS to obtain the third query matrix Q. fusion and the third bond matrix K fusion A linear transformation is performed on the time characteristics of the EEG to obtain the third-value matrix V. fusion ; Q fusion =W T_q F f_out K fusion =W T_k F f_out V fusion =W T_v F e_out In the formula, W T_q W is the fourth weight matrix. T_k W is the fifth weight matrix. T_v This is the sixth weight matrix; S52. Calculate the autocorrelation matrix based on the third query matrix and the third key matrix, and use the autocorrelation matrix and the third value matrix to calculate the dual-modal fusion feature F. st ; In the formula, d fusion For K fusion The size of the dimension, W fusion The autocorrelation matrix is expressed as follows: S53. Convert the dual-modal fusion features into a one-dimensional vector, input it into the classification layer, and calculate the classification probability of the classification result output by the classification layer using the softmax function. The classification result with the highest probability is taken as the classification result for the brain cognitive task. The specific expression for the loss function L of the classification layer is as follows: In the formula, N is the number of samples, C is the number of categories, ω(·) is the exponential function, and y n For the predicted label of the nth sample, l c For the true label of the c-th category, p c This represents the conditional probability of the c-th category in the output.
Citation Information
Patent Citations
Bimodal signal fusion method based on adaptive space-time convolution attention network
CN118626940A
Cognitive state classification method based on EEG-fNIRS deep fusion of feature decoupling
CN119397339A