Universal brain-computer interface decoding method

By constructing an Enhanced-DNN that combines TRCA with causal convolution and dilated convolution, the problem of insufficient generalization ability of brain-computer interface decoding algorithms under multiple paradigms is solved, achieving efficient and lightweight EEG signal decoding and reducing visual fatigue.

CN121209697APending Publication Date: 2025-12-26GUOKEXINNAO (SHANGHAI) INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511353225.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing brain-computer interface decoding algorithms are insufficient in cross-paradigm adaptability and generalization ability. In particular, when faced with multiple SSVEP paradigms, it is difficult to maintain high classification accuracy, and traditional methods are prone to causing visual fatigue.

Method used

We construct an enhanced deep neural network (Enhanced-DNN), combine it with a spatial filter trained by TRCA for data augmentation, and introduce causal convolution, dilated convolution, and multi-layer fully connected networks for EEG signal decoding in multiple paradigms.

Benefits of technology

It exhibits superior classification performance under various SSVEP paradigms, reduces visual fatigue, adapts to different time window lengths, and achieves efficient decoding with a lightweight network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121209697A_ABST
    Figure CN121209697A_ABST
Patent Text Reader

Abstract

The invention discloses a universal brain-computer interface decoding method, and relates to the technical field of neural network algorithms. In order to solve the technical problems that an existing brain-computer interface decoding algorithm is poor in universality, performance is reduced under a complex normal form, and light weight and high classification performance are difficult to balance, the method comprises the steps that original electroencephalogram signals are preprocessed and enhanced through TRCA (Task Related Component Analysis); and decoding is realized in combination with an enhanced deep neural network containing causal convolution, expansion convolution and a multi-layer full-connection layer. The method is verified by experiments under three normal forms of steady-state visual evoked potential SSVEP, high-frequency SSVEP and SSPVEP, not only maintains the light weight of the network to reduce the over-fitting risk, but also shows the classification performance superior to that of a traditional algorithm and an existing deep learning algorithm when the normal form difficulty is improved and the data length is increased, has strong universality and practicability, and is suitable for popularization and application. The method can be widely applied to brain-computer interface related fields such as medical rehabilitation and virtual reality.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of neural network algorithm, and particularly relates to a general-purpose brain-computer interface decoding method. BACKGROUND

[0002] Brain-computer interface aims to establish direct communication between brain and external devices, and converts brain signals into external instructions. At present, many BCI-related products have been applied to human life, and rapidly developed in the fields of medical rehabilitation, virtual reality and brain-to-brain communication. In order to meet different application requirements, a large number of BCI paradigms have been generated. Commonly used brain-computer interface paradigms include motor imagery (MI), P300 and steady-state visual evoked potential (SSVEP). Among them, SSVEP is the most mature and has the highest robustness. By detecting and analyzing the frequency characteristics of SSVEP signals, the BCI system can quickly identify the target of user attention, so as to realize efficient and natural human-computer interaction.

[0003] One of the main challenges faced by SSVEP is that visual stimulation, such as high-frequency flashing LEDs or screens, can cause visual fatigue, headache or discomfort. This problem is particularly evident when the flashing occurs at lower frequencies (e.g. 8-15 Hz). These frequencies are more easily perceived by the human eye and can interfere with the user's daily experience. To alleviate this problem, researchers have created many SSVEP derivative paradigms, such as high-frequency SSVEP, SSPVEP, etc. However, the decoding algorithms of these paradigms rarely consider their generality. There is an urgent need for a general algorithm tool in the field of brain-computer interface, which can be used to evaluate the feasibility of paradigm improvement, verify the effectiveness of new paradigms, and systematically design hybrid paradigms.

[0004] Traditional SSVEP decoding algorithms include CCA and TRCA. The core idea of CCA (canonical correlation analysis) is to determine the best projection direction of electroencephalogram signals through frequency matching, and to identify the target frequency using the maximum correlation coefficient. Due to its high efficiency and ease of implementation, CCA has become a classic decoding algorithm in the field of SSVEP-BCI. TRCA (task-related component analysis) enhances the repeatability of SSVEP in multiple trials through spatial filtering, thereby enhancing electroencephalogram signals. When a subject performs the same task multiple times, the evoked components of the same nature should be included in the electroencephalogram signals. In theory, TRCA is applicable to any BCI paradigm involving decoding of evoked signals with stable waveform characteristics. In terms of generality, CCA has better adaptability across subjects, while TRCA shows stronger generalization ability across paradigms. However, when the paradigm classification is difficult, the performance of TRCA will decrease.

[0005] In recent years, the progress of deep learning has fundamentally changed the decoding algorithm of BCI. The data-driven end-to-end deep learning method provides a powerful classification solution for SSVEP and its derivative paradigms. When faced with paradigms whose signal features are not easy to distinguish, deep learning models can still effectively complete the classification task with their strong feature extraction ability. There have been a number of deep learning-based SSVEP classification studies that have demonstrated state-of-the-art performance. For example, Shubin Zhang et al. used dynamic graph convolutional neural networks to classify multi-channel electroencephalogram data, achieving international leading level; Jianbo Chen et al. used a deep learning model based on Transformer to classify two low-frequency SSVEP data sets, demonstrating its potential to alleviate the calibration process in practical applications; Wenlong Ding et al. used FBtCNN combined with time-domain convolutional neural networks and filter banks to improve performance in short time window scenarios. However, these deep learning methods have not been widely verified in multiple paradigms.

[0006] Few deep learning algorithms consider cross-paradigm generalization ability, and academia is still constantly proposing paradigms for specific tasks, highlighting the urgent need for a robust electroencephalogram classification algorithm with broad applicability. Considering the generalization ability of TRCA and the feature extraction ability of deep learning networks, combining the two is expected to achieve high classification performance in multiple paradigms. Given the large room for improvement of deep learning networks, designing a neural network that can more effectively integrate with TRCA may be able to achieve high classification accuracy in multiple paradigms. SUMMARY

[0007] The application provides a general-purpose brain-computer interface decoding method, constructs an enhanced deep neural network (Enhanced-DNN), performs data enhancement on original electroencephalogram signals through a spatial filter obtained by TRCA training, and then classifies by using an enhanced convolutional neural network; the lightweight neural network is more suitable for fusion with TRCA; in order to make up for the deficiency of the lightweight network in the feature extraction capability, dilated convolution, causal convolution and multi-layer fully connected network are introduced. In order to verify the effectiveness of the proposed model, performance evaluation is carried out under various paradigms and different time window lengths; the generalization ability of the enhanced deep neural network (Enhanced-DNN) is verified through three different paradigms; as the decoding difficulty of the paradigm increases, the Enhanced-DNN shows better performance than other advanced algorithms; at the same time, as the length of the experimental data increases, the performance gap between the Enhanced-DNN and other algorithms is further widened; TRCA is introduced in the data preprocessing stage, causal convolution and dilated convolution are integrated into the DNN, and a multi-layer fully connected structure is added, so that the network is lightweight, and leading performance is obtained; the above solves the problems in the background art.

[0008] To solve the above technical problems, the application is realized by the following technical solutions:

[0009] The general-purpose brain-computer interface decoding method of the application comprises the following steps:

[0010] 1. Determine the stimulation paradigm and experimental design

[0011] Three typical SSVEP paradigms are selected to verify the decoding performance, and the parameters of each paradigm are as follows:

[0012] SSVEP paradigm: 35 healthy subjects (17 females, age 17-34 years, average 22 years) are used in the Benchmark dataset; the experiment is divided into 6 blocks, and each block has 40 trials (corresponding to 8-15.8 Hz, 0.2 Hz interval flashing characters); each trial lasts 5 seconds (including 0.5 second visual prompt); the electroencephalogram signal is sampled at 1000 Hz and down-sampled to 250 Hz; 7 visual-related electrode channels of PO5, PO3, PO4, PO6, O1, Oz and O2 are selected.

[0013] High-frequency SSVEP paradigm: 30 healthy subjects (16 females, age 21-35 years, average 26.5 years) are used in the SSVEPs open dataset (after cropping); 31-60 Hz (1 Hz interval) high modulation depth stimulation conditions are retained; the experiment is divided into 4 time stages, each stage has 12 blocks, and each block has 30 trials; each trial lasts 7 seconds (including 1 second attention guide, 5 seconds flashing stimulation, and 1 second rest); the electroencephalogram signal is sampled and the channel selection is the same as the SSVEP paradigm.

[0014] SSPVEP paradigm: self-experiment data, 10 healthy subjects (3 females, age 22-28 years, mean 25 years); 5x8 visual stimulation matrix (40 targets) was used, stimulation frequency 8-15.6 Hz (interval 0.4 Hz), odd targets were flickering stimuli, even targets were interspersed; the stimulation was generated by the formula s = 0.5 + 0.5 x sin(2pft + f), (f is frequency, f is phase, t is sampling point, s is screen brightness); the display screen resolution is 1920x1080 pixels, refresh rate 60 Hz; the electroencephalogram signal is sampled at 1250 Hz, down-sampled to 250 Hz, de-interference by 50 Hz notch filter, electrode impedance is controlled below 10 kQ, channel selection is the same as the previous two paradigms.

[0015] 2. Raw EEG signal acquisition

[0016] The electroencephalogram acquisition device (such as Neuroscan SynAmps264-256 electrode amplifier) is used, and the user wears the electroencephalogram cap to collect the scalp surface electroencephalogram signal; the signal is transmitted to the computer after being processed by the amplifier, and only the signals of the visual related brain area electrode channels (such as PO5, PO3, PO4, PO6, O1, Oz, O2) are retained.

[0017] 3. TRCA pre-processing enhancement

[0018] The spatial filter is trained through the TRCA algorithm to enhance the original electroencephalogram signal, and the steps are as follows:

[0019] Let the electroencephalogram signal of a single experiment in the training stage be X i (dimension: number of channels x time points), and the number of trials is N t , the superposition weight coefficient w of each electrode channel is solved by formula (1):

[0020]

[0021] In formula (1), X i is the EEG signal (dimension: number of channels x time points) of a single experiment in the first stage (data training stage). The weight coefficient can be obtained by formula (2). N t is the number of trials.

[0022]

[0023] Cov(X i ,X j ) represents the signal similarity between different channels between the i-th and j-th trials. Finally, the eigenvector corresponding to the maximum eigenvalue of Q -1 S is the weight coefficient of each brain area channel

[0024] The original brain electrical signal is filtered through the TRCA spatial filter to obtain a filtered signal Yn=WnXn (Wn is the spatial filter of the nth stimulus); the signal is divided into Nk subbands, and the Nf x Nk filtered signals of the Nf stimuli are rearranged: the signals of different spatial filters are arranged as new channels along the x axis, and the subband signals are arranged along the z axis to form a multi-channel input signal (for example, under the SSPVEP paradigm, after 7 original channels are processed by 40 spatial filters, 40 new channels are formed, and the signal size is Nf x Ns x Nk, Ns is the signal length).

[0025] 4. Enhanced-DNN decoding classification

[0026] An enhanced deep neural network (Enhanced-DNN) is designed, which includes the following levels in turn, and the functions and parameters of each level are as follows:

[0027] Subband fusion layer: 1x1 convolution kernel is used to fuse the harmonic and second harmonic information of the three subbands, and the target stimulus frequency characteristics are fully utilized (especially suitable for low-frequency SSVEP / SSPVEP).

[0028] Spatial information extraction layer: 120 N x 1 convolution kernels (N is the number of target stimuli) are used to extract and accumulate the signal information of different spatial filters.

[0029] Downsampling layer: 120 1x2 convolution kernels are used, and the step is set to 2 to reduce the network complexity and meet the lightweight demand.

[0030] Causal convolution layer: The size of the convolution kernel is adjusted according to the signal time window (1x20 for 1 second, 1x40 for 2 seconds, and 1x60 for 3 seconds and above); the left side is padded with zeros to ensure that the output sequence length is consistent with the input, so that the current time step output only depends on the current and historical input, which meets the physiological characteristics of brain signals and adapts to real-time processing.

[0031] Dilated convolution layer: 120 1x10 convolution kernels are used, and the dilation factor is set to 2; the receptive field is expanded under low computational complexity to capture the long-term dependence characteristics of brain signals, avoid using large-size convolution kernels, and reduce the risk of overfitting.

[0032] Multi-layer fully connected layer: 2 layers of fully connected layers are set to improve the decoding ability and optimize the classification accuracy.

[0033] 5. Experimental process control

[0034] The decoding process is divided into two stages:

[0035] Training stage: User fixate 4 targets (each for several seconds) for 10 times; System collect EEG data, complete TRCA spatial filter training and Enhanced-DNN parameter calibration.

[0036] Using stage: No need to retrain, stimulation flicker at fixed frequency / phase; User fixate target, stimulation stop, system output corresponding control instruction or character.

[0037] The present application has the following beneficial effects relative to the prior art:

[0038] (1) The previous EEG signal decoding method mostly only focuses on achieving high performance in a single paradigm, while we verify the generalization ability of Enhanced-DNN through three different paradigms; It is verified to be effective under the three paradigms of SSVEP, high-frequency SSVEP and SSPVEP, solving the "single paradigm adaptation" problem of existing algorithms, and supporting multi-paradigm BCI system design and optimization.

[0039] (2) The experimental results show that, as the decoding difficulty of the paradigm increases (such as from SSVEP to SSPVEP), Enhanced-DNN shows better performance compared to other traditional CCA / TRCA and existing deep learning algorithms; At the same time, as the length of the experimental data increases (1 second to 4 seconds), the performance gap between Enhanced-DNN and other algorithms is further widened;

[0040] (3) TRCA is introduced in the data preprocessing stage, and causal convolution and dilated convolution are integrated into the DNN, and a multi-layer fully connected structure is added, so as to realize network lightweight while obtaining leading performance;

[0041] (4) Reduce DNN data volume through TRCA preprocessing, combined with causal / dilation convolution and simplified network structure, while ensuring lightweight (parameter size is smaller than traditional deep learning model), reduce overfitting risk, adapt to actual BCI device deployment;

[0042] (5) Support decoding of low visual fatigue paradigms such as high-frequency SSVEP / SSPVEP, reduce user discomfort, and improve long-term use feasibility.

[0043] Of course, any product implementing the present application does not necessarily need to achieve all the advantages described above at the same time. BRIEF DESCRIPTION OF DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed for the description of the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.

[0045] Figure 1 The overall flowchart of the general-purpose brain-computer interface decoding method of the present application;

[0046] Figure 2 The overall flowchart of the general-purpose brain-computer interface decoding method of the present application;

[0047] Figure 3 The schematic diagram of the multi-layer fully connected network structure;

[0048] Figure 4 The SSVEP paradigm stimulation interface diagram;

[0049] Figure 5 The high-frequency SSVEP paradigm data acquisition design diagram;

[0050] Figure 6 The SSVEP paradigm stimulation frequency-phase distribution diagram Figure 7 The SSVEP paradigm stimulation frequency-phase matrix diagram. DETAILED DESCRIPTION

[0051] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort fall within the scope of protection of the present application.

[0052] The main content of the present application is to design an algorithm that can still maintain excellent performance under multiple paradigms, and we construct Enhanced-DNN. As shown in Figure 1 , Enhanced-DNN first performs data enhancement on the original electroencephalogram signal through the spatial filter obtained by TRCA training, and then uses an enhanced convolutional neural network for classification. Studies have shown that lightweight neural networks are more suitable for fusion with TRCA. In order to make up for the shortcomings of lightweight networks in feature extraction capability, we introduce dilated convolution (as shown in Figure 2 ), causal convolution (as shown in Figure 3 ) and multi-layer fully connected network (as shown in Figure 4 ). In order to verify the effectiveness of the proposed model, we perform performance evaluation under multiple paradigms and different time window lengths

[0053] Stimulus paradigm:

[0054] Few algorithms in previous studies have been validated by multiple paradigms. The present application uses three paradigms to verify performance: SSVEP, high-frequency SSVEP, and SSPVEP (as shown in FIGS. Figure 5 、 6 , 7).

[0055] Detection algorithm:

[0056] The enhanced neural network model proposed in this study balances the lightweight demand while improving the feature extraction capability. In view of the risk of overfitting of complex neural networks, this study starts from a lightweight neural network. Considering that the DNN using harmonic, channel and time sub-band convolution has achieved the best performance on two benchmark datasets, this method well meets the demand for lightweight and has high performance, but its feature extraction and classification ability still needs to be improved when facing more complex classification tasks. The Enhanced-DNN proposed in this study introduces causal convolution and dilated convolution into the core idea of time convolution network to enhance the feature extraction capability, and uses two fully connected layers to improve the decoding ability.

[0057] The specific implementation of the user-friendly brain-computer interface proposed in the present application is described from the following aspects: stimulus paradigm, experimental procedure, signal acquisition, and signal processing.

[0058] Stimulus paradigm

[0059] The SSVEP paradigm uses the Benchmark dataset, which records data from 35 healthy subjects (17 females, age range 17-34 years, average age 22 years). The experiment is divided into 6 blocks, each containing 40 trials corresponding to 40 flashing characters randomly presented, with a flashing frequency range of 8-15.8 Hz (interval of 0.2 Hz). Figure 5 The stimulus interface is displayed. Each trial lasts 5 seconds, and the target character is presented through visual cues for 0.5 seconds before the trial starts, guiding the subject to focus attention. The electroencephalogram data is collected by a SynAmps2 (Neuroscan Inc.) device at a sampling rate of 1000 Hz, and then down-sampled to 250 Hz. The average visual delay of the subjects in this dataset is about 140 milliseconds. This dataset uses 64-channel electroencephalogram signals. During the experiment, the subjects are required to focus on the target character and avoid blinking to ensure data quality. To reduce visual fatigue, a few minutes of rest time are arranged between every two blocks. The data length set in this study is 1 second, 2 seconds and 3 seconds. Seven electroencephalogram channels (PO5, PO3, PO4, PO6, O1, Oz and O2) are used in our experiment.

[0060] The open datasets used by high-frequency SSVEPs are simply referred to as SSVEPs. For example... Figure 6 As shown, this dataset records experimental data from 30 healthy subjects (16 women, aged 21–35 years, mean age 26.5 years), using frequencies ranging from 1–60 Hz with 1 Hz intervals. The experiment included 120 stimulus conditions (60 frequencies × 2 modulation depths), divided into four time phases, completed on different dates. Each time phase contained 12 blocks, each block containing 30 trials, randomly presenting 15 combinations of frequencies and two modulation depths, with each stimulus condition repeated 12 times. Each trial lasted 7 seconds, including 1 second for attentional guidance, 5 seconds for a flickering stimulus phase where subjects were instructed to concentrate and avoid blinking, and a short rest period of 1 second. To avoid visual fatigue, rest periods of 1–10 minutes were arranged between every two blocks.

[0061] For the high-frequency SSVEP dataset used in this study, the focus is not on the impact of modulation depth on SSVEP response and user experience, but rather on the decoding performance of the algorithm under high-frequency stimulus conditions. Therefore, this study pruned the dataset, retaining only stimulus frequencies with high modulation depth and within the 31–60 Hz range. The data lengths set in this study were 1 second, 2 seconds, 3 seconds, and 4 seconds.

[0062] The SSPVEP paradigm used self-collected data for experiments. Ten healthy subjects participated in our experiment (3 women, aged 22-28 years, mean age 25 years). All participants were offline and had no history of brain disease. All subjects were required to sign informed consent forms before the experiment. EEG data acquisition was conducted at the Shanghai University Bioinformatics Interactive Brain-Computer Interface Laboratory. All experiments were approved by the Shanghai University Ethics Committee (ECSHU2024-072). All subjects were required to complete a questionnaire before the experiment to measure their comfort. EEG data was recorded using seven electrodes (PO3, PO4, PO5, PO6, O1, O2, OZ) using the Synamp2 system. The sampling rate was 1250 Hz, and the acquisition device was a Neuroscan 64-electrode EEG cap. Electrode impedance was kept below 10 kΩ. All EEG data were downsampled to 250 Hz to improve data processing speed. A 50 Hz notch filter was used to reduce power line interference. The sample sinusoidal stimulation method proposed by Manyakov et al. (2013) was used to present visual flicker stimuli encoded using joint frequency-phase modulation on a 24-inch liquid crystal display (LCD). The visual stimulus was generated by the following formula:

[0063] s = 0.5 + 0.5 × sin(2πft + φ)

[0064] where f is the frequency of the stimulus, φ is the phase, t is the sample point, and s is the luminance of the screen. The sin(·) generates a sinusoidal wave, and the two constants 0.5 make the generated sinusoidal stimulus range between 0 and 1, where 0 represents complete darkness and 1 represents the brightest state. The resolution of the display screen is 1920 x 1080 pixels, and the refresh rate is 60 Hz. The user interface is a 5 x 8 visual stimulus matrix, which contains 40 targets in total. Specifically, 20 of the stimulus targets are labeled with linearly increasing frequencies, with the increment of each frequency being proportional to the stimulus index (as shown in Figure 7 Figure 7 Odd-numbered targets are encoded with flickering stimuli, while even-numbered targets are arranged between the odd-numbered targets. The stimulus frequency ranges from 8 Hz to 15.6 Hz, with an interval of 0.4 Hz. Each stimulus is presented in a disc (with a radius of 70 pixels). All of the design parameters are shown in Figure 7

[0065] Experimental procedure

[0066] When a user uses the brain-computer interface system proposed in the present application, the output of each control instruction is divided into two stages. The first stage is called the training stage. In the training stage, the user needs to first fixate on each target for a few seconds (the number of targets in the present application is 4), and repeat the fixation of all targets about 10 times. The computer will calibrate the user's electroencephalogram data in this stage and complete the training process of the algorithm. The second stage is the use stage of the user. In this stage, the user does not need to retrain the data. All stimuli start to flicker with the same frequency and phase. The user can choose to fixate on the target corresponding to the control instruction or character that he or she wants to output. The computer will output the corresponding control instruction or character when each stimulus stops flickering. The time of the two stages for each target output can be different for the specific system algorithm and the number of control instructions implemented.

[0067] Signal acquisition

[0068] ​​The present application does not make specific limitations on the device for collecting brain electrical signals. The brain electrical signals collected by most brain electrical signal collection devices on the market can meet the paradigm proposed by the present application. For the need of listing examples, Neuroscan SynAmps264-256 brain electrical amplifier is used here. This device supports the connection of a maximum of 64 electrode brain caps and supports the parallel connection of four devices to realize the collection of 256 electrode brain electrical signals. In use, the user needs to wear a collection head set (brain cap). The brain cap directly collects the brain electrical signals on the scalp surface and sends them to the signal amplifier for processing. Finally, the processed signals are sent to the computer end. In this patent, only the electrodes corresponding to the brain regions related to vision are selected.

[0069] Signal processing (use of TRCA algorithm, not limited to using this algorithm)

[0070] The TRCA algorithm obtains the superposition coefficients of each brain region electrode channel by solving formula (1).

[0071]

[0072] In formula (1), X i is the EEG signal (dimension: channel number x time point) of one experiment in the first stage (data training stage). The weight coefficient can be obtained by formula (2). N t is the number of trials.

[0073]

[0074] Cov(X i ,X j ) represents the signal similarity between different channels between the i-th and j-th trials. Finally, the eigenvector corresponding to the maximum eigenvalue of Q -1 S is the weight coefficient

[0075] Signal decoding

[0076] TRCA enhances the data features before passing them to the neural network. During the training process, TRCA applies a spatial filter to extract the task-related components from each trial. Subsequently, the TRCA filtered signals are rearranged to form a new multi-channel signal and used as input for the DNN model for classification. Figure 1 The flowchart of the proposed method is shown. The calibration signal of the stimulus Xn is divided into three sub-band signals by a filter bank.

[0077] The three sub-band signals are used to train the TRCA model and obtain spatial filters for each sub-band. Then Wn x Xn obtains the TRCA filtered signal Yn. For the nth stimulus, each sub-band can obtain the corresponding TRCA filtered signal. Since there are Nf stimuli and Nk sub-bands, a total of Nf x Nk TRCA filtered signals are obtained. We rearrange the TRCA filtered signals according to the following rules. The TRCA filtered signals from Nf spatial filters are rearranged. The filtered signals processed by different spatial filters have complementary EEG signal information, just like the signals collected from electrodes at different positions on the scalp. Therefore, the TRCA filtered signals from different spatial filters will be reorganized along the x-axis into new channels. For self-collected data with 40 stimulus targets under the SSPVEP paradigm, the original EEG signal has 7 channels, and after filtering by 40 spatial filters, the TRCA filtered signal will have 40 new channels. The multi-channel signal from the sub-band signal will be reorganized along the z-axis to take advantage of the harmonic contribution. For a paradigm containing Nf stimuli and lasting for Ns, the size of the reorganized TRCA filtered signal is Nf x Ns x Nk.

[0078] For each layer of the proposed neural network, we will explain its function and provide formal definitions. In addition, we also introduce the training strategy and further implementation details. All parameters are based on the case where the data length is 3 seconds under the SSPVEP mode. For parameter changes for different data lengths, we will discuss them below.

[0079] The first layer uses a 1 x 1 convolution kernel to fuse the three sub-bands, with the goal of fusing information from the three sub-bands to fully utilize the harmonics and second harmonics of the target stimulus frequency, which is particularly important for SSVEP and SSPVEP where the target stimulus frequency is in the low frequency band.

[0080] The second layer uses 120 N x 1 convolution kernels to extract and accumulate information from different spatial filters, where N represents the number of target stimuli. The N spatial filters are generated when we process the data using the TRCA method, and these spatial filters generate the corresponding filtered data in the preprocessing stage.

[0081] The third layer uses 120 1x2 convolutional kernels with a stride of 2 for downsampling and reducing the network complexity. We hope to build a lightweight network with strong feature extraction capability. Early on, we considered two ideas: combining DNN with long short-term memory network (LSTM) or adding an attention mechanism to DNN. However, the combination of DNN and TRCA itself has good performance, and adding other modules will lead to an overly complex network structure and increase the risk of overfitting. Therefore, we used some methods in TCN (temporal convolutional network), namely using dilated convolution and causal convolution. In our lightweight network, we no longer introduce residual blocks.

[0082] The fourth layer uses causal convolution. In causal convolution, the output of the current time step only depends on the input of the current and previous time steps, but not on the future input. In order to keep the output sequence length consistent with the input sequence length (or a specific length), causal convolution usually performs appropriate zero padding at the beginning of the input sequence, so that the convolution kernel does not go out of bounds during calculation. The sliding range of the convolution kernel is designed not to access future input data, and by padding zeros on the left, the convolution kernel can only access current and past input data. As shown in Figure 5 When the input sequence length is 5 and the convolution kernel size is (1, 3), two zeros are padded on the left of the input sequence to achieve causal convolution. The blue squares on the bottom of the figure represent time series data, with time advancing from left to right, and the white squares are padded zeros. The colored part above represents the convolution process. The output of these colored blocks comes from the current and past input values. Causal convolution preserves temporal causality, making it suitable for real-time processing and more consistent with the physiological characteristics of brain signals. We adjust the size of the convolution kernel according to the length of the time window. For larger time windows (e.g., 3 seconds, 4 seconds, or 4.6 seconds), because the amount of data is sufficient, we use 120 convolution kernels with a size of 1x60; when the time window is short, such as 2 seconds, the convolution kernel is set to 1x40; if the time window is very short, such as 1 second, we set the convolution kernel size to 1x20. We have verified that these parameters can achieve better results in all paradigms.

[0083] The fifth layer uses 120 1x10 convolutional kernels to extract information, and we use dilated causal convolution. As shown in Figure 6As shown, the bottom grid represents the input sequence, the dark blue squares represent the convolution kernels, and the top grid represents the convolution output. A convolution kernel of size (2,2) forms a 3×3 receptive field after expansion. Dilated convolution is an operation that introduces intervals during convolution, which can expand the receptive field while maintaining low computational complexity, thereby capturing dependencies over a longer time span. By adding an expansion factor *r* to causal convolution, the convolution operation is not limited to adjacent time points but can also sample across time points, thus capturing dependencies over a longer time span or a larger spatial range. This is particularly suitable for processing data with long-term dependency characteristics. Furthermore, it avoids using very large convolution kernels, meeting the design requirements of lightweight networks. In this study, better performance was achieved when the expansion factor was set to 2 and the convolution kernel size was set to 10.

[0084] Performance verification results

[0085] SSVEP paradigm: With a data length of 1 second, the classification accuracy is 12.3% higher than the traditional TRCA and 18.7% higher than CCA; with a data length of 3 seconds, the accuracy reaches 98.2%, close to the theoretical upper limit.

[0086] High-frequency SSVEP paradigm: With a data length of 4 seconds, the accuracy is improved by 9.5% compared to the Transformer-based model, and the inference speed is improved by 30% (due to the lightweight network structure).

[0087] The SSPVEP paradigm achieved an accuracy of 95.8% with a data length of 3 seconds, which is 7.2% higher than FBtCNN. Furthermore, it reduced visual fatigue scores in subjects by 40% compared to the SSPVEP paradigm (based on a 5-point questionnaire).

[0088] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A general-purpose brain-computer interface decoding method, characterized in that, The method comprises the following steps: S1, determining a stimulation paradigm and an experimental design: selecting at least one of the SSVEP paradigm, the high-frequency SSVEP paradigm and the SSVEP paradigm to verify the decoding performance; S2, collecting original electroencephalogram signals under the stimulation paradigm, and performing TRCA preprocessing on the original electroencephalogram signals: training a spatial filter through a TRCA algorithm, filtering and enhancing the original electroencephalogram signals to obtain a TRCA filtered signal; The TRCA filtered signal is divided into a plurality of sub-bands, and is rearranged according to the rule that the spatial filter signal is along the x-axis as a new channel and the sub-band signal is along the z-axis to form a multi-channel input signal; S3, inputting the multi-channel input signal into an enhanced deep neural network for classification decoding, wherein the enhanced deep neural network comprises a sub-band fusion layer, a spatial information extraction layer, a down-sampling layer, a causal convolution layer, a dilated convolution layer and a multi-layer full connection layer in sequence; wherein the sub-band fusion layer adopts a 1×1 convolution kernel, the spatial information extraction layer adopts an N×1 convolution kernel, N is the target stimulation quantity of the stimulation paradigm, the down-sampling layer adopts a 1×2 convolution kernel with a step of 2, the convolution kernel size of the causal convolution layer is adjusted according to the time window length of the electroencephalogram signal, and the dilated convolution layer has a dilated factor of 2 and adopts a 1×10 convolution kernel.

2. The general-purpose brain-computer interface decoding method of claim 1, wherein, In the S2 step, the TRCA algorithm solves the weight coefficients of the spatial filter by the following formula wherein X i is the electroencephalogram signal of a single experiment in the training stage, with dimensions of channel * time point, N t is the number of trials, Cov(X i , X j ) is the signal similarity between different channels between the ith and jth trials, is the maximum eigenvalue of Q -1 S, and the eigenvector corresponding to the maximum eigenvalue is the weight coefficient of each brain region channel.

3. The general-purpose brain-computer interface decoding method of claim 1, wherein, In the S2 step, the collection parameters of the original electroencephalogram signal are as follows: an electrode amplifier is adopted, and seven visual electrode channels of PO5, PO3, PO4, PO6, O1, Oz and O2 are selected.

4. The general-purpose brain-computer interface decoding method of claim 1, wherein, In the S2 step, the multi-channel input signal has a size of Nf×Ns×Nk; wherein Nf is the target stimulation quantity of the stimulation paradigm, Ns is the time length of the electroencephalogram signal, and Nk is the number of sub-bands, wherein Nk=3.

5. The general-purpose brain-computer interface decoding method of claim 1, wherein, The SSVEP paradigm adopts a Benchmark data set, the electroencephalogram signal is sampled at 1000 Hz and down-sampled to 250 Hz.

6. The general-purpose brain-computer interface decoding method of claim 1, wherein, The high-frequency SSVEP paradigm adopts an SSVEPs open data set, and the high-frequency SSVEP paradigm retains a 31–60 Hz high modulation depth stimulation condition, the electroencephalogram signal is sampled at 1000 Hz and down-sampled to 250 Hz.

7. The general-purpose brain-computer interface decoding method of claim 1, wherein, The SSVEP paradigm adopts a Benchmark data set, the electroencephalogram signal is sampled at 1000 Hz and down-sampled to 250 Hz.

8. The general-purpose brain-computer interface decoding method of claim 1, wherein, In the S3 step, the convolution kernel size of the causal convolution layer is specifically as follows: when the time window of the electroencephalogram signal is 1 second, it is set to 1×20, when it is 2 seconds, it is set to 1×40, and when it is 3 seconds or more, it is set to 1×60; the causal convolution layer makes the output sequence length consistent with the input sequence length through left zero padding.

9. The general-purpose brain-computer interface decoding method of claim 1, wherein, In the S1 step, the parameters of each stimulation paradigm are as follows: SSVEP paradigm: stimulation frequency 8–15.8 Hz, interval 0.2 Hz, each trial lasts 5 seconds, contains 0.5 seconds of visual prompt, experiment is divided into 6 blocks, each block contains 40 trials; High-frequency SSVEP paradigm: stimulation frequency 31–60 Hz, interval 1 Hz, only high modulation depth condition is retained, each trial lasts 7 seconds, contains 1 second of attention guide, 5 seconds of flashing stimulation and 1 second of rest; SSPVEP paradigm: 5x8 stimulus matrix, 40 targets, 8-15.6 Hz stimulation frequency, 0.4 Hz interval, stimulation generated by the formula s=0.5+0.5xsin(2pi ft+phi), display resolution 1920x1080 pixels, refresh rate 60 Hz.

10. The general-purpose brain-computer interface decoding method of claim 1, wherein, Also includes experimental process control steps: decoding process is divided into training stage and use stage; in the training stage, the user gazes at 4 targets, each for several seconds and repeats 1 time, and the system completes the training and calibration of TRCA and enhanced deep neural network.