P300 spelling method based on convolutional mixer network

By using a convolution mixer network to extract spatiotemporal features in P300 waveform detection, and optimizing feature space with supervisory comparison learning, the problem of insufficient accuracy and robustness of P300 waveform detection in the prior art is solved, and higher accuracy and stable spelling performance are achieved.

CN120215699AActive Publication Date: 2025-06-27SOUTH CHINA UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510260279.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The prior art has low signal-to-noise ratio, noise sensitivity and space-time characteristics complexity in P300 waveform detection, resulting in poor detection accuracy and robustness.

Method used

The P300 spelling method based on the convolution mixer network is adopted to extract the spatiotemporal features of the EEG samples through the time convolution mixer and the spatial convolution mixer, and the feature space is optimized using supervised comparison learning to reduce the dependence on a large amount of labeled data.

Benefits of technology

It improves the accuracy and robustness of P300 waveform detection, reduces dependence on a large amount of training data, and achieves more stable and high-precision spelling performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215699A_ABST
    Figure CN120215699A_ABST
Patent Text Reader

Abstract

The invention discloses a P300 spelling method based on a convolutional mixer network. The P300 spelling method comprises the following steps: acquiring an electroencephalogram sample and carrying out preprocessing operation; a time convolution mixer is used for effectively extracting time features of electroencephalogram samples, and the learning ability of the network to time information is improved; compressing and mixing the spatial features by using a spatial mixer, and extracting effective spatial features; carrying out average pooling operation on the spatial features by using a pooling layer; generating the probability that the electroencephalogram sample contains the P300 waveform by using an output layer; and judging the row and the column of the electroencephalogram sample containing the P300 waveform, determining the position of a target character, and realizing the output of the spelling character. According to the method, the deep spatial-temporal characteristic relation of the electroencephalogram signals can be effectively mined through the convolutional mixer network, the valuable characteristic extraction capability of the network is enhanced, the P300 character spelling performance is remarkably improved, and the method has high practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of P300 character spelling, and in particular to a P300 spelling method based on a convolutional mixer network. Background Art

[0002] Brain-computer interface technology provides a way of communication that does not rely on body movement by directly interpreting brain activities and converting them into external device control instructions. Among them, character spelling based on the P300 waveform in event-related potentials is a typical application, which allows users to output specific characters through electroencephalogram (EEG) signals. However, due to the low signal-to-noise ratio, noise sensitivity, and complexity of spatio-temporal characteristics, the accurate detection of the P300 waveform is challenging.

[0003] Currently, the research on P300 detection is mainly divided into two categories: traditional machine learning methods and deep learning methods. Traditional machine learning methods mainly rely on manual feature extraction and classical classifiers. The feature selection and classification models are relatively transparent and have strong interpretability. In the application of P300 signal decoding, methods such as support vector machines, linear discriminant analysis, and stepwise linear discriminant analysis are widely used. However, these traditional methods rely heavily on manual features and are easily affected by inter-subject variability, resulting in poor robustness and generalization ability of the models. With the rise of deep learning, more and more methods based on deep neural networks have been applied to the decoding of P300 signals. Convolutional neural networks are widely used to automatically learn spatio-temporal features in EEG signals, thus effectively improving the detection accuracy of P300 and performing well in character recognition tasks. The EEGNet model optimizes the extraction of spatio-temporal patterns of EEG signals by combining a depthwise separable convolutional structure, further enhancing its generalization ability in complex EEG data. Capsule networks can better handle the hierarchical relationship of EEG signals, significantly improving the decoding effect of P300 signals. These deep learning methods can automatically extract more complex and representative features, so they are significantly superior to traditional methods in terms of performance. However, deep learning models often require a large amount of training data to avoid overfitting, and in a low signal-to-noise ratio environment, the robustness of the models still faces challenges. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a P300 spelling method based on a convolutional mixer network, which can more effectively extract spatio-temporal features of EEG samples, and use label information to enhance the similarity within the same category and expand the differences between different categories, optimize the feature space, thereby reducing the network's dependence on a large amount of labeled data and providing more stable and accurate spelling performance.

[0005] To achieve the above object, the technical solution provided by the present invention is: a P300 spelling method based on a convolutional mixer network, where the convolutional mixer network includes a temporal convolutional mixer, a spatial convolutional mixer, a pooling layer, and an output layer; the temporal convolutional mixer captures important features of EEG samples at the temporal level through convolutional operations; the spatial convolutional mixer is used to extract correlation features between different channels in the spatial dimension to enhance the expression of spatial information; the pooling layer reduces the dimension of features through average pooling operations to highlight important features; the output layer maps the extracted features to a low-dimensional space and outputs the probability of the P300 waveform in the EEG sample, thereby determining the spelled character;

[0006] The specific implementation of the P300 spelling method includes:

[0007] Obtain EEG samples and preprocess the EEG samples to obtain standardized and normalized EEG samples;

[0008] Use the trained convolutional mixer network to perform the following processing on the preprocessed EEG samples:

[0009] Mix the preprocessed EEG samples using the temporal convolutional mixer to effectively extract the temporal features of the EEG samples and obtain a temporal feature map; wherein, the temporal convolutional mixer includes an embedding layer, depthwise convolution, and pointwise convolution, and both the depthwise convolution and the pointwise convolution are followed by an ELU activation function and batch normalization, and there is a residual connection between the embedding layer and the pointwise convolution. The embedding layer divides the EEG samples into smaller time blocks and expands them into multiple feature maps. The depthwise convolution performs individual convolution on each feature map to obtain local temporal features. The pointwise convolution fuses the local temporal features to optimize the global information. The ELU activation function and batch normalization are used to improve the network expression ability and optimize the feature distribution. The residual connection is used to prevent information loss;

[0010] Input the temporal feature map output by the temporal convolutional mixer into the spatial convolutional mixer to extract effective spatial features and obtain a spatial feature map; wherein, the spatial convolutional mixer includes convolution, an ELU activation function, and batch normalization. The convolution compresses and mixes the temporal features at the spatial level, mapping multiple channels into one channel. The ELU activation function and batch normalization are used to improve the network expression ability;

[0011] Perform average pooling operations on the spatial feature map output by the spatial convolutional mixer using the pooling layer to reduce the spatial dimension, reduce the computational complexity, and extract more significant features to obtain pooled features;

[0012] The output layer flattens and fully connects the pooling features output by the pooling layer, and outputs them to a single neuron to obtain the probability that the EEG sample contains the P300 waveform.

[0013] Use the above convolutional mixer network for prediction, obtain the probability that the EEG samples corresponding to multiple flashes in a round of character spelling task contain the P300 waveform, judge the row and column where the EEG sample containing the P300 waveform is located, so as to determine the target character position and realize the output of the spelling character.

[0014] Furthermore, EEG samples are obtained from the P300 speller, which contains a 6×6 character matrix and a total of 36 characters; when the user wants to spell a certain character, this character is the target character. During the period of gazing at the target character, each row and column of the character matrix will flash once. The first six times are column flashes, and the last six times are row flashes, for a total of 12 flashes. 12 EEG samples are recorded during the 12 flashes; when the row or column where the target character is located flashes, the user's brain generates an obvious P300 waveform, and when the row or column that does not contain the target character flashes, the user's brain does not generate a P300 waveform; the row and column of the stimulus are identified by detecting the presence of the P300 waveform in the EEG sample, so as to infer the target character that the user is concerned about; the obtained EEG samples are 2 seconds long and contain data of 32 channels. To obtain the complete waveform of this P300, the EEG data time window from 0 to 610 ms starting from the visual stimulus is intercepted and downsampled to 78 time points; the EEG data of 32 channels are band-pass filtered from 0.5 to 20 Hz, and the data are standardized and normalized; finally, the obtained EEG sample X = {x1, x2, …, x i , …, x N} has a dimension of 78×32, where x i is the i-th EEG sample, and N represents the number of EEG samples input into the convolutional mixer network; the EEG sample X includes two types of EEG samples with and without the P300 waveform. The label Y = {y1, y2, …, y i …, y N} corresponding to each EEG sample indicates whether the EEG sample contains the P300 waveform. If it contains, the label is 1, and if it does not contain, the label is 0, where y i is the label corresponding to the i-th EEG sample.

[0015] Furthermore, the temporal convolutional mixer extracts features in the temporal dimension from the preprocessed EEG samples and outputs a temporal feature map with local and global features, specifically as follows:

[0016] First, use convolution as the embedding layer to divide the preprocessed EEG samples into small time blocks and expand the input EEG samples into 40 feature maps.

[0017] F embedding = Conv(X)

[0018] In the formula, F embedding represents the expanded feature map, and Conv represents the convolution operation;

[0019] Then, depthwise convolution is used to separately convolve each feature map with a specific convolution kernel to generate a depth feature map with 40 channels to extract the local spatial features of each channel; the ELU activation function and batch normalization are applied after the depthwise convolution to optimize the global feature distribution and improve the performance and expressive ability of the network; a linear residual connection is added between the embedding layer and the pointwise convolution to retain the overall trend of time variation and prevent information loss;

[0020] F depth = F embedding + BN(ELU(ConvDepthwise(F embedding )))

[0021] In the formula, F depth represents the depth feature map, BN represents batch normalization, ELU represents the ELU activation function, and ConvDepthwise represents depthwise convolution;

[0022] Since depthwise convolution cannot effectively capture the feature correlations between channels, pointwise convolution is used to fuse the features generated by depthwise convolution from each dimension, ensuring global information fusion; ELU and batch normalization are also applied after the pointwise convolution to obtain the time feature map F temporal :

[0023] F temporal = BN(ELU(ConvPointwise(F depth )))

[0024] In the formula, ConvPointwise represents pointwise convolution.

[0025] Furthermore, the spatial convolution mixer performs spatial feature extraction and fusion on the time feature map, and outputs a spatial feature map, specifically as follows:

[0026] The spatial convolution mixer convolves the time feature map F temporal along the spatial dimension, integrates and extracts valuable features from different channels, and it compresses all channel information into a single channel, thereby achieving global spatial feature extraction and fusion while reducing the data dimension; the spatial convolution mixer also uses the ELU activation function and batch normalization to further improve the training efficiency and performance of the network, and finally obtains the spatial feature map F spatial :

[0027] F spatial = BN(ELU(Conv(F temporal )))。

[0028] Furthermore, the pooling layer performs average pooling operation on the obtained spatial feature map, further compressing and aggregating spatio-temporal information, reducing the number of parameters and minimizing the overfitting risk, to obtain the pooled feature F pooling :

[0029] F pooling = AvgPooling(F spatial )

[0030] In the formula, AvgPooling represents the average pooling operation.

[0031] Furthermore, the output layer first flattens the pooled feature F pooling this high-dimensional data into a one-dimensional representation, then maps the extracted features through a fully connected layer, and finally outputs to a single neuron; the output layer introduces the Sigmoid activation function to limit the output value within the range of [0, 1], representing the probability P that the EEG sample contains the P300 waveform target :

[0032] P target = Sigmoid(Linear(Flatten(F pooling )))

[0033] In the formula, Sigmoid represents the Sigmoid activation function, Linear represents the fully connected layer, and Flatten represents the flattening operation.

[0034] Furthermore, the loss function for training the convolutional mixer network is a hybrid loss function, which consists of the supervised contrast loss L S and the cross-entropy loss L C The supervised contrast loss establishes a better decision boundary in the feature space, enhances the network's understanding of the data, and improves its ability to recognize rare classes. The cross-entropy loss allows explicit classification of EEG samples through probability outputs and directly optimizes the classification accuracy. This hybrid loss function enables the network to extract different feature representations for samples within each class, thus significantly improving the network's feature learning ability and classification accuracy, while also enhancing the network's robustness and generalization ability and reducing the dependence on a large amount of labeled data. Among them, the supervised contrast loss L S is calculated using the pooled feature output by the pooling layer:

[0035]

[0036] In the formula, S + represents the EEG sample xi A set of samples with the same label, where p represents the sample in S + in S, |S + | represents the number of samples in the same batch with the same label as the EEG sample x i S represents the set of other samples in the EEG sample X except the EEG sample x, s represents the sample in S, and u i represents the pooled feature corresponding to the EEG sample x i u represents the pooled feature corresponding to the sample with the same label as the EEG sample x i u represents the pooled feature corresponding to other EEG samples except the EEG sample x, τ represents the temperature, which is used to control the smoothness of training and the influence on difficult samples; p u represents the pooled feature corresponding to the sample with the same label as the EEG sample x i u represents the pooled feature corresponding to the sample with the same label as the EEG sample x s u represents the pooled feature corresponding to other EEG samples except the EEG sample x, τ represents the temperature, which is used to control the smoothness of training and the influence on difficult samples; i The cross-entropy loss L

[0037] is calculated using the probability of the P300 waveform output by the output layer: C In the formula, p

[0038]

[0039] represents the probability of the P300 waveform output by the output layer corresponding to the EEG sample x; i The total loss L of the network is the mixed loss composed of the supervised contrast loss L i and the cross-entropy loss L

[0040] L = λL S +(1 - λ)L C In the formula, λ represents the regularization parameter, which is used to balance the contributions between the supervised contrast loss and the cross-entropy loss.

[0041] L = λL S +(1 - λ)L C

[0042] In the formula, λ represents the regularization parameter, which is used to balance the contributions between the supervised contrast loss and the cross-entropy loss.

[0043] Furthermore, the specific operations of using the convolutional mixer network for prediction and realizing the output of spelling characters are as follows:

[0044] In each round of the spelling task, each row and each column of the character matrix flash once, with the first six flashes being column flashes and the last six flashes being row flashes, for a total of 12 flashes. 12 EEG samples are recorded during the 12 flashes; the P300 waveform only appears when the row and column containing the target character flash; since the target row and column are unique, a probability maximization strategy is adopted to determine the row and column where the target character is located:

[0045]

[0046] Wherein, c represents the column where the target character is located, k represents the row where the target character is located, and P j represents the probability of containing P300 output by the electroencephalogram sample corresponding to the jth flash through the convolutional mixer network; the target character spelled by the user can be determined through the determined row and column where the target character is located, thereby realizing the output of the P300 spelling character.

[0047] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0048] 1. The present invention designs a convolutional mixer network, which divides the input electroencephalogram sample into smaller time blocks, uses methods such as depth convolution and separable convolution to integrate local and global information, then extracts spatio-temporal hybrid features, obtains the spatio-temporal dependence relationship for P300 waveform detection, and realizes P300 character spelling.

[0049] 2. The present invention designs a hybrid loss function, integrates supervised contrast learning into network training, and alleviates the limitations of unclear classification margins and susceptibility to noise when using cross-entropy loss alone. This joint supervision method enables the convolutional mixer network to learn deep features with both inter-class separability and intra-class compactness, thereby improving the prediction accuracy and generalization ability of the network.

[0050] 3. The present invention combines the convolutional mixer network with the idea of supervised contrast learning, can achieve good P300 character spelling performance, and can still be effectively trained with the support of a small amount of training data, thereby saving the cost and time of data collection and annotation, and having broad application prospects.

[0051] In summary, the present invention can effectively mine the deep feature relationship of electroencephalogram signals in the time and space dimensions, make full use of label information, enhance the ability of the network to extract valuable features from electroencephalogram signals, significantly improve the P300 detection ability and character spelling performance of the network, facilitate rapid deployment, and have high practical performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a framework diagram of the method of the present invention.

[0053] Figure 2 is a structural schematic diagram of the convolutional mixer network.

[0054] Figure 3 is a schematic diagram of the P300 character speller; in the figure, S in the upper white square is the target character, and the characters in the black area form a character matrix that can realize spelling.

[0055] Figure 4 is a structural schematic diagram of the time convolutional mixer; in the figure, ELU is the ELU activation function.

[0056] Figure 5 It is a schematic structural diagram of a spatial convolution mixer; in the figure, ELU is the ELU activation function. Specific embodiments

[0057] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0058] This embodiment discloses a P300 spelling method based on a convolutional mixer network. The convolutional mixer network includes a temporal convolutional mixer, a spatial convolutional mixer, a pooling layer, and an output layer; the temporal convolutional mixer captures important features of EEG samples at the temporal level through convolutional operations; the spatial convolutional mixer is used to extract the correlation features between different channels in the spatial dimension and enhance the expression of spatial information; the pooling layer reduces the dimension of the features through average pooling operations and highlights important features; the output layer maps the extracted features to a low-dimensional space and outputs the probability of the P300 waveform in the EEG sample, thereby determining the spelled character. In summary, the present invention designs a temporal convolutional mixer and a spatial convolutional mixer to extract the spatio-temporal features of EEG signals, introduces supervised contrast learning, optimizes the feature space, improves the classification performance, and realizes high-precision P300 character spelling. The method framework of the present invention is as Figure 1 shown, and the structure of the convolutional mixer network is as Figure 2 shown. The specific P300 spelling method includes the following steps:

[0059] Obtain EEG samples and preprocess the EEG samples to obtain standardized and normalized EEG samples;

[0060] Use the trained convolutional mixer network to perform the following processing on the preprocessed EEG samples:

[0061] Mix the preprocessed EEG samples using the temporal convolutional mixer to effectively extract the temporal features of the EEG samples and obtain a temporal feature map. Among them, the temporal convolutional mixer includes an embedding layer, a depth convolution, and a pointwise convolution. After the depth convolution and the pointwise convolution, there are an ELU activation function and batch normalization, and there is a residual connection between the embedding layer and the pointwise convolution; the embedding layer divides the EEG samples into smaller time blocks and expands them into 40 feature maps, the depth convolution performs individual convolution on each feature map to obtain local temporal features, the pointwise convolution fuses the local temporal features to optimize the global information, the ELU activation function and batch normalization are used to improve the network expression ability and optimize the feature distribution, and the residual connection is used to prevent information loss;

[0062] The temporal feature map output by the temporal convolutional mixer is fed into the spatial convolutional mixer to extract effective spatial features, obtaining a spatial feature map; wherein, the spatial convolutional mixer includes convolution, ELU activation function, and batch normalization; the convolution compresses and mixes the temporal features at the spatial level, mapping multiple channels into one channel, and the ELU activation function and batch normalization are used to enhance the network's expressive ability;

[0063] The average pooling operation is performed on the spatial feature map output by the spatial convolutional mixer using a pooling layer to reduce the spatial dimension, reduce the computational complexity, and extract more prominent features, obtaining a pooled feature;

[0064] The output layer flattens and fully connects the pooled features output by the pooling layer, outputs to a single neuron, and obtains the probability that the input EEG sample contains a P300 waveform;

[0065] Using the above convolutional mixer network for prediction, obtaining the probability that the EEG samples corresponding to 12 flashes in a round of character spelling task contain a P300 waveform, determining the row and column where the EEG sample containing the P300 waveform is located, thereby determining the target character position and realizing the output of the spelling character.

[0066] Specifically, obtain EEG samples and preprocess them, and the specific situation is as follows:

[0067] Obtain EEG samples from a traditional P300 speller. This P300 speller contains a 6×6 character matrix, with a total of 36 characters, as Figure 3 shown. When the user wants to spell a certain character, that character is the target character. During the period of gazing at the target character, each row and each column of the character matrix will flash once. The first six times are column flashes, and the last six times are row flashes, for a total of 12 flashes. 12 EEG samples are recorded for 12 flashes; when the row or column where the target character is located flashes, the user's brain generates an obvious P300 waveform, and when the row or column that does not contain the target character flashes, the user's brain does not generate a P300 waveform; by detecting the presence of the P300 waveform in the EEG samples, the row and column of the stimulus are identified, thereby inferring the target character that the user is concerned about; the obtained EEG samples have a duration of 2 seconds and contain data of 32 channels. To obtain the complete waveform of this P300, the EEG data time window from 0 to 610 ms starting from the visual stimulus is intercepted and downsampled to 78 time points; the EEG data of 32 channels is band-pass filtered from 0.5 to 20 Hz, and the data is standardized and normalized; finally, the obtained EEG sample X = {x1, x2, …, x i , …, x N} has a dimension of 78×32, where x iFor the i-th EEG sample, N represents the number of EEG samples input into the convolutional mixer network; the EEG sample X includes two types of EEG samples, those with P300 waveforms and those without P300 waveforms. The label Y = {y1, y2, …, y i …, y N} for each EEG sample indicates whether the EEG sample contains a P300 waveform. If it contains, the label is 1; if not, the label is 0, where y i is the label corresponding to the i-th EEG sample.

[0068] Specifically, the temporal convolutional mixer includes an embedding layer, depthwise convolution, and pointwise convolution, the ELU activation function and batch normalization after the depthwise convolution and pointwise convolution, and a residual connection between the embedding layer and the pointwise convolution, as Figure 4 shown. The preprocessed EEG samples are used with the temporal convolutional mixer for feature extraction in the time dimension, outputting a time feature map with local and global features:

[0069] First, a convolution is used as the embedding layer to divide the preprocessed EEG samples into smaller time blocks and expand the input EEG samples into 40 feature maps.

[0070] F embedding = Conv(X)

[0071] In the formula, F embedding represents the expanded feature map, and Conv represents the convolution operation;

[0072] Then, depthwise convolution is used to separately convolve each feature map with a specific convolution kernel to generate a depth feature map with 40 channels to extract the local spatial features of each channel; the ELU activation function and batch normalization are applied after the depthwise convolution to optimize the global feature distribution and improve the performance and representational ability of the network; a linear residual connection is added between the embedding layer and the pointwise convolution to retain the overall trend of time variation and prevent information loss;

[0073] F depth = F embedding + BN(ELU(ConvDepthwise(F embedding )))

[0074] In the formula, F depth represents the depth feature map, BN represents batch normalization, ELU represents the ELU activation function, and ConvDepthwise represents depthwise convolution;

[0075] Since depth convolution cannot effectively capture the feature correlations between channels, point convolution is used to fuse the features generated by depth convolution from each dimension, ensuring the fusion of global information. ELU and BatchNorm are also applied after pointwise convolution to obtain the temporal feature map F temporal :

[0076] F temporal = BN(ELU(ConvPointwise(F depth )))

[0077] In the formula, ConvPointwise represents pointwise convolution;

[0078] Specifically, the spatial convolution mixer includes convolution, ELU activation function, and batch normalization, as Figure 5 shown. It extracts and fuses spatial features from the temporal feature map, and outputs the spatial feature map:

[0079] The spatial convolution mixer convolves the temporal feature map F temporal along the spatial dimension, integrates and extracts valuable features from different channels. It compresses all channel information into a single channel, thereby achieving global spatial feature extraction and fusion while reducing the data dimension. This module also uses the ELU activation function and batch normalization to further improve the training efficiency and performance of the network, and finally obtains the spatial feature map F spatial :

[0080] F spatial = BN(ELU(Conv(F temporal )))

[0081] Specifically, the pooling layer performs average pooling operation on the obtained spatial feature map, further compresses and aggregates spatio-temporal information, reduces the number of parameters, and minimizes the risk of overfitting, to obtain the pooled feature F pooling :

[0082] F pooling = AvgPooling(F spatial )

[0083] In the formula, AvgPooling represents average pooling operation.

[0084] Specifically, the output layer performs a fully connected operation on the output of the pooling layer, finally obtains the probability of the P300 waveform in the input EEG sample, and conducts network iterative training through a hybrid loss function composed of supervised contrast loss and cross-entropy loss:

[0085] The output layer first takes the pooled feature F poolingThis high-dimensional data is flattened into a one-dimensional representation, and then the extracted features are mapped through a fully connected layer, and finally output to a single neuron. The Sigmoid activation function is introduced to limit the output value within the range of [0, 1], representing the probability P that the target EEG sample contains the P300 waveform. target :

[0086] P target = Sigmoid(Linear(Flatten(F pooling )))

[0087] In the formula, Sigmoid represents the Sigmoid activation function, Linear represents the fully connected layer, and Flatten represents the flattening operation.

[0088] The loss function for network training is a hybrid loss function, which consists of the supervised contrast loss L S and the cross-entropy loss L C . The supervised contrast loss establishes a better decision boundary in the feature space, enhances the network's understanding of the data, and improves its ability to identify rare classes. The cross-entropy loss allows for explicit classification of EEG samples through probability outputs and directly optimizes the classification accuracy. This hybrid loss function enables the network to extract different feature representations for samples within each class, thus significantly improving the model's feature learning ability and classification accuracy. At the same time, it also improves the network's robustness and generalization ability and reduces the dependence on a large amount of labeled data. Among them, the supervised contrast loss L S is calculated using the pooled features output by the pooling layer:

[0089]

[0090] In the formula, S + represents the set of samples with the same label as the EEG sample x i , p represents the sample in S + , |S + | represents the number of samples in the same batch with the same label as the EEG sample x i , S represents the set of other samples in the EEG sample X except the EEG sample x i , s represents the sample in S, u i represents the pooled feature corresponding to the EEG sample x i , u p represents the pooled feature corresponding to the sample with the same label as the EEG sample x i , u s represents the pooled feature corresponding to other EEG samples except the EEG sample x i , and τ represents the temperature, which is set to 0.07 and is used to control the smoothness of training and the impact on difficult samples.

[0091] Cross-entropy loss L C It is calculated using the probability of the P300 waveform contained in the output of the output layer:

[0092]

[0093] In the formula, p i represents the probability of the output of the output layer corresponding to the EEG sample x i containing the P300 waveform;

[0094] The total loss L of the network is the mixed loss composed of the supervised contrast loss L S and the cross-entropy loss L C :

[0095] L = λL S +(1 - λ)L C

[0096] In the formula, λ represents the regularization parameter, which is used to balance the contributions between the supervised contrast loss and the cross-entropy loss.

[0097] Specifically, using the above convolutional mixer network for prediction, the probability of the EEG samples corresponding to 12 flashes in one round of character spelling tasks containing the P300 waveform can be obtained, and the row and column where the EEG sample containing the P300 waveform is located can be determined, so as to determine the target character position and realize the output of the spelling character:

[0098] In each round of spelling tasks, each row and each column of the character matrix flash once. Among them, the first six times are column flashes, and the last six times are row flashes, for a total of 12 flashes. 12 EEG samples are recorded for 12 flashes. The P300 waveform only appears when the row and column containing the target character flash. Since the target row and column are unique, a probability maximization strategy is adopted to determine the row and column where the target character is located:

[0099]

[0100] In the formula, c represents the column where the target character is located, k represents the row where the target character is located, and P j represents the probability of containing P300 output by the convolutional mixer network for the EEG sample corresponding to the jth flash; the target character spelled by the user can be determined through the determined row and column where the target character is located, so as to realize the output of the P300 spelling character.

[0101] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A P300 spelling method based on a convolutional mixer network, characterized in that: The convolution mixer network includes a time convolution mixer, a space convolution mixer, a pooling layer and an output layer; the time convolution mixer captures important features of the EEG sample at the time level through convolution operation; the space convolution mixer is used to extract correlation features between different channels in the spatial dimension and enhance the expression of spatial information; the pooling layer reduces the dimension of features and highlights important features through average pooling operation; the output layer maps the extracted features to a low-dimensional space and outputs the probability of the EEG sample containing a P300 waveform, thereby determining the spelled character; The specific implementation of the P300 spelling method includes: Obtaining EEG samples, and preprocessing the EEG samples to obtain standardized and normalized EEG samples; The preprocessed EEG samples are processed as follows using the trained convolution mixer network: The preprocessed EEG samples are mixed using a time convolution mixer to effectively extract the time features of the EEG samples and obtain a time feature map; wherein the time convolution mixer includes an embedding layer, a deep convolution and a point-by-point convolution, and the deep convolution and the point-by-point convolution both contain an ELU activation function and batch normalization, and a residual connection is provided between the embedding layer and the point-by-point convolution, the embedding layer divides the EEG samples into smaller time blocks and expands them into multiple feature maps, the deep convolution performs a separate convolution on each feature map to obtain local time features, the point-by-point convolution fuses the local time features and optimizes the global information, the ELU activation function and batch normalization are used to improve the network expression ability and optimize the feature distribution, and the residual connection is used to prevent information loss; The temporal feature map output by the temporal convolution mixer is passed into the spatial convolution mixer to extract effective spatial features and obtain a spatial feature map; wherein the spatial convolution mixer includes convolution, ELU activation function and batch normalization, the convolution compresses and mixes the temporal features at the spatial level, maps multiple channels into one channel, and the ELU activation function and batch normalization are used to improve the network expression ability; The pooling layer is used to perform average pooling operation on the spatial feature map output by the spatial convolution mixer to reduce the spatial dimension, reduce the computational complexity and extract more significant features to obtain pooled features; The output layer is used to flatten and fully connect the pooled features output by the pooling layer, and output them to a single neuron to obtain the probability that the EEG sample contains a P300 waveform; The above-mentioned convolutional mixer network is used for prediction to obtain the probability that the EEG samples corresponding to multiple flashes in a round of character spelling tasks contain P300 waveforms, and the rows and columns of the EEG samples containing the P300 waveforms are determined to determine the position of the target character and realize the output of the spelled characters.

2. The P300 spelling method based on a convolutional mixer network according to claim 1, characterized in that: The EEG samples were obtained from a P300 speller, which contained a 6×6 character matrix containing 36 characters in total. When the user wanted to spell a character, the character became the target character. While the user was looking at the target character, each row and column of the character matrix would flash once, with the first six flashes being column flashes and the next six being row flashes, for a total of 12 flashes. A total of 12 EEG samples were recorded for the 12 flashes. When the row or column containing the target character flashed, the user's brain would produce an obvious P300 waveform, and when the row or column not containing the target character flashed, the user's brain would not produce a P300 waveform. Generate a P300 waveform; identify the row and column of the stimulus by detecting the presence of the P300 waveform in the EEG sample, and thus infer the target character that the user is paying attention to; the EEG sample obtained is 2 seconds long and contains a total of 32 channels of data. In order to obtain the complete waveform of the P300, the EEG data time window of 0 to 610ms from the start of the visual stimulus is intercepted and downsampled to 78 time points; the 32-channel EEG data is band-pass filtered at 0.5 to 20 Hz, and the data is standardized and normalized; finally, the obtained EEG sample X = {x1, x2, …, x i ,…,x N } has a dimension of 78×32, where x i is the ith EEG sample, N represents the number of EEG samples input into the convolution mixer network; the EEG sample X includes two types of EEG samples: EEG samples with P300 waveform and EEG samples without P300 waveform. The label Y corresponding to each EEG sample is {y1, y2, …, y i …,y N } indicates whether the EEG sample contains a P300 waveform. If it does, the label is 1, and if it does not, the label is 0. i is the label corresponding to the i-th EEG sample.

3. The P300 spelling method based on a convolutional mixer network according to claim 2, characterized in that: The temporal convolution mixer extracts the features of the preprocessed EEG samples in the time dimension and outputs a temporal feature map with local and global features, as follows: First, convolution is used as an embedding layer to divide the preprocessed EEG samples into small time blocks and expand the input EEG samples into 40 feature maps; F embedding =Conv(X) In the formula, F embedding Represents the feature map after expansion, and Conv represents the convolution operation; Then, a deep convolution is used to perform a separate convolution on each feature map using a specific convolution kernel to generate a deep feature map with 40 channels to extract the local spatial features of each channel; After the ELU activation function and batch normalization are applied to deep convolution, the global feature distribution is optimized, improving the performance and expression ability of the network; a linear residual connection is added between the embedding layer and the point-by-point convolution to retain the overall trend of time changes and prevent information loss; F depth =F embedding +BN(ELU(ConvDepthwise(F embedding ))) In the formula, F depth represents the deep feature map, BN represents batch normalization, ELU represents the ELU activation function, and ConvDepthwise represents the deep convolution; Since deep convolution cannot effectively capture the feature correlation between channels, point-by-point convolution is used to fuse the features generated by deep convolution from various dimensions to ensure global information fusion; ELU and batch normalization are also applied to point-by-point convolution to obtain the temporal feature map F temporal : F temporal =BN(ELU(ConvPointwise(F depth ))) Where ConvPointwise represents point-by-point convolution.

4. The P300 spelling method based on a convolutional mixer network according to claim 3, characterized in that: The spatial convolution mixer extracts and fuses the spatial features of the temporal feature map and outputs the spatial feature map, as follows: The spatial convolution mixer combines the temporal feature map F along the spatial dimension. temporal Convolution is performed to integrate and extract valuable features from different channels. It compresses all channel information into a single channel, thereby reducing the data dimension while achieving global spatial feature extraction and fusion; the spatial convolution mixer also uses the ELU activation function and batch normalization to further improve the training efficiency and performance of the network, and finally obtains the spatial feature map F spatial : F spatial =BN(TWO(Conv(F temporal )))。 5. The P300 spelling method based on a convolutional mixer network according to claim 4, characterized in that: The pooling layer performs an average pooling operation on the obtained spatial feature map to further compress and aggregate the spatiotemporal information, reduce the number of parameters and minimize the risk of overfitting, and obtain the pooled feature F pooling : F pooling =AvgPooling(F spatial ) Where AvgPooling represents the average pooling operation.

6. The P300 spelling method based on a convolutional mixer network according to claim 5, characterized in that: The output layer first pools the features F pooling This high-dimensional data is flattened into a one-dimensional representation, and then the extracted features are mapped through a fully connected layer and finally output to a single neuron; the output layer introduces a Sigmoid activation function to limit the output value to the range of [0, 1], indicating the probability P that the EEG sample contains a P300 waveform target : P target =Sigmoid(Linear(Flatten(F pooling ))) In the formula, Sigmoid represents the Sigmoid activation function, Linear represents the fully connected layer, and Flatten represents the flattening operation.

7. The P300 spelling method based on a convolutional mixer network according to claim 6, characterized in that: The loss function of the convolutional mixer network training is a hybrid loss function, which is composed of the supervised contrast loss L S and the cross entropy loss L C The supervised contrast loss establishes a better decision boundary in the feature space, enhances the network's understanding of the data, and improves its ability to identify rare classes. The cross entropy loss allows explicit classification of EEG samples through probability output, directly optimizing the classification accuracy. This hybrid loss function enables the network to extract different feature representations for samples in each class, thereby significantly improving the network's feature learning ability and classification accuracy, while also improving the network's robustness and generalization ability, and reducing its dependence on a large amount of calibration data. The supervised contrast loss L S The pooling features output by the pooling layer are used for calculation: In the formula, S + Represents the EEG sample x i A set of samples with the same label, p represents S + The samples in |S + |Represents and EEG samples x i The number of samples in the same batch with the same label, S represents the number of samples in EEG sample X except EEG sample x i The set of other samples outside, s represents the sample in S, u i Represents EEG sample x i The corresponding pooling feature, u p Represents the EEG sample x i The pooled features corresponding to samples with the same label, u s Represents the EEG sample x i The pooling features corresponding to other EEG samples are as follows: τ represents the temperature, which is used to control the smoothness of training and the impact on difficult samples; Cross entropy loss L C The probability of the P300 waveform output by the output layer is used for calculation: In the formula, p i Represents EEG sample x i The corresponding output layer outputs the probability of containing the P300 waveform; The total loss L of the network is the supervised contrast loss L S and the cross entropy loss L C The mixed loss is: L=λL S +(1-λ)L C Where λ represents the regularization parameter, which is used to balance the contribution between the supervised contrast loss and the cross entropy loss.

8. The P300 spelling method based on a convolutional mixer network according to claim 7, characterized in that: The specific operations of using the convolutional mixer network to predict and implement spelling character output are as follows: In each round of the spelling task, each row and column of the character matrix flashed once, of which the first six were column flashes and the last six were row flashes, totaling 12 flashes, and 12 EEG samples were recorded for the 12 flashes; the P300 waveform only appeared when the row and column containing the target character flashed; since the target row and column were unique, a probability maximization strategy was used to determine the row and column where the target character was located: In the formula, c represents the column where the target character is located, k represents the row where the target character is located, and P j The probability that the EEG sample corresponding to the j-th flash contains P300 is output by the convolution mixer network; the target character spelled by the user can be determined by the row and column where the target character is determined, thereby achieving the output of the P300 spelling character.

Citation Information

Patent Citations

  • A P300 detection method based on CNN-LSTM network

    CN109389059A

  • Method for improving brain-computer interface performance based on dynamic inverse learning network

    CN112381124A

  • P300 signal detection method based on dilated convolutional neural network

    CN113017645A

  • Electroencephalogram signal classification method based on addition network and supervised contrast learning

    CN114595725A

  • RSVP type unbalanced electroencephalogram signal classification method based on adaptive channel mixed attention mechanism and decoupling learning

    CN119385578A