A method for monitoring error-related potentials in a brain-computer interface based on mutual information

By extracting the time and frequency domain features of EEG signals in the brain-computer interface and using the least squares support vector machine for classification, the problem of low ErrP signal recognition rate is solved, the recognition accuracy of ErrP is improved, and the friendliness of human-computer interaction and the active participation of users are enhanced.

CN115964655BActive Publication Date: 2025-10-10HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310041182.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-13
Publication Date
2025-10-10
Estimated Expiration
2043-01-13

AI Technical Summary

Technical Problem

In existing brain-computer interface technologies, the recognition rate of error-related potential signals is low, resulting in a decline in the performance of human-computer interaction tasks and may even cause accidents or secondary injuries, and users' confidence in ErrP-BCI is gradually reduced.

Method used

A method based on mutual information is used to extract the time and frequency domain features of EEG signals, and classification is performed using a least squares support vector machine to optimize feature screening and improve the recognition accuracy of ErrP signals.

Benefits of technology

Through feature combination and optimized screening, the recognition accuracy of ErrP signals is improved, the friendliness of human-computer interaction and the user's active participation are enhanced, and the occurrence of erroneous operations is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115964655B_ABST
    Figure CN115964655B_ABST
Patent Text Reader

Abstract

The application discloses a method for monitoring error-related potentials in a brain-computer interface based on mutual information, and relates to a method for identifying error-related potentials, and aims at solving the problem of low recognition rate of error-related potential signals in the existing brain-computer interface. The application extracts time domain features by pre-processing original electroencephalogram signals and performing non-overlapping sliding window analysis; extracts frequency domain features from the signals by using the Welch method; combines the time domain features and the frequency domain features; uses mutual information as a measure between features and positive and negative categories; calculates the mutual information of the features and the categories and sorts them; filters out the features with high rankings; uses a least squares support vector machine to classify initial error-related potentials; performs leave-one-out cross-validation on samples; obtains and retains individual optimal models; and obtains the accuracy rate of final error-related potential classification. The application has the beneficial effect of improving the recognition accuracy of error-related potentials.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a method for identifying error-related potentials. BACKGROUND

[0002] Brain-computer interface (BCI) technology reveals human intention and transmits intention to external devices, providing a bright prospect for effective communication between humans and intelligent machines; in recent years, BCI technology has been widely used in teaching demonstrations, technical performances, robot device control, and human-computer interaction (HCI) tasks for specific industrial operations; BCI in teaching mode can enhance students' attention and memory; some interactive games can control virtual objects according to the user's intention, such as the cursor reaching the designated target; intelligent devices such as robots can complete some specific tasks according to human instructions; industrial exoskeleton robotic arms can assist in processing and handling heavy objects and perform repetitive fine operations, greatly saving the physical strength of workers while significantly improving production efficiency; rehabilitation robots based on patients' own brain and muscle electrical signals can replace rehabilitation therapists to assist patients with high-intensity repetitive rehabilitation training; in addition to performing expected tasks through setting programs, whether the controlled object correctly performs the corresponding operation according to the user's intention also affects the effectiveness of HCI performance.

[0003] Although the controlled virtual object or external device can generally move to the specified target or perform the expected operation according to the user's intention to some extent, meaning that the machine intelligence of the controlled virtual object or external device is gradually approaching the biological intelligence that humans have; but sometimes the virtual object or external device will move to the wrong target or perform the wrong non-expected operation, meaning that the machine intelligence has not yet reached the level of human biological intelligence; these errors will interrupt teaching, industrial production, and rehabilitation training, greatly affecting the performance of HCI; in dangerous operation work, some unexpected errors can even cause industrial accidents; in rehabilitation training, the wrong action of the rehabilitation robot can even cause secondary injury to the patient; in addition, due to long-term tasks and excessive dependence on external devices, users are prone to carelessness, which is also not conducive to the further execution of human-computer interaction tasks, so it is necessary to discover errors and promptly stop.

[0004] Although the operation of the supervisor to remind the operator of the wrong execution action helps the smooth progress of the human-computer interaction task, the passive operation of the user still exists, and the active and positive participation cannot be completely achieved; the user is repeatedly reminded, which can easily lead to distraction of attention and increase the labor cost; in order to solve the above problems, the user itself needs to detect the observed error in the HCI task in real time, so as to timely correct and intervene; the error-related potential (ErrP) emerges as the times require; it has been proved that when a human being observes a wrong action, no matter what type of error occurs, the ErrP signal is unconsciously generated in the brain without any prior training; the occurrence of the ErrP signal belongs to the normal reaction of the human being when the operation of the external thing is contrary to the will of the human being; the frequency of the ErrP signal is mainly concentrated in 1-10Hz and tends to be stable, and plays an important role in the supervision of the specific task; the error monitoring technology of the brain-computer interface based on the ErrP can significantly improve the active participation and confidence of the user, and the user will be more focused and cautious when performing the expected task; like the traditional BCI, the ErrP signal usually needs to be detected separately, and can be regarded as a single source input in the BCI system, the current error is found by detecting the ErrP, and the correction measure is taken to prevent the wrong action from being continuously executed, and the role of the biological intelligence of the human being is played in the brain-computer interface; the effective combination of the machine intelligence of the brain-computer interface and the biological intelligence of the human being forms a hybrid enhanced intelligence, and the experience of the user in the human-computer interaction task is significantly improved.

[0005] With the emergence of machine learning technology and powerful computing resources, the classification rate of neural signals has been significantly improved; with the integration of the ErrP automatic error correction mechanism in the BCI, it is expected to be further promoted; however, the generally low ErrP classification rate constitutes a major challenge; due to the low ErrP classification rate, some human-computer interaction tasks such as the P300 speller have hardly improved in the integration of the ErrP, and the participants even find that the automatic correction strategy is more chaotic and unpredictable; the overall performance decline leads some users to prefer to use the P300 speller without correction because they think that there is no benefit, and the confidence of using the ErrP-BCI gradually decreases; therefore, it is urgent to propose an algorithm for processing the error-related potential signal to improve the recognition rate of the ErrP signal and effectively monitor the error of the brain-computer interface. SUMMARY

[0006] The purpose of the present application is to solve the problem of poor recognition rate of the error-related potential signal in the existing brain-computer interface, and a monitoring method for the error-related potential in the brain-computer interface based on mutual information is proposed.

[0007] The monitoring method for the error-related potential in the brain-computer interface based on mutual information comprises the following steps:

[0008] Step one, pre-processing the original electroencephalogram signal to obtain the original error-related potential signal;

[0009] Step two, non-overlapping sliding window analysis is performed on the original error-related potential signal obtained in step one, the average absolute value of each sliding window signal is extracted, and the average absolute value is taken as the time domain feature of the original error-related potential signal;

[0010] Step three, the Welch method is used to extract the features of the original error-related potential signal obtained in step one, and the frequency domain features of the original error-related potential signal are obtained;

[0011] Step four, the time domain features obtained in step two and the frequency domain features obtained in step three are combined to obtain the combined features;

[0012] Step five, mutual information is used as the measure between the combined features in step four and the positive and negative categories, the mutual information amount of the combined features and the positive and negative categories is calculated and sorted in descending order, the best model is obtained and retained, and the error-related potential is classified by using the least square support vector machine based on the best model.

[0013] The application proposes a novel brain-computer interface error monitoring method based on error-related potential (Error-related Potential, ErrP): the time and frequency domain features of the pre-processed user real-time ErrP are extracted and combined, the mutual information amount of the features and the categories is calculated and sorted, the features with larger mutual information amount are preferentially selected and classified by using the least square support vector machine, the samples are subjected to leave-one-out cross-validation, and the individual best model is obtained and retained; the method optimizes the selection of features, improves the recognition accuracy of ErrP, provides direction and strategy for subsequent human intervention, and plays the role of human biological intelligence. The effective combination of machine intelligence and human biological intelligence of the brain-computer interface forms a hybrid enhanced intelligence, and the friendliness of the human-computer interaction process is enhanced. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 A flow chart of a brain-computer interface error-related potential monitoring method based on mutual information amount according to the first embodiment;

[0015] Figure 2 A flow chart of the best feature number screening based on mutual information amount in the first embodiment;

[0016] Figure 3 A subject observation cursor movement direction true or false diagram in the first embodiment;

[0017] Figure 4 A mutual information classification Venn diagram in the first embodiment;

[0018] Figure 5 Fig. 6 is a waveform diagram of the error-related potential of each subject in the first embodiment of the present application;

[0019] Figure 6 Fig. 7 is a comparison chart of the classification accuracy of the FCz channel of each subject in the first embodiment of the present application based on the MAV, Welch power spectrum and double feature combination;

[0020] Figure 7 Fig. 8 is a comparison chart of the classification accuracy of the Cz channel of each subject in the first embodiment of the present application based on the MAV, Welch power spectrum and double feature combination. DETAILED DESCRIPTION

[0021] In combination Figures 1 to 7 In this embodiment, a monitoring method for error-related potentials in a brain-computer interface based on mutual information is described. The monitoring method comprises the following steps:

[0022] Step 1: Preprocess the original electroencephalogram signal to obtain the original error-related potential signal;

[0023] Step 2: Perform non-overlapping sliding window analysis on the original error-related potential signal obtained in step 1 to extract the average absolute value of each sliding window signal, and take the average absolute value as the time domain feature of the original error-related potential signal;

[0024] Step 3: Use the Welch method to extract features from the original error-related potential signal obtained in step 1 to obtain the frequency domain features of the original error-related potential signal;

[0025] Step 4: Combine the time domain features obtained in step 2 with the frequency domain features obtained in step 3 to obtain the combined features;

[0026] Step 5: Use mutual information as the measure between the combined features in step 4 and the correct and incorrect categories, calculate the mutual information between the combined features and the correct and incorrect categories and sort them in descending order, obtain and retain the best model; based on the best model, use the least squares support vector machine to classify the error-related potential.

[0027] In this embodiment, the Dataset 22 - Monitoring error-related potentials data in the public database - EU Brain-Computer Interface Horizon 2020 (BNCI Horizon 2020) database is used. The subjects sit in front of a computer screen and are asked to observe whether the cursor moves towards the correct target or not, while their electroencephalogram signals are recorded. A moving cursor (i.e. a gray square) is displayed on the screen. The black square on the left or right side of the cursor represents the target position, and the dashed square represents the previous position of the cursor, as shown inFigure 1 The cursor moves horizontally according to the position of the target in each trial. Each trial lasts about 2000 milliseconds. Once the target is reached, the cursor will remain in place, and a new target position will appear. During the experiment, the user is asked to monitor the performance of the system and evaluate whether it is working properly. In order to study the signals generated by the erroneous movements, in each trial the cursor has a probability of moving in the wrong direction (i.e. opposite to the target position). When the subject finds that the cursor has moved to the wrong target, an ErrP occurs. The dataset includes 64-channel whole-brain ErrP signals of 6 subjects collected according to the 10-20 international standard, with a sampling rate of 512 Hz. The present invention only selects data with an error probability of 20% and analyzes the ErrP signals of each subject one by one. Each subject is divided into two stages of collection, with different days apart. The present invention selects the data of the first stage for training and building an individual optimal model, and uses the data of the second stage for testing. Due to the large amount of original data samples and the imbalance of the proportion of different categories of samples, the randperm function is used to non-repeatedly rearrange and sample the majority class samples, and the same number of majority class samples as the minority class samples are randomly extracted to balance the data.

[0028] In the preferred embodiment, the raw electroencephalogram signal in step one is preprocessed, and the specific method for obtaining the original error-related potential signal includes: sequentially processing the raw electroencephalogram signal by using common average reference method, median filtering method, FIR band-pass filtering method, independent component analysis, anti-aliasing filtering, downsampling and sliding window; wherein the common average reference method is: removing the electrode common noise in the mean value of the original electroencephalogram signal of 64 channels to obtain the central electrode channel signal of 9 central cortex regions of the error-related potential signal;

[0029] The median filtering method is to remove the baseline drift by using a median filter with a window length of 201.

[0030] The FIR band-pass filtering method is to perform band-pass filtering on the data by using a finite-length unit impulse response filter of 1-10 Hz.

[0031] In the present embodiment, first, the mean value of the signals of 64 channels is removed by the common average reference (CAR) method to eliminate all electrode common noise and achieve spatial filtering. Since the error-related activity is associated with the central cortex region, 9 central electrode channels (FC1, FCz, FC2, C1, Cz, C2, CP1, CPz, CP2) are preferentially selected for further analysis to achieve dimension reduction. This channel selection process also helps to remove electrodes that may be contaminated by blinking and muscle artifacts.

[0032] The median filter with window length of 201 is used to remove baseline drift, and then the data is band-pass filtered by 1-10 Hz Finite Impulse Response (FIR) filter, the Hamming window is selected for window function, and the window length is 50. Finally, the electrooculogram signal is removed by Independent Component Analysis (ICA) to reconstruct the ErrP signal. The FastICA algorithm proposed by Hyvarinen, also known as fixed-point algorithm, is a fast optimization iterative algorithm in batch mode.

[0033] The reconstructed ErrP is filtered by an 82-order FIR filter before down-sampling to prevent frequency aliasing after down-sampling. After down-sampling, a fixed length ErrP window period in each trial is selected for analysis. The Mean Average Value (MAV) of each sliding window signal is extracted as the time domain feature by using the non-overlapping sliding window analysis method.

[0034] In the preferred embodiment, the specific method of step two is to perform non-overlapping sliding window analysis on the original error-related potential signal to extract the average absolute value of each sliding window signal.

[0035] Let x u be the u-th sample point of the original error-related potential signal x, N be the signal length, the average absolute value represent the degree of separation between the signal and the far point in the effective data collected by the signal; the sliding window length is 7, there are 12 sliding windows in each trial, a total of 12 time domain features are obtained, 9 central electrode channels are preferentially selected for analysis, and a total of 9*12=108-dimensional features are extracted from each experimental sample.

[0036] The average absolute value expression is shown in equation (1):

[0037]

[0038] Wherein, MAV is the average absolute value.

[0039] In the embodiment, the length signal is also analyzed by the overlapping sliding window analysis method. Since the ErrP is concentrated in 1-10 Hz, including the delta (0.5-4 Hz), theta (4-8 Hz), and alpha1 (8-10 Hz) bands, when extracting the frequency domain features, the Welch power spectrum features of the above three bands in the ErrP are extracted respectively for classification, and then the Welch power spectrum features in the 1-10 Hz band are extracted for classification, so as to compare the contribution of the extraction of different frequency band frequency domain features to the single-channel single-feature classification rate. The Hamming window is selected as the window function, the window length is 64, and the overlapping length is 32.

[0040] In the preferred embodiment, the specific method for obtaining the frequency domain features of the original error-related potential signal in step three is as follows:

[0041] First, the original error-related potential signal is decomposed into data x(m) with a length of N, n=0, 1, …, N-1, which is divided into L segments, m is the mth sampling point of the data x(m), each segment has M data, and the tth segment of data is represented as:

[0042] x t (m)=x(m+tM-M),0≤m≤M,1≤t≤L (2)

[0043] Wherein, x t (m) is the tth segment of data;

[0044] The window function w(m) is added to each segment of data to obtain the periodogram of each segment of data, and the periodogram of the tth segment of data is:

[0045]

[0046] Wherein, I t (ω) is the periodogram of the tth segment of data; U is a normalization factor; j is an imaginary unit; and ω is a frequency;

[0047]

[0048] The periodograms of each segment of data are considered to be mutually independent, and the power spectral density is:

[0049]

[0050] Wherein, P xx (e jω ) is the power spectral density; and L is the total number of segments of the data x(n);

[0051] Each trial has 257 features, and 9 central electrode channels are preferentially selected for analysis, so that 9*257=2313-dimensional features are extracted from each experimental sample;

[0052] The power spectrum density is the frequency domain feature.

[0053] In this implementation, the MAV and Welch power spectrum feature vectors for each channel were first cross-validated using LS-SVM leave-one-out cross-validation, and the single-channel, single-feature ErrP classification accuracy was obtained and compared. After combining the MAV and power spectrum, each trial had a total of 9*(12+257)=2421 features. LS-SVM leave-one-out cross-validation was then performed, and the single-channel, dual-feature classification accuracy was compared one by one.

[0054] Support Vector Machines (SVMs) are a new pattern recognition method developed in the 1990s based on Statistical Learning Theory (SLT) established by Vapnik et al. Similar to SLT, SVMs are statistical learning methods designed for small samples. They effectively overcome the shortcomings of neural network methods, such as convergence difficulties, unstable solutions, and poor generalization. They are particularly effective when the number of training samples is small, achieving good classification results. Currently, SVMs have been widely used in pattern recognition, signal processing, communications, and other fields.

[0055] In terms of classification learning, traditional pattern recognition methods are to reduce the dimensionality of data, while SVM is just the opposite. For samples in the feature space, SVM uses a mapping method to map them to a high-dimensional space and find the optimal classification hyperplane for each sample in the high-dimensional space.

[0056] In a preferred embodiment, the frequency band classification in step 5 includes delta frequency band, theta frequency band and alpha1 frequency band;

[0057] The delta frequency band ranges from 0.5 Hz to 4 Hz;

[0058] The theta frequency band ranges from 4 Hz to 8 Hz;

[0059] The alpha1 frequency band ranges from 8 Hz to 10 Hz.

[0060] In a preferred embodiment, the specific method of classifying the combined features in step five is: classifying the combined features using an optimal classification function;

[0061] Let the training set be (x i ,y i ), i = 1, 2, ... n, a total of n training samples, where x i Represents the input of the training model, y i Represents the output of the training model; the training set has a linear discriminant function g(x i) = w T x i +b, w is the normal vector of the hyperplane, and b is the intercept of the linear discriminant function; the problem of finding the optimal classification surface is represented as finding the minimum value of ||w||2 under the constraint condition y i (w T x i +b)-1 ≥ 0. 2

[0062] The original Lagrange function is defined as shown in equation (6):

[0063]

[0064] where L(w, b, α) is the function value of the original Lagrange function, α i is the Lagrange coefficient, and α i ≥ 0; w is the weight vector, w T is the transpose of the weight vector; and b is the intercept of the linear discriminant function.

[0065] The partial derivatives of the above equation with respect to w and b are taken, which is converted into a quadratic programming dual problem:

[0066] The maximum value of the optimization function is found under the constraint conditions and α i ≥ 0. In order to distinguish w T and w, w T is written as w is rewritten as where α j is another Lagrange coefficient that is different from α i , y j is another output of the training model that is different from y i , and x j is another input of the training model that is different from x i :

[0067]

[0068] where Q(α) is the optimization function.

[0069] The above equation has a unique optimal solution under the inequality constraint condition, and satisfies:

[0070] α i [y i (w T x i +b)-1] = 0 i = 1, 2,..., n (8)

[0071] If α is the optimal solution of equation (7), then​ For most samples is 0, Samples whose values ​​are not 0 are support vectors;

[0072] Then the optimal classification function is:

[0073]

[0074] Among them, b * is the optimal solution for the intercept of the linear discriminant function;

[0075] In the nonlinear case, using The linearly inseparable samples x i Mapping to high-order space achieves linear separability. The optimization function is:

[0076]

[0077] in, for and The inner product of

[0078] Therefore, the optimal classification function is:

[0079]

[0080] When classifying in high-dimensional space, kernel functions are used to reduce the dimensionality of the optimal classification function;

[0081] The kernel function K(·) is defined as:

[0082]

[0083] Then the optimization function is transformed into:

[0084]

[0085] The corresponding optimal classification function is:

[0086]

[0087] In a preferred embodiment, the kernel function includes:

[0088] (a) Polynomial kernel function

[0089]

[0090] At this time, SVM is a q-order polynomial classifier;

[0091] (b) Radial basis kernel function

[0092] K(x i ,xj ) = exp(-||x i -x j || 2 / 2μ 2 ) (16)

[0093] ||x i -x j || 2 The square Euclidean distance between two feature vectors can be seen as a free parameter μ. In this case, the SVM is a radial basis function classifier, which differs from the traditional radial basis function in that the center of each basis function corresponds to a support vector;

[0094] (c) Sigmoid kernel function

[0095]

[0096] where β0is the slope and β1is the intercept.

[0097] In addition, the kernel function also includes the number-shaped radial kernel function, Fourier series, spline function and B-spline function.

[0098] In the preferred embodiment, the specific process of using mutual information as the measure between the combined feature and the optimal classification in step six is as follows:

[0099] Mutual information measures the degree of association between the feature item and the class. The mutual information of two discrete variables X and Y is defined as:

[0100]

[0101] where p(x, y) is the joint probability distribution function of X and Y, and p(x) and p(y) are the marginal probability distribution functions of X and Y, respectively; the mutual information of two random variables is a measure of the mutual dependence between the variables; the mutual information is equivalently expressed as:

[0102]

[0103] where H(X) and H(Y) are the marginal entropies, H(X|Y) and H(Y|X) are the conditional entropies, and H(X, Y) is the joint entropy of X and Y; it is expressed as shown in the Venn diagram Figure 4 .

[0104] The mutual information amount of the features and the categories is continuously calculated and arranged in descending order, the features with larger mutual information amount are preferentially reserved, the single-channel joint classification accuracy is obtained after leave-one-out cross validation by using the LS-SVM, the channels with higher classification results are screened out, the features thereof are combined and screened by using the mutual information, the individual best model is reserved after leave-one-out cross validation by using the LS-SVM, and the final ErrP classification accuracy is obtained. Specifically, N features are arranged in descending order according to the mutual information amount from large to small, the first to k features are selected one by one to be classified by using the least square support vector machine on the basis of the individual best model, leave-one-out cross validation is performed on the samples, the classification accuracy A k is reserved, as shown in Figure 3 This helps to convey the corresponding instructions, terminate the erroneous movement of the cursor, and facilitate timely intervention.

[0105] In the preferred embodiment, the specific method for obtaining and reserving the best model in step six is as follows:

[0106] Specifically, N features are arranged in descending order according to the mutual information amount from large to small, the first to k features are selected one by one to be classified by using the least square support vector machine on the basis of the individual best model, leave-one-out cross validation is performed on the samples, the classification accuracy A k is reserved when the k value is the highest; as shown in Figure 2 This helps to convey the corresponding instructions, terminate the erroneous movement of the cursor, and facilitate timely intervention.

[0107] In the preferred embodiment, the specific method for classifying the error-related potential by using the least square support vector machine in step five is as follows: the least square support vector machine adopts a least square linear system as a loss function, and changes the quadratic programming method in the classical support vector machine algorithm into a solution of a linear equation set.

[0108] The target optimization function of the least square support vector machine algorithm is as follows:

[0109]

[0110] The constraint condition is as follows:

[0111]

[0112] wherein w is a weight vector; γ is a regularization parameter; b is an intercept of a linear discriminant function; e i is an error;

[0113] The Lagrange function is updated as follows:

[0114]

[0115] w, b, and ei and α i Taking partial derivative of the parameters, we have:

[0116]

[0117] The formula (23) is arranged as:

[0118]

[0119] wherein, l = [1, 1,..., 1] T , i, j = 1, 2,..., n, I is a unit matrix, α i = [α1, α2, …, α n ] T , y = [y1, y2,..., y n ] T ;

[0120] Let Solve the matrix equation:

[0121]

[0122] α i = A -1 (y-bl) (26)

[0123] The prediction value of the least squares support vector machine for the unknown sample x is:

[0124]

[0125] In each classification, the kernel function is selected first, and then the optimal regular parameter γ and the optimal kernel parameter σ are trained by using the training data and the corresponding sample, and then the Lagrange coefficient α i and the linear discriminant function intercept b are trained; the individual optimal model is generated according to the Lagrange coefficient α i and the linear discriminant function intercept b, the individual optimal model is obtained and retained, and the error-related potential is classified through the individual optimal model.

[0126] Experimental results:

[0127] Taking the FCz channel as an example, Figure 5 the error-related potential waveforms and the total average waveforms of all subjects are shown. The thin line represents the ErrP of six subjects, and the thick line is the total average waveform. The ErrP has high temporal locking: after about 200 ms of observing the error, the initial positive peak value appears first, followed by a larger negative deflection near 255 ms, and a second larger positive peak value appears at about 340 ms.

[0128] After the restructured ErrP is down-sampled to 128 Hz, the ErrP signal period between -50 ms and 700 ms in each trial is selected for analysis. In the training model, the RBF kernel function is selected, and the optimal regular parameter γ and the optimal kernel parameter σ are trained by using the training data and the corresponding samples, and then the Lagrange coefficient α and the linear discriminant function intercept b are trained. According to the above four key parameters, the individual optimal model is formed to predict the label of the test data.

[0129] After comparison, it is found that compared with extracting delta, theta, and alpha1, the classification accuracy of extracting the Welch power spectrum feature in the 1-10 Hz frequency band is higher. In addition, it is found that among the 9 channels, the classification rate of FCz and Cz channels is higher. Therefore, FCz and Cz two central electrode channels are further analyzed. Figure 6 、 7 The single-channel average absolute value, single-channel Welch power spectrum, and recognition rate after combining the single-channel average absolute value and Welch power spectrum feature of FCz and Cz are shown respectively.

[0130] It can be seen that the classification effect of FCz channel is generally better than that of Cz; the classification accuracy based on the time domain feature MAV is higher than that of the frequency domain feature Welch power spectrum. By comparing the classification effects of single features and feature combinations, it is found that the more features are not necessarily more conducive to classification, and some features are redundant and even hinder the classification rate.

[0131] The time-frequency features of FCz and Cz are combined for double-channel double-feature joint classification, and there are 2*(12+257)=538 in each trial. Finally, the mutual information of the features and categories is calculated and arranged in descending order, and the first k features with larger mutual information are selected one by one to use the least squares support vector machine classification, and the leave-one-out cross-validation is performed on the samples. The k value when the classification accuracy is the highest is retained. The k values of the six subjects are 8, 11, 7, 6, 8, and 9 respectively, and the results are shown in Table 1. It can be seen that similar to the combination of single-channel time-frequency features, the superposition of multi-channel feature numbers does not significantly improve the classification rate. On the contrary, the operation time is significantly prolonged, and the calculation amount and complexity are increased. After the mutual information is sorted and screened, the joint classification rate of the ErrP signal is significantly improved, and the ErrP recognition contributes a major role.

[0132] Table 1 Comparison of classification accuracy of double-channel double-feature combination and mutual information screening of each subject

[0133]

[0134] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for monitoring error-related potentials in a brain-computer interface based on mutual information, characterized in that: The monitoring method comprises the following steps: Step 1: pre-process the original EEG signal to obtain the original error-related potential signal; Step 2: Perform non-overlapping sliding window analysis on the original error correlation potential signal obtained in step 1, extract the average absolute value of each sliding window signal, and use the average absolute value as the time domain feature of the original error correlation potential signal; Step 3: Using the Welch method to extract features from the original error correlation potential signal obtained in step 1, and obtain the frequency domain features of the original error correlation potential signal; Step 4: Combine the time domain features obtained in step 2 with the frequency domain features obtained in step 3 to obtain combined features; Step 5: Use mutual information as a measure between the combined features in step 4 and the true and false categories, calculate the mutual information between the combined features and the true and false categories and sort them in descending order, obtain and retain the best model; based on the best model, use the least squares support vector machine to classify the error correlation potentials.

2. The method for monitoring error-related potentials in a brain-computer interface based on mutual information according to claim 1, characterized in that: The raw EEG signal in step 1 is preprocessed to obtain the raw error-related potential signal, including: sequentially processing the raw EEG signal using a common average reference method, a median filter method, an FIR bandpass filter method, an independent component analysis, an anti-aliasing filter, downsampling, and a sliding window; wherein the common average reference method removes the electrode common noise from the mean of the raw EEG signal of 64 channels to obtain the central electrode channel signal of 9 central cortical regions related to the error-related potential signal; The median filter method uses a median filter with a window length of 201 to remove baseline drift; The FIR bandpass filtering method is to perform bandpass filtering on the data through a 1-10 Hz finite length unit impulse response filter.

3. The method for monitoring error-related potentials in a brain-computer interface based on mutual information according to claim 1, characterized in that: In step 2, the original error-related potential signal is subjected to non-overlapping sliding window analysis, and the specific method for extracting the average absolute value of each sliding window signal is as follows: Let x u is the u-th sample point of the original error-related potential signal x, N is the signal length, and the average absolute value expression is shown in (1): Here, MAV is the mean absolute value.

4. The method for monitoring error-related potentials in a brain-computer interface based on mutual information according to claim 3, characterized in that: The specific method for obtaining the frequency domain characteristics of the original error-related potential signal in step 3 is: First, the original error-related potential signal is decomposed into data x(m) of length N, n = 0, 1, ..., N-1, divided into L segments, m is the mth sampling point of the data x(m), each segment has M data, and the tth segment data is expressed as: x t (m)=x(m+tM-M),0≤m≤M,1≤t≤L (2) Among them, x t (m) is the data of the tth segment; Add the window function w(m) to each segment of data and find the periodogram of each segment of data. The periodogram of the t-th segment of data is: Among them, I t (ω) is the periodogram of the t-th segment data; U is the normalization factor; j is the imaginary unit; ω is the frequency; The periodograms of each segment of data are considered to be unrelated, and the power spectrum density is: Among them, P xx (e jω ) is the power spectrum density; L is the total number of segments of data x(n); The power spectrum density is the frequency domain feature.

5. The method for monitoring error-related potentials in a brain-computer interface based on mutual information according to claim 4, characterized in that: In step 5, the frequency band classification includes delta band, theta band and alpha1 band; The delta frequency band ranges from 0.5 Hz to 4 Hz; The theta frequency band ranges from 4 Hz to 8 Hz; The alpha1 frequency band ranges from 8 Hz to 10 Hz.

6. The method for monitoring error-related potentials in a brain-computer interface based on mutual information according to claim 5, characterized in that: The specific method of classifying the combined features in step five is: classifying the combined features using the optimal classification function; Let the training set be (x i ,y i ), i=1,2,…n, a total of n training samples, where x i Represents the input of the training model, y i Represents the output of the training model; the training set has a linear discriminant function g(x i )=w T x i +b, w is the normal vector of the hyperplane, b is the intercept of the linear discriminant function; the problem of finding the optimal classification surface is expressed as i (w T x i +b)-1≥0, find ||w|| 2 / 2 minimum value; Define the original Lagrange function as shown in (6): Among them, L(w,b,α) is the function value of the original Lagrange function, α i is the Lagrange coefficient, and α i ≥0; w is the weight vector, w T is the transpose of the weight vector; b is the intercept of the linear discriminant function; Take partial derivatives of the above formula with respect to w and b respectively, and transform it into the dual problem of quadratic programming: Under the constraints and α i ≥0, find the maximum value of the optimization function, in order to distinguish w T and w, w T writing w rewrite Among them, α j is different from α i Another Lagrange coefficient, y j Is different from y i The output of another trained model, x j Is different from x i Another input for training the model: Among them, Q(α) is the optimization function; The above formula has a unique optimal solution under the inequality constraints and satisfies: α i [y i (w T x i +b)-1]=0 i=1,2,...,n (8) like is the optimal solution of formula (7), then For most samples is 0, Samples whose values ​​are not 0 are support vectors; Then the optimal classification function is: Among them, b * is the optimal solution for the intercept of the linear discriminant function; In the nonlinear case, using The linearly inseparable samples x i Mapping to high-order space achieves linear separability; the optimization function is: in, for and The inner product of Therefore, the optimal classification function is: When classifying in high-dimensional space, kernel functions are used to reduce the dimensionality of the optimal classification function; The kernel function K(·) is defined as: Then the optimization function is transformed into: The corresponding optimal classification function is:

7. The method for monitoring error-related potentials in a brain-computer interface based on mutual information according to claim 6, characterized in that: The kernel function includes: (a) Polynomial kernel function At this time, SVM is a q-order polynomial classifier, where q is the order of the polynomial classifier; (b) Radial basis kernel function K(x i ,x j )=exp(-||x i -x j || 2 / 2μ 2 ) (16) ||x i -x j || 2 It can be regarded as the squared Euclidean distance between two eigenvectors, and μ is a free parameter. In this case, SVM is a radial basis function classifier. The difference from the existing radial basis function is that the center of each basis function corresponds to a support vector. (c) S-shaped kernel function Where β0 is the slope and β1 is the intercept.

8. The method for monitoring error-related potentials in a brain-computer interface based on mutual information according to claim 7, characterized in that: The specific process of using mutual information as a measure between the combined features and the optimal classification in step 5 is: Mutual information measures the degree of association between feature items and categories. The mutual information of two discrete variables X and Y is defined as: where p(x,y) is the joint probability distribution function of X and Y, and p(x) and p(y) are the marginal probability distribution functions of X and Y, respectively. The mutual information of two random variables is a measure of the interdependence between the variables. Mutual information is equivalently expressed as: I(X;Y)=H(X)-H(X|Y)=H(Y)-H(Y|X) =H(X)+H(Y)-H(X,Y) (19)=H(X,Y)-H(X|Y)-H(Y|X) Among them, H(X) and H(Y) are marginal entropies, H(X|Y) and H(Y|X) are conditional entropies, and H(X,Y) is the joint entropy of X and Y.

9. The method for monitoring error-related potentials in a brain-computer interface based on mutual information according to claim 8, characterized in that: The specific method of using the least squares support vector machine to classify the error-related potential in step 5 is as follows: the least squares support vector machine uses the least squares linear system as the loss function, and changes the quadratic programming method in the classic support vector machine algorithm to solving a linear equation system; The objective optimization function of the least squares support vector machine algorithm is: Constraints: Among them, w is the weight vector; γ is the regularization parameter; b is the intercept of the linear discriminant function; e i is the error; Update the Lagrange function to: For w, b, e respectively i and α i Taking partial derivatives of equal parameters, we get: Arranging formula (23) yields: where l = [1, 1, ..., 1] T , i,j=1,2,...,n, I is the identity matrix, α i =[α1,α2,……,α n ] T ,y=[y1,y2,...,y n ] T ; make Solving the matrix equation yields: α i =A -1 (y-bl) (26) For unknown sample x, the prediction value of the least squares support vector machine is: Each time classification is performed, the kernel function is selected first, and then the training data and corresponding samples are used to first train the optimal regularization parameter γ and the optimal kernel parameter σ, and then the Lagrange coefficient α is trained. i and the linear discriminant function intercept b; according to the Lagrange coefficient α i and the linear discriminant function intercept b to generate an individual best model, obtain and retain the individual best model, and classify the error-related potentials using the individual best model.

Citation Information

Patent Citations

  • Multiclass classification method for the estimation of EEG signal quality

    CA3105269A1

  • Classification system of epileptic EEG signals based on non-linear dynamics features

    US20210000426A1