A domain-adaptive method for addressing subject variability in motor imagery brain-computer interfaces
By addressing subject differences in motor imagery brain-computer interfaces through a domain-adaptive approach, and utilizing pre-aligned and manifold embedding distribution-aligned feature maps, combined with the principle of structural risk minimization, a classifier is constructed. This solves the problem of low classification accuracy caused by subject differences in the MI-BCI system, and improves recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510063312.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-01-15
AI Technical Summary
In motor imagery brain-computer interfaces, the low classification accuracy of the MI-BCI system due to differences in EEG signals among subjects and the poor performance of traditional machine learning algorithms have become obstacles to the further development of MI-BCI technology.
A domain-adaptive approach is adopted, which involves performing multiple motor imagery tasks through an online brain-computer interface system to collect EEG signal data. A classifier is constructed to improve classification accuracy by using a pre-alignment strategy and a domain-adaptive manifold embedding distribution-aligned feature mapping method, combined with the principle of structural risk minimization and feature fusion.
It significantly improves the recognition accuracy and robustness of motor imagery tasks, effectively avoids the problem of low classification accuracy caused by differences in EEG signals among subjects, and improves the performance of the MI-BCI system.
Smart Images

Figure CN119987544B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of transfer learning technology in motor imagery brain-computer interfaces, specifically relating to a domain adaptation method for addressing subject differences in motor imagery brain-computer interfaces. Background Technology
[0002] Brain-computer interfaces (BCIs) are powerful communication tools between users and systems, enhancing the human brain's ability to communicate and interact directly with its environment. Representing a technological innovation and an indispensable frontier, BCIs have a significant impact on the future development of society, enabling direct communication between the brain and external devices without the need for muscles or nerves. BCIs can be categorized based on the type of electroencephalogram (EEG) signals used. Among them, the motor imagery brain-computer interface (MI-BCI), based on motor imagery electroencephalogram (MI-EEG), is widely used due to its active and motor-related nature. However, significant individual differences exist in MI-BCI. In this case, using the same model structure and parameters with different subjects often leads to a decrease in the classification accuracy of the MI-BCI paradigm. Traditional machine learning algorithms perform poorly in handling subject differences; therefore, this issue has become one of the main obstacles restricting the further development of MI-BCI technology. Summary of the Invention
[0003] The purpose of this invention is to address the problem of low classification accuracy of the MI-BCI system in motor imagery brain-computer interfaces due to differences in EEG signals among subjects, and to propose a domain adaptation method to solve the subject differences in motor imagery brain-computer interfaces.
[0004] The technical solution adopted by this invention to solve the above-mentioned technical problems is: a domain adaptation method for solving subject differences in motor imagery brain-computer interfaces, the method specifically including the following steps:
[0005] Step 1: The ID-position subjects used an online brain-computer interface system to perform multiple motor imagery tasks. The motor imagery tasks included leftward motor imagery tasks and rightward motor imagery tasks. During each motor imagery task, the subjects' EEG signal data were collected. The collected EEG signal data of the ID-position subjects was used as the training set data, and the EEG signal data collected by the subjects under test during the motor imagery tasks was used as the target domain data.
[0006] Step 2: Extract the EEG signal data corresponding to the time period of motor imagery from each group of EEG signal data in the training set and the target domain, and preprocess the extracted EEG signal data to obtain the preprocessed EEG signal data of each group.
[0007] Step 3: Use a pre-alignment strategy to process each group of preprocessed EEG signal data in the training set and the target domain to obtain the processed EEG signal data.
[0008] Step 4: Based on each group of processed EEG signal data, calculate the covariance matrix of each group of EEG signal data, and then use the tangent space feature mapping method to project the covariance matrix of each group of EEG signal data onto the tangent space to obtain the tangent space vector of each group of EEG signal data.
[0009] Step 5: Initialize id = 1;
[0010] Step 6: Use the domain adaptive manifold embedding distribution alignment feature mapping method to project and map the tangent space vectors of each group of EEG signal data of the id-th subject in the training set and the tangent space vectors of each group of EEG signal data in the target domain, so as to obtain the feature vectors of the id-th subject and the tangent space vectors of the target domain after projection mapping.
[0011] Step 7: Construct a classifier CLS using the feature vectors of the id-th subject and the projected target domain. id ;
[0012] Step 8: Determine if id = ID;
[0013] If the conditions are met, then the constructed ID classifiers CLS1, CLS2, ..., CLS are obtained. ID After projecting and mapping the target domain, the feature vectors are passed through classifiers CLS1, CLS2, ..., CLS. ID The classification results Rst1, Rst2, ..., Rst are obtained respectively. ID Then proceed to step nine;
[0014] If the condition is not met, then set id = id + 1 and return to step six.
[0015] Step 9: Perform feature fusion on the ID classification results obtained in Step 8 to obtain the final classification result for the target domain data.
[0016] Furthermore, the specific process of step one is as follows:
[0017] For any given subject, the subject is positioned facing the computer screen. At t=0, a fixed cross symbol appears on the computer screen. At t=2 seconds, a motion imagery prompt appears on the computer screen to prompt the subject to perform motion imagery. The motion imagery prompt is an arrow prompt, which includes two types: left-pointing arrow prompts and right-pointing arrow prompts.
[0018] For a total of 4 seconds from t=2 to t=6, the arrow prompt on the computer screen remained unchanged, and the subject continued to perform the motor imagery prompted by the current prompt; then rest for 2 seconds; at t=8 seconds, the second round of motor imagery experiment began, and the above process was repeated, and the subject was subjected to multiple motor imagery experiments; and during each motor imagery task, the subject's EEG signal data were collected from the beginning to the end of the task.
[0019] Similarly, EEG signal data of each subject was collected during the motor imagery task, and the collected EEG signal data of the ID subject was used as the training set. The training set data is the source domain data.
[0020] The EEG signal data collected during the subject's motor imagery task is used as the target domain data.
[0021] Furthermore, the specific process of step two is as follows:
[0022] For each set of collected EEG signal data, the motor imagery period was extracted, that is, the EEG signal data within 2.5 seconds to 5.5 seconds was extracted, and the extracted EEG signal data was processed by power frequency notch filtering and bandpass filtering.
[0023] Furthermore, the specific process of step three is as follows:
[0024] The covariance matrix is calculated based on the preprocessed EEG signal data of the ID-bit subjects and the target domain in the training set. Then, the mean of the covariance matrix in Riemann space is calculated. Based on the mean, the preprocessed EEG signal data of each group is aligned to obtain the processed EEG signal data of each group.
[0025] Furthermore, the specific process of step six is as follows:
[0026] Step 61: Calculate the intra-class scattering matrix and inter-class scattering matrix for the subject at position id in the training set. Then, the projection matrix, based on the constraints of the intra-class scattering matrix and the inter-class scattering matrix, is as follows:
[0027]
[0028] Where A is the projection matrix, A T It is the transpose of A, tr(·) is the trace of the matrix, S wLet be the intraclass scattering matrix of the EEG signal data of the subject at position id. It is a matrix composed of the tangent space vectors corresponding to the data belonging to category c in the EEG signal data of the subject at position id. Let C represent a real number, where C is the number of types of EEG signal data from the id-th subject. Let be the number of data groups of category c in the EEG signal data of subject id, and D be the dimension of the tangent space vector corresponding to each EEG signal data group. For the data center matrix of category c, S is the identity matrix. b Let be the inter-class scattering matrix of the EEG signal data of the subject at position id. Let be the mean of the tangent space vectors corresponding to the EEG signal data of each group in category c of subject id. Let be the tangent space vector corresponding to the i-th group of EEG signal data in category c of subject id. Let be the mean of the tangent space vectors corresponding to all EEG signal data of the subject at position id. X i Let n represent the tangent space vector corresponding to the i-th group of EEG signal data of the id-th subject. s This represents the total number of EEG signal data sets for the subject at position id.
[0029] Step 62: Calculate the EEG signal data scattering matrix V of the target domain. t Then the projection matrix is constrained by the data scattering matrix as follows:
[0030] m B axtr(B T V t B)
[0031] Where B represents the projection matrix, X t This represents the matrix composed of the tangent space vectors corresponding to each group of processed EEG signal data in the target domain. n t H represents the total number of EEG signal data sets in the target domain. t The center matrix of the target domain, It is the identity matrix;
[0032] Step 63: Projection matrices A and B satisfy the following constraints:
[0033]
[0034] Among them, ||·|| F This indicates the calculation of the F-norm;
[0035] Step 64: Solve the objective function according to the constraints of Steps 61 to 63 to obtain the projection matrices A and B;
[0036] Next, the tangent space vectors of each group of EEG signal data in the target domain are projected and mapped according to projection matrices A and B to obtain the projected vectors; then the projected vectors are input into the SVM classifier to obtain the classification results of each group of EEG signal data in the target domain, and the obtained classification results are used as the initial classification results.
[0037] Step 65: Initialize the iteration count l = 1;
[0038] Step 66: Calculate the intra-class and inter-class scattering matrices of each group of EEG signal data in the target domain. The projection matrix, based on the constraints of the intra-class and inter-class scattering matrices, is as follows:
[0039]
[0040] Among them, T w The intraclass scattering matrix of each group of EEG signal data in the target domain. It is a matrix composed of the tangent space vectors corresponding to the EEG signal data belonging to category c in the target domain. This represents the number of EEG signal data sets belonging to category c in the target domain. The center matrix of EEG signal data for category c, T is the identity matrix. b The inter-class scattering matrix of the target domain EEG signal data. This represents the mean of the tangent space vectors corresponding to EEG signal data belonging to category c in the target domain. Let represent the tangent space vector corresponding to the i-th group of EEG signal data belonging to category c in the target domain. This represents the mean of the tangent space vectors corresponding to each group of EEG signal data in the target domain. X i is the tangent space vector corresponding to the i-th group of EEG signal data in the target domain;
[0041] Step 67: Solve the objective function according to the constraints of Steps 61 to 63 and Step 66 to obtain the projection matrices A and B;
[0042] Step 68: Determine if the number of iterations has reached the set maximum number of iterations L;
[0043] If not achieved, the tangent space vectors of each group of EEG signal data in the target domain are projected and mapped according to projection matrices A and B. The projected and mapped vectors are input into the SVM classifier to obtain the classification results of each group of EEG signal data in the target domain. Let l = l + 1, and use the obtained classification results to return to step six.
[0044] If this is achieved, the tangent space vectors of the id-th subject and each group of EEG signal data in the target domain within the training set are projected and mapped according to projection matrices A and B to obtain the projection mapping result, which is the feature vector of the id-th subject and the target domain after projection mapping.
[0045] Furthermore, the objective function in step six-four is:
[0046]
[0047] Where I is the identity matrix, and ω, μ, α and λ are coefficients;
[0048] By constructing a Lagrange function and setting its derivative to 0, the objective function can be expressed as:
[0049]
[0050] Among them, let Φ is a matrix A diagonal matrix composed of eigenvalues, Φ = diag(Φ1, Φ2, ..., Φ k ), Φ1,Φ2,...,Φ k Represents a matrix The eigenvalues are sorted in descending order, and the 1st, 2nd, ..., kth eigenvalues are denoted as W1, W2, ..., Wk. k The eigenvalues are Φ1, Φ2, ..., Φ k The corresponding eigenvectors, matrix W = [W1,...,W k W = [A; B].
[0051] Furthermore, the specific process of step seven is as follows:
[0052] Step 71: Initialize the iteration count l′ = 1;
[0053] Step 72, the principle of minimizing SRM is expressed as:
[0054]
[0055] In the formula, Let f represent a set of classifiers in the kernel space, where f is the classifier. Let represent the loss function on the matrix x consisting of the projected mapping results (projected feature vectors) of the subject at position id in the training set and the EEG signal data of each group in the target domain, y represent the classification label of x, η represent the hyperparameters of R(f), and R(f) represent... The square norm of f in;
[0056] The expansion of f is as follows:
[0057]
[0058] In the formula, n s n represents the number of EEG signal data sets of the id-th subject in the training set. t The number of EEG signal data sets in the target domain is represented by β, where β represents the weight vector. This indicates the weight of each group of EEG signal data. K(x) is a real number. i (,x) represents x i Correlation with x, x i This represents the projection mapping result of the id-th subject in the training set and any set of EEG signal data in the target domain;
[0059] Expanding f further, we get:
[0060]
[0061] In the formula, f(·) represents the output of classifier f, K represents the kernel matrix, and K(x) represents the kernel matrix. i ,x j ) represents x i With x j The correlation, K(x) i ,x j Let y be the element in the i-th row and j-th column of the kernel matrix K. i x represents i Category tags, Let Y represent the label vector, Y be the one-hot encoding matrix of y, and A′ represent the diagonal matrix. If i belongs to the training set data, then the i-th diagonal element A′ in the diagonal matrix is... i ′ i =1; otherwise, A i ′ i =0;
[0062] Step 73: Calculate the maximum average difference between the projected mapping result of the EEG signal data of the id-th subject in the training set and the projected mapping result of the EEG signal data in the target domain:
[0063]
[0064] In the formula, Indicates the maximum average difference, M0 represents the marginal MMD matrix, and K i It is the i-th column vector of K. It is the (j+n)th digit of K. s column vectors, H K Represents the Hilbert kernel space;
[0065]
[0066] In the formula, This represents the set of projection mapping results of the EEG signal data of each group of subjects at position id in the training set. M0 represents the set of projection mapping results for each group of EEG signal data in the target domain. ij This represents the element in row i and column j of M0;
[0067] Step 74: Calculate Conditional Metric Constraints
[0068]
[0069] in, This represents the matrix consisting of the projection mapping results of the data belonging to category c in the projection mapping results of the subject at position id in the training set; (x t ) (c) This represents the matrix consisting of the projection mapping results corresponding to data belonging to category c in the projection mapping results of the target domain; (x) (c) This represents the matrix consisting of the projection mapping results of data belonging to category c in the projection mapping results of the subject at position id in the training set and each group of EEG signal data in the target domain; This indicates the category of the projected mapping results corresponding to the id-th subject in the training set and each group of EEG signal data in the target domain. The matrix composed of the projection mapping results corresponding to the data; E[·] represents a category that does not belong to category c; E[·] represents the expectation. For intra-class MMD matrices; For inter-class MMD matrix;
[0070]
[0071] in, express The element in the i-th row and j-th column, Let represent the set of data belonging to category c in the projection mapping results of the subject at position id in the training set. This represents the number of projection mapping results belonging to category c among the projection mapping results of the subject at position id in the training set. This represents the set of data belonging to category c in the projection mapping result of the target domain. This represents the number of mapping results belonging to category c in the projection mapping results of the target domain;
[0072]
[0073] In the formula, express The element in the i-th row and j-th column, Let n represent the set of projected mapping results belonging to category c in the projected mapping results of the id-th subject in the training set and the projected mapping results of the target domain. (c) This represents the number of projection mapping results belonging to category c in both the projection mapping results of the id-th subject in the training set and the projection mapping results of the target domain. This represents the number of projection mapping results that do not belong to category c in the projection mapping results of the subject at position id in the training set and the projection mapping results of the target domain.
[0074] Step 75: Calculate the mean H of the within-class difference center matrix:
[0075]
[0076] in, when When the data at the corresponding position belongs to category c, the element at the corresponding position is 1; otherwise, the element at the corresponding position is 0. Let n = n s +n t ;
[0077] Step 76: Calculate matrix S:
[0078]
[0079] In the formula, S ij Let represent the element in the i-th row and j-th column of matrix S, and sim(·,·) represent the distance between two points. Representing point x i Belongs to x j The neighborhood set;
[0080] Laplace regularization Represented as:
[0081]
[0082] In the formula, L is the Laplace matrix, L ij Let L represent the element in the i-th row and j-th column of the Laplace matrix, where L = DS, D is a diagonal matrix, and
[0083] Step 77: Based on steps 72 to 76, obtain the final classifier:
[0084]
[0085] In the formula, ε, ρ, and γ are all coefficients;
[0086] Based on the final classifier in the above formula, establish the objective function:
[0087]
[0088] By taking the derivative of the objective function and setting it to zero, we obtain the weight vector β of the classifier. * :
[0089] β * =((A′+εM0+(1-ε)M) c +ρL+γH)K+ηI) -1 A′Y T
[0090] Steps 7 and 8: Determine whether the number of iterations l′ has reached the set maximum number of iterations L′;
[0091] If the target is not reached, let l′ = l′ + 1, use the classifier obtained in step 77 to obtain the pseudo label of the target domain, and then return to execute step 72.
[0092] If this is achieved, the classifier obtained in the last iteration will be used as the classifier CLS obtained based on the subject at position id in the training set. id .
[0093] Furthermore, the specific process of step nine is as follows:
[0094] Step 91: Initialize the subject id in the source domain to 1;
[0095] Step 92: Pass the tangent space vectors of all EEG signal data in the target domain and the tangent space vectors of all EEG signal data of the subject at position id in the source domain through a linear classifier, and calculate the edge distance based on the output of the linear classifier.
[0096] The projection mapping results corresponding to the tangent space vectors of all EEG signal data in the target domain and the projection mapping results corresponding to the tangent space vectors of all EEG signal data of the subject at position id in the source domain are passed through a linear classifier, and the edge distance is calculated based on the output of the linear classifier.
[0097] Step 93: Pass the tangent space vector of the c-th type EEG signal data of the target domain and the source domain of the id-th subject through a linear classifier, and calculate the conditional distance of the tangent space vector of the c-th type EEG signal data based on the output of the linear classifier.
[0098] The projection mapping results of the tangent space vector of the id-th subject's EEG signal data of type c in the target domain and the source domain are passed through a linear classifier. The conditional distance of the projection mapping result corresponding to the tangent space vector of the id-th subject's EEG signal data is calculated based on the output of the linear classifier. c = 1, 2, ..., C;
[0099] Step 94: Adjust edge distance and conditional distance The integration is performed to obtain the integration distance corresponding to the subject at position id in the source domain;
[0100] The specific process for calculating the integration distance is as follows:
[0101]
[0102] Where, d M and d c All are intermediate variables;
[0103] Then the integration distance d corresponding to the id-th subject in the source domain id,Mc for:
[0104]
[0105] Step 95: Determine if id = ID;
[0106] If id = ID, then proceed to step nine six;
[0107] If id = ID is not satisfied, then set id = id + 1 and return to step nine two.
[0108] Step 96: Based on the integration distance d corresponding to each subject in the source domain id,Mc We calculate the weight factor of the classifier for each subject, and denote the weight factor of the classifier for the subject at position id in the source domain as wt. id ;
[0109] Then proceed to step 97;
[0110] Step 97: For any set of EEG signal data in the target domain, use the projection mapping results corresponding to the set of EEG signal data as the input of the classifier for each subject in the source domain.
[0111] Then, the output of the classifier corresponding to the subject at position id in the source domain is combined with the weight factor wt. id Multiply to obtain the result, which includes the probability that the set of EEG signal data belongs to each category, and id = 1, 2, ..., ID;
[0112] The elements in the multiplication results of each classifier are added together. Each element in the summation result represents the final probability of the EEG signal data belonging to each category. The maximum value is then selected from the final probabilities of each category, and the category corresponding to the maximum value is the category to which the EEG signal data belongs.
[0113] Furthermore, the method for calculating the edge distance is as follows:
[0114] For edge distance The tangent space vectors of all EEG signal data in the target domain and the tangent space vectors of all EEG signal data of the subject at position id in the source domain are passed through a linear classifier. The linear classifier outputs whether the EEG signal data corresponding to each tangent space vector belongs to the source domain or the target domain.
[0115]
[0116] Where h represents the linear classifier, and ε(h) represents the inaccuracy rate of the linear classifier's output.
[0117] Furthermore, the specific process of step ninety-six is as follows:
[0118]
[0119] Where, wt id represents the weight factor of the classifier corresponding to the subject at position id in the source domain, and e is the base of the natural logarithm.
[0120] The beneficial effects of this invention are:
[0121] This invention fully utilizes information from both the source domain (training set) and target domain samples, combining systematic data processing and classification algorithms to significantly improve the recognition accuracy and robustness of motor imagery tasks. Through efficient data preprocessing and feature extraction methods in domain-adaptive manifold embedding, the consistency of feature mapping between the training set and the target domain is ensured. Furthermore, classifier optimization based on the principle of structural risk minimization further enhances classification performance. Through feature fusion and voting mechanisms, the reliability of label classification is effectively improved. This invention effectively avoids the problem of low classification accuracy in the MI-BCI system due to differences in EEG signals among subjects, laying a solid foundation for applications in related fields. Attached Figure Description
[0122] Figure 1 A flowchart illustrating the experimental paradigm for acquiring electroencephalogram (EEG) signals;
[0123] Figure 2 The flowchart shows the domain-adaptive manifold embedding distribution alignment method. Detailed Implementation
[0124] Specific implementation method one: Combining Figure 2 This embodiment describes a domain adaptation method for addressing subject-specific differences in motor imagery brain-computer interfaces. The specific process of this method is as follows:
[0125] Step 1: The ID-position subjects used an online brain-computer interface system to perform multiple motor imagery tasks. The motor imagery tasks included leftward motor imagery tasks and rightward motor imagery tasks. During each motor imagery task, the subjects' EEG signal data were collected. The collected EEG signal data of the ID-position subjects was used as the training set data, and the EEG signal data collected by the subjects under test during the motor imagery tasks was used as the target domain data.
[0126] Step 2: Extract the EEG signal data corresponding to the time period of motor imagery from each group of EEG signal data in the training set and the target domain, and preprocess the extracted EEG signal data to obtain the preprocessed EEG signal data of each group.
[0127] Step 3: Use a pre-alignment strategy to process each group of preprocessed EEG signal data in the training set and the target domain to obtain the processed EEG signal data.
[0128] Step 4: Based on each group of processed EEG signal data, calculate the covariance matrix of each group of EEG signal data, and then use the tangent space feature mapping method to project the covariance matrix of each group of EEG signal data onto the tangent space to obtain the tangent space vector of each group of EEG signal data. The dimension of a single trial obtained after projection is D.
[0129] Step 5: Initialize id = 1;
[0130] Step 6: Use the domain adaptive manifold embedding distribution alignment feature mapping method to project and map the tangent space vectors of each group of EEG signal data of the id-th subject in the training set and the tangent space vectors of each group of EEG signal data in the target domain, so as to obtain the feature vectors of the id-th subject and the tangent space vectors of the target domain after projection mapping.
[0131] Step 7: Based on Structural Risk Minimization (SRM) and Regularization Theory (LR), construct a classifier CLS using the feature vector after projection mapping of the id-th subject and the target domain. id ;
[0132] Step 8: Determine if id = ID;
[0133] If the conditions are met, then the constructed ID classifiers CLS1, CLS2, ..., CLS are obtained. ID After projecting and mapping the target domain, the feature vectors are passed through classifiers CLS1, CLS2, ..., CLS.ID The classification results Rst1, Rst2, ..., Rst are obtained respectively. ID Then proceed to step nine;
[0134] If the condition is not met, then set id = id + 1 and return to step six.
[0135] Step 9: Perform feature fusion on the ID classification results obtained in Step 8 to obtain the final classification result for the target domain data.
[0136] Specific Implementation Method Two: Combining Figure 1 This embodiment is described below. The difference between this embodiment and specific embodiment one is that the specific process of step one is as follows:
[0137] For any given subject, the subject is positioned facing the computer screen. At t=0, a fixed cross symbol appears on the computer screen. At t=2 seconds, a motion imagery prompt appears on the computer screen to prompt the subject to perform motion imagery. The motion imagery prompt is an arrow prompt, which includes two types: left-pointing arrow prompts and right-pointing arrow prompts.
[0138] For 4 seconds, from t=2 to t=6, the arrow prompt on the computer screen remained unchanged, and the subject continued to perform the motor imagery prompted by the current prompt. After a two-second rest, at t=8, the second round of motor imagery experiment began, and the above process was repeated. The subject underwent multiple motor imagery experiments. During each motor imagery task, the subject's EEG signal data was collected from the beginning to the end of the task (sampling rate set to 100Hz). The number of times the left arrow prompt appeared was the same as the number of times the right arrow prompt appeared, without any experimental bias. 200 sets of experimental data were collected for each subject.
[0139] Similarly, EEG signal data of each subject during the motor imagery task were collected separately, and the collected EEG signal data of the ID subject was used as the training set. The training set data is the source domain data (and the data in the training set were numbered according to the order of the experiment to ensure the standardization of data processing and the repeatability of the experiment, which is used to assist in the classification of the EEG data of the subjects in the test set).
[0140] The EEG signal data collected during the subject's motor imagery task is used as the target domain data (the category of the subject's EEG signal data is unknown beforehand; the purpose of this invention is to classify the subject's EEG signal data).
[0141] The other steps and parameters are the same as in Specific Implementation Method 1.
[0142] Specific Implementation Method Three: This implementation method differs from Specific Implementation Method One or Two in that the specific process of step two is as follows:
[0143] For each set of collected EEG signal data, the motor imagery period was extracted, that is, the EEG signal data within 2.5 seconds to 5.5 seconds was extracted, and the extracted EEG signal data was processed by power frequency notch filtering (filtering frequency of 50Hz) and bandpass filtering (filtering frequency of 8H to 30Hz).
[0144] Other steps and parameters are the same as in specific implementation method one or two.
[0145] Specific Implementation Method Four: This implementation method differs from Specific Implementation Methods One to Three in that the specific process of step three is as follows:
[0146] The covariance matrix is calculated based on the preprocessed EEG signal data of the ID-bit subjects and the target domain in the training set. Then, the mean of the covariance matrix in Riemann space is calculated. Based on the mean, the preprocessed EEG signal data of each group is aligned to obtain the processed EEG signal data of each group.
[0147] The other steps and parameters are the same as those in one of the specific implementation methods one to three.
[0148] After data alignment, the distribution of all trial data is more similar, making it more suitable for subsequent classification.
[0149] Specific Implementation Method Five: This implementation method differs from Specific Implementation Methods One to Four in that the specific process of step six is as follows:
[0150] Step 61: Calculate the intra-class scattering matrix and inter-class scattering matrix for the subject at position id in the training set. Then, the projection matrix, based on the constraints of the intra-class scattering matrix and the inter-class scattering matrix, is as follows:
[0151]
[0152] Where A is the projection matrix, A T It is the transpose of A, tr(·) is the trace of the matrix, S w Let be the intraclass scattering matrix of the EEG signal data of the subject at position id. It is a matrix composed of the tangent space vectors corresponding to the data belonging to category c in the EEG signal data of the subject at position id. Let C represent a real number, where C is the number of EEG signal data types for the subject at position id (since this invention includes leftward and rightward movement cues, C is set to 2). Let be the number of data groups of category c in the EEG signal data of subject id, and D be the dimension of the tangent space vector corresponding to each EEG signal data group. For the data center matrix of category c, S is the identity matrix. b Let be the inter-class scattering matrix of the EEG signal data of the subject at position id. Let be the mean of the tangent space vectors corresponding to the EEG signal data of each group in category c of subject id. Let be the tangent space vector corresponding to the i-th group of EEG signal data in category c of subject id. Let be the mean of the tangent space vectors corresponding to all EEG signal data of the subject at position id. X i Let n represent the tangent space vector corresponding to the i-th group of EEG signal data of the id-th subject. s This represents the total number of EEG signal data sets for the subject at position id.
[0153] The intra-class scattering matrix and the inter-class scattering matrix can preserve the discriminative information within the domain, ensuring that data samples of the same class are close to each other and data samples of different classes are far apart from each other.
[0154] Step 62: Calculate the category-independent EEG signal data scattering matrix V of the target domain. t Then the projection matrix is constrained by the data scattering matrix as follows:
[0155]
[0156] Where B represents the projection matrix, X t This represents the matrix composed of the tangent space vectors corresponding to each group of processed EEG signal data in the target domain. n t H represents the total number of EEG signal data sets in the target domain. t The center matrix of the target domain, It is the identity matrix;
[0157] Due to the uncertainty of the predicted labels, relying solely on discriminative information is insufficient to prevent irrelevant features from entering irrelevant dimensions. Therefore, we continue to use target domain variance maximization to avoid information brought by irrelevant dimensions.
[0158] Step 63: To achieve better generalization performance, it is necessary to ensure that A and B do not contain extreme values, that is, the projection matrices A and B satisfy the following constraints:
[0159]
[0160] Among them, ||·|| F This indicates the calculation of the F-norm;
[0161] The requirement that projection matrices A and B satisfy the above conditions can reduce the difference between the source domain and the target domain, and further narrow the difference between the source domain and the target domain.
[0162] Step 64: Solve the objective function according to the constraints of Steps 61 to 63 to obtain the projection matrices A and B;
[0163] Next, the tangent space vectors of each group of EEG signal data in the target domain are projected and mapped according to projection matrices A and B to obtain the projected vectors; then the projected vectors are input into the SVM classifier to obtain the classification results of each group of EEG signal data in the target domain, and the obtained classification results are used as the initial classification results.
[0164] Step 65: Initialize the iteration count l = 1;
[0165] Step 66: Calculate the intra-class and inter-class scattering matrices of each group of EEG signal data in the target domain. The projection matrix, based on the constraints of the intra-class and inter-class scattering matrices, is as follows:
[0166]
[0167]
[0168] Among them, T w The intraclass scattering matrix of each group of EEG signal data in the target domain. It is a matrix composed of the tangent space vectors corresponding to the EEG signal data belonging to category c in the target domain. This represents the number of EEG signal data sets belonging to category c in the target domain. The center matrix of EEG signal data for category c, T is the identity matrix. b The inter-class scattering matrix of the target domain EEG signal data. This represents the mean of the tangent space vectors corresponding to EEG signal data belonging to category c in the target domain. Let represent the tangent space vector corresponding to the i-th group of EEG signal data belonging to category c in the target domain. This represents the mean of the tangent space vectors corresponding to each group of EEG signal data in the target domain. X i is the tangent space vector corresponding to the i-th group of EEG signal data in the target domain;
[0169] By using a classifier trained in the source domain to help predict pseudo-labels, information in the target domain is preserved, resulting in intra-class and inter-class scattering matrices that are close to the true matrices. Minimizing T... w Maximize T at the same time b This makes it easier to categorize.
[0170] Step 67: Solve the objective function according to the constraints of Steps 61 to 63 and Step 66 to obtain the projection matrices A and B;
[0171] Step 68: Determine if the number of iterations has reached the set maximum number of iterations L;
[0172] If not achieved, the tangent space vectors of each group of EEG signal data in the target domain are projected and mapped according to projection matrices A and B. The projected and mapped vectors are input into the SVM classifier to obtain the classification results of each group of EEG signal data in the target domain. Let l = l + 1, and use the obtained classification results to return to step six.
[0173] If this is achieved, the tangent space vectors of the id-th subject and each group of EEG signal data in the target domain within the training set are projected and mapped according to projection matrices A and B to obtain the projection mapping result, which is the feature vector of the id-th subject and the target domain after projection mapping.
[0174] The other steps and parameters are the same as those in one of the specific implementation methods one to four.
[0175] To obtain a more stable projection matrix, steps 66 to 68 require multiple iterations. After solving for projection matrices A and B, the tangent space vectors are projected to obtain the feature vectors of the projected training set and the target domain. Step 6 is repeated for ID subjects in the source domain training set and the current target domain to obtain different projection matrices for the ID groups.
[0176] Specific Implementation Method Six: This implementation method differs from Specific Implementation Methods One to Five in that the objective function in step six-four is:
[0177]
[0178] Where I is the identity matrix, ω, μ, α and λ are all coefficients (the specific values can be selected according to the actual situation), ω corresponds to the parameter for preserving the discriminative information of the source domain, μ corresponds to the parameter for maximizing the variance of the target domain, λ corresponds to the parameter for minimizing the subspace difference and regularization, and α corresponds to the parameter for preserving the discriminative information of the target domain.
[0179] By constructing a Lagrange function and setting its derivative to 0, the objective function can then be expressed as:
[0180]
[0181] The above equation can be solved using generalized eigenvalue decomposition;
[0182] Among them, let Φ is a matrix A diagonal matrix composed of eigenvalues, Φ = diag(Φ1, Φ2, ..., Φ k ), Φ1,Φ2,…,Φ k Represents a matrix The eigenvalues are sorted in descending order, and the 1st, 2nd, ..., kth eigenvalues are denoted as W1, W2, ..., Wk. k The eigenvalues are Φ1, Φ2, ..., Φ k The corresponding eigenvectors, matrix W = [W1,...,W k W = [A; B].
[0183] The other steps and parameters are the same as those in one of the specific implementation methods one to five.
[0184] In step six-four, when calculating the objective function, let T... b The matrix is zero. In each subsequent iteration, T is calculated according to the value obtained in step six of the current iteration. b That's all.
[0185] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Methods One to Six in that the specific process of step seven is as follows:
[0186] Step 71: Initialize the iteration count l′ = 1;
[0187] Step 72: Based on the SRM principle, first calculate the squared loss in recursive least squares and the squared norm of CLS in the Hilbert kernel space; the principle of minimizing SRM is expressed as:
[0188]
[0189] In the formula, Let f represent a set of classifiers in the kernel space (which may include, but is not limited to, the rbf kernel space), where f is the classifier. Let y represent the loss function on the matrix x consisting of the projection mapping results (projected feature vectors) of the subject at position id in the training set and the EEG signal data of each group in the target domain, y represent the classification label of x, η represent the hyperparameters of R(f) (those skilled in the art can set them according to experience), and R(f) represent The square norm of f in;
[0190] The expansion of f is as follows:
[0191]
[0192] In the formula, n s n represents the number of EEG signal data sets of the id-th subject in the training set. tThe number of EEG signal data sets in the target domain is represented by β, where β represents the weight vector. This indicates the weight of each group of EEG signal data. K(x) is a real number. i (,x) represents x i Correlation with x, x i This represents the projection mapping result of the id-th subject in the training set and any set of EEG signal data in the target domain;
[0193] Expanding f further, we get:
[0194]
[0195] In the formula, f(·) represents the output of classifier f, K represents the kernel matrix (the kernel space that can be used includes, but is not limited to, the rbf kernel space), K(x i ,x j ) represents x i With x j The correlation, K(x) i ,x j Let y be the element in the i-th row and j-th column of the kernel matrix K. i x represents i Category tags, Let Y represent the label vector, Y be the one-hot encoding matrix of y, and A′ represent the diagonal matrix. Since this part only focuses on the source domain SRM, the diagonal matrix in the equation... In other words, if i belongs to the training set data, then the i-th diagonal element A in the diagonal matrix i ′ i =1; otherwise, A i ′ i =0;
[0196] Step 73: Calculate the maximum mean difference (MMD) between the projected mapping results of the EEG signal data of the id-th subject in the training set and the projected mapping results of the EEG signal data in the target domain.
[0197]
[0198] In the formula, Indicates the maximum average difference, M0 represents the marginal MMD matrix, and K i It is the i-th column vector of K. It is the (j+n)th digit of K. s column vectors, H K Let Hilbert kernel space be represented, and the above formula calculates the 2-norm in Hilbert kernel space;
[0199]
[0200] In the formula, This represents the set of projection mapping results of the EEG signal data of each group of subjects at position id in the training set. M0 represents the set of projection mapping results for each group of EEG signal data in the target domain. ij This represents the element in row i and column j of M0;
[0201] This step minimizes the distance between the source and target domains, measured by the maximum average difference in projection, through edge metric constraints, thereby further reducing the global difference between the two domains.
[0202] Step 74: Calculate Conditional Metric Constraints
[0203]
[0204] in, This represents the matrix consisting of the projection mapping results of the data belonging to category c in the projection mapping results of the subject at position id in the training set; (x t ) (c) This represents the matrix consisting of the projection mapping results corresponding to data belonging to category c in the projection mapping results of the target domain; (x) (c) This represents the matrix consisting of the projection mapping results of data belonging to category c in the projection mapping results of the subject at position id in the training set and each group of EEG signal data in the target domain; This indicates the category of the projected mapping results corresponding to the id-th subject in the training set and each group of EEG signal data in the target domain. The matrix composed of the projection mapping results corresponding to the data; E[·] represents a category that does not belong to category c; E[·] represents the expectation. For intra-class MMD matrices; For inter-class MMD matrix;
[0205]
[0206] in, express The element in the i-th row and j-th column, Let represent the set of data belonging to category c in the projection mapping results of the subject at position id in the training set. This represents the number of projection mapping results belonging to category c among the projection mapping results of the subject at position id in the training set. This represents the set of data belonging to category c in the projection mapping result of the target domain. This represents the number of mapping results belonging to category c in the projection mapping results of the target domain;
[0207]
[0208] In the formula, express The element in the i-th row and j-th column, Let n represent the set of projected mapping results belonging to category c in the projected mapping results of the id-th subject in the training set and the projected mapping results of the target domain. (c) This represents the number of projection mapping results belonging to category c in both the projection mapping results of the id-th subject in the training set and the projection mapping results of the target domain. This represents the number of projection mapping results that do not belong to category c in the projection mapping results of the subject at position id in the training set and the projection mapping results of the target domain.
[0209] This step uses conditional metric constraints to reduce intra-class differences between the source and target domains and increase inter-class differences to help with better classification. When constraining the target domain, the first iteration uses a linear discriminator to predict pseudo-labels for the target domain. After multiple iterations, pseudo-labels that are close to the true labels are obtained. In each subsequent iteration of the target domain, the pseudo-labels are obtained using the classifier obtained in the previous iteration.
[0210] Step 75: Calculate the mean H of the within-class difference center matrix:
[0211]
[0212] in, when When the data at the corresponding position belongs to category c, the element at the corresponding position is 1; otherwise, the element at the corresponding position is 0. Let n = n s +n t ;
[0213] This step addresses the excessive constraints generated in the balancing condition metric constraints by using metric correction to minimize certain inter-class differences; similarly, projection MMD is used to minimize inter-class differences.
[0214] Step 76: Calculate matrix S:
[0215]
[0216] In the formula, S ij Let represent the element in the i-th row and j-th column of matrix S, and sim(·,·) represent the distance between two points (cosine distance is used in this invention). Representing point x i Belongs to x j The neighborhood set (the neighborhood set includes points x) j The nearest p points (where p is a pre-defined parameter);
[0217] Laplace regularization Represented as:
[0218]
[0219] In the formula, L is the Laplace matrix, L ij Let L represent the element in the i-th row and j-th column of the Laplace matrix, where L = DS, D is a diagonal matrix, and
[0220] This step uses Laplace regularization to introduce constraints in the parameter space. These constraints form a smooth hyperplane, causing the optimized solution to tend towards a region with less variation.
[0221] Step 77: Based on steps 72 to 76, obtain the final classifier:
[0222]
[0223] In the formula, ε, ρ, and γ are all coefficients; ε corresponds to the marginal metric constraint parameter, 1-ε is the conditional metric constraint parameter, ρ represents the Laplace regularization parameter, and γ represents the metric correction parameter.
[0224] Based on the final classifier in the above formula, establish the objective function:
[0225]
[0226] By taking the derivative of the objective function and setting it to zero, we obtain the weight vector β of the classifier. * :
[0227] β * =((A′+εM0+(1-ε)M) c +ρL+γH)K+ηI) -1 A′Y T
[0228] Then β * Substitution Obtain the classifier f;
[0229] Steps 7 and 8: Determine whether the number of iterations l′ has reached the set maximum number of iterations L′;
[0230] If the target is not reached, let l′ = l′ + 1, use the classifier obtained in step 77 to obtain the pseudo label of the target domain, and then return to execute step 72.
[0231] If this is achieved, the classifier obtained in the last iteration will be used as the classifier CLS obtained based on the subject at position id in the training set. id .
[0232] The other steps and parameters are the same as those in one of the specific implementation methods one to six.
[0233] Specific Implementation Method Eight: This implementation method differs from Specific Implementation Methods One to Seven in that the specific process of step Nine is as follows:
[0234] Step 91: Initialize the subject id in the source domain to 1;
[0235] Step 92: Pass the tangent space vectors of all EEG signal data in the target domain and the tangent space vectors of all EEG signal data of the subject at position id in the source domain through a linear classifier, and calculate the edge distance based on the output of the linear classifier.
[0236] The projection mapping results corresponding to the tangent space vectors of all EEG signal data in the target domain and the projection mapping results corresponding to the tangent space vectors of all EEG signal data of the subject at position id in the source domain are passed through a linear classifier, and the edge distance is calculated based on the output of the linear classifier.
[0237] Step 93: Pass the tangent space vector of the c-th type EEG signal data of the target domain and the source domain of the id-th subject through a linear classifier, and calculate the conditional distance of the tangent space vector of the c-th type EEG signal data based on the output of the linear classifier.
[0238] The projection mapping results of the tangent space vector of the id-th subject's EEG signal data of type c in the target domain and the source domain are passed through a linear classifier. The conditional distance of the projection mapping result corresponding to the tangent space vector of the id-th subject's EEG signal data is calculated based on the output of the linear classifier. c = 1, 2, ..., C;
[0239] Step 94: Adjust edge distance and conditional distance The integration is performed to obtain the integration distance corresponding to the subject at position id in the source domain;
[0240] The specific process for calculating the integration distance is as follows:
[0241]
[0242] Where, d M and d c All are intermediate variables;
[0243] Then the integration distance d corresponding to the id-th subject in the source domain id,Mc for:
[0244]
[0245] By integrating edge distance and conditional distance, the integrated overall distribution distance can be obtained. The difference in overall distribution distance indicates the overall difference between the source domain and the current target domain. The range of the overall distribution distance difference after integration remains unchanged, still within the [0,2] interval. The larger the distribution distance difference, the lower the credibility of the current source domain label. The source domain with the smallest distribution distance to the target domain will be considered as having the largest weight factor in the classification result.
[0246] Step 95: Determine if id = ID;
[0247] If id = ID, then proceed to step nine six;
[0248] If id = ID is not satisfied, then set id = id + 1 and return to step nine two.
[0249] Step 96: Based on the integration distance d corresponding to each subject in the source domain id,Mc We calculate the weight factor of the classifier for each subject, and denote the weight factor of the classifier for the subject at position id in the source domain as wt. id ;
[0250] Then proceed to step 97;
[0251] Step 97: For any set of EEG signal data in the target domain, use the projection mapping results corresponding to the set of EEG signal data as the input of the classifier for each subject in the source domain.
[0252] Then, the output of the classifier corresponding to the subject at position id in the source domain is combined with the weight factor wt. id Multiply to obtain the result, which includes the probability that the set of EEG signal data belongs to each category, and id = 1, 2, ..., ID;
[0253] Then, the elements in the multiplication results of each classifier are added together (the first element in each multiplication result is added together, and the sum is the first element in the sum result; similarly, each element in the sum result is obtained). The elements in the sum result are the final probabilities of the EEG signal data belonging to each category. Then, the maximum value is selected from the final probabilities of each category, and the category corresponding to the maximum value is the category to which the EEG signal data belongs.
[0254] The other steps and parameters are the same as those in any of the specific implementation methods one to seven.
[0255] Specific Implementation Method Nine: This implementation method differs from Specific Implementation Methods One to Eight in that the method for calculating the edge distance is as follows:
[0256] For edge distance The tangent space vectors of all EEG signal data in the target domain and the tangent space vectors of all EEG signal data of the subject at position id in the source domain are passed through a linear classifier (the linear classifier that can be used in this invention includes, but is not limited to, SVM). The linear classifier outputs whether the EEG signal data corresponding to each tangent space vector belongs to the source domain or the target domain.
[0257]
[0258] Where h represents the linear classifier, and ε(h) represents the error rate of the linear classifier's output (in this invention, the label for all EEG signal data in the target domain is 0, and the label for all EEG signal data in the source domain is 1. The error rate can be calculated based on the label and the output of the linear classifier).
[0259] The other steps and parameters are the same as those in one of the specific implementation methods one to eight.
[0260] Similarly, and All calculations were performed using the method described in this embodiment.
[0261] Specific Implementation Method Ten: This implementation method differs from Specific Implementation Methods One to Nine in that the specific process of step Nine Six is as follows:
[0262]
[0263] Where, wt id represents the weight factor of the classifier corresponding to the subject at position id in the source domain, and e is the base of the natural logarithm.
[0264] The other steps and parameters are the same as those in any of the specific implementation methods one to nine.
[0265] The above examples of the present invention are merely illustrative of the computational model and process of the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is impossible to exhaustively list all possible implementations here. Any obvious variations or modifications derived from the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A domain adaptation method for solving subject difference in motor imagery brain-computer interface, characterized in that, The method specifically comprises the following steps: Step one, the ID bit subject uses an online brain-computer interface system to perform a plurality of motor imagery tasks, wherein the motor imagery task comprises a left motor imagery task and a right motor imagery task, and the brain electrical signal data of the subject is collected during each motor imagery task, and the collected brain electrical signal data of the ID bit subject is taken as training set data, and the brain electrical signal data collected from the subject to be detected during the motor imagery task is taken as target domain data; Step two, the brain electrical signal data corresponding to the motor imagery time period is extracted from each set of brain electrical signal data of the training set and the target domain, and the extracted brain electrical signal data is preprocessed to obtain each set of preprocessed brain electrical signal data; Step three, the pre-alignment strategy is used to process each set of preprocessed brain electrical signal data in the training set and the target domain to obtain each set of processed brain electrical signal data; Step four, the covariance matrix of each set of brain electrical signal data is calculated according to each set of processed brain electrical signal data, and the cut space vector of each set of brain electrical signal data is obtained by projecting the covariance matrix of each set of brain electrical signal data to the cut space by using the cut space feature mapping method; Step five, id is initialized as 1; Step six, the cut space vector of the id bit subject in the training set and the cut space vector of each set of brain electrical signal data in the target domain are projected and mapped by using the domain self-adaptive manifold embedding distribution alignment feature mapping method to obtain the feature vector of the cut space vector of the id bit subject and the target domain after projection and mapping; Step seven, construct the classifier CLS by using the projection mapped eigenvectors of the id-th subject and the target domain id ; Step eight, whether id satisfies ID is judged; If yes, the ID constructed classifiers CLS1, CLS2, …, CLS are obtained ID ; the projection mapped feature vectors of the target domain are respectively passed through the classifiers CLS1, CLS2, …, CLS ID , and classification results Rst1, Rst2, …, Rst are respectively obtained ID , and step nine is executed again; If not, id is set as id+1, and step six is returned to be executed; Step nine, the ID classification results obtained in step eight are fused to obtain the final classification result of the target domain data.
2. The domain adaptation method for solving subject difference in motor imagery brain-computer interface according to claim 1, characterized in that, The specific process of step one is as follows: For any subject, the subject faces the computer screen, a fixed ten-character pattern appears on the computer screen at t=0, and a motor imagery prompt appears on the computer screen at t=2 seconds to prompt the subject to perform motor imagery; wherein the motor imagery prompt is an arrow prompt, and the arrow prompt comprises a left arrow prompt and a right arrow prompt; During t=2 to t=6, the arrow prompt on the computer screen remains unchanged, and the subject continuously performs the current prompted motor imagery; then rest for two seconds; at t=8 seconds, the second round of motor imagery experiment is started again, and the above process is repeated for the subject; and during each motor imagery task, the brain electrical signal data of the subject from the beginning of the task to the end of the task is collected; Similarly, the brain electrical signal data of each subject during the motor imagery task is collected, and the collected brain electrical signal data of the ID bit subject is taken as the training set, and the training set data is taken as the source domain data; The brain electrical signal data collected from the subject to be detected during the motor imagery task is taken as the target domain data.
3. The domain adaptation method for solving subject difference in motor imagery brain-computer interface according to claim 2, characterized in that, The specific process of step two is as follows: For each set of collected brain electrical signal data, the motor imagination period is extracted, i.e. the brain electrical signal data within 2.5-5.5 seconds is extracted, and the extracted brain electrical signal data is subjected to power frequency notch filtering and band-pass filtering processing.
4. The domain adaptation method for solving subject difference in motor imagery brain-computer interface according to claim 1, characterized in that, The specific process of the step three is: According to the covariance matrix of the ID bit subject in the training set and the preprocessed brain electrical signal data of the target domain, the mean of the covariance matrix in the Riemann space is calculated, and the data alignment processing is performed on each group of preprocessed brain electrical signal data according to the mean, to obtain the processed brain electrical signal data.
5. The domain adaptation method for solving subject difference in motor imagery brain-computer interface according to claim 1, wherein, The specific process of the step six is: Step six one, the within-class scatter matrix and the between-class scatter matrix of the ID bit subject in the training set are calculated, and the constraint of the projection matrix based on the within-class scatter matrix and the between-class scatter matrix is: where A is a projection matrix, A T is the transpose of A, tr(·) is the trace of a matrix, S w is the within-class scatter matrix of the electroencephalogram data of the ith subject, is the matrix composed of the cut space vectors corresponding to the data belonging to class c in the electroencephalogram data of the ith subject, represents a real number, C is the number of classes of the electroencephalogram data of the ith subject, is the number of groups of data of class c in the electroencephalogram data of the ith subject, D is the dimension of the cut space vector corresponding to each group of electroencephalogram data, is the center matrix of data of class c, is an identity matrix, S b is the between-class scatter matrix of the electroencephalogram data of the ith subject, is the mean of the cut space vectors corresponding to each group of electroencephalogram data in class c of the ith subject, is the cut space vector corresponding to the ith group of electroencephalogram data in class c of the ith subject, is the mean of the cut space vectors corresponding to all groups of electroencephalogram data of the ith subject, X i represents the cut space vector corresponding to the ith group of electroencephalogram data of the ith subject, n s represents the total number of groups of electroencephalogram data of the ith subject; Step six two, calculate the target domain electroencephalogram signal data scattering matrix V t The projection matrix is based on the constraint of the data scattering matrix: Wherein, B represents a projection matrix, X t A matrix composed of the cut space vectors corresponding to each group of processed electroencephalogram signal data in the target domain, n t H represents the total number of groups of electroencephalogram signal data in the target domain, t is the center matrix of the target domain, is the unit matrix; Step six three, the projection matrix A and the projection matrix B satisfy the constraint: Among them, ||·|| F This indicates the calculation of the F-norm; Step six four, the objective function is solved according to the constraints of step six one to step six three, to obtain the projection matrix A and B; According to the projection matrix A and B, the tangent space vectors of each group of brain electrical signal data of the target domain are projected and mapped, to obtain the projected and mapped vectors; then the projected and mapped vectors are input into the SVM classifier, to obtain the classification results of each group of brain electrical signal data of the target domain, and the obtained classification results are taken as the initial classification results; Step six five, the iteration number l is initialized as 1; Step six six, the within-class scatter matrix and the between-class scatter matrix of each group of brain electrical signal data of the target domain are calculated, and the constraint of the projection matrix based on the within-class scatter matrix and the between-class scatter matrix is: wherein T w is an intra-class scatter matrix of the brain electrical signal data of each group in the target domain, is a matrix composed of the tangent space vectors corresponding to the brain electrical signal data belonging to class c in the target domain, denotes the number of groups of the brain electrical signal data belonging to class c in the target domain, is a center matrix of the brain electrical signal data of class c, is an identity matrix, T b is an inter-class scatter matrix of the brain electrical signal data in the target domain, denotes the mean of the tangent space vectors corresponding to the brain electrical signal data belonging to class c in the target domain, denotes the tangent space vector corresponding to the i-th group of brain electrical signal data belonging to class c in the target domain, denotes the mean of the tangent space vectors corresponding to each group of brain electrical signal data in the target domain, X i is the tangent space vector corresponding to the i-th group of brain electrical signal data in the target domain; Step six seven, the objective function is solved according to the constraints of step six one to step six three and step six six, to obtain the projection matrix A and B; Step six eight, it is judged whether the iteration number reaches the set maximum iteration number L; If not, the tangent space vectors of each group of brain electrical signal data of the target domain are projected and mapped according to the projection matrix A and B, the projected and mapped vectors are input into the SVM classifier, to obtain the classification results of each group of brain electrical signal data of the target domain, l is set as l+1, and the obtained classification results are returned to step six six for execution; If yes, the tangent space vectors of each group of brain electrical signal data of the target domain and the ID bit subject in the training set are projected and mapped according to the projection matrix A and B, to obtain the projected and mapped results, i.e. the feature vectors of the tangent space vectors of the ID bit subject and the target domain after the projection and mapping.
6. The domain adaptation method for solving subject difference in motor imagery brain-computer interface according to claim 5, characterized in that, The objective function in the step six four is: Wherein, I is a unit matrix, ω, μ, α and λ are all coefficients; By constructing the Lagrange function, the derivative of the Lagrange function is set as 0, and the objective function is represented as: wherein let Φ is a diagonal matrix composed of eigenvalues of matrix Φ = diag(Φ1, Φ2, …, Φ k ), Φ1, Φ2, …, Φ k represent the 1st, 2nd, …, kth eigenvalues of eigenvalues of matrix sorted from large to small, W1, W2, …, W k are eigenvectors corresponding to eigenvalues Φ1, Φ2, …, Φ k respectively, matrix W = [W1, …, W k ], W = [A; B].
7. The domain adaptation method for solving subject difference in motor imagery brain-computer interface according to claim 6, characterized in that, The specific process of the step seven is: Step seven one, the iteration number l' is initialized as 1; Step seven two, the minimization principle of SRM is represented as: In the formula, represents a set of classifiers in the nuclear space, f is a classifier, represents a loss function on the matrix x composed of the projection mapping results (feature vectors after projection mapping) corresponding to each set of electroencephalogram data of the id subject in the training set and in the target domain, y represents the classification label of x, η represents the hyperparameter of R(f), and R(f) represents the square norm of f in the formula . The expansion form of f is as follows: In the formula, n s n represents the number of EEG signal data sets of the id-th subject in the training set. t The number of EEG signal data sets in the target domain is represented by β, where β represents the weight vector. This indicates the weight of each group of EEG signal data. K(x) is a real number. i (,x) represents x i Correlation with x, x i This represents the projection mapping result of the id-th subject in the training set and any set of EEG signal data in the target domain; f is continuously expanded as: where f(·) denotes the output of the classifier f, K denotes the kernel matrix, K(x i ,x j ) denotes the correlation of x i and x j , K(x i ,x j ) is the element of the i-th row and j-th column of the kernel matrix K, y i denotes the classification label of x i , represents the label vector, Y is the one-hot encoding matrix of y, A' denotes a diagonal matrix, if i belongs to the training set data, the i-th diagonal element of the diagonal matrix A i ' i = 1; otherwise ,A i ′ i =0; Step seven three, the maximum average difference between the projection mapping results of the brain electrical signal data of the ID bit subject in the training set and the projection mapping results of the brain electrical signal data of the target domain is calculated: wherein denotes the maximum mean discrepancy, M0represents the edge MMD matrix, K i is the i-th column vector of K, is the j + n s column vector of K, H K denotes the Hilbert kernel space; In the formula, denotes a set of projection mapping results of each group of electroencephalogram data of the ith subject in the training set, denotes a set of projection mapping results of each group of electroencephalogram data in the target domain, (M0) ij denotes an element in the ith row and the jth column in M0; Step seven four, calculating conditional metric constraints wherein, (x s ) (c) denotes a matrix composed of projection mapping results corresponding to data belonging to class c in the projection mapping results of the ith subject in the training set;(x t ) (c) denotes a matrix composed of projection mapping results corresponding to data belonging to class c in the projection mapping results of the target domain;(x) (c) denotes a matrix composed of projection mapping results corresponding to data belonging to class c in the projection mapping results corresponding to each set of EEG signal data of the ith subject in the training set and the target domain; denotes a matrix composed of projection mapping results corresponding to data belonging to class in the projection mapping results corresponding to each set of EEG signal data of the ith subject in the training set and the target domain; denotes a class other than class c; E[·] denotes expectation; is an intra-class MMD matrix; is an inter-class MMD matrix; wherein, denotes denotes an element in the i-th row and the j-th column of the matrix, denotes a data set belonging to the class c in the projection mapping result of the i-th subject in the training set, denotes the number of projection mapping results belonging to the class c in the projection mapping result of the i-th subject in the training set, denotes a data set belonging to the class c in the projection mapping result of the target domain, denotes the number of mapping results belonging to the class c in the projection mapping result of the target domain. In the formula, denotes the element in the ith row and jth column of the matrix, denotes the projection mapping result of the idth subject in the training set and the projection mapping result set belonging to the class c in the projection mapping result of the target domain, n (c) denotes the number of projection mapping results belonging to the class c in the projection mapping result of the idth subject in the training set and the projection mapping result of the target domain, denotes the number of projection mapping results not belonging to the class c in the projection mapping result of the idth subject in the training set and the projection mapping result of the target domain. Step seven five, the mean H of the within-class difference center matrix is calculated: wherein When The element in the corresponding position is 1 when the data in the corresponding position in the vector belongs to the category c, otherwise the element in the corresponding position is 0. The vector represents elements all of which are 1, n = n s + n t ; Step seven six, the matrix S is calculated: In the formula, S ij Let represent the element in the i-th row and j-th column of matrix S, and sim(·,·) represent the distance between two points. Representing point x i Belongs to x j The neighborhood set; laplace regularization is represented as: where L is the Laplacian matrix, L ij denotes the element in the ith row and jth column of the Laplacian matrix, L = D - S, D is a diagonal matrix, and Step seven, based on steps seven two to seven six, the final classifier is obtained: wherein ε, p and γ are coefficients; The target function is established according to the final classifier of the above formula: Derive the target function and let the derivative of the target function be 0 to obtain the weight vector β of the classifier * : β * = ((A' + εM0+ (1 - ε)M c + ρL+ γH)K+ ηI) -1 A'Y T Step seven eight, judge whether the iteration number l' reaches the set maximum iteration number L'; If not, let l'=l'+1, obtain the pseudo label of the target domain by using the classifier obtained in step seven seven, and return to execute step seven two; If so, the classifier obtained at the last iteration is taken as the classifier CLS obtained from the ith subject of the training set id .
8. The domain adaptation method for solving subject difference in motor imagery brain-computer interface according to claim 7, characterized in that, The specific process of the step nine is: Step nine one, initialize the subject id=1 in the source domain; Step nine, the tangent space vector of the target domain of all brain electrical signal data and the tangent space vector of the first id subject in the source domain of all brain electrical signal data are subjected to a linear classifier, and an edge distance is calculated according to the output of the linear classifier The tangent space vector of the target domain of all brain electrical signal data and the tangent space vector of the first id subject in the source domain of all brain electrical signal data are subjected to a linear classifier, and an edge distance is calculated according to the output of the linear classifier Step nine three, the target domain and the source domain of the first id subjects of the first c type of brain electrical signal data cut space vector through the linear classifier, according to the output of the linear classifier to calculate the first c type of brain electrical signal data cut space vector condition distance the projection mapping result of the tangent space vector of the cth type of brain electrical signal data of the idth subject in the target domain and the source domain is subjected to a linear classifier, and a conditional distance of the projection mapping result corresponding to the tangent space vector of the cth type of brain electrical signal data is calculated according to an output of the linear classifier Step nine four, edge distance And conditional distance Integrating to get the integration distance corresponding to the idth subject in the source domain; The specific process of calculating the integrated distance is: wherein d M and d c are intermediate variables; Then the integration distance d corresponding to the id subject in the source domain id,Mc is: Step nine five, judge whether id=ID is satisfied; If id=ID is satisfied, execute step nine six; If id=ID is not satisfied, let id=id+1, return to execute step nine two; Step 96: Based on the integration distance d corresponding to each subject in the source domain id,Mc We calculate the weight factor of the classifier for each subject, and denote the weight factor of the classifier for the subject at position id in the source domain as wt. id ; Then execute step nine seven; Step nine seven, for any one set of electroencephalogram signal data in the target domain, the projection mapping result corresponding to the set of electroencephalogram signal data is taken as the input of the classifier corresponding to each subject in the source domain respectively; Then, the output of the classifier corresponding to the idth subject in the source domain is multiplied by the weight factor wt id to obtain a multiplication result, wherein the multiplication result includes the probability that the set of electroencephalogram signal data belongs to each category, and id = 1, 2, …, ID. And the elements in the multiplication result of each classifier are added correspondingly, each element contained in the addition result is the final probability that the set of electroencephalogram signal data belongs to each category, and the maximum value is selected from the final probability of each category, and the category corresponding to the maximum value is the category to which the set of electroencephalogram signal data belongs.
9. The domain adaptation method for solving subject difference in motor imagery brain-computer interface according to claim 8, characterized in that, The calculation method of the edge distance is: For edge distance The cut space vector of the target domain of all brain electrical signal data and the cut space vector of the idth subject in the source domain of all brain electrical signal data are subjected to the linear classifier, and the linear classifier outputs whether the brain electrical signal data corresponding to each cut space vector belongs to the source domain or the target domain. Wherein, h represents a linear classifier, and ε(h) represents the incorrect rate of the output result of the linear classifier.
10. The domain adaptation method for solving subject difference in motor imagery brain-computer interface according to claim 9, characterized in that, The specific process of the step nine six is: wherein wt id represents the weight factor of the classifier corresponding to the id subject in the source domain, and e is the base of the natural logarithm.
Citation Information
Patent Citations
Brain-computer interface transfer learning method based on manifold embedding distribution alignment
CN111723661A
Motor imagery electroencephalogram signal adaptive classification method based on covariance alignment
CN115640539A