A principal transport migration component identification method based on riemannian manifold tangent space alignment
By employing a master transport migration component identification method aligned with the optimal Riemannian manifold tangent space, the problem of feature extraction from EEG signals with large individual variability and nonstationary/nonlinear characteristics is solved, improving the accuracy and efficiency of EEG signal recognition and enhancing the model's adaptability.
Patent Information
- Application Number
- CN202211482033.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-11-24
AI Technical Summary
Existing methods for recognizing motor imagery EEG signals struggle to effectively extract features when faced with highly individualized and non-stationary/nonlinear EEG signals, resulting in low model generalization performance and time-consuming and inconvenient calibration processes.
The method of identifying master transport migration components with optimal Riemannian manifold tangent space alignment is adopted. Through covariance centroid alignment, tangent space feature extraction, optimal transport domain adaptation and migration component analysis, individual differences are reduced and domain-invariant shared features are extracted.
It improves the accuracy and efficiency of EEG signal recognition, reduces computation time, enhances the model's generalization ability, and adapts to the extraction of EEG signal features from different individuals.
Smart Images

Figure CN115935140B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a motor imagery electroencephalogram signal recognition method, in particular to a main transport transfer component recognition method (OTCA) based on optimal Riemannian manifold tangent space alignment. TECHNICAL BACKGROUND
[0002] For decades, motor imagery (MI) based brain-computer interfaces (BCI) have attracted considerable attention because it can help subjects to directly manipulate electronic devices using brain activity triggered by imagined movements without involving traditional muscle-dependent pathways. Electroencephalography (EEG) is the most commonly used paradigm in BCI because it is safe and easy to obtain, and has a wide range of application prospects, such as rehabilitation, entertainment and drowsiness detection. There are several reasons for the success of EEG-based BCI, such as low price, non-invasiveness and high temporal resolution compared with other neurophysiological models. However, there are still many challenges in the application of brain-computer interfaces, which mainly depend on the characteristics of electroencephalogram signals. On the one hand, although there are many available methods, such as general spatial patterns, fast Fourier transform, particle swarm optimization and deep network theory, the non-stationary, nonlinear and non-Gaussian electroencephalogram signals of the brain still cause difficulties in distinguishing feature extraction and signal analysis. On the other hand, the characteristics of electroencephalogram signals are greatly affected by various individual differences such as age and psychology, and there may be differences in electroencephalogram signals of different individuals in the same event, making it difficult to extract effective information related to a specific task. Therefore, when calibrating a BCI system, it is usually necessary to conduct many training experiments for specific subjects, and this calibration process is both time-consuming and inconvenient for users. The above difficulties need to develop more powerful and effective methods in further research. SUMMARY
[0003] The present application aims to solve the problem that part of the transfer learning model does not consider the probability coupling of source data and target data in the space of electroencephalogram signals, and the shared features obtained cannot solve the problem of large individual differences in EEG signals and low model generalization performance, respectively introduces the optimal transport domain adaptation algorithm and the transfer component analysis, proposes a main transport transfer component recognition method based on Riemannian manifold tangent space alignment, which is more universal, and then re-extracts the domain-invariant shared features of the motor imagery dataset.
[0004] According to the summary, the specific steps include the following steps:
[0005] Step 1, covariance centroid alignment, extracting motor imagery electroencephalogram signals, and aligning the covariance matrices of the source domain and the target domain of the motor imagery electroencephalogram signals, so that the marginal distribution of the covariance matrices of the source domain and the target domain is closer.
[0006] Step 2, cut space feature extraction, mapping the aligned covariance matrix to a cut space vector.
[0007] Step 3, optimal transport domain adaptation, learning the optimal transport mapping in the extracted cut space vector, and obtaining the optimal transport source domain by optimal transport of the source domain, the dimension after mapping is unchanged.
[0008] Step 4, transfer component analysis, learning similarity transformation in reproducing and Hilbert space (RKHS), and then transforming the optimal transport source domain and target domain into a similar subspace, so that the distributions of the two are more similar.
[0009] The present application considers the calculation accuracy and the calculation time-consuming, and selects a suitable domain adaptation method for the data set to obtain the optimal transfer component.
[0010] As preferred, in the field adaptation link, optimal transport domain adaptation can be used to obtain the optimal source domain to reduce the dimension of the feature, and reduce the difference between the domains.
[0011] Step 3-1: Assuming that there are two different joint probability distributions P s (x t ,y) and P s (x s ,y) in the measurable space Ω t and Ω t , μ s and μ t are the marginal probability distributions thereon. A series of probability couplings P(Ω s ×Ω t ) of the marginal distributions μ s and μ t can be obtained. Let there be a measurable mapping T:Ω s →Ω t , such that T#μ s =μ t , then T is the optimal transport mapping found. At this time, T can be called the transport mapping or propulsion from μ s to μ t . The Kantorovitch problem to be solved is to find a coupling matrix γ∈Π between Ω s and Ω t .
[0012] It is difficult to search T in all possible transformation spaces, and some restrictions need to be imposed. Here, the selection of T needs to satisfy the minimization of the following loss function:
[0013]
[0014] Wherein the loss function is: C(T) is a distance function over the metric space Ω. C(T) can be understood as the energy required to move the probability density from x to T(x). While this sounds simple, searching for T across all measurable spaces is by no means easy. In this work, the Monge method for the optimal transport problem involves moving from μ... s to μ t We seek an optimal mapping T0 from all possible transmissions as a solution to the minimization problem.
[0015]
[0016] When μ s and μ t When obtained only from discrete samples, the corresponding empirical distribution can be written as:
[0017]
[0018] in Is The Dirac function for position, and p i t It is the probability density of the i-th sample. The Kantorovich formula for the optimal transportation problem is directly applied to the discrete case. Let Π represent the probability coupling set between two empirical distributions, defined as:
[0019]
[0020] Among them 1 d Let represent a d-dimensional vector of all 1s. Then the optimal transmission can be written as:
[0021]
[0022] In this case, γ can be understood as the marginal distribution μ s and μ t The joint probability measure. γ0 is also a transfer matrix. Then we can obtain μ. s and μ t p-order Wasserstein distance between them:
[0023]
[0024] Step 3-2: Assume that the source domain has N s Tag Examples in It is the i-th characteristic matrix of the source domain, y s,i∈{1,...,l} are the corresponding labels, where h,t and l represent the number of channels, time-domain samples, and categories. Similarly, assume the target domain has N... t Example in For domain adaptation, once the optimal transmission scheme γ0 is obtained, the source domain samples can be transmitted to the target domain. This paper defines γ0 as a transport mapping T. For each... This mapping can be conveniently represented as the following loss-based centroid mapping:
[0025]
[0026] in This refers to the new transmission source domain after optimal transmission, i.e., the optimal transmission source domain, which greatly increases the distribution similarity with the target domain.
[0027] Preferably, step 4 includes:
[0028] Step 4: The key assumption proposed in the above domain adaptation setting is that the source domain and the target domain have different edge distributions, but have the same conditional distribution in the new subspace, i.e., P(X s )≠P(X t But P(Y) s |φ(X s ))=P(Y t |φ(X t This is based on minimizing the maximum mean difference (MMD) distance MMD(X). s ′,X t This is achieved through ′). Where X is... s ′={x s ′ ,i}={φ(x s,i )}, X t ′={x t ′ ,i}={φ(x t,i )}, expecting to have P(X) s ′)≈P(X t Then the distance between the two domains can be expressed as:
[0029]
[0030] Therefore, the desired nonlinear mapping φ can be found by minimizing the above equation. However, φ is typically highly nonlinear, and directly optimizing this function may lead to poor local minima. Therefore, Pan proposed transforming this problem into a kernel problem. Using kernel tricks, such as k(x... i ,x j )=φ(x i )′φ(xj ), in formula (8), the distance between the empirical mean of two fields can be written as:
[0031] Dist(X s ′,X t ′)=tr(KL)(9)
[0032] Wherein:
[0033]
[0034] It is a kernel matrix. Wherein K s,s ,K t,t ,K t,s Gram matrix of source data, target data and cross data in projection space respectively, N s And N t The number of samples of source domain and target domain. L is the multiplier of K, L can be expressed as:
[0035]
[0036] By decomposing the matrix K and defining the result kernel, there is a matrix Map the feature to the m-dimensional space. Therefore, the objective function of TCA can be expressed as:
[0037]
[0038] Wherein Is the unit matrix, Is the center matrix. Is the all-1 column vector. The solution of W is the dominant eigenvector (I+mu KLK) -1 KHK.
[0039] Compared with the prior art, the present application has the following beneficial effects: the BCI system is disturbed by external interference in the signal acquisition process, and the data will drift. Data drift makes the EEG signal have huge individual difference, so that the movement intention recognition becomes a challenging task. The main transmission transfer component recognition method based on Riemannian manifold tangent space alignment proposed in the present application can find the best transmission mapping matrix and transmit the source domain features to the target domain, so as to reduce the data drift problem of the probability distribution between the domains. BRIEF DESCRIPTION OF DRAWINGS
[0040] Figure 1 The method flowchart of the embodiment of the present application.
[0041] Figure 2 The classification stage flowchart in the embodiment of the present application.
[0042] Figure 3Time consuming statistics for the six classification methods on the data sets MI1 and MI2.
[0043] Figure 4 Distributed data visualization based on the t-SNE method. DETAILED DESCRIPTION
[0044] The application will be further described below in connection with specific embodiments. The following description is only exemplary and explanatory in nature and not intended to limit the application in any form. The embodiments of the application provide a principal transport migration component identification method based on Riemannian manifold tangent space alignment, as shown in Figure 1 The implementation steps are as follows:
[0045] Step 1, covariance centroid alignment as a preprocessing step can reduce the difference between the marginal probability distribution of the source domain and the target domain. Let be the covariance matrix of the ith source domain, and the reference matrix M ref -1 / 2 , where M can be Riemannian mean, Euclidean mean, etc. Therefore, the aligned covariance matrix can be expressed as:
[0046] P′ s,i ref P s,i M ref ,i=1,...,N s (1)
[0047] Step 2, after obtaining the centroid-aligned covariance matrix, each covariance matrix is mapped to a tangent space vector.
[0048]
[0049] For different subjects, the reference matrix M ref of the specific subject is used. Next, the new feature matrix and m=h(h+1) / 2 is the dimension of the tangent space.
[0050] Step 3, constructing the optimal transport source domain, the optimal transport (OT) algorithm uses the optimal transport theory to calculate the optimal coupling between two probability distributions in an economical and efficient way. Based on the theory of optimal transport, some scholars have proposed the optimal transport domain adaptation (OTDA) algorithm, which mainly includes learning a mapping matrix using the probability density functions of the source domain and the target domain, while limiting the same class of samples in the source domain to remain close during transport and in the target domain. This can reduce the distribution of the source domain and the target domain, thereby learning domain-invariant features to achieve better recognition performance. The specific steps are as follows:
[0051] Step 3-1: Assume x s and x t represent the subject covariance matrix of the source domain and the target domain mapped into the tangent space vector in the Riemannian tangent space, where the loss function is the distance function on the metric space Ω s × Ω t , then there exists a measurable mapping T: Ω s → Ω t such that T#μ s = μ t , T is the optimal transport mapping to be found, and the problem to be solved is to find a general coupling matrix γ ∈ Π between Ω s and Ω t , the selection of T needs to satisfy the minimization of the following loss function, and Π represents the probability coupling set between two empirical distributions, then the form of optimal transport can be written as:
[0052]
[0053] Step 3-2: For domain adaptation, once the optimal transport matrix γ0 is obtained, the source domain samples can be transported to the target domain. Define γ0 as a transport mapping T. For each x i s This mapping can be conveniently represented as a loss-based barycentric mapping:
[0054]
[0055] Step 4, find the required nonlinear mapping φ by minimizing equation (5)
[0056]
[0057] After obtaining the nonlinear mapping φ, calculate the mapped source domain and target domain features X s ′ and X tFinally, the sLDA linear discriminator is used for prediction.
[0058] Two MI datasets are used to evaluate the proposed method, which are MI dataset 1 (MI1) and MI dataset 2 (MI2) respectively. For the two MI datasets, the experimental paradigm is as follows: a subject sits in front of a computer screen. At the beginning of the experiment, a fixed cross appears on the black screen, prompting the subject to be ready. After a while, an arrow pointing in a certain direction appears as a visual cue for a few seconds, during which the subject performs a specific MI task. Then, the visual cue disappears, and the next experiment starts after a short rest. The EEG signal is recorded during the experiment and used to classify the MI performed by the user. For the first dataset, the EEG signals of 7 healthy subjects are recorded at a frequency of 100 Hz, with 59 channels, and each subject performs 100 left-hand MIs and 100 right-hand MIs. For the second dataset, the EEG signals of 9 healthy subjects are recorded at a frequency of 250 Hz, with 22 channels, and each subject performs 72 left-hand MIs and 72 right-hand MIs.
[0059] For the two MI datasets, a 50-order finite impulse response (FIR) band-pass filter is used to remove muscle artifacts and DC current drift, thereby obtaining cleaner MI signals. Next, the EEG signal between [0.5, 3.5] seconds after the visual cue appears is extracted as the experiment.
[0060] The proposed method OTCA is compared with five state-of-the-art baseline algorithms to verify the effectiveness of the optimal transport algorithm. Figure 2 is a system flowchart of the proposed method, Figure 2 is a flowchart of the classification stage.
[0061] 1) CSP-LDA: used as a baseline method to evaluate the accuracy of MI classification.
[0062] 2) CSP-OT-LDA: adds optimal transport OT to CSP-LDA to verify the effectiveness of the OT algorithm.
[0063] 3) CA-LDA: uses centroid alignment to reduce the probability distribution of the source domain and the target domain before using the LDA classifier for classification.
[0064] 4) MEKT: manifold embedding knowledge transfer method is used for comparison.
[0065] 5) OTDA: a baseline method based on optimal transport domain adaptation.
[0066] 6) OTCA: the method proposed in the present application.
[0067] Table 1 summarizes the classification accuracy of all these methods on the two datasets. The results show that the proposed OTCA method outperforms other methods on BCI Competition IV Dataset 1 and Dataset 2a, with average accuracy of 72.57% and 69.19%, respectively. For Dataset 1, the best average accuracy obtained by OTCA is 1.57% higher than the second best MEKT and 5.08% higher than the third best OTDA, and the classification accuracy of each subject is higher than other methods. For Dataset 2, the best average accuracy obtained by OTCA is 1.46% higher than the second best MEKT and 2.18% higher than the third best OTDA, and only subjects 2, 4, 5, and 7 are lower than other methods.
[0068] Table 1. Classification accuracy of each method on MI1
[0069]
[0070] Table II. Classification accuracy of each method on MI2
[0071]
[0072]
[0073] To emphasize the computational cost of different algorithms, Figure 3 The time-consuming of 6 transfer learning methods is shown. To eliminate interference, the algorithm will be repeated 10 times and the average value will be calculated to obtain the final calculation time-consuming. For CA-LDA, since the input of CA-LDA classifier is SPD matrix, and the number of channels of dataset dataset 1 is 22, and the number of channels of dataset dataset 2 is 64. According to Figure 3 It can be found that the calculation time-consuming of the two is nearly 6 times different, that is, there is a huge correlation between the calculation amount and the number of channels of EEG signal. By comparing CSP-LDA and CSP-OT-LDA algorithms, it is found that after adding optimal transport, the classification accuracy of the dataset is increased by 11.42% (MI1) and 7.19% (MI2). At the same time, it also reduces the calculation time-consuming, which shows the effectiveness of OT algorithm. Among the last four algorithms, the calculation amount of OTCA is the smallest, while OTDA and CA-LDA are basically the same, and the calculation amount of MEKT is the largest. Considering the classification accuracy, the classification progress of OTCA is the highest, followed by MEKT, which highlights the superior performance of OTCA algorithm, that is, OTCA algorithm realizes the best trade-off between classification accuracy and calculation cost.
[0074] In Figure 4In the table, the source domain and target domain of the first row are selected from the 1st and 3rd subjects in the MI1 dataset, respectively. The source domain and target domain of the third row are selected from the 1st and 5th subjects in the MI1 dataset, respectively. The source domain and target domain of the fifth row are selected from the 1st and 7th subjects in the MI1 dataset, respectively. Figure 4 (a), (c) and (e) represent the original EEG feature distribution. Figure 4 (b), (d) and (f) represent the feature distribution after OTCA. The solid circle represents the source domain belonging to class 1, the pentagram represents the source domain belonging to class 2, the triangle represents the target domain belonging to class 1, and the rectangle represents the target domain belonging to class 2.
[0075] Optimal transport (OT) aims to find an optimal transport mapping T that can reduce the distribution difference between the source domain and the target domain after mapping. In order to further verify its effectiveness, the present application visualizes the feature distribution in different states. Here, the t-SNE (random neighborhood embedding) technique is used to reduce the dimension of the features. The t-SNE technique can reduce high-dimensional data to 2-3 dimensions, and can visualize the reduced features. The meanings of the x-axis and y-axis are to reduce the data to one-dimensional and two-dimensional feature values, respectively, using the t-SNE dimension reduction technique. From Figure 4 As can be seen in (a), the data distribution of the source domain (solid circle and pentagram) and the target domain (triangle and rectangle) is more concentrated in a certain position, which means that there is a difference in the marginal probability distribution between the target domain and the source domain. In Figure 4 In b, after OTCA, the data of the source domain is more sparse, and the similarity with the target domain is increased. This shows the effectiveness of OTCA in making the two data distributions more similar.
Claims
1. A principal transport migration component identification method based on Riemannian manifold tangent space alignment, characterized in that, Comprise the following steps: Step 1, covariance centroid alignment, extract motor imagination electroencephalogram, and align the covariance matrix of the source domain and the target domain of the motor imagination electroencephalogram; Step 2, cut space feature extraction, map the aligned covariance matrix to a cut space vector; Step 3, optimal transport domain adaptation, learn the optimal transport mapping in the extracted cut space vector, and obtain the optimal transport source domain by optimal transport of the source domain; Step 4, transfer component analysis, learn similarity transformation in the reproducing and Hilbert space, and then transform the optimal transport source domain and the target domain into a similar subspace, so that the distributions of the two are more similar.
2. The method of claim 1, wherein, The dimension of the source domain after mapping is unchanged.
3. The method of claim 1, wherein, The learning method of the optimal transport mapping is: Assume x s and x t denote the tangent space vectors in the tangent space of the Riemannian quotient space, where the loss function c: is the distance function on the metric space Ω s × Ω t , then there exists a measurable mapping T: Ω s → Ω t such that T#μ s = μ t , T is the optimal transport mapping to be found, and the problem is to find a coupling matrix γ ∈ Π between Ω s and Ω t , the selection of T needs to satisfy the minimization of the following loss function, and Π represents the set of probability couplings between two empirical distributions, then the optimal transport can be written as: In this case, γ can be understood as the joint probability measure of the marginal distributions μ s and μ t , γ0is an optimal transport matrix, then the p-order Wasserstein distance between μ s and μ t can be obtained:
4. The method of claim 3, wherein, The method for obtaining the optimal transmission source domain through optimal transmission of the source domain is as follows: for domain adaptation, obtaining an optimal transmission matrix γ0, then the sample of the source domain can be transmitted to the target domain, defining γ0 as a transmission mapping T, for each This mapping can be conveniently expressed as the following loss-based barycenter mapping: wherein is the new transmission source domain, i.e., the optimal transmission source domain, wherein N s and N t represent the number of samples of the source domain and the target domain, respectively, represents the tangent space of a d-dimensional Riemannian manifold, and x is a feature in .
Citation Information
Patent Citations
Semi-supervised optimal transmission method for heterogeneous field adaptation
CN111062406A
Electroencephalogram signal domain adaptation method based on Riemannian manifold coordinate alignment
CN112580436A
Cited By
A stroke cross-subject electroencephalogram decoding method with pathological awareness and timing calibration
CN122527771A