A method for selecting the source domain of electroencephalogram across subjects based on domain similarity

Through the improved MMD and Copula function models, the problem of data migration selection in cross-participants' emotions recognition is solved, which improves accuracy and efficiency and reduces calculation time.

CN115034296BActive Publication Date: 2025-07-08HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210620691.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-07-08
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

In the emotional recognition across subjects, it is difficult for the prior art to effectively use the similarity between the source domain and the target domain to select suitable transfer learning data, resulting in a decrease in the accuracy of emotion recognition and an increase in computer computing time.

Method used

The improved maximum mean difference (MMD) method is used to calculate the inter-class spacing and Copula function modeling, and the source domain data with high similarity to the target domain are selected, and transfer learning is performed through Kendall rank correlation coefficient and adaptive distribution alignment method to eliminate negative transfer data.

Benefits of technology

It improves the accuracy of emotion classification, reduces computer computing time, and enhances the efficiency of emotional recognition across subjects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115034296B_ABST
    Figure CN115034296B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for selecting a source domain across subjects based on domain similarity of electroencephalogram. Aiming at the problems of the prior art, the technical solution proposed is as follows: First, the inter-class distance MMD(xi) within each subject is obtained according to the improved MMD formula and the confidence level is given; the Kendall rank correlation coefficient of the selected Copula function is calculated and the confidence level of the former is superimposed, and a threshold is set to select approximately 1 / 3 of the source domains as the migration objects for transfer learning; then, the distribution balance is adaptively adjusted to balance the conditional distribution and the marginal distribution; finally, the soft labels of the target domain are updated, and finally the classifier is returned for three-class classification to output the classification accuracy. The present invention uses source domain selection to improve the efficiency of distribution alignment based on manifold embedding, and finally uses the source domain suitable for transfer learning for learning, which improves the accuracy and greatly reduces the operation time compared with the traditional method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for selecting the source domain of emotional electroencephalogram (EEG) signals, specifically a method for selecting a source domain suitable for transfer learning by comparing the similarity between the source domain and the target domain. Background Art

[0002] Human emotions are psychological phenomena such as the needs and desires of human beings as the main body. In recent years, researchers have been classifying human emotions. In modern times, the theme of emotion recognition has gradually entered the field of vision of human beings and plays an important role in different fields. For example, in the medical industry, rapid diagnosis can be achieved by monitoring the emotions of patients, and emotion is used for image retrieval, etc.

[0003] In the field of cross-subject pattern recognition, transfer learning comes into view due to its advantages of using previously obtained knowledge to improve the efficiency and accuracy of learning in a certain field and not requiring a large amount of labeled data. Its method is to obtain better results by transferring data annotation, transferring models, adaptive learning, and transferring knowledge in similar fields using a small number of data labels. For different individuals, it is difficult to use a single-subject model to recognize the emotions of each individual. By measuring the domain similarity between the source domain and the target domain, strongly associated source domains with stronger similarity are selected, and weakly associated or even domains with poor similarity are excluded, which greatly improves the accuracy of cross-subject emotion recognition. Summary of the Invention

[0004] In the problem of EEG cross-subject emotion recognition, the present invention mainly proposes to use the Copula function to model the non-linear correlation relationship between EEG signals among cross-subjects, and calculate the weight of the within-class distance by using the calculation method of the superimposed improved maximum mean discrepancy (MMD), and screen the source domain data so that the selected data can be better transferred.

[0005] The object of the present invention can be achieved by the following technical solutions:

[0006] The present invention selects EEG data based on domain similarity and performs unsupervised classification on the selected data, specifically including the following steps:

[0007] Step 1. First, calculate the between-class distance MMD(xi) within each source domain subject according to the improved MMD formula and give the confidence level;

[0008] Step 2. To determine the type of Copula function used and facilitate the comparison of the distributions of the source domain and the target domain, first divide the selected source domain data. In each different emotion category of each subject's each experiment, each category of emotion is divided and a joint distribution is constructed with the overall mean of the target domain. Select and analyze the kernel distribution estimates of the two to select the Frank-Copula function in the Copula function and perform parameter estimation, and obtain the parameter estimation value of 18.1323;

[0009] Step 3. Calculate the Kendall rank correlation coefficient of the Frank-Copula function and superimpose the confidence level in Step 1. Set the threshold to 1.5 to select some source domains as the transfer objects for transfer learning;

[0010] Step 4. Adaptively adjust the distribution balance to balance the conditional distribution and the marginal distribution, then update the target domain soft labels, and finally return the classification accuracy of the classifier output.

[0011] In Step 1, the MMD distance (Maximum mean discrepancy) is a measure of the distance between two distributions in the reproducing kernel Hilbert space. The formula means to find the mean distance of two sets of data in the kernel space. In the selection of the source domain for EEG emotion recognition, this paper experimented with a large amount of data and improved the formula based on MMD to calculate the mean difference of the inter-class distances within the source domain and sum them up:

[0012] Its mathematical expression:

[0013]

[0014] C is the category within the source domain, MMD(xi) is the intra-class spacing of a certain subject, D s Corresponds to the source domain, n s Corresponds to the n samples within a single source domain, and f is the mapping function corresponding to the sample; because it is infinite-dimensional and cannot be directly solved in the original space, the formula is squared, simplified to obtain the inner product, and solved using the kernel function. It is understood that the sample distributions within a single source domain are respectively mapped to the corresponding points in the reproducing Hilbert space, and the distance between these two distributions is represented by the inner product of two points, and finally the sum of the distances is obtained.

[0015] Based on traditional transfer learning, first use the mean of the inter-class means within the source domain to screen different source domains, then perform transfer learning and the classification recognition accuracy, and use the Pearson function to perform a correlation test on the results to obtain a correlation coefficient of 0.61, which is significantly correlated. The greater the intra-class distance, the relatively higher the accuracy.

[0016] In Step 2, the source domain is segmented, analyzed, and a Copula function is selected for parameter estimation. The specific steps are as follows:

[0017] 2-1. In different emotion categories of each subject's experiment each time, divide an experiment into 15 segments according to the label in time periods. To facilitate fitting as the marginal distribution, calculate the mean of each segment. When processing the target domain, calculate the mean of the entire target domain segment for similarity measurement.

[0018] 2-2. Estimate the approximate overall distribution type according to the kernel distribution of the empirical distribution function. It can be obtained that the tails of the EEG emotion distribution images of different subjects are asymmetric. Then, the Archimedean-Copula is selected to describe the distribution of cross-subject EEG emotion signals.

[0019] 2-3. Measure according to the Euclidean distance between the empirical Copula function and the Copula function. The smaller the distance, the better the fitting degree of the Copula function.

[0020] The distance of the Frank-Copula in the Archimedean-Copula function is the smallest at 0.0467. Therefore, it is selected for modeling.

[0021] In step 3 described above, the Kendall correlation coefficient τ is a statistic for categorical variables and is an index used to reflect the correlation of categorical variables. It is applicable to the case where both categorical variables are ordinal categories, especially for the classification of EEG signals.

[0022] Its mathematical expression:

[0023]

[0024] where (x i -x j ), i, j = 1, 2,..., d are the observed data, sign is the sign function. The larger the τ value, the more significant the correlation between variables.

[0025] In step 4 described above, to avoid feature distortion and quantitatively evaluate the importance of marginal distribution and conditional distribution, first, to deal with the elimination of degenerate features, the original data learns the manifold feature functional g(.) in the Grassmann manifold space and introduces the geodesic flow kernel (GFK) to promote its domain adaptation, and quantitatively evaluates the importance of marginal distribution and conditional distribution in domain adaptation through dynamic distribution alignment.

[0026] Its mathematical expression:

[0027] μ ∈ [0, 1] is the adaptive factor, D f is the marginal distribution, is the conditional distribution, C is the category, and the adaptive μ value is balanced by calculating the A-distance to balance the importance of the two distributions. Its mathematical expression:

[0028]

[0029] d M is the marginal A-distance, d cFor the A distance of a certain type, after summarizing the above manifold feature learning and dynamic distribution alignment through the principle of structural risk minimization (SRM), the following loss function expression is obtained:

[0030] The beneficial effects of the present invention are as follows:

[0031] Regarding the problem that the accuracy of cross-disciplinary emotion recognition in EEG signal transfer learning inevitably decreases due to the negative transfer of some data in the source domain, a new method is proposed to dynamically select data suitable for transfer learning and eliminate the source domain data that may cause negative transfer; using emotional EEG signals as the feature extraction object, improving the method of manifold embedding distribution alignment, and selecting domain similarity based on the Copula function and the improved MMD distance method. The data of each class in each source domain is segmented, and the source domain closer to the target domain is found through non-linear similarity analysis and the inter-class distance within a single domain is superimposed, largely screening and eliminating the data that may cause negative transfer. Theoretically, this method improves the accuracy of emotion classification and greatly reduces the computer operation time. Description of the Drawings

[0032] Figure 1 It is an image related to the sum of the intra-class distances in the source domain for some classification results;

[0033] Figure 2 It is a comparison of frequency histograms for some cross-subject cases;

[0034] Figure 3 It is a comparison of frequency histograms for different time periods of the same subject;

[0035] Figure 4 It is a flow chart based on domain similarity selection. Detailed Embodiment

[0036] The present invention will be further described below in conjunction with specific embodiments. The following description is only for demonstration and explanation, and does not impose any formal restrictions on the present invention.

[0037] Embodiment 1

[0038] The dataset used in this paper is the SEED dataset provided by the BCMI Laboratory led by Professor Lv Baoliang of Shanghai Jiao Tong University. In the experiment, 15 Chinese movie clips were selected as stimuli for positive, neutral, and negative emotions. There were 15 trials in each experiment. There was a 5-second prompt before each clip, 45 seconds for self-assessment after the clip, and 15 seconds for rest.

[0039] A method for selecting cross-subject source domains of EEG based on domain similarity includes the following steps:

[0040] Step 1. Using the SEED dataset, the collected EEG signals are downsampled to 200 Hz. To further filter out noise and remove artifacts, the signals are passed through a band-pass filter with a frequency range of 0 - 75 Hz, and differential entropy features are extracted. According to the improved MMD formula, the between-class distance MMD(xi) within each subject is calculated and a confidence level is given, with the confidence interval being [0, 1].

[0041] Step 2. In different emotion categories of each subject's each experiment, one experiment is segmented into 15 segments according to the labels at different time intervals, and the mean value of each segment is calculated. When processing the target domain, the mean value of the entire target domain segment is calculated for similarity measurement. The Archimedean - Copula is selected to describe the distribution of cross - subject EEG emotion signals, which includes the following three models: Gumbel - Copula, Clayton - Copula, and Frank - Copula. The parameter values of the three functions are solved, and the leave - one - out method is used to calculate the parameter values, and the average of all parameters is taken:

[0042]

[0043] According to the Euclidean distance d between the empirical Copula function and the Copula function 2 for measurement, the smaller the distance, the better the fitting degree of the Copula function:

[0044]

[0045] Step 3. Calculate the Kendall rank correlation coefficient of the Frank - Copula function and superimpose the confidence level in Step 1. Set the threshold to 1.5 to select some source domains as the transfer objects for transfer learning;

[0046] Step 4. Input specific parameters, map the source domain and the target domain to the manifold space for distribution alignment, output the classifier, update the soft labels of the target domain, and finally return the classification accuracy of the classifier output.

[0047] In the above - mentioned Step 1, as Figure 1 shown, based on traditional transfer learning, first, the between - class distances within the source domain are used to screen different source domains, and then transfer learning is carried out and the classification recognition accuracy is obtained. The Pearson function is used to conduct a correlation test on the results, and a correlation coefficient of 0.61 is obtained, which belongs to a significant correlation. The greater the between - class distance, the relatively higher the accuracy.

[0048] In the above - mentioned Step 2, as Figure 2 、 3As shown, for different subjects, the difference in the frequency histogram is more obvious to the naked eye compared with different data of the same subject. Through the skewness, kurtosis and normality tests of different domains, it is obviously not normally distributed, and it is necessary to further determine its distribution type and select a suitable Copula function.

[0049] Parameter estimation is carried out for the Copula function, and the unknown parameters in the alternative Copula functions are estimated. Commonly used parameter estimation methods include the maximum likelihood estimation method (ML estimation), the iterative estimation method (IFM estimation) and the semi-parametric estimation method (CML estimation). Among them, the semi-parametric estimation (CML estimation) is further divided into the standard maximum likelihood estimation based on the empirical distribution function and the maximum likelihood estimation based on the non-parametric kernel density. In this paper, kernel distribution estimation is adopted. It can be seen from the results that the parameter estimation values of the Copula function are all within the parameter value range. And according to the Euclidean distance between the empirical Copula function and the Copula function, it can be seen that the fitting effect of Frank-Copula is the best.

[0050] In step 3 described above, the Kendall correlation coefficient is selected as the correlation estimation value and the inter-class distance within the source domain is superimposed for source domain screening. The specific steps are as follows:

[0051] 3-1. Substitute the Copula parameter estimation value into the Kendall correlation coefficient equation for solution;

[0052] 3-2. Add the obtained result to the result of the inter-class distance within the source domain in step 1, and perform a reverse order sorting on the combination;

[0053] 3-3. Select the data with a threshold of more than 1.5 as the source domain for data transfer learning.

[0054] In step 4 described above, based on the method of transfer learning, the source domain and the target domain are projected into the manifold space for distribution alignment. The specific steps are as follows:

[0055] 4-1. First, use a 20-dimensional subspace to model the domain, and then embed it into the Grassmann manifold space. The source domain and the target domain are respectively projected into the PCA subspace;

[0056] 4-2. Each subspace can be regarded as a point in the manifold space, and the domain displacement is eliminated through the geodesic flow;

[0057] 4-3. By calculating the marginal A distance, denoted as d m , and the A distance between conditional distributions is denoted as d c , then the μ value is adaptively calculated through the above formula, so as to balance the conditional distribution and the marginal distribution;

[0058] 4-4. Summarize the loss function according to the principle of structural risk minimization and dynamic distribution alignment, and then output the classifier to obtain the classification accuracy.

[0059] The influence of different numbers of source domains on the classification accuracy in source domain selection is as Figure 3 shown. It can be seen that when the threshold remains unchanged, expanding the source domain data also significantly improves the accuracy for each target subject, but the corresponding calculation period will increase significantly.

[0060] The flowchart of all work is as Figure 4 shown. Here, 15 subjects are used as experimental objects, with 3 experiments for each subject. Each experiment records 62 channels, which are recorded according to the international 10-20 standard system. After that, the collected signals are preprocessed to extract different frequency components. The original data is cut with a time window length of 1 s, and the differential entropy features are calculated in five frequency bands (δ: 1-3 Hz, θ: 4-7 Hz, α: 8-13 Hz, β: 14-30 Hz, γ: 31-50 Hz) for each 1 s segment.

Claims

1. A method for selecting a source domain of electroencephalogram across subjects based on domain similarity, characterized in that It includes the following steps: Step 1. First, calculate the between-class distance MMD(x i ) within each source domain subject according to the improved MMD formula and give the confidence level; The MMD distance measures the distance between two distributions in the reproducing kernel Hilbert space and is used to calculate and sum the mean difference of the inter-class distances within the source domain in EEG emotion recognition source domain selection. Its mathematical expression: Among them: C is the category in the source domain, MMD(x i ) is the within-class distance of a certain subject, D s corresponds to the source domain, n s corresponds to n samples in a single source domain, and f is the mapping function corresponding to the samples; For the SEED dataset, based on traditional transfer learning, after screening different source domains using the within-source-domain inter-class means and sums, transfer learning is then carried out and the classification accuracy is obtained. The Pearson function is used to conduct a correlation test on the results, obtaining a correlation coefficient of 0.61, which belongs to a significant correlation. The greater the within-class distance, the relatively higher the accuracy; Step 2. To determine the type of Copula function to be used and facilitate the comparison of the distributions of the source domain and the target domain, the selected source domain data is segmented. In each different emotion category of each subject's each experiment, each type of emotion is segmented and combined with the overall mean of the target domain to construct a joint distribution. By analyzing the kernel distribution estimates of the two sets of data, a suitable Copula function is obtained. Through verification, the Frank Copula function is found to be the most suitable, and it is used for parameter estimation; Step 3. Calculate the Kendall rank correlation coefficient of the Frank-Copula function and superimpose the confidence level of Step 1. Set the threshold to 1.5 to select some of the source domains as the transfer objects for transfer learning; Step 4. Adaptively adjust the distribution balance to balance the conditional distribution and the marginal distribution, then update the soft labels of the target domain, and finally return the classification accuracy output by the classifier.

2. The method for selecting a source domain of electroencephalogram across subjects based on domain similarity according to claim 1, wherein In Step 2 described above, the source domain segmentation is selected, analyzed, and a Copula function is selected for parameter estimation. The specific steps are as follows: Step 2-1. Segment an experiment into 15 segments according to the labels in different time periods. To facilitate fitting as the marginal distribution, the mean value of each segment is calculated. When processing the target domain, the mean value of the entire target domain segment is calculated for similarity measurement; Step 2-2. According to the kernel distribution estimate of the empirical distribution function to approximate the overall distribution type, it is obtained that the tails of the EEG emotion distributions of different subjects are asymmetric. Select Archimedean-Copula to describe the cross-subject EEG emotion signal distribution; Step 2-3. Measure according to the Euclidean distance between the empirical Copula function and the Copula function. The smaller the distance, the better the fitting degree of the Copula function. Select the one with the smallest Frank-Copula distance for modeling.

3. A method for cross-subject source domain selection of EEG based on domain similarity according to claim 1, characterized in that In Step 3 described above, the Kendall correlation coefficient τ is a statistic for categorical variables and is an index used to reflect the correlation of categorical variables, which is applicable to the case of EEG emotion classification. Its mathematical expression: where: (x i - x j ), i, j = 1, 2, ..., d are the observed data, sign is the sign function, and the larger the value of τ, the more significant the correlation between variables.

4. A method for selecting a source domain of electroencephalogram across subjects based on domain similarity according to claim 1, characterized in that, In Step 4 described above, to avoid feature distortion and quantitatively evaluate the importance of the marginal distribution and the conditional distribution, To address the elimination of degenerate features, the original data learns the manifold feature functional g(.) in the Grassmann manifold space and introduces the geodesic flow kernel GFK to facilitate its domain adaptation, and quantitatively evaluates the importance of marginal distribution and conditional distribution in domain adaptation through dynamic distribution alignment; Its mathematical expression: $\mu\in[0,1]$ is the adaptive factor, and $D$ f is the marginal distribution, is the conditional distribution, $C$ is the category. The value of the adaptive factor $\mu$ is estimated by calculating the A-distance to balance the weights of the two distributions, and its mathematical expression is: d M is the edge A distance, d c is the A distance of a certain type. After summarizing the above manifold feature learning and dynamic distribution alignment through the principle of structural risk minimization, the following loss function expression is obtained:

Citation Information

Patent Citations

  • Electroencephalogram emotion migration model training method and system and electroencephalogram emotion recognition method and device

    CN112690793A

  • Electroencephalogram signal recognition method based on metric transfer learning

    CN114330559A