Electroencephalogram signal privacy protection method and device based on de-identity feature

By using an identity-de-associated separation model and leveraging a multi-domain, multi-level time-frequency fusion network and attention feature selection, task features with identity information removed are generated. This resolves the contradiction between identity privacy protection and task decoding accuracy in existing methods, achieving efficient protection of EEG signal privacy and improvement of decoding accuracy.

CN120995496APending Publication Date: 2025-11-21ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511098477.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing methods for protecting the privacy of EEG signals often interfere with task features and affect the accuracy of decoding tasks while protecting identity privacy, and generative algorithms have the risk of being identifiable.

Method used

An identity-de-correlated separation model is adopted, which extracts features through a multi-domain, multi-level time-frequency fusion network. Attention feature selection and gradient inversion loss are used to generate task features without identity information. Combined with task feature compensation and adversarial training, the task decoding accuracy is ensured.

Benefits of technology

It effectively removes identity information, significantly improves the decoding accuracy of EEG tasks, reduces the risk of privacy leaks, and maintains the accuracy of task features and cross-domain generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995496A_ABST
    Figure CN120995496A_ABST
Patent Text Reader

Abstract

The invention discloses an electroencephalogram signal privacy protection method and device based on de-identity features. The protection method comprises the following steps of: for electroencephalogram signals which are respectively generated during multiple imagination after data alignment, generating task characteristics Ftotal for removing identity information by adopting an identity removal correlation separation model; according to the model, computer features F are extracted from electroencephalogram signals through a multi-domain multi-stage time-frequency fusion network, identity features Fid and task features Fmix-task are extracted from F through attention feature selection, correlation coefficients rho and task features Ftask with gradient inversion loss only containing a small amount of identity information are obtained through calculation according to Fid and Fmix-task, task feature compensation items are fused in Ftask, and Ftotal is obtained. The model constraint conditions are as follows: in the formula, identity-level task classification loss, gradient inversion loss and identity classification loss, rho is a correlation coefficient, inferior task feature loss and dominant task feature loss, and lambda1-lambda1 respectively represent tradeoff coefficients of corresponding losses. .
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of electroencephalogram signal analysis, and particularly relates to an electroencephalogram signal privacy protection method based on de-identity features and an electroencephalogram signal privacy protection device based on de-identity features. BACKGROUND

[0002] Traditional electroencephalogram source data protection mainly relies on encryption, including homomorphic encryption (HE), secure multi-party computation (SMC) and secure processors. Data can be encrypted through HE protocols. Outsourcing parties can securely compute data through SMC, and both operations can be implemented on secure processors. However, once the encryption is cracked, the source data will be completely unprotected. Therefore, in recent years, a privacy protection method that emphasizes saving data locally and only transferring relevant model parameters or API interfaces to perform electroencephalogram decoding tasks has been more widely applied, and federated learning and source-free transfer learning are two representative algorithms. Federated learning (FL) is a distributed machine learning method that allows model training on multiple edge devices or data sources without centralizing data on a single server. This method can protect user privacy because data is always kept locally, and only model updates (such as gradient information) are sent to the central server for aggregation and optimization.

[0003] Source-free transfer learning aims to address the data distribution shift between the target domain (test objects) and the source domain (training objects). Unlike traditional transfer learning, source-free transfer learning adapts without access to source domain data, usually by transferring a pre-trained source model to ensure privacy protection of source data. However, although there are fewer attacks on models in the electroencephalogram field, the privacy of such methods is still worth considering. Protection of electroencephalogram identity privacy, as there is private information in electroencephalogram, such as identity, users do not want to expose such information and other electroencephalogram data that may be associated with it, and some algorithms specifically designed for identity protection in electroencephalogram have received attention. Current mainstream algorithms include generative electroencephalogram identity protection and interference electroencephalogram identity protection. Generative electroencephalogram identity protection synthesizes EEG data through a generative adversarial network (GAN), and these generated data are similar to real data but cannot identify individual identity. This algorithm emphasizes the protection of electroencephalogram identity information, but often ignores the accuracy of electroencephalogram decoding tasks, and in addition, the protection of electroencephalogram identity achieved still has certain recognizability. SUMMARY

[0004] To reduce the risk of privacy leakage and reduce the interference of task features, the application provides an electroencephalogram signal privacy protection method based on de-identity features and an electroencephalogram signal privacy protection device based on de-identity features.

[0005] This invention is achieved using the following technical solution: a method for protecting the privacy of EEG signals based on de-identification features, which is as follows:

[0006] The EEG signals generated during multiple imagery sessions after data alignment were used to generate identity-removed task features F using an identity-relevance separation model. total ;

[0007] The model utilizes a multi-domain, multi-level time-frequency fusion network to extract computer features F from EEG signals, and uses attention feature selection to incorporate identity features F. id and task characteristics F mixed-task Extracted from computer feature F, based on F id and F mixed-task The correlation coefficient rho and gradient reversal loss l were calculated. GRL Task characteristics F containing only a small amount of identity information task Integrate task feature compensation terms into task feature F task In the process, the fused task features F are obtained. total The model constraints are:

[0008] min(λ1l cls +λ2l GRL +λ3l dcls -λ4rho-λ5l acls +λ6l tcls )

[0009] In the formula, l cls For identity-level task classification loss, l GRL For gradient reversal loss, l dcls For identity classification loss, rho is the correlation coefficient, and l acls For the loss of disadvantageous task characteristics, l tcls The loss represents the advantageous task characteristics, and λ1 to λ1 represent the trade-off coefficients of the corresponding losses.

[0010] As a further improvement to the above scheme, the method for extracting computer feature F includes the following steps:

[0011] EEG signals Perform time-domain branching to obtain the time-domain local discriminant features F. 1 and temporal global discriminant features F 2 Also to Perform frequency domain branching to obtain local frequency domain discriminative features. and frequency domain global discriminant features

[0012] Then, first fuse F 1 and Re-fusion F 2 and Output multi-level, multi-domain feature information F:

[0013]

[0014] Further, the method for processing the time domain branch is:

[0015] Five parallel two-dimensional convolution layers Extract time features f of different scales 21 k k represents the serial number 1-5 branches;

[0016] After five spatial domain convolutions, the output time splicing signal f 22 is spliced;

[0017] Then, after the squeeze and excitation (SE) operation, the output F 1 ;

[0018] After time domain convolution and squeeze and excitation, output F 2 .

[0019] Further, the method for processing the frequency domain branch is:

[0020] Five parallel wavelet convolution pairs Extract frequency features f corresponding to the frequency band 11 k ;

[0021] After five spatial domain convolutions, the output frequency splicing signal f 12 is spliced;

[0022] Then, after the squeeze and excitation operation, the output

[0023] After time domain convolution and squeeze and excitation, output

[0024] As a further improvement of the above scheme, the correlation coefficient rho, the gradient reversal loss l GRL , and the task feature F containing only a small amount of identity information task are respectively:

[0025]

[0026] l GRL = l CE (GRL(F mixed-task ,y))

[0027]

[0028] In the formula, h idFor identity feature F id The identity feature vector obtained by linear dimensionality reduction, h mixed-task To determine the task features F mixed-task The task feature vector obtained by linear dimensionality reduction, f(·) represents the mean function, f(h) mixed-task ) and f(h id ) represent h respectively mixed-task and h id The mean of g(h), g(·) represents the variance function, g(h) mixed-task ) and g(h id ) represent h respectively mixed-task and h id The variance;

[0029] l CE Cross-entropy loss, GRL is the gradient inversion function, and y is the task label;

[0030] R m×n It is a mathematical range concerning EEG features, where m is the batch size and n is the feature dimension.

[0031] As a further improvement to the above scheme, the identity feature vector F id Perform secondary task classification to extract lost task decoding features f id-task Then, the weights ω are fused to obtain the task feature compensation term ωf. id-task .

[0032] As a further improvement to the above scheme, the identity feature vector F is adjusted according to different expectation functions. id By performing different cross-entropy calculations, the corresponding identity classification loss l is obtained. dcls And task classification loss cls :

[0033] l dcls =E (x,s)~D [l CE (F id ,s)]

[0034] l cls =E (x,y)~C [l CE (f id-task ,y)]

[0035] In the formula, E (x,s)~D Let represent the expectation function of the identity classifier D, s represent the function used for labels, and E represent the expectation function. (x,y)~C Let C represent the expectation function of the first task classifier C.

[0036] As a further improvement to the above scheme, based on the fused task characteristics F total The attention-based adversarial feature selection network is used to compute the dominant feature F.TD and disadvantageous features F TI :

[0037]

[0038] denotes multiplication, denotes addition, and A is an adversarial loss weight;

[0039] According to the advantageous features F TD and the disadvantageous features F TI respectively, and calculate the cross-entropy, and the corresponding advantageous task feature loss l tcls and the disadvantageous task feature loss l acls :

[0040]

[0041] In the formula, is the expected function of the third task classifier C2, is the expected function of the second task classifier C1.

[0042] Further,

[0043] In the formula, epsilon is Gumbel noise used to introduce randomness, tau is a temperature coefficient to control the smoothness of Softmax, U(0, 1) is a normal distribution with 0 as the mean value and 1 as the standard deviation, and alpha is an adversarial intensity factor to make the model more robust.

[0044] The application also provides an electroencephalogram privacy protection device based on de-identity features, which adopts any of the above electroencephalogram privacy protection methods based on de-identity features, and the electroencephalogram privacy protection device comprises:

[0045] A data alignment module is configured to generate electroencephalogram signals generated during multiple times of imagination respectively after data alignment;

[0046] A de-identity correlation separation model is configured to generate task features F total without identity information from the electroencephalogram signals after data alignment;

[0047] In the formula, the model extracts computer features F from the electroencephalogram signals by using a multi-domain multi-level time-frequency fusion network, extracts identity features F id and task features F mixed-task from the computer features F by using attention feature selection, calculates a correlation coefficient rho, a gradient reversal loss l GRL , and task features F task containing only a small amount of identity information according to F id and F mixed-task , and fuses a task feature compensation term in the task features Ftask In the formula, the fused task feature F is obtained total The model constraint condition is:

[0048] min(λ1l cls +λ2l GRL +λ3l dcls -λ4rho-λ5l acls +λ6l tcls )

[0049] In the formula, l cls is an identity level task classification loss, l GRL is a gradient reversal loss, l dcls is an identity classification loss, rho is a correlation coefficient, l acls is a disadvantaged task feature loss, l tcls is an advantaged task feature loss, λ1-λ1 respectively represent the weighting coefficients of the corresponding losses.

[0050] In order to reduce the risk of privacy leakage and reduce the interference of the task feature, the EEG privacy protection method based on identity removal (ID-RemovalNet) is proposed, and the core is to construct an identity correlation separation model, that is, an EEG signal decomposition framework. Unlike the overall EEG protection method, the EEG is decomposed into task features, identity privacy features and noise three components, and the identity privacy feature protection is focused on. The identity correlation separation model uses a multi-domain multi-level fusion feature extraction network (MDMLNet) to extract high-quality and rich feature information from the EEG signal, and on the other hand, designs an identity decorrelation separation network (IDSNet) (because the existing generative and subtle disturbance methods interfere with the task classification features while protecting the identity privacy), directly removes the identity privacy features irrelevant to the task, which will not affect the subsequent task classification work. Finally, EEG privacy protection and decoding enhancement are taken into account, and a task feature enhancement module is introduced. The task feature enhancement module stimulates the selection of dominant features through an attention-based feature selection network (AAF), improves the decoding accuracy and cross-domain generalization ability, and designs a loss-guided identity level task feature re-fusion module (TFRF) to further make up for the loss of task features after removing the identity information, and ensure the accuracy of task decoding.

[0051] The contribution points of the present application are as follows:

[0052] 1. The identity-removal model (ID-RemovalNet) proposed in this invention designs an identity-removal separation (IDS) module to remove identity information and protect EEG identity privacy. At the same time, it designs a feature enhancement and classifier (AAFS & LITFR) module. LITFR effectively compensates for the partial loss of EEG task features during the removal process, and AAFS stimulates the selection of dominant features of the EEG task, thereby improving and generalizing the decoding of the EEG task.

[0053] 2. The ID-RemovalNet proposed in this invention designs a multi-domain, multi-level fusion feature extraction (MDML) module. By designing the fusion of global and local features in the time and frequency domains, it extracts rich EEG features, thereby improving the decoding task of EEG tasks.

[0054] 3. The identity-de-associated separation model proposed in this invention removes identity information to 0.43% on four EEG datasets of two different paradigms, while significantly improving the decoding accuracy of EEG tasks by 3.28%, and achieving state-of-the-art recognition rate in "cross-subject" EEG experiments. Attached Figure Description

[0055] Figure 1 This is a flowchart of the EEG signal privacy protection method based on de-identification features according to the present invention.

[0056] Figure 2 To achieve Figure 1 A schematic diagram of a module for a brainwave signal privacy protection device based on de-identified features.

[0057] Figure 3 for Figure 2 A flowchart of the method for extracting computer features F based on MDML.

[0058] Figure 4 for Figure 2 A flowchart illustrating the feature selection network processing method for attention adversarial approaches. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0060] This invention discloses a method for protecting the privacy of EEG signals based on de-identification features, aiming to improve decoding performance while focusing on the privacy and security of EEG identity. The method is used to process EEG signals x generated during multiple imagery sessions. iAn identity-relevance separation model is used to generate task features F after removing identity information. total The identity-decorrelated separation model effectively removes identity information from EEG signals through multi-domain, multi-level feature extraction and identity decorrelation techniques, reducing the risk of privacy breaches and minimizing its interference with task features. Furthermore, loss-guided feature refusion and attention-based adversarial feature selection improve the accuracy of EEG task decoding. Experiments validated the method on four different EEG datasets, showing a low identity removal rate of 0.43% and a significant improvement in decoding accuracy of 3.28%, representing an average improvement of 6.00% compared to baseline methods in subject-dependent experiments.

[0061] Please see Figure 1 and Figure 2 , Figure 1 This is a flowchart of the EEG signal privacy protection method based on de-identification features according to the present invention. Figure 2 To achieve Figure 1 A schematic diagram of a module for a de-identified EEG signal privacy protection device, which is a method for protecting the privacy of EEG signals. The EEG signal privacy protection device is used to implement... Figure 1 This is a specific example of a method for protecting the privacy of EEG signals. The method mainly includes two steps:

[0062] Step 1: EEG signal data alignment: Aligning the EEG signals x generated during multiple visualizations. i ;like Figure 2 The area marked "A)" is shown in the image.

[0063] A) EEG signal data alignment

[0064] The step of aligning EEG signal data can be performed by the data alignment module, which will align the EEG signal x i Data alignment yields aligned EEG signals If it is possible to detect brainwave signals x i Data alignment can effectively reduce subject differences between different sessions and improve the accuracy of EEG decoding.

[0065] Given an EEG dataset in This refers to the EEG signal generated during the i-th imagery session, such as motor imagery EEG. In each trial, the left and right hands are imagined. Event-related potentials (ERPs) are collected at the time of the event in each trial, where c represents the sample channel, t represents the sample time, and y... i ∈Y={1,...,K} is the label corresponding to the task, u i ∈U={1,...,U} are the user labels for the i-th trial, and N is the number of EEGs.

[0066] The EA (data) alignment effectively reduces the subject difference between different sessions, and for N EEG tests in a specific domain, EA first calculates the Euclidean arithmetic mean of all N spatial covariance matrices R:

[0067]

[0068] Then, the alignment is performed, and the average spatial covariance matrix becomes a unit matrix, and the EEG signal after data alignment is obtained

[0069]

[0070] Step two, generate identity information removed task features F by using the identity de-correlation separation model total : 1) identity de-correlation, as shown in the "B) marked area in the Figure 2 ; 2) feature enhancement and classification, as shown in the "C) marked area in the Figure 2 .

[0071] The identity de-correlation separation model of the present application learns to express the following optimization problem:

[0072] min(λ1l cls +λ2l GRL +λ3l dcls -λ4rho-λ5l acls +λ6l tcls )

[0073] In the formula, l cls is the identity level task classification loss, l GRL is the gradient reversal loss, l dcls is the identity classification loss, rho is the correlation coefficient, l acls is the inferior task feature loss, l tcls is the superior task feature loss, and λ1-λ1 respectively represent the weighting coefficients before the above losses.

[0074] The corresponding EEG signal privacy protection device comprises a data alignment module for performing the "A)" step, and an identity de-correlation separation model for performing the second step. The identity de-correlation separation model comprises an identity de-correlation module for performing the "B)" step and a feature enhancement and classification module for performing the "C)" step.

[0075] B) Identity de-correlation

[0076] The identity decorrelation method includes the following steps: I to IV. The corresponding identity decorrelation modules include a multi-domain, multi-level fusion feature extraction module, an attention feature selection module, a computation module, and a task feature fusion module. Features are extracted through the multi-domain, multi-level fusion feature extraction (MDML) module. Then, the attention feature selection module implements a gradient inversion layer and distance constraints to separate task features and identity features. Based on the task features and identity features, the computation module calculates the correlation coefficient rho and the gradient inversion loss l. GRL Task characteristics F containing only a small amount of identity information task The task feature fusion module integrates task feature compensation items into task feature F. task In the process, the fused task features F are obtained. total This allows for the removal of identity features through feature subtraction.

[0077] I utilize multi-domain, multi-level time-frequency fusion network (MDML) to analyze EEG signals Extract computer features F.

[0078] Please combine Figure 3 , Figure 3 for Figure 2 A flowchart illustrating the method for extracting computer-generated features F. The multi-domain, multi-level fusion feature extraction module is used to implement... Figure 3 A specific example of a method for extracting computer feature F.

[0079] The method for extracting computer feature F includes the following steps: analyzing electroencephalogram (EEG) signals... Perform time-domain branching to obtain the time-domain local discriminant features F. 1 and temporal global discriminant features F 2 It also affects brain signals. Perform frequency domain branching to obtain local frequency domain discriminative features. and frequency domain global discriminant features Then, the temporal local discriminant features F are first fused. 1 and frequency domain local discrimination features Re-fusion of temporal global discriminative features F 2 and frequency domain global discriminant features Output multi-level, multi-domain feature information F.

[0080]

[0081] Correspondingly, the computer feature extraction module includes a time-domain branch processing unit, a frequency-domain branch processing unit, and a fusion unit. The time-domain branch processing unit is used to process EEG signals. Perform time-domain branching to obtain the time-domain local discriminant features F. 1 and temporal global discriminant features F 2 The frequency domain branching processing unit is used to process EEG signals x.i Branch processing in frequency domain to obtain frequency domain local discriminative features and frequency domain global discriminative features The fusion unit is used to first fuse the time domain local discriminative features F 1 and the frequency domain local discriminative features The time domain global discriminative features F 2 and the frequency domain global discriminative features Output multi-level and multi-domain feature information F.

[0082] The time domain branch processing method of the time domain branch processing unit is five parallel two-dimensional convolution layers to extract time features f 21 k of different scales, k represents the sequence number 1-5 branches, and after five spatial domain convolutions, the time domain splicing signal f 22 is output, then after the squeeze and excitation (SE) operation, the time domain local discriminative features F 1 are output, and after the time domain convolution and the squeeze and excitation, the time domain global discriminative features F 2 are output, which represent different levels of time dependence.

[0083] The frequency domain branch processing method of the frequency domain branch processing unit is based on the research in the literature, and according to the multi-rhythm of the electroencephalogram signal, five parallel learnable continuous wavelet convolutions are used to extract frequency features f 11 k of corresponding frequency bands, after five spatial domain convolutions, f 12 is output, and after the squeeze and excitation, the frequency domain local discriminative features F are output, and after the time domain convolution and the squeeze and excitation, the frequency domain global discriminative features F

[0084] The fusion unit is different from the fusion of only focusing on different domain features, but adopts a fusion strategy of first fusing the time domain local discriminative features and the frequency domain local discriminative features, and then fusing the time domain global discriminative features and the frequency domain global discriminative features, to effectively output multi-level and multi-domain feature information F.

[0085] II Using attention feature selection to extract identity features F id and task features F mixed-task from F:

[0086] F = F mixed-task + F id + F noise

[0087] In the formula, F mixed-task represents the task features in the electroencephalogram signal , F id represents the electroencephalogram signal identity feature in F noise representing electroencephalogram signal noise feature in F.

[0088] This step can be performed by an attention feature selection module, which utilizes attention feature selection to extract F mixed-task and F id from F.

[0089] III. Calculate the correlation coefficient rho, the gradient reversal loss l id and the task feature F mixed-task containing only a small amount of identity information from the identity feature F GRL and the task feature F task :

[0090]

[0091] l GRL = l CE (GRL(F mixed-task , y))

[0092]

[0093] where h id is the identity feature vector obtained by performing linear dimension reduction on the identity feature F id , h mixed-task is the task feature vector obtained by performing linear dimension reduction on the task feature F mixed-task , f(·) represents the mean function, f(h mixed-task ) and f(h id ) represent the mean of h mixed-task and h id respectively, g(·) represents the variance function, g(h mixed-task ) and g(h id ) represent the variance of h mixed-task and h id respectively;

[0094] l CE is the cross-entropy loss, GRL is the gradient reversal function, and y is the task label.

[0095] R m×n is the mathematical range of the electroencephalogram feature, where m is the batch size and n is the feature dimension.

[0096] This step can be performed by a calculation module for calculating the correlation coefficient rho, the gradient reversal coefficient l id and the task feature F mixed-task containing only a small amount of identity information from the identity feature F GRL and the task feature F task .

[0097] Due to the complexity of the brain electrical characteristics, the task features and the identity features extracted by the MDML have coupling correlation, therefore, the application introduces a de-correlation regularization term to calculate the correlation rho between the two, and by minimizing the coefficient, the effective separation of the task features and the identity features is realized. mixed-task and the identity features F id are linearly reduced in dimension to obtain vectors h task and h id , wherein f(h task ) and f(h id ) represent the mean value processing of the vectors, so as to center them, g(h task ) and g(h id ) represent the calculation of the variance of the vectors, and finally the correlation coefficient rho is obtained.

[0098] In addition, a gradient reversal layer (GRL) is introduced to realize the further separation of the features by inter-domain adversarial training. The GRL reverses the gradient direction, promotes the network to optimize the features in the training process, so that the task-related features F mixed-task and the identity-related features F id are distinguished as much as possible. Wherein, the output of the task classifier is f(x), and the loss of the task classifier is L d .

[0099]

[0100] At this time, the task features and the identity features are highly de-correlated, and finally the feature subtraction is performed, so as to obtain the task features F task containing only a small amount of identity information.

[0101] IV The task feature compensation term ωf id-task is fused in the task features F task to obtain the fused task features F total , that is, the final task features:

[0102] F total =F task +ωf id-task

[0103] In the formula, f id-task represents the lost task decoding features, because the simple removal of the identity information may cause the loss of the task decoding features, and ω represents the feature fusion weight, which can adopt an empirical value. The source of the task feature compensation term ωf id-task is described below, and it is best to be optimized in the continuous learning of the model until the ωf id-taskFor the best task feature compensation term.

[0104] This step can be performed by a task feature fusion module for fusing the task feature compensation term ωf id-task in the task feature F task , to obtain the fused task feature F total .

[0105] C) Feature enhancement and classification

[0106] The feature enhancement method includes the following steps I-V. The corresponding feature enhancement and classification module includes: a task feature enhancement module, an identity classifier D, three task classifiers C, C1, C2, an attention-based adversarial feature selection network (AAFS), and an identity-level task feature re-fusion module (TFRF). The task feature enhancement module is LITFR, which guides the identity-level task feature through the loss to supplement the task feature. The attention-based adversarial feature selection network is AAFS, which generates attention weights for adversarial training.

[0107] I. Different cross-entropy calculations are performed on the identity feature vector F id according to different expected functions, to obtain the corresponding identity classification loss l dcls and task classification loss l cls :

[0108] l dcls = E (x,s)~D [l CE (F id ,s)]

[0109] l cls = E (x,y)~C [l CE (f id-task ,y)]

[0110] In the formula, E (x,s)~D represents the expected function of the identity classifier D, s represents the label, and E (x,y)~C represents the expected function of the first task classifier C.

[0111] This step can be performed by the identity classifier D and the first task classifier C.

[0112] II. The identity feature vector F id is subjected to secondary task classification to further extract residual task features, i.e., lost task decoding features f id-task , and then fused with a weight ω to obtain ωf id-task .

[0113] This step can be performed by a task feature enhancement module, which can be LITFR. Because the task feature and the identity feature are not completely linearly related, simple removal of identity information may result in loss of task decoding features. Thus, subsequently, the identity feature loss to the task can be calculated by a loss-guided identity-level task feature re-fusion module (TFRF), and the feature fusion weight is dynamically adjusted to balance the removal of identity information and the integrity of task decoding.

[0114] III According to the fused task feature F total , an advantage feature F TD and a disadvantage feature F TI are calculated by a feature selection network (AAF S) :

[0115]

[0116] represents multiplication, represents addition, and A is the adversarial loss weight.

[0117] Please refer to Figure 4 , the attention-adversarial feature selection network (AAF S) is used to: decode the fused feature task F total to obtain an output adversarial loss initial weight a, and optimize a to the most weight adversarial loss weight A; then on the one hand, the advantage feature F TD is calculated according to the adversarial loss weight A, and on the other hand, the disadvantage feature F TI is calculated for 1-A.

[0118]

[0119] In the formula, ε is Gumbel noise used to introduce randomness, τ is a temperature coefficient to control the smoothness of Softmax, U(0,1) is a normal distribution with 0 as the mean value and 1 as the standard deviation, and α is an adversarial intensity factor to make the model more robust.

[0120] This step can be performed by AAF S. During the training process of the neural network, the model usually activates the main features related to the label. However, when facing unseen test data, these features may not work, resulting in performance degradation. In order to solve this problem, inspired by the self-challenge mechanism, a feature selection module based on attention-adversarial is designed to strengthen the influence of key features, force the model to better mine the features most useful for decoding in the task, and at the same time enhance the robustness of the model, thereby improving its generalization ability on unseen data. As shown in formula 8, the attention weight a is calculated, and the most weight A is obtained through the above formula.

[0121] IV According to the advantage feature F TD and the disadvantage feature FTI respectively, and the cross-entropy is calculated, and the corresponding superior task feature loss l tcls and inferior task feature loss l acls :

[0122]

[0123] In the formula, is the expected function of the third task classifier C2, is the expected function of the second task classifier C1.

[0124] This step can be performed by the second task classifier C1 and the third task classifier C2. F TD and F TI are respectively sent into the main classifier, i.e., C2, and the auxiliary classifier, i.e., C1, for training. Under the supervision of the cross-entropy loss, the backbone feature generator and the classifier are trained to predict the correct label.

[0125] VAccording to the identity classification loss l dcls , the identity-level task classification loss l cls , the superior task feature loss l tcls , the inferior task feature loss l acls , the gradient reversal loss l GRL , and the correlation coefficient rho, the learning expression of the de-identity correlation separation model is as follows:

[0126] min(λ1l cls +λ2l GRL +λ3l dcls -λ4rho-λ5l acls +λ6l tcls )

[0127] In the formula, λ1-λ1 respectively represent the weighting coefficients before the above losses.

[0128] This step can be performed by the identity-level task feature re-fusion module (TFRF). Here, λ1-λ1 are weighting coefficients, and by minimizing the constraint of the above loss function, the task feature F total without identity information can be generated, and the obtained F total not only can hardly identify the identity but also can robustly decode across sessions.

[0129] Next, experiments are used to verify the beneficial effects of the present application.

[0130] Experiment

[0131] 1. Database

[0132] The following four public databases are used in this experiment, as shown in Table 1.

[0133] Table 1: Detailed information of the four datasets used in the experiments

[0134]

[0135] 1) Four-class motor imagery database (MI4): derived from BCI Competition IV1 dataset 2a. The data were recorded at 250 Hz with 22 EEG channels. The data within 0-4 seconds after each imagination cue were extracted and band-pass filtered at [8-32] Hz.

[0136] 2) P300 evoked potential (P300): P300 evoked potential data from 4 disabled subjects and 4 healthy subjects were collected, with classification of target and non-target, recorded at 32 channels, at 2048 Hz. We re-referenced the data to remove mastoid channels and filtered with a 1-40 Hz band-pass filter, followed by down-sampling to 128 Hz. The length of each EEG signal epoch was 0-1 second.

[0137] 3) Two-class motor imagery database (MI2): derived from BCI Competition IV1 dataset 2b. The first two sessions without visual feedback were used in this paper, recorded at 250 Hz with 3 EEG channels. The data within 0-4 seconds after each imagination cue were extracted and band-pass filtered at [1-40] Hz.

[0138] 4) Error-related negativity (ERN): derived from the 2015 IEEE Conference on Neural Engineering Competition, the training set from 16 users was used in this paper. The data were recorded at 56 channels, at 200 Hz, followed by down-sampling to 128 Hz. EEG signal epochs of 0-1.3 seconds were extracted.

[0139] 2、Evaluation index

[0140] To balance the unbalanced problem of some samples in the experiment, we used balanced classification accuracy (BCA) to evaluate the performance of the task classifier.

[0141]

[0142] To measure the degree of identity information mining from EEG data, we used identity recognition accuracy to evaluate the performance of the identity classifier.

[0143]

[0144] For each database, leave-one-session cross-validation was performed, i.e. one session was used to train the task classifier and the identity recognizer, and the remaining sessions were used to test the task classification accuracy and the identity recognition results of the task features, ensuring that all sessions became a training set once.

[0145] We repeat each experiment three times and report the average.

[0146] 3. Baseline models

[0147] We use the following three CNN models as baseline feature extractors and keep the last fully connected layer of each model as the task classifier while using two fully connected layers as the identity discriminator: (1) EEGNet: a compact CNN architecture designed for EEG classification tasks with depthwise separable convolutions. (2) ShallowCNN: a shallow version of DeepCNN with only one convolutional block and larger kernel and different pooling methods. (3) DeepCNN: contains four convolutional blocks, the first one is designed for EEG input, and the rest are standard convolutional blocks.

[0148] 4. Hyperparameter settings

[0149] In the experiment, we use a batch size of 128, an initial learning rate of 0.01, and adjust the learning rate to 0.001 after 50 rounds. The model is trained for a total of 100 times, and the best model is selected for testing. We use balanced classification accuracy (BCA) to evaluate the performance of the task classifier and use identity accuracy (UIA) to evaluate the performance of the identity classifier. For each database, we perform a leave-one-session cross-validation and report the average of three experiments. For each loss function of multi-task joint learning, we test the hyperparameters in the range of [0.01-1], and use the step-by-step adjustment method, that is, from small to large adjustment of each hyperparameter while keeping other hyperparameters unchanged until the optimal value is found. After determining the optimal value of each hyperparameter, we fix the value to ensure that each hyperparameter can achieve the best effect. In addition, according to experience, we design a general parameter setting for each database: 1.00, 0.05, 1.00, 0.05, 0.01, 1.00. See Table 2 for detailed experimental parameter settings.

[0150] Table 2: Parameter settings for the four datasets used in the experiment

[0151]

[0152]

[0153] 4. Results

[0154] 4.1. Task features with identity decorrelation for identity protection

[0155] To evaluate the performance of the proposed model in EEG identity privacy protection and task decoding, we conducted experiments on four data using four models. The data is shown in Table 2, which shows the EEG decoding rate and user identification rate of the original data, and the task feature EEG decoding rate and user identification rate after identity feature removal.

[0156] (1) The higher UIA value on raw EEG indicates that the original data contains a higher degree of identity information, which means that through the original EEG data, it is easy to identify the attribution of a segment of EEG features.

[0157] (2) The experiment proves that ID-RemovalNet performs excellently in identity protection. For MI4C, the UIA value of Decorrelated EEG decreases by up to 92.12%; for MI2C, the UIA value of Decorrelated EEG decreases by up to 73.62%; for ERN, the UIA value of Decorrelated EEG decreases by up to 80.04%; for P300, the UIA value of Decorrelated EEG decreases by up to 97.01%.

[0158] (3) The experiment proves that ID-RemovalNet can enhance the decoding of EEG tasks. For MI4C, the BCA value of Decorrelated EEG increases by up to 7.79%; for MI2C, the BCA value of Decorrelated EEG increases by up to 4.27%; for ERN, the BCA value of Decorrelated EEG increases by up to 3.63%; for P300, the BCA value of Decorrelated EEG increases by up to 2.82%.

[0159] (4) Whether on raw EEG or on Decorrelated EEG, the BCA value of the proposed MDML remains at the optimal result. At the same time, the UIA value is at a low average level.

[0160] In addition, we can obtain clean task features through any network model in the training or testing stage. This makes our method highly adaptable and can achieve the desired purpose in any database and feature extractor.

[0161] Table 3: In MI4C, MI2C, ERN, P300, the task recognition accuracy and identity recognition accuracy of three basic models and the proposed CWTnet on raw EEG signals and identity-removed EEG features

[0162]

[0163]

[0164] 4.2, Ablation Experiment

[0165] We conducted an ablation experiment by adding each component in the method step by step, and the results of the ablation experiment are listed in Table 4.

[0166] (1) The baseline data is the result of the original EEG after task decoding and user identification without adding any operation.

[0167] (2) We added the alignment processing, and the BCA value was greatly improved, and the UIA value decreased in the motor imagery MI4C and MI2C databases. This indicates that alignment is more conducive to task decoding of EEG data.

[0168] (3) The third and fourth columns of each database in Table 4 are the results of adding the identity-related separation network, and the data shows that after adding the de-correlation separation component, the task information and identity information in the EEG are separated, and the user identification result of the EEG task feature at this time is reduced to 5%-20%. The BCA has little effect. After adding the identity component, the BCA remains almost unchanged, but the UIA at this time again decreases greatly to below 2.5%.

[0169] (4) We added the IDSNet component, Dec. represents the feature de-correlation component, and Sub. represents the feature subtraction component. The data of the four databases shows that after adding Dec., the task information and identity information in the EEG are separated, and the user identification result of the EEG task feature at this time is reduced to 5%-20%. The BCA has little effect. After adding Sub., the BCA remains almost unchanged, but the UIA at this time again decreases greatly to below 2.5%.

[0170] (5) We continue to add the feature enhancement component, which includes attention-based adversarial feature selection (AAF S.) and loss-guided identity-level task feature re-fusion (LITFR.), and the data shows that after adding only the attention-based adversarial feature selection component, the UIA value appears a small fluctuation while maintaining a low value, but the BCA value at this time appears a large increase, and after adding the loss-guided identity-level task feature re-fusion component, a small increase appears on different databases.

[0171] (6) The fifth and sixth columns of each database in Table 4 are the results of adding the loss-guided identity-level task feature re-fusion and attention-based adversarial feature selection feature enhancement network, and the data shows that after adding only the attention-based adversarial feature selection component, the UIA value appears a small fluctuation while maintaining a low value, but the BCA value at this time appears a large increase, and to avoid removing too many task features, after adding the loss-guided identity-level task feature re-fusion component, a small increase appears on different databases.

[0172] Table 4: The task recognition accuracy and user recognition accuracy achieved by three basic models and the proposed CWTNet in the ablation study on MI4C, MI2C, ERN, P300. The third and fourth rows in the table represent the IDSNet results, and the fifth and sixth rows represent the feature enhancement (AAF&LITFR) results

[0173]

[0174] 4.3 Ablation experiment of multi-domain and multi-level fusion network

[0175] To verify the effectiveness of the proposed CWTNet in extracting high-quality EEG task features and identity features, we conducted an ablation experiment by gradually adding each branch (TD&FD) and fusion method (GC&FC) in the method. The ablation experiment results are shown in Table 5.

[0176] As shown in Table 5, except for P300, multi-scale frequency domain convolution performs better than multi-scale time domain convolution. On motor imagery MI4C and MI2C, globally fused features are better for recognition, but on event-related potentials ERN and P300, local features perform better.

[0177] Experiments show that the CWTNet with complete structure performs the best BCA on the four databases, with an average of about 1.00% higher than the rest of the results, and the lowest UIA value on the ERN and P300 databases.

[0178] Table 5: The task recognition accuracy and user recognition accuracy achieved by CWTNet in the ablation study on MI4C, MI2C, ERN, P300

[0179]

[0180] 5. Discussion

[0181] 5.1 Cross-subject comparison experiment

[0182] Electroencephalogram (EEG) cross-subject comparison test is an important challenge in brain-computer interface (BCI) research. However, due to physiological differences between individuals, there is great variability in EEG data between subjects, which brings many problems to cross-subject decoding. But ID-RemoveNet proves to be effective in removing identity features, protecting privacy while reducing its interference with task features. We reproduce the code in the existing paper, and conduct cross-subject experiments on databases BCI 2a, BCI 2b, ERN, P300. In order to maintain the fairness of the experiment, we use the cross-subject comparison test, which is a model trained by one session data applied to the remaining sessions, and use the test data features to do cross-subject experiments. DRDA and BDAN-SPD use the original paper data.

[0183] ID-RemovalNet with baseline modelTable 6 shows that we select the baseline model EEGNet, ShallowCNN, DeepCNN for experiments, and each subject becomes a target domain. The experiment shows that the cross-subject performance of EEGNet, ShallowCNN, and DeepCNN using ID-RemovalNet is better than the original data. Specifically, for BCI2a, EEGNet improves from 69.8% to 77.0%, ShallowCNN improves from 71.3% to 76.4%, and DeepCNN improves from 65.6% to 71.8%. At the same time, the improvement of each subject is achieved. In addition, Table 7 further demonstrates the applicability of ID-RemovalNet. The results of all baseline models using ID-RemovalNet on ERN, P300, and BCI 2b databases are still better than the original data.

[0184] ID-RemovalNet and othersIn addition, the accuracy of the proposed ID-RemovalNet on BCI2a, ERN, P300, and BCI2b is as high as 83.9%, 86.9%, 87.4%, and 73.0%, respectively, which surpasses all other experimental results we reproduced, indicating that the task features extracted by our model effectively remove the interference of identity information and the superiority of the feature enhancement module, and also proving that MDML extracts discriminative features, further enhancing the classification performance.

[0185] Table 6: BCI2a cross-subject comparison test

[0186]

[0187] Table 7: ERN, P300, BCI 2b cross-subject comparison test

[0188]

[0189] 6、Summary

[0190] The present application proposes a new electroencephalogram privacy protection framework ID-RemoveNet, aiming to solve the challenge of maintaining decoding performance while protecting the privacy of electroencephalogram data. The framework effectively extracts high-quality electroencephalogram features through a multi-domain multi-level fusion feature extraction module, and uses an identity de-correlation separation module to strip identity features, both protecting personal privacy and reducing the interference of identity information on task features. In addition, ID-RemoveNet further optimizes task feature selection by feature enhancement, makes up for feature loss, and effectively enhances decoding accuracy. Experiments show that ID-RemoveNet removes identity information to 0.43% on four electroencephalogram datasets of two different paradigms, while significantly improving the electroencephalogram task decoding accuracy by 3.28%, and achieving the best in the cross-subject experiment. In future research, we will focus on the remaining privacy information in electroencephalogram and use multi-feature removal network to further protect electroencephalogram.

[0191] In the present application, we propose a new electroencephalogram privacy protection framework ID-RemoveNet, aiming to solve the challenge of maintaining decoding performance while protecting the privacy of electroencephalogram data. The framework effectively extracts high-quality electroencephalogram features through a multi-domain multi-level fusion feature extraction network, and effectively protects electroencephalogram identity privacy through an identity de-correlation separation network. At the same time, the feature enhancement module optimizes task feature selection, makes up for feature loss, and effectively enhances decoding accuracy. In the evaluation of four databases, ID-RemoveNet significantly improves decoding accuracy, greatly reduces identity information, and performs well in the defined cross-subject. In future research, we will study the remaining privacy information in electroencephalogram and use multi-feature removal network to further protect electroencephalogram.

[0192] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for protecting the privacy of EEG signals based on de-identified features, characterized in that, It is: The EEG signals generated during multiple imagery sessions after data alignment were used to generate identity-removed task features F using an identity-relevance separation model. total ; The model utilizes a multi-domain, multi-level time-frequency fusion network to extract computer features F from EEG signals, and uses attention feature selection to incorporate identity features F. id and task characteristics F mixed-task Extracted from F, based on F id and F mixed-task The correlation coefficient rho and gradient reversal loss l were calculated. GRL Task characteristics F containing only a small amount of identity information task Integrate task feature compensation items into F task In the middle, we get F total The model constraints are: min(λ1l cls +λ2l GRL +λ3l dcls -λ4rho-λ5l acls +λ6l tcls ) In the formula, l cls For identity-level task classification loss, l GRL For gradient reversal loss, l dcls For identity classification loss, rho is the correlation coefficient, and l acls For the loss of disadvantageous task characteristics, l tcls The loss represents the advantageous task characteristics, and λ1 to λ1 represent the trade-off coefficients of the corresponding losses.

2. The method for protecting the privacy of EEG signals based on de-identified features as described in claim 1, characterized in that: The method for extracting computer feature F includes the following steps: EEG signals Perform time-domain branching to obtain the time-domain local discriminant features F. 1 and temporal global discriminant features F 2 Also to Perform frequency domain branching to obtain local frequency domain discriminative features. and frequency domain global discriminant features Then, first fuse F 1 and Re-fusion F 2 and Output multi-level, multi-domain feature information F:

3. The method for protecting the privacy of EEG signals based on de-identification features as described in claim 2, characterized in that: The method for handling time-domain branching is as follows: Five parallel two-dimensional convolutional layers Extracting time features f at different scales 21 k k represents the number of branches 1-5; After five spatial convolutions, the output time-concatenated signal f is spliced ​​together. 22 ; Then, after the squeezing and excitation (SE) operation, the output F is... 1 ; After temporal convolution, compression, and excitation, the output F is obtained. 2 .

4. The method for protecting the privacy of EEG signals based on de-identified features as described in claim 2, characterized in that: The method for frequency domain branching is as follows: Five parallel wavelet convolution pairs are used Extract the frequency features f of the corresponding frequency band 11 k ; After five spatial convolutions, the output frequency spliced ​​signal f is concatenated. 12 ; Then, after compression and excitation, the output is... Then, after temporal convolution, compression, and excitation, the output is...

5. The method for protecting the privacy of EEG signals based on de-identified features as described in claim 1, characterized in that: Correlation coefficient rho, gradient reversal loss l GRL Task characteristics F containing only a small amount of identity information task They are respectively: l GRL =l CE (GRL(F mixed-task ,y)) In the formula, h id For identity feature F id The identity feature vector obtained by linear dimensionality reduction, h mixed-task To determine the task features F mixed-task The task feature vector obtained by linear dimensionality reduction, f(·) represents the mean function, f(h) mixed-task ) and f(h id ) represent h respectively mixed-task and h id The mean of g(h), g(·) represents the variance function, g(h) mixed-task ) and g(h id ) represent h respectively mixed-task and h id The variance; l CE Cross-entropy loss, GRL is the gradient inversion function, and y is the task label; R m×n It is a mathematical range concerning EEG features, where m is the batch size and n is the feature dimension.

6. The method for protecting the privacy of EEG signals based on de-identified features as described in claim 1, characterized in that: The identity feature vector F id Perform secondary task classification to extract lost task decoding features f id-task Then, the weights ω are fused to obtain the task feature compensation term ωf. id-task .

7. The method for protecting the privacy of EEG signals based on de-identified features as described in claim 1, characterized in that: Based on different expectation functions, the identity feature vector F id By performing different cross-entropy calculations, the corresponding identity classification loss l is obtained. dcls And task classification loss cls : l dcls =E (x,s)~D [l CE (F id ,s)] l cls =E (x,y)~C [l CE (f id-task ,y)] In the formula, E (x,s)~D Let represent the expectation function of the identity classifier D, s represent the function used for labels, and E represent the expectation function. (x,y)~C Let C represent the expectation function of the first task classifier C.

8. The method for protecting the privacy of EEG signals based on de-identification features as described in claim 1, characterized in that: Based on the fused task characteristics F total The attention-based adversarial feature selection network is used to compute the dominant feature F. TD and disadvantageous characteristics F TI : Indicates multiplication. Indicates addition, where A is the weight for adversarial loss; Based on the dominant characteristic F TD and disadvantageous characteristics F TI Perform adversarial training separately and calculate cross-entropy to obtain the corresponding advantageous task feature loss l. tcls and disadvantageous task feature loss l acls : In the formula, It is the expectation function of the third task classifier C2. It is the expectation function of the second task classifier C1.

9. The method for protecting the privacy of EEG signals based on de-identification features as described in claim 8, characterized in that: In the formula, ε is Gumbel noise used to introduce randomness, τ is the temperature coefficient that controls the smoothness of Softmax, U(0,1) is a normal distribution with a mean of 0 and a standard deviation of 1, and α is the adversarial strength factor used to make the model more robust.

10. A device for protecting the privacy of EEG signals based on de-identification features, comprising the method for protecting the privacy of EEG signals based on de-identification features as described in any one of claims 1 to 9, characterized in that, The EEG signal privacy protection device includes: The data alignment module is used to align the EEG signals generated during multiple visualizations. The identity-relevance separation model generates task features F from the aligned EEG signals, removing identity information. total ; The model utilizes a multi-domain, multi-level time-frequency fusion network to extract computer features F from EEG signals, and uses attention feature selection to incorporate identity features F. id and task characteristics F mixed-task Extracted from computer feature F, based on F id and F mixed-task The correlation coefficient rho and gradient reversal loss l were calculated. GRL Task characteristics F containing only a small amount of identity information task Integrate task feature compensation terms into task feature F task In the process, the fused task features F are obtained. total The model constraints are: min(λ1l cls +λ2l GRL +λ3l dcls -λ4rho-λ5l acls +λ6l tcls ) In the formula, l cls For identity-level task classification loss, l GRL For gradient reversal loss, l dcls For identity classification loss, rho is the correlation coefficient, and l acls For the loss of disadvantageous task characteristics, l tcls The loss represents the advantageous task characteristics, and λ1 to λ1 represent the trade-off coefficients of the corresponding losses.