EEG-based frequency stripe high-order map emotion recognition method

By constructing an EEG-based high-order image emotion recognition method, the problem of inaccurate feature extraction and low recognition in EEG-enzyme emotional recognition is solved, and a more accurate emotion recognition effect is achieved.

CN120234673AActive Publication Date: 2025-07-01TAIYUAN UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202510393207.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-01
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The prior art has problems in the emotional recognition of EEG signals, such as inaccurate feature extraction and low recognition, and lacks a general EEG emotional recognition model.

Method used

The EEG-based high-order graph emotion recognition method is adopted, and dynamic time-varying characterization is extracted through EEG signal acquisition and preprocessing, and a higher-order graph based on the minimum spanning tree is constructed. The GIN encoding module and the node clustering graph pooling module are used to perform emotion recognition in combination with the deep learning model.

Benefits of technology

A more accurate classification detection of emotion recognition is achieved, the ability to represent emotional states is enhanced, and the limitations of traditional methods in time-domain and frequency-domain feature extraction are overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234673A_ABST
    Figure CN120234673A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of EEG (electroencephalogram) emotion recognition, and particularly relates to an EEG-based frequency stripe high-order map emotion recognition method, which comprises the following steps: data acquisition and preprocessing, a dynamic time-varying representation extraction module, an MST-based high-order map construction module, a GIN coding module, a node clustering map pooling module and an emotion classification module. The SEED data set is used for model training and evaluation. Firstly, the periodicity of a time sequence is extracted by using a dynamic time-varying representation extraction module; then, a high-order graph construction module based on MST is used for dynamically capturing space-time information of the EEG signals; thirdly, a GIN coding module and a node clustering graph pooling module are adopted to learn the complex relation between the electrode nodes, and the spatial features of the electrode nodes are aggregated, so that effective graph representation is obtained; and finally, performing sentiment classification through a dynamic linear layer. The method provides a novel thought for dynamic analysis of emotion response, and has a wide prospect in the application of emotion analysis and emotion regulation technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electroencephalogram (EEG) signal emotion recognition, and in particular relates to an EEG-based ripple high-order graph emotion recognition method. Background Art

[0002] Emotion recognition is an important module in the fields of human-computer interaction, diagnosis and rehabilitation of mental illness, transportation, etc. The study of emotions is an advanced stage of artificial intelligence, which will greatly promote the development of emotional robots, anthropomorphic control theory, etc. The EEG-based emotion recognition method has the advantages of being difficult to disguise and having a higher temporal resolution than other physiological signals. In recent years, more and more researchers have used EEG for emotion recognition research. The accuracy and reliability of emotion recognition are closely related to the selection of EEG features and the construction of deep learning models.

[0003] EEG data is a non-stationary time-varying signal that is difficult to analyze directly. It is necessary to extract multivariate features from the EEG data. In terms of EEG feature selection, the more common traditional EEG features include time domain and frequency domain features, but in current research, there are some limitations to using only one feature. Time domain features reflect the timing characteristics of the EEG, but there are some empirical components in the extraction process; frequency domain features analyze the distribution of energy values ​​contained in the EEG, which can better express the physiological significance of the EEG, but the recognition of frequency domain features will be low. The feature information representation ability of a single domain is not strong, and how to accurately extract effective and non-redundant time-frequency-space multivariate EEG emotional representation has not yet been solved.

[0004] Existing studies have shown that traditional neural networks cannot directly process non-Euclidean data. EEG is discrete and discontinuous in the spatial domain. It is more advantageous to construct EEG emotion graph structures based on graph theory knowledge and use graph neural networks to process information in the graph domain, which can better describe the intrinsic relationship between channels. Feature extraction of EEG signals and construction of brain networks, combined with deep learning model recognition methods, can reflect the interaction between different brain regions under different emotional states and show the specific state of the brain, which can help understand the working principles of the brain and obtain the brain's emotional cognitive mechanism. However, most existing studies focus on brain structural characteristics and combine functional connectivity indicators to model EEG as a graph structure, ignoring the graph structure representation that integrates physiological significance. There is no general EEG emotion recognition model for obtaining the brain's emotional cognitive mechanism. Summary of the invention

[0005] In response to the technical problems existing in the above-mentioned traditional neural networks, the present invention provides an EEG-based ripple high-order graph emotion recognition method to achieve more accurate classification detection for emotion recognition.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0007] A method for emotion recognition based on EEG ripple high-order graphs, comprising the following steps:

[0008] S1. EEG signal acquisition and preprocessing. A bandpass filter bank is used to extract the α, β, γ, δ, and θ bands of the original EEG signal, and a sliding window is performed on each frequency band to extract DE features.

[0009] S2, dynamic time-varying representation extraction module, is used to decompose complex time changes into multiple intra-cycle and inter-cycle changes, so that it can more effectively capture the complexity and periodicity of time series in 2D space;

[0010] S3, MST-based high-order graph construction module, the module consists of an initial graph construction layer and an MST-based high-order graph construction layer;

[0011] S4, GIN encoding module, is used to learn the complex relationship between various electrode nodes and aggregate their spatial features to obtain an effective graph representation;

[0012] S5, node clustering graph pooling module, selectively downsamples nodes from the graph structure to reduce the number of nodes in the graph and retain the most representative structural information

[0013] S6, emotion classification module, in each stage, each subject's sample is used as the test set, and the remaining

[0014] The data standards used in S1 include: having complete scale information.

[0015] The dynamic time-varying representation extraction module in S2 integrates a data conversion layer, a dynamic multi-scale feature extraction layer and an adaptive feature aggregation layer;

[0016] The data conversion layer uses fast Fourier transform to transform the input signal X 1D Perform a frequency domain transformation and find the period and the changes between periods, revealing the strength of its frequency components; perform a 1D FFT on each time step:

[0017] X f =FFT(X 1D )

[0018]

[0019] Where: X 1D represents the input one-dimensional time domain signal, FFT represents the fast Fourier transform, which converts the time domain signal into the frequency domain signal; X f The input signal X 1D Frequency domain representation; Padding(X 1D ) indicates padding the input signal; Denotes reshaping the padded signal into a two-dimensional tensor with dimensions p i ×f i , where p i is the number of time steps and f i is the number of frequency points. Each branch i ∈ {1,..., k} corresponds to different p i and f i , achieving multi-scale analysis; Denotes the two-dimensional tensor generated by the i-th branch;

[0020] The dynamic multi-scale feature extraction layer uses multi-scale one-dimensional convolution operations to perform feature extraction on the obtained 2D tensor and learn its rich time information;

[0021]

[0022] Where: Denotes a one-dimensional convolution operation with a kernel size of t; LeakyReLU denotes a leaky rectified linear unit activation function; AvgPool denotes an average pooling operation; Denotes the multi-scale features extracted by the i-th branch;

[0023] Finally, the learned 2D tensor is converted back to 1D space for dimensional reduction and reshaping, which is defined as:

[0024]

[0025] Where: Denotes flattening the two-dimensional features into one dimension with dimensions 1 × p i ×f i ; Trunc denotes a truncation operation to adjust the signal length to be consistent with the original input; Denotes the one-dimensional feature representation output by the i-th branch;

[0026] The adaptive feature aggregation layer fuses the output results of the feature extraction layer, fusing k one-dimensional representations and provides them to the next module; its definition is:

[0027]

[0028] Where: A f Denotes learnable attention weight parameters used to measure the importance of different scale features; Softmax(A f ) denotes normalizing A f to generate weight coefficients that sum to 1; Denotes the final output after aggregation.

[0029] The MST-based high-order graph construction module in S3 consists of an initial graph adjacency matrix construction layer, an initial graph edge embedding construction layer, and an MST-based high-order graph construction layer;

[0030] In the initial graph adjacency matrix construction layer, each electrode of the input signal is regarded as a single node in the constructed brain network, and the learned dynamic time-frequency representation of a single electrode is regarded as the node attribute; according to the characteristic that the relationship between two electrode nodes is mutual, the adjacency matrix of the undirected initial graph assuming global connection is defined as A initial ∈R C×C , where C is the number of electrodes, that is, the total number of nodes in the graph; the embedding of the edge in the initial graph edge embedding construction layer is calculated by concatenating the embeddings of the source node and the target node;

[0031]

[0032] Where: represents the concatenation result of the source node and target node features; W3, b3, W4, b4 are respectively learnable weight matrices and bias parameters for linear transformation; σ ReLU represents the ReLU activation function; σ Sigmoid represents the Sigmoid activation function, which compresses the output to the range [0,1][0,1], representing the weight or importance of the edge; e st represents the embedding representation of the edge (s,t);

[0033] In the MST-based high-order graph construction layer, let represent the graph constructed based on EEG data, V represents the set of nodes, represents the one-dimensional feature representation output by the i-th branch, E represents the set of edges, and the minimum spanning tree is calculated for the graph G. Its goal is to find a subgraph that contains all nodes, such that the subgraph is connected and the sum of the edge weights is minimized, that is, to find a subgraph such that connectivity, all nodes in the graph T are connected; acyclicity, the graph T does not contain any cycles; minimum weight, the sum of the edge weights in the graph T is the minimum.

[0034] The GIN encoding module in S4 for the high-order graph learns the complex relationships between each electrode node through GIN, aggregates its spatial features, and thus obtains an effective graph representation; GIN uses a message passing process to combine the information of neighbor nodes to update the representation of the nodes; the message passing process can be referred to the following formula: for the k-th layer in GIN, the formula for node update in the GIN model is expressed as:

[0035]

[0036] Wherein: represents the set of adjacent nodes of a specific node v ∈ V, and f (k) is a function that converts adjacent node features and edge weights into an aggregated vector, and g (k) is a trainable function that maps the current node representation form and the aggregated vector to a new representation form;

[0037] For the feature corresponding to the k-th layer node v in the graph it will be updated according to its neighbor nodes and its own information, and finally a node representation with more global information is obtained.

[0038] In the S5, the node clustering graph pooling module uses HGP-SL, combines with the graph neural network, and through the graph pooling function of the HGP-SL operator, maximally retains node information, removes redundant noise, aggregates information from less important nodes to more important nodes, so as to learn a more refined graph structure.

[0039] Aggregate the outputs corresponding to each frequency band and feed them to the fully connected layer to obtain the final classification result:

[0040]

[0041] Wherein: GeLU(·) is an activation function, which helps the gradient descent optimization algorithm to converge more easily; Dropout(·) randomly sets the outputs of some neurons to zero during the training process, thereby reducing the over-dependence of the neural network on certain specific neurons and improving the generalization ability of the model; Linear(·) is the fully connected layer to obtain the final prediction result.

[0042] Compared with the prior art, the beneficial effects of the present invention are:

[0043] 1. The present invention extracts signals of five frequency bands of α, β, γ, δ, and θ through a band-pass filter bank and calculates the dynamic entropy features of each frequency band, fully capturing the specific contributions of different frequency bands in emotion recognition. Compared with single-frequency band analysis, multi-frequency band joint modeling can more comprehensively reflect the time-frequency characteristics of brain activities and enhance the representation ability of emotional states.

[0044] 2. The present invention converts the time-domain signal into the frequency domain through the fast Fourier transform (FFT) to reveal the periodicity and frequency component intensity of the signal, overcoming the limitations of traditional time-domain methods in modeling complex time patterns. The multi-scale Reshape operation decomposes the signal into two-dimensional tensors with different time-frequency resolutions, capturing the synergistic effects of short-term local fluctuations and long-term periodic changes.

[0045] 3. The multi-scale one-dimensional convolution (Conv1D) of the present invention is combined with the LeakyReLU activation function to extract local features under different time spans and enhance the model's adaptability to temporal dynamics. The adaptive feature aggregation layer dynamically fuses multi-scale features through the Softmax attention mechanism, suppresses redundant information and highlights key ripple patterns, thereby improving the robustness of feature expression.

[0046] 4. The present invention uses electrodes as nodes and dynamic time-frequency features as node attributes to construct an initial graph of global connections and retain the potential correlation between electrodes. The edge embedding calculation learns edge weights through the ReLU-Sigmoid dual activation structure, combines node feature splicing, quantifies the functional connection strength between electrodes, and enhances the interpretability of edge weights.

[0047] 5. The present invention aggregates node features and neighbor node information through multi-layer MLP to achieve layer-by-layer update of node features and effectively model the nonlinear interaction relationship between electrodes. Combined with weighted aggregation of edge weights, it highlights the contribution of important connections to emotion recognition and improves the discriminability of spatial feature learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.

[0049] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.

[0050] Figure 1 A schematic diagram of a method for emotion recognition based on EEG ripple high-order graphs according to an embodiment of the present invention;

[0051] Figure 2 Schematic diagram of the dynamic time-frequency feature extraction process in an embodiment of the present invention;

[0052] Figure 3 A schematic diagram of the principle of constructing an initial graph in an embodiment of the present invention;

[0053] Figure 4 Schematic diagram of the GIN encoding and node clustering graph pooling process in an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. These descriptions are only to further illustrate the features and advantages of the present invention, rather than to limit the claims of the present invention. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0055] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0056] The embodiment of the present invention relates to an EEG-based ripple high-order graph emotion recognition method, the flow chart is as follows: Figure 1 As shown, the following steps are included:

[0057] The method for emotion recognition based on EEG ripple high-order graph is characterized by comprising:

[0058] Step S1: EEG signal acquisition and preprocessing. A bandpass filter bank is used to extract the α, β, γ, δ, and θ bands of the original EEG signal, and a sliding window is performed on each frequency band to extract DE features.

[0059] Step S2: Dynamic time-varying representation extraction module is used to decompose complex time changes into multiple intra-cycle and inter-cycle changes, so that it can more effectively capture the complexity and periodicity of the time series in 2D space.

[0060] Step S3: Construct a module based on the MST-based high-order graph. The module consists of the initial graph and the MST-based high-order graph.

[0061] Step S4: GIN encoding module. It is used to learn the complex relationship between various electrode nodes and aggregate their spatial features to obtain an effective graph representation.

[0062] Step S5: Node clustering graph pooling module. Selectively downsample nodes from the graph structure to reduce the number of nodes in the graph and retain the most representative structural information.

[0063] Step S6: Emotion classification module. In each stage, the samples of each subject are taken as the test set in turn, and the samples of the remaining subjects are combined into the training set.

[0064] The specific steps are as follows:

[0065] In step S2, ifFigure 2 As shown, the dynamic time-varying representation extraction module integrates a data conversion layer, a dynamic multi-scale feature extraction layer, and an adaptive feature aggregation layer.

[0066] In the data conversion layer, the Fast Fourier Transform (FFT) is used to perform a frequency-domain conversion on the input signal X 1D and find the changes between periods and periods, revealing the intensity of its frequency components. Specifically, a 1D FFT is performed on each time step:

[0067] X f = FFT(X 1D )

[0068] where FFT(·) represents performing an FFT transformation on the signal along the time dimension of the EEG feature, and the resulting frequency-domain signal X f is a complex number. In this layer, we focus on its amplitude.

[0069] Next, the mean value of the signal amplitude at each frequency in the frequency domain is calculated for each channel:

[0070] A f = Avg(Amp(X f ))

[0071] In the formula, Amp(·) is used to calculate the amplitude value of each frequency, and Avg(·) is used to take its average to obtain the mean amplitude value A f of each frequency component. In particular, the DC component is set to zero.

[0072] According to the calculated frequency mean, the frequency indices of the top k maximum amplitudes are selected. For ease of representation, the selected frequency indices are defined as f i , and the k most significant frequency components {f1, f2..., f k} in each electrode channel and frequency band are selected.

[0073] f i = arg Topk(A f )

[0074] where f i is the set of frequency indices corresponding to the top k maximum amplitudes selected. According to the frequency index f i , the period of each selected frequency can be calculated. Assuming that each selected frequency is {f1, f2..., f k} (i ∈ {1,..., k}), then the period p i (i ∈ {1,..., k}) of each frequency can be calculated by the following formula:

[0075]

[0076] where p i is the period corresponding to the frequency f i , representing the repeating pattern of the signal in the time domain, and L is the length of the time series.

[0077] According to the conjugate property of the frequency domain, only the frequencies within are considered during the calculation. Therefore, the X 1D time series can be reconstructed into multiple X 2D tensors, as shown in the formula:

[0078]

[0079] For the period corresponding to each most significant frequency component (main frequency component), zero-padding is performed along the time dimension through the Padding(·) operation, making its length a multiple of the period, which is convenient for conversion in the 2D space. And the padded data is reshaped into a 2D form using the Reshape(·) operation to obtain where p i , f i is the period of each frequency component, and the two-dimensional tensor represents the two-dimensional variations at k different times from different periods.

[0080] The dynamic multi-scale feature extraction layer uses multi-scale one-dimensional convolution operations to perform feature extraction on the obtained 2D tensor and learn its rich time information.

[0081]

[0082] As shown in the above formula, dynamic features are obtained by applying 1D convolution kernels of different sizes to the samples one by one. Among them, (1, t) is the size of the convolution kernel, is the input sample, and Conv1D() is the one-dimensional convolution operation on the input sample with a convolution kernel of (1, t) and a convolution stride of (1, 1). Secondly, the LeakyRe LU(·) activation function is used in the convolution operation, and the feature map is downsampled through the average pooling function AvgPool(·) to reduce the influence of noise and feature dimensions on the signal quality.

[0083] Finally, the learned 2D tensor is converted back to the 1D space for dimensional reduction and reshaping, which is defined as:

[0084]

[0085] Among them, since the Padding(·) operation is performed from a 1D space to a 2D space, Trunc(·) is used to truncate the length to the original length. During the data conversion process, due to different periods, the shapes of the obtained 2D tensors are also different. Therefore, when using Conv1D(·) for feature extraction, the same-sized convolutional kernels can be used in the 2D tensors corresponding to different epochs to achieve weight sharing and thus improve efficiency.

[0086] The adaptive feature aggregation layer fuses the output results of the feature extraction layer into k one-dimensional representations and provides them to the next module. The amplitude A calculated in the data conversion layer f can reflect the relative importance of the selected significant frequencies and the corresponding periods. Therefore, the Softmax(·) function is applied to each frequency, and then it is multiplied by the corresponding and summed. It is defined as:

[0087]

[0088] In the formula, the period weights are applied to the output of each period. Subsequently, the outputs of all k periods are adaptively aggregated. The amplitude A f is used as the weight coefficient after the Softmax(·) operation for the aggregation operation. Taking the aggregated tensor after residual connection as the input again, multi-layer extraction can be performed to obtain deeper information, or it can be directly transmitted to the next layer.

[0089] In step S3, as Figure 3 shown, it consists of an initial graph adjacency matrix construction layer, an initial graph edge embedding construction layer, and an MST-based high-order graph construction layer.

[0090] Each electrode of the input signal in the initial graph adjacency matrix construction layer is regarded as a single node in the constructed brain network, and the learned dynamic time-frequency representation of a single electrode is regarded as the node attribute. The size of the initial graph depends on how many electrode channels the input EEG signal contains (depending on the EEG device used during acquisition). In this embodiment, the dot product between the electrode node attributes (features) within each frequency band is calculated to reflect the relationship between the electrode nodes. It should be noted that the similarity adjacency matrix is dynamic and instance-specific.

[0091] According to the property that the relationship between two electrode nodes is mutual. Define the adjacency matrix A of the undirected initial graph assuming global connection initial ∈R C×C as:

[0092]

[0093] The above formula constructs the adjacency matrix of the graph by calculating the similarity between nodes, where is the dot product, h i , i ∈ {1, 2,... C} are the feature vectors generated by each electrode node after passing through the dynamic time-varying feature extraction module.

[0094] The initial graph edge embedding construction layer calculates the embedding of the edge by concatenating the embeddings of the source node and the target node. First, a graph convolutional network is used to update the node embeddings, and here the message passing process of the Graph Isomorphism Network (GIN) is adopted. Each node feature h i in the graph will be updated according to the information of its neighbor nodes and itself. Suppose the embeddings of the source node and the target node are obtained as h s and h t , respectively. Then the edge embedding e st is obtained by concatenating the embeddings of the source node and the target node:

[0095] e st = g(Concat(h s , h t ))

[0096] That is, the learned node features are concatenated to create the embedded edge features Here, the aggregation function g(·) uses a multi-layer perceptron MLP, which consists of a linear layer with trainable weights , a ReLU activation function σ ReLU , and a linear layer with a weight of b4 ∈ R, and the Sigmoid function is used to convert the result to the range (0, 1):

[0097]

[0098] In the high-order graph construction layer based on MST, let represent the graph constructed based on the EEG data, where V = {v i : v = 1, 2,..., C} represents the set of nodes. A initial = [a st : s, t ∈ V] ∈ {0, 1} C×C is the adjacency matrix describing its connection information, and a st = 1 indicates that there is an edge between nodes, otherwise there is no connection. is the set of node attributes, representing the attributes corresponding to each node v i , and E = {e st : s, t ∈ V} ∈ R C×Cis a set of edge weights, representing the connectivity strength between nodes, and C represents the number of EEG electrode channels.

[0099] Calculate the Minimum Spanning Tree (MST) for graph G. Its goal is to find a subgraph that contains all nodes, such that this subgraph (the minimum spanning tree) is connected and the sum of the edge weights is minimized. That is, find a subgraph such that for connectivity, all nodes in graph T are connected; for acyclicity, graph T does not contain any cycles; for minimum weight, the sum of the edge weights in graph T is the minimum.

[0100] In step S4, as Figure 4 shown, for high-order graphs use GIN to learn the complex relationships between each electrode node, aggregate its spatial features, and thus obtain an effective graph representation. GIN uses a message-passing process to update the representation of nodes by combining the information of neighboring nodes. The message-passing process can be referred to the following formula:

[0101]

[0102] where represents the set of adjacent nodes of a specific node v ∈ V, and f (k) is a function that converts the adjacent node features and edge weights into an aggregated vector, and g (k) is a trainable function that maps the current node representation form and the aggregated vector to a new representation form. Here, f (k) is the weighted sum of the node features and edge weights, and g (k) is an MLP layer that updates the representation of nodes through a non-linear mapping. Therefore, for the k-th layer in GIN, the formula for node update (message passing) in the GIN model can be expressed as:

[0103]

[0104] That is, for the feature corresponding to the k-th layer node v in the graph will be updated according to the information of its neighboring nodes

[0105] and its own information, and finally obtain a node representation with more global information.

[0105] The above process can be written in matrix form:

[0106]

[0107] where H (k) is the matrix of behavior and denotes the Hadamard product (element - wise multiplication), 1 is an M - dimensional vector consisting of all 1s, and is a trainable weight matrix, where, d (0) = 3 and d (1) = d, and BN(·) represents the batch normalization operation.

[0108] In step 5, in order to perform selective down - sampling on nodes from the graph structure to reduce the number of nodes in the graph and retain the most representative structural information, this module uses HGP - SL, combined with a graph neural network. Through the graph pooling function of the HGP - SL operator, it maximally retains node information, removes redundant noise, aggregates information from less important nodes to more important nodes, and learns a more refined graph structure.

[0109] Specifically, the importance score of each node is calculated, and scores from multiple sources are used here: the degree centrality of the node, the dot product of the feature and the weight, and the PageRank score. For node v i , the importance score s i can be represented by the following formula:

[0110]

[0111]

[0112] The above formula calculates the node degree centrality score, the node feature importance score, and the node PageRank score.

[0113] Among them, σ Sigmoid (·) is the sigmoid activation function, deg(v i ) represents the degree centrality of node v i , α is the degree centrality weighting coefficient, controlling the influence degree of the degree centrality on the node score, β is used as an offset to adjust the sensitivity of node selection, ε is a small positive constant in logarithm (e.g., 10 - 6), and H (k) is the node representation matrix of the k - th layer.

[0114] These scores are combined into a vector s to obtain the scores of all nodes, and w 1 , w 2 , w 3 are trainable weights:

[0115]

[0116] Based on the score s, select the nodes that the pooling operator should retain. Here, the top-rank (·) operation is used to implement this. Retaining nodes with relatively large node information scores better preserves the core information of the graph and obtains a refined graph structure. Specifically, first reorder the nodes in the graph according to the node information score, and then select the top-ranked node subset. The node ratio selected by pooling is r, that is, r×C nodes are retained:

[0117]

[0118] top-rank(·) means return The index function of and Indicates row or (and) column extraction to form the node representation matrix and adjacency matrix of the subgraph. Finally, we get and Represents the node features and graph structure information of the next layer.

[0119] In step S4, it includes:

[0120] Aggregate the outputs corresponding to each frequency band and feed them into the fully connected layer to get the final classification result:

[0121]

[0122] Where GeLU(·) is the activation function, which helps the gradient descent optimization algorithm to converge more easily.

[0123] Dropout(·) randomly sets the output of some neurons to zero during the training process, thereby reducing the excessive dependence of the neural network on certain specific neurons and improving the generalization ability of the model. Linear(·) is a fully connected layer to obtain the final prediction result.

[0124] The emotion classification model of the present invention is evaluated by using the average accuracy of leave-one-out cross-validation at the subject level.

[0125] The present invention provides an EEG-based ripple high-order graph emotion recognition method. Experimental results show that the classification accuracy of the model on the SEED dataset reaches 94.67%, proving the effectiveness and innovation of the method in emotion recognition. Evaluation on the SEED dataset shows that the EEG-based ripple high-order graph emotion recognition method is superior to the current state-of-the-art method and shows stronger effects in emotion recognition. The model has key modules such as dynamic time-frequency feature extraction and high-order graph construction. It can adapt to the data heterogeneity of different subjects, has strong generalization ability, and shows application potential in EEG signal analysis and emotion recognition.

[0126] The above only elaborates in detail on the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention, and all such changes should be included within the protection scope of the present invention.

Claims

1. A method for emotion recognition based on EEG ripple high-order graph, characterized in that: The following steps are involved: S1. EEG signal acquisition and preprocessing. The α, β, γ, δ and θ bands of the original EEG signal are extracted using a bandpass filter bank, and each frequency band is windowed to extract DE features; S2, dynamic time-varying representation extraction module, is used to decompose complex time changes into multiple intra-cycle and inter-cycle changes, so that it can more effectively capture the complexity and periodicity of time series in 2D space; S3, MST-based high-order graph construction module, the module consists of an initial graph construction layer and an MST-based high-order graph construction layer; S4, GIN encoding module, is used to learn the complex relationship between various electrode nodes and aggregate their spatial features to obtain an effective graph representation; S5, node clustering graph pooling module, selectively downsamples nodes from the graph structure to reduce the number of nodes in the graph and retain the most representative structural information S6, emotion classification module, in each stage, each subject’s sample is taken as the test set in turn, and the samples of the remaining subjects are combined into the training set.

2. The method for emotion recognition based on EEG ripple high-order graph according to claim 1, characterized in that: The data standards used in S1 include: having complete scale information.

3. The method for emotion recognition based on EEG ripple high-order graph according to claim 1, characterized in that: The dynamic time-varying representation extraction module in S2 integrates a data conversion layer, a dynamic multi-scale feature extraction layer and an adaptive feature aggregation layer; The data conversion layer uses fast Fourier transform to transform the input signal X 1D Perform a frequency domain transformation and find the period and the changes between periods, revealing the strength of its frequency components; perform a 1D FFT on each time step: X f =FFT(X1D) Where: X 1D represents the input one-dimensional time domain signal, FFT represents the fast Fourier transform, which converts the time domain signal into the frequency domain signal; X f The input signal X 1D Frequency domain representation; Padding(X 1D ) indicates padding the input signal; Reshape the padded signal into a two-dimensional tensor with dimension p i ×f i , p i is the number of time steps, f i is the number of frequency points, each branch i∈{1,...,k} corresponds to a different p i and f i , achieving multi-scale analysis; Represents the two-dimensional tensor generated by the i-th branch; The dynamic multi-scale feature extraction layer uses a multi-scale one-dimensional convolution operation to obtain a 2D tensor Perform feature extraction and learn its rich temporal information; in: Represents a one-dimensional convolution operation with a convolution kernel size of t; LeakyReLU represents a leaky rectified linear unit activation function; AvgPool represents an average pooling operation; Represents the multi-scale features extracted by the i-th branch; Finally, the learned 2D tensor is converted back to 1D space Perform dimension reduction and reshaping, which is defined as: in: Indicates flattening the two-dimensional feature into one dimension, with a dimension of 1×p i ×f i ;Trunc represents the truncation operation, which adjusts the signal length to be consistent with the original input; Represents the one-dimensional feature representation of the output of the i-th branch; The adaptive feature aggregation layer combines the output of the feature extraction layer with k one-dimensional representations. and provided to the next module; it is defined as: Among them: A f Represents a learnable attention weight parameter, which is used to measure the importance of features at different scales; Softmax(A f ) indicates that A f Normalize and generate weight coefficients to satisfy the sum of 1; Represents the final output after aggregation.

4. The method for emotion recognition based on EEG ripple high-order graph according to claim 1, characterized in that: The MST-based high-order graph construction module in S3 consists of an initial graph adjacency matrix construction layer, an initial graph edge embedding construction layer, and an MST-based high-order graph construction layer; Each electrode of the input signal in the initial graph adjacency matrix construction layer is regarded as a single node in the constructed brain network, and the learned dynamic time-frequency representation of a single electrode is regarded as a node attribute; according to the characteristic that the relationship between two electrode nodes is mutual; define the adjacency matrix of the undirected initial graph assuming global connection as A initial ∈R C×C , C is the number of electrodes, i.e., the total number of nodes in the graph; the embedding of the initial graph edge embedding construction layer edge is calculated by concatenating the embeddings of the source node and the target node; in: Represents the concatenation result of the source node and the target node features; W3, b3, W4, b4 are respectively the learnable weight matrix and bias parameters used for linear transformation; σ ReLU represents the ReLU activation function; σ Sigmoid represents the Sigmoid activation function, which compresses the output to the range of [0,1][0,1], indicating the weight or importance of the edge; e st represents the embedding representation of edge (s, t); In the MST-based high-order graph construction layer, let represents a graph constructed based on EEG data, V represents a node set, represents the one-dimensional feature representation of the output of the i-th branch, E represents the set of edges, and the minimum spanning tree is calculated for the graph G. Its goal is to find a subgraph containing all nodes so that the subgraph is connected and the sum of the edge weights is the smallest, that is, to find a subgraph To ensure connectivity, all nodes in graph T are connected; acyclicity, graph T does not contain any cycles; minimum weight, the sum of the weights of the edges in graph T is the smallest.

5. The method for emotion recognition based on EEG ripple high-order graph according to claim 1, characterized in that: The GIN encoding module in S4 is used for high-order graphs GIN is used to learn the complex relationship between various electrode nodes and aggregate their spatial features. In this way, an effective graph representation is obtained; GIN uses a message passing process to update the representation of a node in combination with the information of neighboring nodes; the message passing process can be referred to the following formula: For the kth layer in GIN, the formula for node update in the GIN model is expressed as: in: represents the set of neighboring nodes of a specific node v∈V, f (k) is a function that converts adjacent node features and edge weights into an aggregation vector, g (k) is a trainable function that maps the current node representation and the aggregate vector to a new representation; For the feature corresponding to the k-th layer node v in the graph According to its neighbor nodes and its own information to update, and finally obtain a node representation with more global information.

6. The method for emotion recognition based on EEG ripple high-order graph according to claim 1, characterized in that: The node clustering graph pooling module in S5 uses HGP-SL, which is combined with a graph neural network. Through the graph pooling function of the HGP-SL operator, the node information is retained to the greatest extent, redundant noise is removed, and information is aggregated from less important nodes to more important nodes to learn a more refined graph structure.

7. The method for emotion recognition based on EEG ripple high-order graph according to claim 1, characterized in that: The output corresponding to each frequency band is aggregated and fed into the fully connected layer to obtain the final classification result: Among them: GeLU(·) is an activation function, which helps the gradient descent optimization algorithm converge more easily; Dropout(·) is to randomly set the output of some neurons to zero during the training process, thereby reducing the excessive dependence of the neural network on certain specific neurons and improving the generalization ability of the model; Linear(·) is a fully connected layer to obtain the final prediction result.

Citation Information

Patent Citations

  • Variable-scale symbolized compensation transfer entropy-based emotion-induced electroencephalogram signal analysis method

    CN112244880A

  • Electroencephalogram emotion recognition method of dynamic graph attention network model based on multi-branch feature extraction and staged fusion

    CN118885868A

  • System and method for neuroenhancement to enhance emotional response

    WO2019133997A1

Cited By

  • Prediction entropy feedback electroencephalogram data clustering model closed-loop iteration updating device and method

    CN121834257A