Multi-modal emotion recognition method based on heart and brain coupling and graph neural network

By constructing a multimodal emotion recognition method based on heart-to-brain coupling, using graph neural network and domain adversarial strategies, the problem of insufficient generalization caused by neglect of central brain signal association and individual differences in the existing technology is solved, and more efficient emotion recognition and generalization are achieved.

CN120217050APending Publication Date: 2025-06-27JILIN UNIVERSITY
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510314800.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art ignores the physiological association of heart and brain signals in emotional recognition, resulting in insufficient modeling of spatiotemporal representations and individual autonomic nerve differences, resulting in insufficient generalization of physiological signals across subjects.

Method used

The multimodal emotion recognition method based on heart-brain coupling and graph neural network is adopted to extract the frequency domain and time domain characteristics of EEG signals and ECG signals through data preprocessing, and a frequency domain functional connection diagram and a time domain data-driven connection diagram are constructed. Multimodal representation is captured using multi-view diagram convolution network and fusion diagram network, and the generalization of the model is enhanced by combining domain adversarial strategies.

Benefits of technology

It improves the accuracy and generalization of emotion recognition, can more effectively model the space-time dynamic correlation of heart and brain emotional states, and overcomes the impact of individual differences on model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217050A_ABST
    Figure CN120217050A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal emotion recognition method based on heart and brain coupling and a graph neural network, and belongs to the field of artificial intelligence. Comprising the steps of data preprocessing, graph representation construction, multi-view graph convolutional network construction, fusion graph network construction and cross-domain joint optimization and sentiment classification. The method has the advantages that an adaptive adjacency matrix optimization strategy based on a triple constraint mechanism is proposed to solve the modal alignment and deviation problems represented by a multi-modal diagram in a data-driven branch, redundant noise is eliminated by adopting global regularization constraint, and unique feature representation in a modal is enhanced through modal specificity; a deep association rule is mined in combination with a cross-modal interaction module, the modeling ability of a heart and brain emotional state is improved, a multi-view image convolutional network is further designed, global features and local features are extracted, features of a cognitive heuristic branch and a data driven branch are combined by adopting an attention mechanism-based image fusion network, a domain confrontation strategy is introduced, and a cognitive network is constructed. And the generalization of the method is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and relates to technologies such as emotion computing and signal processing. Specifically, it is a multi-modal emotion recognition method based on heart-brain coupling and graph neural network. Background Art

[0002] Emotion computing is an important research direction in human-computer interaction, and has broad application prospects in the auxiliary diagnosis of various mood disorders such as depression, post-traumatic stress, and anxiety disorders. The emotion model theory has experienced a paradigm evolution from discrete to continuous. Traditional discrete emotion models simplify emotions into labeled classifications such as happiness, anger, sadness, and joy, but it is difficult to characterize the gradual change and complexity of emotions. Therefore, emotion models based on continuous dimensions such as valence and arousal have gradually become the mainstream.

[0003] With the rapid development of the field, emotion recognition based on electroencephalogram (EEG) signals has become a research hotspot in emotion computing. However, due to the high time-variability of EEG signals and the fact that a single modality can only reflect local dimensions of emotional responses, there are representation blind spots. Therefore, combining peripheral physiological signals can construct complementary representations. Among them, electrocardiogram (ECG) signals have attracted much attention due to their easy acquisition and clear physiological significance. The cardiac autonomic nervous system and the central nervous system form a bidirectional regulation through the brain-heart coupling mechanism: the activity of the prefrontal cortex regulates heart rate variability through the vagus nerve, and the change of cardiac rhythm will in turn affect the energy distribution of EEG frequency bands. This dynamic synergistic effect has a key impact on emotion generation.

[0004] The core of multi-modal sentiment analysis lies in constructing a highly discriminative fused representation. Existing research has limitations: only shallow feature concatenation is performed, and the physiological prior knowledge of brain-heart coupling is not embedded, resulting in the lack of modeling of spatio-temporal dynamic associations; secondly, the differences in individual autonomic regulation mechanisms lead to significant inter-subject variability in physiological signals, while mainstream algorithms assume that the training / test data are identically distributed, causing a sharp decline in model performance in the cross-subject scenario. Xefteris et al. combined EEG-based features with peripheral physiological signals for emotion recognition in the paper "Graph theoretical analysis of eeg functional connectivity patterns and fusion with physiological signals for emotion recognition". The extracted graph features were fused with the statistical features of the peripheral signals and input into a traditional classifier to achieve emotion recognition. However, this method does not explicitly model the brain-heart interaction mechanism, limiting the discriminability of the representation. Gagliardi et al. proposed BHI-XAI-CNN in the paper "Fine-grained emotion recognition using brain-heart interplay measurements and explainable convolutional neural networks", which achieved 9-level valence recognition by encoding brain-heart interaction as a sparse image and used an explainable convolutional neural network for classification. However, this method does not fully consider the specificity of physiological signals, which limits the generalization of the model. Summary of the Invention

[0005] The present invention provides a multi-modal emotion recognition method based on brain-heart coupling and graph neural network to solve the problems of insufficient spatio-temporal representation modeling caused by ignoring the physiological association of brain-heart signals and insufficient generalization of physiological signals across subjects due to individual autonomic differences, thereby improving the accuracy and generalization of emotion recognition.

[0006] The technical solution adopted by the present invention includes the following steps:

[0007] (1) Data preprocessing, obtaining electroencephalogram signals and electrocardiogram signals from a multi-modal emotion database, extracting frequency-domain feature X f and time-domain feature X t , and constructing a functional connectivity matrix A using the spectral coherence coefficient f to represent the brain-heart coupling strength;

[0008] (2) Constructing a graph representation, constructing a frequency-domain functional connectivity graph (X f , A f), and the time-domain data-driven connection graph (X t , A t ) of the two-stream representation, capturing the inherent features and dynamic correlation features, forming a complementary and collaborative representation;

[0009] (3) Construct a multi-view graph convolutional network

[0010] The multi-view graph convolutional network adopts a double-branch isomorphic architecture. Although the two branches have the same network design but independent parameters: one branch uses the frequency-domain functional connection graph (X f , A f ) as the input, and the other branch processes the time-domain data-driven graph (X t , A t ). Each branch decouples the physiological and anatomical subgraphs from the global graph (X a , A a ) through the microscopic aggregation module and the emotion-induced subgraph while retaining the topological structure of the global graph, where a ∈ {t, f}, t represents the time-domain data-driven branch, and f represents the frequency-domain functional connection branch. On this basis, discriminative spatio-temporal features are extracted from the two subgraphs and the global graph through the graph convolutional network and the cascaded readout function respectively. Finally, the spatio-temporal features of the subgraph and the global graph are concatenated to form the branch representation Z a ;

[0011] (4) Construct a fusion graph network

[0012] The fusion graph network uses the attention mechanism to merge the branch representation Z a , calculates the attention weights of the specific branch features through the shared parameter layer, and realizes content-aware graph fusion;

[0013] (5) Cross-domain joint optimization and emotion classification

[0014] Adopt the domain adversarial method to alleviate the data distribution shift caused by individual differences. The core architecture consists of a feature extractor, a domain classifier, and a label predictor. A gradient reversal layer is introduced between the feature extractor and the domain classifier to achieve dynamic adversarial training. The Adam optimization algorithm is used to train the network composed of the feature extractor, the domain classifier, and the label predictor; in the test phase, the features of the test data are extracted and fed into the trained network to obtain the emotion recognition result.

[0015] In step (1) of the present invention, the EEG signals and ECG signals are divided into several non-overlapping time segments, and the frequency-domain features X f and the time-domain features X tExtraction: In frequency-domain feature extraction, differential entropy (DE) of the θ, α, β, and γ frequency bands is calculated for EEG signals. For ECG signals, Coiflet4 wavelet is used for eight-scale decomposition, and the band energies of the third to sixth components are extracted. Together with the differential entropy of EEG signals, they form the joint frequency-domain feature X. f ;

[0016] In time-domain feature extraction, the signal is further split using non-overlapping sliding windows. For ECG signals, time-domain indices of heart rate variability are calculated, including the mean, median, standard deviation, root mean square of RR intervals, and the percentage of RR interval differences greater than 50 milliseconds. For EEG signals, Hjorth parameters are fused with the mean and kurtosis of statistical features to form the time-domain feature X. t ; Construct the functional connectivity matrix A. f , The EEG signals are divided into four frequency bands: θ, α, β, and γ frequency bands. The ECG signals directly use the full-band signals to participate in the calculation, and the coupling strength is quantitatively characterized by the spectral coherence coefficient (ERCoh).

[0017] In step (2) of the present invention, the frequency-domain functional connectivity graph (X f , A f ) is directly represented by the predefined functional connectivity matrix A f and the frequency-domain feature X f , forming a fixed topological structure. The core of the time-domain data-driven graph is to dynamically generate the adjacency matrix A t of the time-domain heart-brain signals from the time-domain feature X t . A triple-constraint mechanism is used to achieve the collaborative optimization of the adjacency matrix A t , ensuring a better representation of the emotional correlation degree of the time-domain channels.

[0018] In the present invention, the multi-modal adjacency matrix optimization in step (2) adopts a triple-constraint mechanism: a global sparsity loss L norm = ||A t ||1 is imposed through L1 regularization to eliminate redundant connections and improve the generalization ability of the model; the L2,1 norm is designed to act on the cross-modal matrices A cross and A' cross to calculate the cross-modal connection loss L cro = ||A cross || 2,1 + ||A cross' || 2,1 , and the deep-seated correlation rules are mined using row sparsity to improve the modeling ability of the heart-brain emotional state; finally, the Hilbert-Schmidt independence criterion is introduced to evaluate the modal independence loss L pri , strengthening the unique features within the modality. The calculation formula is as follows:

[0019]

[0020] Among them, I is the identity matrix, and e o is the column vector of 1. The number of samples n represents the data scale. The electroencephalogram signal kernel matrix and the electrocardiogram signal kernel matrix are constructed through a non-positive definite inner product kernel function with parameter s. The matrix trace operation Tr() measures the statistical dependence between modalities by quantifying the interaction trace of the two centralized kernel matrices. This multi-granularity optimization strategy enables the adjacency matrix of time-domain heart-brain signals to represent both the global sparse topology and distinguish the emotional coupling mechanisms within and between modalities.

[0021] In step (3) of the present invention, the step of forming the branch representation Z a is as follows:

[0022] 3.1 Microscopic aggregation module

[0023] Based on the prior knowledge of the physiological and anatomical structure and the distribution of emotion-induced functional regions of the heart-brain coupling mechanism, two locally subgraph structures with biological interpretations are decoupled from the global graph (X a , A a ). The two types of subgraphs are generated by quantifying the connection matrix e. The specific calculation formula of the connection matrix e is as follows:

[0024]

[0025] where W represents the learnable weight matrix, represents matrix transpose. LeakyReLu preserves the gradient propagation in the negative value interval by introducing a small negative slope. The weight coefficient of the local subgraph is calculated by summing each row of the connection matrix e Subsequently, the feature matrix and the adjacency matrix The parameter locate ∈ {3, 8} represents the subgraph node scale: when locate = 3, it represents the emotion-induced subgraph, and when locate = 8, it represents the physiological and anatomical subgraph:

[0026]

[0027] where softmax represents the activation function. The physiological structure subgraph and the emotion-induced subgraph are respectively denoted as and where the superscript number represents the subgraph node scale, and F a is the dynamic feature dimension: in the frequency-domain functional connection branch, it corresponds to the frequency-domain feature dimension F, and in the time-domain driving branch, it is mapped to the time-domain feature dimension M. The two are independently learned through branch-specific parameters;

[0028] 3.2 Graph convolution module

[0029] The graph convolution module adopts a three-way collaborative architecture to process the global graph (X a , A a ), the physiological structure sub-graph and the emotion induction sub-graph A two-layer K-order Chebyshev graph convolution network is used for neighborhood feature aggregation, and its calculation process is as follows:

[0030]

[0031] Among them, is the normalized Laplacian matrix, Q k is the Chebyshev polynomial basis function, are learnable parameters. After passing through the Chebyshev process, the global graph and the two sub-graphs respectively obtain node-level features Z a , Through the concatenated summation readout function R sum and the recurrent neural network R GRU The readout function converts the node features into graph-level features:

[0032]

[0033] Among them, represents matrix concatenation. Finally, the graph-level feature Gra a of the global graph and the graph-level features of the two sub-graphs are concatenated to form the branch representation Z a , and the specific formula is:

[0034]

[0035] Therefore, specifically, the frequency-domain functional connectivity branch representation is Z f , and the time-domain data-driven branch representation is Z t .

[0036] In step (4) of the present invention, the fusion graph network adopts an attention mechanism to merge the frequency-domain functional connectivity branch representation Z f and the time-domain data-driven branch representation Z t , calculates the attention weights of the branch features through a shared parameter layer, and the calculation formula is as follows:

[0037] (λ f , λ t ) = Att(Z f , Z t ) (6)

[0038] Among them, λ f , λ tRepresents the attention weights of the frequency domain functional connection branch and the time domain data driven branch. Att represents the same linear transformation of the branch representation (shared weight matrix and bias), generates attention weights for each feature, and realizes content-aware graph fusion. Subsequently, the attention weights are normalized by the softmax function to obtain the attention score of each feature. For the branch representation Z a The lth feature in , the attention score The calculation method is as follows:

[0039]

[0040] The attention score With branch representation Z a Weighted fusion is performed to obtain the final graph representation Z, and the calculation formula is as follows:

[0041] Z=ω f *Z f +ω t *Z t (8).

[0042] The robustness of the model is enhanced by adopting a domain adversarial approach in step (5) of the present invention. Its core architecture is composed of a feature extractor, a domain classifier and a label predictor. The adversarial optimization of the feature extractor and the domain classifier is dynamically coordinated through a gradient reversal layer. The feature extractor has learned the branch representation Z through steps (2) to (4), and the graph representation Z is respectively input into the domain classifier and the label predictor to obtain the prediction results of the label predictor and the domain classifier. A gradient reversal layer is introduced between the feature extractor and the domain classifier to realize dynamic adversarial training. The feature gradient is kept normally transmitted during the forward propagation process, and the domain classifier gradient is reversed during the back propagation, thereby forcing the feature extractor and the domain classifier to form an optimization dynamic that is mutually adversarial.

[0043] The advantages of the present invention are: based on the theory of mind-brain coupling, cognitive inspiration branches and data-driven branches are constructed to capture inherent features and dynamic correlation features to form complementary and synergistic representations. An adaptive adjacency matrix optimization strategy based on a triple constraint mechanism is proposed to solve the modal alignment and deviation problems of multimodal graph representation in data-driven branches: global regularization constraints are used to eliminate redundant noise, and unique feature representations within the modality are strengthened through modal specificity. In combination with cross-modal interaction modules, deep-level correlation rules are mined to enhance the modeling capabilities of the mind-brain emotional state. A multi-view graph convolutional network is further designed to realize the extraction of global and local features. A graph fusion network based on an attention mechanism is used to merge the features of cognitive inspiration branches and data-driven branches. A domain adversarial strategy is introduced to enhance the generalization of the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1It is a flowchart of the emotional analysis method for heart and brain signals;

[0045] Figure 2 It is a diagram of the adaptive adjacency matrix optimization method based on a triple constraint mechanism;

[0046] Figure 3 It is a diagram of local subgraph partitioning based on physiological anatomy: The diagram shows the distribution of 32-lead EEG electrodes of the international 10-20 standard and 3-lead ECG signals. According to anatomy, it is divided into prefrontal, parietal, occipital, and temporal lobe regions. The temporal lobe electrodes are further subdivided into smaller local subregions. The multi-lead ECG signals are aggregated to a single electrode for characterization, and finally divided into 8 subregions;

[0047] Figure 4 It is a diagram of local subgraph partitioning based on emotion induction: The diagram shows the distribution of 32-lead EEG electrodes of the international 10-20 standard and 3-lead ECG signals. According to the asymmetry of emotional brain regions, the brain is divided into left and right hemispheres. At the same time, the multi-lead ECG signals are mapped to a single electrode for characterization, and finally divided into 3 subregions;

[0048] Figure 5 It is a structural diagram of a multi-view graph convolutional network;

[0049] Figure 6 It is a structural diagram of a domain adversarial network. Specific implementation manner

[0050] The present invention provides a multi-view graph convolutional emotion recognition method based on heart-brain coupling. First, based on the heart-brain coupling relationship, a frequency-domain functional connectivity graph and a time-domain data-driven connectivity graph with complementary characteristics are constructed to form a basis for bimodal representation. Subsequently, a multi-view graph convolutional network is constructed to micro-aggregate the graph and use the graph neural network to encode deep emotional information. The advantageous information of the time-domain branch and the frequency-domain branch features is dynamically integrated through an attention mechanism. In addition, to break through the limitation of individual physiological differences on model generalization, a domain generalization module for cross-subject scenarios is designed to extract highly generalizable emotional representations by decoupling subject-specific information. Specifically, as Figure 1 shown, it is carried out according to the following steps:

[0051] (1) Data preprocessing

[0052] Obtain EEG signals and ECG signals from a multi-modal emotion database, divide them into several non-overlapping time segments (for example, each segment has a duration of 6 s), and extract frequency-domain features X f and time-domain features X t respectively. In the extraction of frequency-domain features, for EEG signals, calculate the differential entropy (DE) of the θ (4 - 8 Hz), α (8 - 14 Hz), β (14 - 30 Hz), and γ (30 - 45 Hz) frequency bands;

[0053] The differential entropy (DE) is defined as:

[0054] DE = -∫ X p(x) log(p(x)) dx (1)

[0055] where p(x) is the probability density function of the EEG signal x. For the ECG signal, Coiflet4 wavelet is used for eight-scale decomposition, and the band energy of the 3rd to 6th components (corresponding to the 2.81 - 45 Hz frequency band) is extracted, which together with the differential entropy of the EEG signal forms the joint frequency-domain feature N represents the number of channels, and F represents the dimension of the frequency-domain feature;

[0056] In time-domain feature extraction, non-overlapping sliding windows are used to further split the signal (for example, the window duration is 2 s). For the ECG signal, time-domain indexes of heart rate variability (average RR, median, standard deviation, root mean square, and percentage of RR interval difference greater than 50 milliseconds) are calculated, while for the EEG signal, the Hjorth parameters and statistical features (mean, kurtosis) are fused to form time-domain features T represents the number of windows, and M represents the dimension of the time-domain feature, represents the set of real numbers;

[0057] Subsequently, a functional connectivity matrix A is constructed f , specifically, the EEG signal is divided into four frequency bands: θ, α, β, and γ bands, while the ECG signal directly uses the full-band signal to participate in the calculation. The coupling strength is quantitatively characterized by the spectral coherence coefficient (ERCoh), and its mathematical definition is:

[0058]

[0059] where f represents the frequency, G x,x (f) and G y,y (f) represent the auto-power spectra of the EEG signal x and the ECG signal y at a given frequency, while G x,y (f) represents the cross-power spectrum of the two signals. For the EEG signal with N1 leads and the ECG signal with N2 leads, the calculated functional connectivity matrix is expressed as (where N = N1 + N2);

[0060] (2) Constructing the graph representation

[0061] The graph data is defined as an undirected graph G = (V, E), where the node set V represents the channels of the heart-brain signals, the edge set E represents the emotional coupling strength between channels, the topological relationship of the emotional coupling strength between channels is quantified by the adjacency matrix A, and the node feature matrix is X which characterizes the signal characteristics of each channel; Based on this, a two-stream graph representation framework is designed, and the frequency-domain functional connectivity graph (X f , A f) and the time-domain data-driven connection graph (X t , A t ), the frequency-domain functional connection graph (X f , A f ) is directly represented by the predefined functional connection matrix A f and the frequency-domain feature X f , forming a fixed topological structure. The core of the time-domain data-driven graph lies in dynamically generating the adjacency matrix A t of the time-domain heart-brain signals from the time-domain feature X t , and its generation process can be expressed as:

[0062] A t = ReLU((PX t + B)Qω) (3)

[0063] where, represents the left multiplication matrix, represents the bias matrix, represents the right multiplication matrix that summarizes the time information, represents the projection matrix of the learnable parameters. ReLU is the activation function. The finally constructed time-domain data-driven graph (X t , A t ) encodes the emotional coupling relationship between heart-brain signals through the dynamic adjacency matrix, captures the cross-channel dynamic emotional association pattern, and overcomes the limitations of the fixed graph structure;

[0064] To better characterize the emotional association degree of the time-domain channels, through the collaborative optimization method of the adjacency matrix based on the triple constraint mechanism, as Figure 2 shown, the adjacency matrix A t of the time-domain heart-brain signals can be decomposed into two parts: the within-modal connection relationship and the inter-modal connection relationship. The former describes the internal coupling of the same type of signal channels, and respectively quantify the dynamic emotional associations within the EEG channels and within the ECG channels. The latter models the synergistic effect of the heart-brain cross-modal channels through A cross and A' cross ;

[0065]

[0066] First, apply the global sparsity loss L norm = ||A t ||1 to impose the global sparsity constraint on the adjacency matrix, eliminate redundant connections, and improve the generalization ability of the model. Secondly, design the L2,1 norm to act on the cross-modal matrices A cross and A' cross to calculate the cross-modal connection loss L cro = ||A cross|| 2,1 +||A cross' || 2,1 , leveraging row sparsity to mine deep collaborative patterns and enhancing the modeling ability of cardio-cerebral emotional states; finally, introducing the Hilbert-Schmidt Independence Criterion to evaluate the modal independence loss L pri , ensuring the specificity of the intra-modal connectivity relationships between electrocardiogram and electroencephalogram signals. The calculation formula is as follows:

[0067]

[0068] where I is the identity matrix, e o is the column vector of 1, the number of samples n represents the data scale, and the electroencephalogram signal kernel matrix and the electrocardiogram signal kernel matrix are constructed through a non-positive definite inner product kernel function (parameter s). The matrix trace operation Tr() measures the statistical dependence between modalities by quantifying the interaction trace of the two centralized kernel matrices. This multi-granularity optimization strategy enables the adjacency matrix of time-domain cardio-cerebral signals to represent both the global sparse topology and distinguish the emotional coupling mechanisms within and between modalities;

[0069] Therefore, the frequency-domain functional connectivity graph (X f ,A f ) and the time-domain data-driven graph (X t ,A t ) can be uniformly represented as a global graph (X a ,A a ), a ∈ {t, f}, where t represents the time-domain data-driven graph and f represents the frequency-domain functional connectivity graph;

[0070] (3) Construct a multi-view graph convolutional network

[0071] The multi-view graph convolutional network adopts a double-branch isomorphic architecture. Although the two branches have the same network design but independent parameters: one branch takes the frequency-domain functional connectivity graph (X f ,A f ) as the input, and the other branch processes the time-domain data-driven graph (X t ,A t ). Each branch decouples the physiological and anatomical subgraph a ,A a ) from the global graph (X and the emotion-induced subgraph Among them, the superscript number represents the subgraph node scale, and the subscript represents the branch. Meanwhile, the topological structure of the global graph is retained. a ∈ {t, f}, where t represents the time-domain data-driven branch and f represents the frequency-domain functional connection branch. On this basis, discriminative spatio-temporal features are extracted from the two local sparse subgraphs and the global graph through the graph convolutional network and the cascaded readout function respectively. Finally, the spatio-temporal features of the subgraph and the global graph are concatenated to form the branch output Z a , such as Figure 5 shown. While realizing the extraction of global dynamic features, local feature perception is ensured through microscopic aggregation;

[0072] 3.1 Microscopic Aggregation Module

[0073] Based on the prior knowledge of the physiological and anatomical structure and the distribution of emotion-induced functional regions in the heart-brain coupling mechanism, two local subgraph structures with biological interpretations are decoupled from the global graph (X a , A a ), as Figure 3 and Figure 4 shown. The two types of subgraphs are generated by quantifying the connection matrix e. The specific calculation formula of the connection matrix e is as follows:

[0074]

[0075] where W represents the learnable weight matrix, represents the matrix transpose, and LeakyReLu retains the gradient propagation in the negative value interval by introducing a small negative slope. The weight coefficient of the local subgraph is calculated by summing each row of the connection matrix e. Subsequently, the feature matrix and the adjacency matrix of the local subgraph are calculated. The parameter locate ∈ {3, 8} represents the subgraph node scale: when locate = 3, it represents the emotion-induced subgraph, and when locate = 8, it represents the physiological and anatomical subgraph:

[0076]

[0077] where softmax represents the activation function. The physiological structure subgraph and the emotion-induced subgraph are denoted as and respectively. Among them, the superscript number represents the subgraph node scale, and F a is the dynamic feature dimension: corresponding to the frequency-domain feature dimension F in the frequency-domain functional connection branch, and mapped to the time-domain feature dimension M in the time-domain driven branch. The two are independently learned through branch-specific parameters;

[0078] 3.2 Graph Convolution Module

[0079] The graph convolution module adopts a three-way collaborative architecture to process the global graph (X a , Aa ) Physiological structure sub - graph and emotion - inducing sub - graph Use a two - layer K - order Chebyshev graph convolutional network for neighborhood feature aggregation. Its calculation process is as follows:

[0080]

[0081] Among them, is the normalized Laplacian matrix, Q k is the Chebyshev polynomial basis function, are learnable parameters. The global graph and the two sub - graphs obtain node - level features Z a , Through the cascaded summation read - out function R sum and the recurrent neural network R GRU The read - out function converts the node features into graph - level features:

[0082]

[0083] Among them, represents matrix concatenation. Finally, the graph - level feature Gra a of the global graph and the graph - level features of the two sub - graphs are concatenated to form the branch representation Z a , and the specific formula is:

[0084]

[0085] Therefore, the frequency - domain functional connectivity branch representation is Z f , and the time - domain data - driven branch representation is Z t ;

[0086] (4) Construct a fusion graph network

[0087] The fusion graph network uses an attention mechanism to merge the frequency - domain functional connectivity branch representation Z f and the time - domain data - driven branch representation Z t , and calculates the attention weights of the branch features through a shared - parameter layer. The calculation formula is as follows:

[0088] (λ f ,λ t ) = Att(Z f ,Z t ) (11)

[0089] Among them, λ f ,λ tDenote the attention weights of the frequency-domain functional connection branch and the time-domain data-driven branch. Att represents the same linear transformation (shared weight matrix and bias) on the branch representations, generating attention weights for each feature, achieving content-aware graph fusion. Subsequently, the attention weights are normalized through the softmax function to obtain the attention scores for each feature. For the l-th feature in the graph representation Z a in, the attention score is calculated as follows:

[0090]

[0091] Multiply the attention score with the graph branch representation Z a for weighted fusion to obtain the final graph representation Z. The calculation formula is as follows:

[0092] Z = ω f *Z f + ω t *Z t (13)

[0093] (5) Cross-domain joint optimization and sentiment classification

[0094] To alleviate the problem of data distribution deviation caused by individual differences, domain adversarial methods are used to enhance the robustness of the model. Its core architecture consists of a feature extractor, a domain classifier, and a label predictor. The gradient reversal layer is used to dynamically coordinate the adversarial optimization between the feature extractor and the domain classifier, as shown in Figure 6 . Among them, the feature extractor has realized the learning of the final graph representation Z through steps (2) to (4). The final graph representation Z is respectively input into the domain classifier and the label predictor. The specific calculation formulas are as follows:

[0095]

[0096] where z i is the graph representation of the i-th sample, and are the prediction results of the label predictor and the domain classifier respectively. φ y , γ y , φ d , γ d are learnable parameters. Both the label predictor and the domain classifier use cross-entropy as the loss function:

[0097]

[0098] where L y represents the label loss, L d represents the domain classification loss, L is the number of samples, R y and R dare the number of categories and the number of domains respectively, y is the true label, and is the predicted value of the model. d is the true domain, is the predicted value of the model;

[0099] A gradient reversal layer is introduced between the feature extractor and the domain classifier to implement dynamic adversarial training. During the forward propagation process, the feature gradient is kept normally transmitted, while during the backward propagation, the gradient of the domain classifier is reversed, thereby forcing the feature extractor and the domain classifier to form an optimized dynamic of mutual confrontation. The loss function of domain adversarial is defined as:

[0100] L DG =-L y +τL d (16)

[0101] where τ, as the adversarial intensity coefficient, controls the balance of the two types of losses, and the Adam optimization algorithm is used to train the network composed of the feature extractor, the domain classifier, and the label predictor;

[0102] In the test phase, the features of the test data are extracted and fed into the trained network to obtain the emotion recognition result.

[0103] The following further illustrates the present invention through experimental examples.

[0104] In this experimental example, the MAHNOB-HCI (A Multimodal Database for Affect Recognition and Implicit Tagging) multimodal emotion database is adopted. In practice, other emotion databases including electroencephalogram (EEG) and electrocardiogram (ECG) signals can also be used. The MAHNOB-HCI database provides 32-channel EEG signals with a sampling rate of 256Hz and 3-channel ECG signals. The EEG signals are processed by a band-pass filter with a frequency range of 4 - 45Hz, and the ECG signals are processed by a low-pass filter with a frequency of 1Hz. Samples with both EEG and ECG signals are selected for four-class recognition in the arousal-valence space.

[0105] First step, data preprocessing: Associate the EEG signals, ECG signals, and label data. The signals are segmented into 6s segments and divided according to the subjects. The data of the first 26 subjects are used as the training set, and the data of the last subject are used as the test set. Frequency-domain feature extraction is performed. Four-band differential entropy features are extracted from the EEG signals, and the frequency band energy of four wavelet components is extracted from the ECG signals. Calculate the spectral coherence coefficient of the EEG signals and the ECG signals. Time-domain feature extraction is performed. First, the signals are further split into 2s, and then the Hjorth parameters and statistical features (mean, kurtosis) are calculated for the EEG signals, and the time-domain heart rate variability features (RR mean, median, standard deviation, root mean square, and the percentage of RR interval difference greater than 50 milliseconds) are calculated for the ECG signals;

[0106] The following are the detailed formulas for feature extraction:

[0107] The wavelet transform coefficient W x is expressed as (a, b):

[0108]

[0109] where ψ is the mother wavelet function, a is the scale parameter, b is the translation parameter, and * represents complex conjugate.

[0110] For a given signal x(t), its energy in the frequency band [w1, w2] can be expressed as:

[0111]

[0112] where X(w) is the Fourier transform of the signal x(t);

[0113] The Hjorth parameters include the following three parts: activity (variance of the signal), mobility (ratio of the standard deviation of the first-order difference of the signal to the standard deviation of the signal), and complexity (ratio of the standard deviation of the second-order difference of the signal to the standard deviation of the first-order difference).;

[0114] Assume R i is the i-th RR interval, is the average value of each RR interval, and n is the number of RR intervals. The unit is milliseconds. Mean of RR intervals:

[0115]

[0116] Standard deviation of RR intervals (SDNN):

[0117]

[0118] Root mean square of the differences between adjacent RR intervals (RMSSD):

[0119]

[0120] Percentage of consecutive RR intervals with differences greater than 50 milliseconds (pNN50), where NN50 is the number of intervals with differences between consecutive RR intervals greater than 50 milliseconds:

[0121]

[0122] In the second step, construct a graph representation: Based on the frequency-domain feature X f and the functional connectivity matrix A f construct a frequency-domain functional connectivity graph (X f , A f ), from the time-domain feature X tDynamically generate the adjacency matrix A t , construct the time-domain data-driven graph (X t , A t );

[0123] Step 3: Construct a multi-view graph convolutional network, which mainly includes the following parameter settings:

[0124] Physiological and anatomical perspective division: EEG electrodes are first anatomically divided into four major regions: prefrontal lobe, parietal lobe, occipital lobe, and temporal lobe. Among them, the temporal lobe electrodes are further subdivided into smaller local sub-regions. Multi-lead ECG signals are aggregated to a single electrode for characterization;

[0125] Emotion induction perspective division: According to the hemispheric lateralization characteristics of brain regions for emotion processing, the brain is divided into left and right hemispheres. At the same time, multi-lead ECG signals are mapped to single-electrode characterization;

[0126] The polynomial order of the Chebyshev graph convolutional network is set to 2;

[0127] Use the multi-view graph convolutional network to encode the frequency-domain functional connectivity graph and the time-domain data-driven graph to obtain the graph branch representations Z f and Z t ;

[0128] Step 4: Construct a fusion graph network: Use the attention mechanism to fuse and extract the multi-modal emotion fusion representation Z for the graph branch representations Z f and Z t ;

[0129] Step 5: Cross-domain joint optimization and emotion classification: Predict the fusion representation Z through the domain classifier and the label predictor, introduce the gradient reversal layer to achieve adversarial training, and jointly optimize the network parameters by minimizing the emotion loss and the domain adversarial loss. Repeat steps 1 to 4, input the test data into the trained network for classification and recognition to obtain the emotion category, and finally output one of the four emotion states: high arousal and high valence, high arousal and low valence, low arousal and high valence, and low arousal and low valence. After 100 iterations, the accuracy rate and the loss value gradually tend to be stable, and the accuracy rate for the test data is 93.09%.

Claims

1. A multimodal emotion recognition method based on heart-brain coupling and graph neural network, characterized in that: The following steps are involved: (1) Data preprocessing: obtaining EEG and ECG signals from the multimodal emotion database and extracting frequency domain features X f and the time domain feature X t , the spectral coherence coefficient is used to construct the functional connectivity matrix A f Characterize the strength of heart-brain coupling; (2) Constructing a graph representation, constructing a frequency domain functional connectivity graph (X f ,A f ) and time domain data driven connection diagram (X t ,A t ) to capture both inherent features and dynamic correlation features, forming a complementary and synergistic representation; (3) Constructing a multi-view graph convolutional network The multi-view graph convolutional network adopts a dual-branch isomorphic architecture. The two branches have the same network design but independent parameters: one branch is a frequency domain functional connection map (X f ,A f ) as input, and the other branch processes the time domain data driven graph (X t ,A t ), each branch is aggregated from the global graph (X a ,A a ) to decouple the physiological and anatomical subgraphs and the emotion-inducing subgraph At the same time, the topological structure of the global graph is retained, where a∈{t,f}, t represents the time domain data-driven branch, and f represents the frequency domain functional connection branch. On this basis, the graph convolutional network and the cascaded readout function are used to extract discriminative spatiotemporal features from the two subgraphs and the global graph respectively, and finally the spatiotemporal features of the subgraphs and the global graph are spliced ​​to form a branch representation Z a ; (4) Constructing a Fusion Graph Network The fusion graph network uses an attention mechanism to merge branch representations Z a , by calculating the attention weights of specific branch features through a shared parameter layer, content-aware graph fusion is achieved; (5) Cross-domain joint optimization and sentiment classification The domain adversarial approach is used to alleviate the data distribution deviation caused by individual differences. The core architecture consists of a feature extractor, a domain classifier, and a label predictor. A gradient reversal layer is introduced between the feature extractor and the domain classifier to implement dynamic adversarial training. The Adam optimization algorithm is used to train the network composed of the feature extractor, domain classifier, and label predictor. In the testing phase, the features of the test data are extracted and fed into the trained network to obtain the emotion recognition results.

2. According to claim 1, a multimodal emotion recognition method based on heart-brain coupling and graph neural network is characterized in that: In step (1), the EEG signal and the ECG signal are divided into several non-overlapping time segments, and the frequency domain features X are respectively calculated. f and the time domain feature X t Extraction, in the frequency domain feature extraction, the differential entropy DE of the θ (4-8Hz), α (8-14Hz), β (14-30Hz), and γ (30-45Hz) frequency bands are calculated for the EEG signal. For the ECG signal, the Coiflet4 wavelet is used for eight-scale decomposition to extract the frequency band energy of the 3rd to 6th components, and together with the differential entropy of the EEG signal, form the joint frequency domain feature X f ; In the time domain feature extraction, non-overlapping sliding windows are used to further split the signal. The time domain indicators of heart rate variability are calculated for the ECG signal, including the RR mean, median, standard deviation, root mean square, and the percentage of RR interval differences greater than 50 milliseconds. The Hjorth parameter is integrated with the statistical feature mean and kurtosis of the EEG signal to form the time domain feature X t ; Constructing the functional connectivity matrix A f The EEG signal is divided into four frequency bands: θ, α, β and γ bands, while the ECG signal directly uses the full-band signal to participate in the calculation, and the coupling strength is quantitatively characterized by the spectral coherence coefficient ERCoh.

3. According to claim 1, a multimodal emotion recognition method based on heart-brain coupling and graph neural network is characterized in that: In step (2), the frequency domain functional connection diagram (X f ,A f ) directly through the predefined functional connection matrix A f and frequency domain feature X f Indicates that a fixed topological structure is formed, and the core of the time domain data-driven graph is to extract the time domain feature X t The adjacency matrix A of the time-domain heart and brain signals is dynamically generated by the neural network t , using the triple constraint mechanism to implement the adjacency matrix A t The collaborative optimization ensures a better representation of the emotional correlation of time domain channels.

4. According to claim 3, a multimodal emotion recognition method based on heart-brain coupling and graph neural network is characterized in that: The multimodal adjacency matrix optimization in step (2) adopts a triple constraint mechanism: applying a global sparsity loss L through L1 regularization norm =||A t ||1 to eliminate redundant connections and improve the generalization ability of the model; design the L2,1 norm to act on the cross-modal matrix A cross and A′ cross Calculate the cross-modal connection loss L cro =||A cross || 2,1 +||A cross' || 2,1 , using row sparsity to explore deep-level association rules and improve the modeling ability of the emotional state of the heart and brain; finally, the Hilbert-Schmidt independence criterion is introduced to evaluate the modal independence loss L pri , strengthen the unique features within the mode, and the calculation formula is as follows: in, I is the identity matrix, e o is a column vector of 1, the number of samples n represents the data scale, and the EEG signal kernel matrix ECG signal core matrix Constructed through a non-positive definite inner product kernel function with parameter s, the matrix trace operation Tr() measures the statistical dependence between modalities by quantifying the interaction traces of two centralized kernel matrices. This multi-granularity optimization strategy enables the adjacency matrix of time-domain cardio-brain signals to both characterize the global sparse topology and distinguish the emotional coupling mechanisms within and between modalities.

5. The multimodal emotion recognition method based on heart-brain coupling and graph neural network according to claim 1 is characterized in that: The branch representation Z is formed in step (3) a The steps are as follows: 3.1 Micro-aggregation module Based on the prior knowledge of the physiological and anatomical structure of the heart-brain coupling mechanism and the regional distribution of emotion-induced functions, from the global map (X a ,A a ) decouples two local subgraph structures with biological interpretations. The two types of subgraphs are generated by quantizing the connection matrix e. The specific calculation formula of the connection matrix e is as follows: Where W represents the learnable weight matrix, Represents matrix transpose. LeakyReLu preserves the gradient propagation in the negative range by introducing a smaller negative slope. The weight coefficient of the local subgraph is calculated by summing the connection matrix e row by row. Then, the feature matrix of the local subgraph is calculated and the adjacency matrix The parameter locate∈{3,8} represents the size of the subgraph nodes: when locate=3, it represents the emotion-induced subgraph, and locate=8 represents the physiological and anatomical subgraph: Among them, softmax represents the activation function, and the physiological structure subgraph and the emotion induced subgraph are respectively denoted as and The superscript numbers represent the size of the subgraph nodes, F a It is a dynamic feature dimension: it corresponds to the frequency domain feature dimension F in the frequency domain functional connection branch and is mapped to the time domain feature dimension M in the time domain driving branch. The two are learned independently through branch-specific parameters. 3.2 Graph Convolution Module The graph convolution module adopts a three-way collaborative architecture to process the global graph (X a ,A a ), physiological structure subgraph and the emotion-inducing subgraph A two-layer K-order Chebyshev graph convolutional network is used for neighborhood feature aggregation, and the calculation process is as follows: in, is the normalized Laplace matrix, Q k is the Chebyshev polynomial basis function, is a learnable parameter. The global graph and the two subgraphs are processed by Chebyshev to obtain node-level features Z a , Reading out the function R by cascading summation sum R with Recurrent Neural Networks GRU The readout function converts node features into graph-level features: in, Represents matrix concatenation, and finally the graph-level features Gra a And the graph-level features of the two subgraphs Splice to form branch representation Z a , the specific formula is: Therefore, specifically, the frequency domain functional connection branch is represented as Z f , the time domain data driven branch is characterized as Z t .

6. The multimodal emotion recognition method based on heart-brain coupling and graph neural network according to claim 1 is characterized in that: In step (4), the fusion graph network uses an attention mechanism to merge the frequency domain functional connection branch representation Z f and time domain data driven branch representation Z t , the attention weight of the branch feature is calculated through the shared parameter layer, and the calculation formula is as follows: (λ f ,λ t )=Att(Z f ,WITH t ) (6) Among them, λ f ,λ t Represents the attention weights of the frequency domain functional connection branch and the time domain data driven branch. Att represents the same linear transformation of the branch representation (shared weight matrix and bias), generates attention weights for each feature, and realizes content-aware graph fusion. Subsequently, the attention weights are normalized by the softmax function to obtain the attention score of each feature. For the branch representation Z a The lth feature in , the attention score The calculation method is as follows: The attention score With branch representation Z a Weighted fusion is performed to obtain the final graph representation Z, and the calculation formula is as follows: Z=ω f *Z f +ω t *Z t (8)。 7. The multimodal emotion recognition method based on heart-brain coupling and graph neural network according to claim 1 is characterized in that: The domain adversarial approach in step (5) is used to enhance the robustness of the model. Its core architecture consists of a feature extractor, a domain classifier, and a label predictor. The adversarial optimization between the feature extractor and the domain classifier is dynamically coordinated through a gradient reversal layer. The feature extractor has learned the branch representation Z through steps (2) to (4), and input the graph representation Z into the domain classifier and the label predictor respectively to obtain the prediction results of the label predictor and the domain classifier. A gradient reversal layer is introduced between the feature extractor and the domain classifier to realize dynamic adversarial training. The feature gradient is transmitted normally during the forward propagation process, and the domain classifier gradient is reversed during the back propagation, thereby forcing the feature extractor and the domain classifier to form an optimization dynamic that is mutually adversarial.

Citation Information

Cited By

  • Coastal city storm surge disaster toughness evaluation method based on graph neural network

    CN120471458A

  • A Graph Neural Network-Based Method for Assessing Storm Surge Resilience in Coastal Cities

    CN120471458B

  • Cross-individual emotion online monitoring system based on electroencephalogram and graph deep learning

    CN121265056A

  • Central autonomic nervous system stability monitoring method and system

    CN122604392A

  • Adaptive multi-modal data fusion prediction method and system based on dynamic graph convolution

    CN122654992A