Multimodal emotion data prediction method, device and related medium based on EEG

By extracting the differential entropy features of EEG data at different resolutions and building a domain adaptive neural network, combining deep convolutional network to extract audio-visual features and using hypergraph segmentation, the individual differences and noise interference problems of the emotion prediction model in the prior art are solved, achieving higher emotion prediction accuracy.

CN114118165BActive Publication Date: 2025-05-23SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111465384.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-05-23
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

Existing EEG-based emotion prediction models have stability and generalization challenges in individual differences and noise interference, and traditional audio-visual features cannot fully express discriminant features related to emotions.

Method used

By extracting the differential entropy features of EEG data at different resolutions, a domain adaptive neural network is constructed to conduct prediction voting; at the same time, deep convolutional network is used to extract the deep visual and auditory features of audio-visual content, and the hidden emotional prediction tag data is obtained through hypergraph segmentation, and finally the label data of EEG and audio-visual features are weighted and fused to improve the accuracy of emotional prediction.

Benefits of technology

Improve the accuracy of emotion prediction, provide more complementary information through multimodal fusion, enable more accurate emotion modeling and fully express discriminant characteristics related to emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114118165B_ABST
    Figure CN114118165B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal emotion data prediction method, device and related media based on electroencephalogram data. The method includes: predicting and voting on electroencephalogram data based on a domain adaptive neural network to obtain individual emotion prediction label data; extracting deep visual features and deep auditory features from preset audio-visual content through a deep convolutional network model, and fusing the deep visual features and deep auditory features into deep audio-visual fusion features; constructing a hypergraph based on the deep visual features, deep auditory features and deep audio-visual fusion features, and obtaining latent emotion prediction label data corresponding to the deep visual features, deep auditory features and deep audio-visual fusion features through hypergraph segmentation; giving weights to individual emotion prediction label data and latent emotion prediction label data and fusing them, and using the fused result as the emotion data prediction result. The present invention combines electroencephalogram data and audio-visual features to perform multimodal prediction, thereby improving the accuracy of emotion prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer software technology, and in particular to a method and device for predicting multimodal emotion data based on electroencephalogram (EEG), and related media. Background Art

[0002] EEG provides a more natural way to record human brain activity and is also widely used in emotional intelligence research. Existing literature shows that deep neural network learning methods provide an effective way to extract deep feature information from EEG signals and achieve good results in EEG-based emotion prediction. However, due to the problem of individual differences, the stability and generalization of EEG-based emotion prediction models are great challenges. Specifically, EEG is a very weak signal and is easily disturbed and affected by external noise, making it difficult to distinguish individual characteristics and meaningful EEG features from noise.

[0003] Visual content and auditory content are the main factors that induce emotions. The same objective content is transmitted to individuals, inducing different emotions in different individuals. Therefore, the emotion prediction model based on individual physiological signals has problems of missing information and individual differences, and cannot accurately model emotions. Compared with the single-modal emotion prediction model, the multimodal fusion method can provide more complementary information that is missing in the single modality for emotion prediction, and can achieve more accurate modeling. Existing methods for extracting audio-visual features are based on traditional audio-visual features. Due to the existence of the "semantic gap" (or "emotional gap"), traditional audio-visual features cannot fully express the discriminative features related to emotions. Summary of the invention

[0004] The embodiments of the present invention provide a method, device and related medium for predicting multimodal emotion data based on electroencephalogram data, aiming to improve the accuracy of emotion prediction.

[0005] In a first aspect, an embodiment of the present invention provides a method for predicting multimodal emotion data based on electroencephalogram data, comprising:

[0006] Extracting differential entropy features of EEG data for training for different sub-bands at different resolutions, and constructing a domain adaptive neural network based on the differential entropy features;

[0007] Based on the domain adaptive neural network, predictive voting is performed on the EEG data of the target user to obtain individual emotion prediction label data;

[0008] Extracting deep visual features and deep auditory features from preset audio-visual content through a deep convolutional network model, and fusing the deep visual features and deep auditory features into deep audio-visual fusion features;

[0009] Constructing a hypergraph based on the deep visual features, deep auditory features and deep audio-visual fusion features, and obtaining latent emotion prediction label data corresponding to the deep visual features, deep auditory features and deep audio-visual fusion features through hypergraph segmentation;

[0010] The individual emotion prediction label data and the latent emotion prediction label data are weighted and fused, and the fused result is used as the emotion data prediction result.

[0011] In a second aspect, an embodiment of the present invention provides a multimodal emotion data prediction device based on EEG data, comprising:

[0012] A network construction unit, used for extracting differential entropy features of EEG data for training for different sub-bands at different resolutions, and constructing a domain adaptive neural network based on the differential entropy features;

[0013] A first prediction unit is used to perform prediction voting on the EEG data of the target user based on the domain adaptive neural network to obtain individual emotion prediction label data;

[0014] A feature extraction unit, configured to extract deep visual features and deep auditory features from preset audio-visual content through a deep convolutional network model, and fuse the deep visual features and deep auditory features into deep audio-visual fusion features;

[0015] A second prediction unit is used to construct a hypergraph based on the deep visual features, the deep auditory features and the deep audio-visual fusion features, and obtain latent emotion prediction label data corresponding to the deep visual features, the deep auditory features and the deep audio-visual fusion features through hypergraph segmentation;

[0016] The label fusion unit is used to assign weights to the individual emotion prediction label data and the latent emotion prediction label data and fuse them, and use the fused result as the emotion data prediction result.

[0017] In a third aspect, an embodiment of the present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the multimodal emotion data prediction method based on EEG data as described in the first aspect is implemented.

[0018] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the multimodal emotion data prediction method based on EEG data as described in the first aspect is implemented.

[0019] The embodiment of the present invention provides a method, device and related medium for predicting multimodal emotion data based on EEG data. The method includes: extracting differential entropy features of EEG data for training from different sub-bands at different resolutions, and constructing a domain adaptive neural network based on the differential entropy features; predicting and voting on the EEG data of the target user based on the domain adaptive neural network to obtain individual emotion prediction label data; extracting deep visual features and deep auditory features from preset audio-visual content through a deep convolutional network model, and fusing the deep visual features and deep auditory features into deep audio-visual fusion features; constructing a hypergraph based on the deep visual features, deep auditory features and deep audio-visual fusion features, and obtaining latent emotion prediction label data corresponding to the deep visual features, deep auditory features and deep audio-visual fusion features through hypergraph segmentation; giving weights to the individual emotion prediction label data and the latent emotion prediction label data and fusing them, and using the fused result as the emotion data prediction result. The embodiment of the present invention combines EEG data and audio-visual features to perform multimodal prediction, which can improve the accuracy of emotion prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.

[0021] Figure 1 A schematic flow chart of a multimodal emotion data prediction method based on EEG data provided by an embodiment of the present invention;

[0022] Figure 2 A schematic diagram of a sub-process of a multimodal emotion data prediction method based on EEG data provided by an embodiment of the present invention;

[0023] Figure 3 A schematic diagram of another sub-process of a multimodal emotion data prediction method based on EEG data provided by an embodiment of the present invention;

[0024] Figure 4 A schematic diagram of the overall network structure of a multimodal emotion data prediction method based on EEG data provided by an embodiment of the present invention;

[0025] Figure 5 A schematic diagram of the network structure of a domain adaptive neural network in a multimodal emotion data prediction method based on EEG data provided by an embodiment of the present invention;

[0026] Figure 6A schematic block diagram of a multimodal emotion data prediction device based on EEG data provided by an embodiment of the present invention;

[0027] Figure 7 A sub-schematic block diagram of a multimodal emotion data prediction device based on EEG data provided by an embodiment of the present invention;

[0028] Figure 8 Another sub-schematic block diagram of a multimodal emotion data prediction device based on EEG data provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0030] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0031] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0032] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0033] See below Figure 1 , Figure 1 A flowchart of a multimodal emotion data prediction method based on EEG data provided by an embodiment of the present invention specifically includes: steps S101 to S105.

[0034] S101, extracting differential entropy features of EEG data for training for different sub-bands at different resolutions, and constructing a domain adaptive neural network based on the differential entropy features;

[0035] S102, performing prediction voting on the EEG data of the target user based on the domain adaptive neural network to obtain individual emotion prediction label data;

[0036] S103, extracting deep visual features and deep auditory features from preset audio-visual content through a deep convolutional network model, and fusing the deep visual features and deep auditory features into deep audio-visual fusion features;

[0037] S104, constructing a hypergraph based on the deep visual features, the deep auditory features, and the deep audio-visual fusion features, and obtaining latent emotion prediction label data corresponding to the deep visual features, the deep auditory features, and the deep audio-visual fusion features through hypergraph segmentation;

[0038] S105 , assigning weights to the individual emotion prediction label data and the latent emotion prediction label data, and fusing them, and using the fused result as the emotion data prediction result.

[0039] In this embodiment, a multi-resolution domain adversarial neural network, namely the domain adaptive neural network (Multi-scale Domain Adversarial Neural Network, MsDANN), is first constructed based on the domain adversarial neural network to enhance the generalization ability of cross-individual EEG feature representation and the performance of model individualized prediction. In order to reduce the impact of individual differences in EEG signals, this embodiment uses audio-visual features to predict the potential emotional information therein as complementary information in emotion prediction. Due to the "semantic gap" between traditional features and emotions, traditional features cannot fully express the discriminative information related to emotions. Therefore, this embodiment proposes a hypergraph clustering method based on deep audio-visual features (DeepAudio-Visual Feature based Hypergraph Clustering Method, DAVFHC) for extracting discriminative high-level audio-visual features. The final emotion prediction result is determined by the decision layer fusion model, which is mainly achieved by giving the same weight to the individualized emotion prediction label data of EEG and the latent emotion prediction label data of audio-visual features, so that the complementary information of different modalities is used for emotion prediction.

[0040] This embodiment combines EEG data and audio-visual features to perform multimodal prediction, providing more complementary information missing in single-modality for emotion prediction, enabling more accurate modeling. At the same time, it can also fully express the discriminative features related to emotions, thereby improving the accuracy of emotion prediction.

[0041] Combination Figure 4As shown, this embodiment includes an individualized emotion prediction module based on EEG, a latent emotion prediction module based on audiovisual features, and a multimodal fusion module. In the individualized emotion prediction module based on EEG, the multi-resolution representation of the EEG signal is first extracted, and then the feature extractor network in the domain adaptive neural network (i.e., the multi-resolution domain adversarial neural network) is used to extract features, and then the extracted features are classified and discriminated by the task classifier network and the discriminator network, so as to obtain individual emotion prediction label data corresponding to the individualized emotion. In the latent emotion prediction module based on audiovisual features, fragment-based visual features and fragment-based auditory features are extracted at the visual level and the auditory level, respectively, and then the latent emotion prediction label data corresponding to the latent emotion is obtained by hypergraph clustering. The individual emotion prediction label data and the latent emotion prediction label data are fused by the multimodal fusion module to obtain the final emotion data prediction result.

[0042] In one embodiment, if Figure 2 As shown, the step S101 includes: steps S201 to S205.

[0043] S201, obtaining EEG data with emotion labels in a training set and setting it as a source domain; and obtaining EEG data without emotion labels in a test set and setting it as a target domain;

[0044] S202, respectively obtaining a source domain feature representation and a target domain feature representation of the source domain and the target domain;

[0045] S203, inputting the source domain feature representation and the target domain feature representation into the generator, and sequentially passing through the first fully connected layer, the first ELU layer, the second fully connected layer, the second ELU layer, the third fully connected layer, and the third ELU layer in the generator to obtain training features and test features accordingly;

[0046] S204, inputting the training features and the corresponding training labels into a classifier, and performing classification prediction through a fourth fully connected layer in the classifier;

[0047] S205, inputting the training features and the test features into the discriminator, and performing discrimination prediction in sequence through the fifth fully connected layer, the RELU layer and the sixth fully connected layer in the discriminator.

[0048] In this embodiment, first, the differential entropy (DE) features of EEG data are extracted from different sub-bands at different resolutions (e.g., 1 Hz, 0.5 Hz, and 0.25 Hz, etc.). Then, these differential entropy features are used to construct a domain adaptive neural network based on transfer learning, MsDANN, and the domain adaptive neural network is trained by the domain adversarial method, so as to solve the problem of individual differences in the process of EEG-based emotion prediction. Specifically, the EEG data of different individuals are regarded as different domains, the source domain refers to the information of existing individuals, and the target domain refers to the information of newly added individuals. Based on the input features of different resolutions, the feature extractor network, the task classification network, and the discriminator are respectively designed to extract features with discriminativeness and domain invariance in the source domain and the target domain, and make the feature distributions of the source domain and the target domain similar or close, so that the source domain and the target domain can be predicted on the same prediction model.

[0049] Combination Figure 5 , the network structure of the domain adaptive neural network (MsDANN) mainly includes three parts: a generator (feature extractor network) for extracting deep features, a classifier (task classification network) for predicting emotion labels, and a discriminator (discriminator) for identifying true and false data. The generator and the classifier can be regarded as standard forward structures. The generator and the discriminator are trained by the method of reverse gradient layer to ensure that the feature distribution of the two domains is as difficult to distinguish as possible. In this embodiment, the EEG data with emotion labels is regarded as the source domain for training the generator, classifier and discriminator; and the EEG data without emotion labels is regarded as the target domain for training the generator and the discriminator. Through this multi-resolution deep framework, a series of transferable features related to emotional information are extracted, so that cross-domain differences can be interoperable; at the same time, the classification performance of the source domain and the target domain can be effectively improved. Here, since the data sample may come from the source domain or the target domain, the role of the discriminator is to determine whether the data sample belongs to the source domain or the target domain.

[0050] In one embodiment, the step S102 includes:

[0051] Extract high-resolution feature representation, medium-resolution feature representation and low-resolution feature representation of the target user's EEG data respectively;

[0052] The high-resolution feature representation is sequentially input into the first generator and the first classifier to obtain a high-resolution label; the medium-resolution feature representation is sequentially input into the second generator and the second classifier to obtain a medium-resolution label; the low-resolution feature representation is sequentially input into the third generator and the third classifier to obtain a low-resolution label;

[0053] Voting is performed on the high-resolution labels, medium-resolution labels, and low-resolution labels, and the voting results are used as individual emotion prediction label data.

[0054] In this embodiment, combined with Figure 5 When using the domain adaptive neural network to classify and predict EEG data, high-resolution feature representation, medium-resolution feature representation and low-resolution feature representation are first extracted from the EEG data, and then the generator and classifier are used to classify the high-resolution feature representation, medium-resolution feature representation and low-resolution feature representation respectively, and the corresponding high-resolution labels, medium-resolution labels and low-resolution labels are obtained, and then the obtained resolution labels are voted to obtain the final individual emotion prediction label data.

[0055] In one embodiment, the multimodal emotion data prediction method based on EEG data further includes:

[0056] The domain adversarial training objective function E of the domain adaptive neural network is constructed according to the following formula:

[0057]

[0058] In the formula, and denote the source domain and the target domain respectively, x l is the EEG data with emotion labels, z l for In the unlabeled EEG data, θ, σ and μ are all parameters;

[0059] The binary cross-entropy loss function of the discriminator is constructed according to the following formula:

[0060]

[0061] In the formula, r θ and d μ Represent the generator and discriminator respectively;

[0062] The loss function of the classifier is constructed as follows:

[0063]

[0064] In the formula, is the classification loss of the source domain.

[0065] In this embodiment, in order to learn the feature space shared by the source domain and the target domain and ensure that the learned features contain enough information to reveal the emotional state, the loss objective function is designed as follows. Assume that the source domain and the target domain are and In domain learning, The EEG data with emotion labels is x l = and and is the feature of the EEG input data represented by the lth frequency domain resolution, y i yes The corresponding emotion label. is x l On the other hand, The unlabeled EEG data in express, is the feature of the EEG input data represented by the lth frequency domain resolution, Yes l This embodiment uses r with parameters θ, σ and μ. θ 、c σ and d μ Represent the generator, classifier and discriminator respectively. In order to ensure r θ The features learned from the source domain or the target domain are indistinguishable, and the domain adversarial training objective function is as follows:

[0066]

[0067] Here, is the binarized cross entropy loss of the discriminator, which is used to train the distinction and The definition is as follows:

[0068]

[0069] Here, For the classifier part, this embodiment adds another new loss function based on the above formula: As the loss function of the classifier, it is as follows:

[0070]

[0071] Here, is the classification loss of the source domain, given by ,λ is the balance parameter in the learning process and is defined as follows:

[0072]

[0073] Here, γ and p are the constant and factor respectively in each pass of the algorithm.

[0074] Among them, the loss function of the classifier is the final objective function of MsDANN model training.

[0075] In one embodiment, if Figure 3As shown, the step S103 includes: steps S301 to S306.

[0076] S301, extracting all frames of visual information from preset audio-visual content, and inputting each frame of visual information into a VGG16 network;

[0077] S302, using each convolutional layer in the VGG16 network to extract a feature map of each frame of visual information, and calculating a corresponding average feature map under the feature map of each convolutional layer;

[0078] S303, based on the average feature map of each convolutional layer, using an adaptive method to extract key frame features of each convolutional layer;

[0079] S304, concatenating the key frame features corresponding to the last two convolutional layers into the deep visual features;

[0080] S305, dividing the auditory information in the preset audio-visual content into multiple auditory segments without overlapping, using each convolutional layer in the VGGish network to calculate the average feature map corresponding to each of the auditory segments, and splicing the average feature maps corresponding to the last two convolutional layers into the deep auditory feature;

[0081] S306: Fuse the deep visual feature and the deep auditory feature into the deep visual and auditory fusion feature.

[0082] In this embodiment, deep visual features and deep auditory features are extracted through a pre-trained VGG16 network and a VGGish network, respectively.

[0083] The VGG16 network structure contains 13 convolutional layers and 3 fully connected layers. The number of convolution kernels in each convolution layer is 64, 64, 128, 128, 256, 256, 256, 512, 512, 512, 512, 512, 512, and the size of the convolution kernel is 3×3.

[0084] There are four steps to extract the deep visual features:

[0085] ① Extract frame visual features. Each frame of the video is input into the VGG16 network, and the feature maps corresponding to each frame in each convolution layer are extracted. For each convolution layer, the corresponding average feature map is calculated as the feature vector of the layer.

[0086] ② Extract visual features of the segment. This embodiment uses an adaptive method to extract key frames in each audiovisual segment to represent the video segment. Specifically, the video is segmented into 1 second segments without overlap. Assuming that each segment contains k frames, ι=1,…N, represents the ιth convolutional layer, and each frame is extracted through the VGG16 network. The steps of key frame extraction are as follows:

[0087] B ι All frames of are clustered into one category using clustering method;

[0088] Find the center point c of the cluster ι ;

[0089] Calculate each frame and cluster center c ι The distance is expressed as

[0090] Select the frame with the smallest distance from the center point as the key frame of the segment, denoted as

[0091] The corresponding key frame features are regarded as the features of the video clip.

[0092] ③ Visual feature fusion of video clips. In this embodiment, the visual features of the last two convolutional layers (ι=12, 13) are fused in a splicing manner as the deep visual feature Ψ obtained by the DAVFHC method. V .

[0093] For extracting deep auditory features, this embodiment uses the pre-trained convolutional neural network model VGGish for extraction. The network structure has 6 convolutional layers, the number of convolution kernels is 64, 128, 256, 256, 512 and 512 respectively, and the convolution kernel size is 3×3. First, the auditory information in the video content is divided into 10 audio clips without overlap according to the duration of 1 second, and then the convolutional features of each convolutional layer of each audio clip are extracted using the pre-trained VGGish network, and then the auditory features of the last two convolutional layers (ι=5, 6) are fused in a splicing manner as the deep auditory feature Ψ obtained by the DAVFHC method. A .

[0094] The deep visual feature Ψ V and the corresponding deep auditory feature Ψ A The deep audio-visual fusion feature is obtained by fusion, and there is a deep audio-visual fusion feature Ψ M =[Ψ V Ψ A ].

[0095] In one embodiment, the step S104 includes:

[0096] The audiovisual content segments corresponding to the deep visual features, deep auditory features, and deep audiovisual fusion features are set as vertices of the hypergraph, and the similarity between any two vertices is calculated according to the following formula, and then the hypergraph is constructed based on this:

[0097]

[0098] In the formula, and For any two vertices, N M is the feature dimension;

[0099] Segmenting the hypergraph into a plurality of clusters corresponding to the emotional states by a spectral hypergraph segmentation method;

[0100] The clusters are normalized, the normalized clusters are optimally segmented using a real value optimization method, and the optimal segmentation results are used as the latent emotion prediction label data.

[0101] In this embodiment, deep visual features, deep auditory features, and deep audio-visual fusion features are used to construct a hypergraph in the Valence and Arousal dimensions based on the principle of hypergraph partition to perform unsupervised prediction of the latent emotions of each clip. The complex relationship between each video clip is constructed through a hypergraph, which is regarded as a method for describing complex hidden data relationships. In a traditional graph, only two paired vertices can be connected, which will lead to information leakage. In a hypergraph, an edge (also called a hyperedge in a hypergraph) can connect more than two vertices, and the relationship between vertices can be well described. In the embodiment, it is assumed that the hypergraph is G={V,E}, E={e 1 ,e 2 ,e 3 ,…,e |E|} is the set of hyperedges, V = {v 1 ,v 2 ,v 3 ,…,v |V|} is a set of vertices. Belongs to the hyperedge e k The vertex set of ∈E is denoted as To define the relationship between vertices and hyperedges, any two vertices (emotion-inducing video clips) and (N M The similarity between (is the feature dimension) is defined as:

[0102]

[0103] and is the distance between two vertices, calculated by the following formula:

[0104]

[0105] Based on the calculated similarity matrix (N is the sample size), the association matrix can be calculated as H∈|V|×|E|, and the relationship between the vertex V and the hyperedge E is expressed as follows:

[0106]

[0107] The weight matrix W of the hypergraph is a diagonal matrix that represents the weights of all hyperedges E in the hypergraph G. k ∈E weight w(e k ) is based on the same hyperedge e k The similarity matrix between the vertices is calculated as follows:

[0108]

[0109] is the vertex v i and v j τ is the value of the similarity connected to the hyperedge e k The number of vertices. k ) is a measure of the similarity relationship between all vertices belonging to the same hyperedge. The larger w(e k ) value indicates that vertices with similar attributes belonging to the same hyperedge have a stronger connection relationship, while a small w(e k ) value indicates that the vertices belonging to the same hyperedge have weak connections, which means that these vertices have fewer similar attributes. In other words, the hypergraph structure can well describe the attribute relationship between audiovisual clips. The order matrix of the vertex (D v ) is a diagonal matrix representing the order of all vertices in the hypergraph G. A vertex v k The order of ∈V is the sum of the weights of all hyperedges to which the vertex belongs, defined as follows:

[0110]

[0111] The rank matrix of the hyperedge (D e ) is also a diagonal matrix, representing the order of all hyperedges in the hypergraph G. A hyperedge e k The order of ∈E refers to the sum of the orders of all vertices connected to the hyperedge, calculated as follows:

[0112]

[0113] The hypergraph problem can be solved by segmenting the constructed hypergraph into several clusters corresponding to the emotional state (high or low) through the spectral hypergraph segmentation method. Therefore, this is a bilateral hypergraph segmentation problem, which can be expressed by the following formula:

[0114]

[0115] Here, S and are the partition sets of vertex V. For the partition of two sides, is the complement of S. θS is the boundary of the segmentation, defined as d(e) is the order of the hyperedge. To prevent unbalanced partitioning, is normalized to:

[0116]

[0117] vol(S) and They are S and The volume of v∈S d(v) and The rule for segmentation is to find the partitioning rules that make S and The weakest connection between the two partitions and the close connection within each partition (large hyperedge weight value). Finding the weakest connection between two partitions is an NP-complete problem that can be solved by real-valued optimization methods. The optimal partition is calculated by the following formula:

[0118]

[0119]

[0120] Here, Θ is:

[0121]

[0122] I is the identity matrix with the same number of rows and columns as W. The Laplacian matrix of the hypergraph is defined as:

[0123] Δ=I-Θ。

[0124] The optimal solution to this problem is transformed into finding the eigenvector of the minimum eigenvalue of Δ. In other words, the optimal hypergraph segmentation result is to find the vector corresponding to the minimum non-zero eigenvalue of Δ to form a new feature space, and this feature space is used for subsequent K-means-based clustering. Through this method, all vertices are clustered into two classes, and the emotional state corresponding to each class is determined by the emotional state of the majority of vertices within the class. If the emotional state of most vertices in the class belongs to a high emotional level, then the class is designated as a high emotional level. If the emotional state of most vertices in the class belongs to a low emotional level, then the class is designated as a low emotional level. In practice, in order to prevent information leakage, the emotional state within the class is determined only by training samples.

[0125] In one embodiment, the step S105 includes:

[0126] The individual emotion prediction label data and the latent emotion prediction label data are weighted and fused according to the following formula:

[0127]

[0128] In the formula, Predict label data for individual emotions, is the potential emotion prediction label data, w EEG and w MUL are the weights of individual emotion prediction label data and latent emotion prediction label data in the fusion process, This is the final multimodal fusion emotion prediction result.

[0129] In this embodiment, based on the above steps, the prediction labels of deep visual features, deep auditory features and deep audio-visual fusion features (i.e., the latent emotion prediction label data) and the corresponding individualized prediction labels of EEG features (i.e., the individual emotion prediction label data) are used to carry out decision-level fusion and calculate the final prediction label of each segment. In other words, the EEG data and audio-visual information are mainly fused by giving them the same weights.

[0130] In one embodiment, the emotion data prediction result is evaluated according to the following formula:

[0131]

[0132]

[0133] In the formula, Accuracy and F1-score are both evaluation indicators, n TN and n TP is the sample of correct prediction, n FN and n FP is the sample with incorrect prediction, P pre and P sen They are accuracy and sensitivity.

[0134] The individual-based true label refers to the different labels that each subject gives in the Valence and Arousal dimensions when watching the video. The cross-individual-based true label means that all subjects have the same emotional label when watching the same video. Accuracy is an indicator to measure the overall prediction performance, while F1-score is the harmonic mean of precision and sensitivity, which is not easily affected by the imbalanced classification problem.

[0135] In one embodiment, evaluation is performed on the Valence and Arousal dimensions based on individual and cross-individual true labels, respectively, and the results are shown in Tables 1 and 2 below.

[0136]

[0137] Table 1

[0138] In Table 1, EEG represents the predicted label of EEG signal in MsDANN network; Fusion represents the predicted label of deep audio-visual fusion feature in hypergraph segmentation method; Visual represents the predicted label of deep visual feature in hypergraph segmentation method; Audio represents the predicted label of deep auditory feature in hypergraph segmentation method.

[0139]

[0140] Table 2

[0141] In Table 2, EEG represents the predicted label of EEG signal in MsDANN network; Fusion represents the predicted label of deep audio-visual fusion feature in hypergraph segmentation method; Visual represents the predicted label of deep visual feature in hypergraph segmentation method; Audio represents the predicted label of deep auditory feature in hypergraph segmentation method.

[0142] The higher the values ​​in Table 1 and Table 2, the better the prediction performance. At the same time, this shows that in the two dimensions of Valence and Arousal, the emotion prediction accuracy of the method provided by the embodiment of the present invention that integrates EEG, visual features and auditory features is better than the emotion prediction accuracy of EEG and visual features or auditory features.

[0143] The effectiveness of the domain adversarial network model was evaluated based on individual and cross-individual true labels in the Valence and Arousal dimensions, respectively. The results are shown in Tables 3 and 4.

[0144]

[0145] Table 3

[0146] In Table 3, EEG represents the predicted label of EEG signal in MsDANN / MsNN network; Fusion represents the predicted label of deep audio-visual fusion feature in hypergraph segmentation method; Visual represents the predicted label of deep visual feature in hypergraph segmentation method; Audio represents the predicted label of deep auditory feature in hypergraph segmentation method.

[0147]

[0148] Table 4

[0149] In Table 4, EEG represents the predicted label of EEG signal in MsDANN or MsNN network; Fusion represents the predicted label of deep audio-visual fusion feature in hypergraph segmentation method; Visual represents the predicted label of deep visual feature in hypergraph segmentation method; Audio represents the predicted label of deep auditory feature in hypergraph segmentation method.

[0150] The data in Tables 3 and 4 are the comparison of the results of the decision fusion of the labels generated by the two network models MsDANN and MsNN (Multi-scale Neural Network, multi-resolution neural network without deep domain adaptation) with the label decision fusion of the deep features of the video content. First, in the Valence and Arousal dimensions, the decision fusion results of the EEG prediction labels generated by the MsDANN network model with the deep audio-visual fusion feature labels, deep visual feature labels, and deep auditory feature labels are better than the decision fusion results of the EEG prediction labels generated by the MsNN network model with the deep audio-visual fusion feature labels, deep visual feature labels, and deep auditory feature labels, which shows that the domain adversarial training method of the MsDANN network can effectively reduce individual differences in EEG data, which is conducive to emotion prediction modeling based on EEG data, thereby improving emotion prediction performance. Secondly, in the Valence and Arousal dimensions, the decision fusion results of the EEG prediction labels and deep audio-visual fusion feature labels generated by the two network models MsDANN and MsNN are better than the decision fusion results of the EEG prediction labels and deep visual features or deep auditory feature labels, which fully demonstrates that multimodal decision fusion can provide more discriminative information for emotion prediction, thereby improving the accuracy of emotion prediction.

[0151] Figure 6 A schematic block diagram of a multimodal emotion data prediction device 600 based on EEG data provided by an embodiment of the present invention, the device 600 includes:

[0152] A network construction unit 601 is used to extract differential entropy features of EEG data for training for different sub-bands at different resolutions, and to construct a domain adaptive neural network based on the differential entropy features;

[0153] A first prediction unit 602 is used to perform prediction voting on the EEG data of the target user based on the domain adaptive neural network to obtain individual emotion prediction label data;

[0154] A feature extraction unit 603 is used to extract deep visual features and deep auditory features from preset audio-visual content through a deep convolutional network model, and fuse the deep visual features and deep auditory features into deep audio-visual fusion features;

[0155] A second prediction unit 604 is used to construct a hypergraph based on the deep visual features, the deep auditory features and the deep audio-visual fusion features, and obtain latent emotion prediction label data corresponding to the deep visual features, the deep auditory features and the deep audio-visual fusion features through hypergraph segmentation;

[0156] The label fusion unit 605 is used to assign weights to the individual emotion prediction label data and the latent emotion prediction label data and fuse them, and use the fused result as the emotion data prediction result.

[0157] In one embodiment, if Figure 7 As shown, the network construction unit 601 includes:

[0158] The domain setting unit 701 is used to obtain EEG data with emotion labels in the training set and set it as the source domain; and obtain EEG data without emotion labels in the test set and set it as the target domain;

[0159] A representation acquisition unit 702, configured to respectively acquire a source domain feature representation and a target domain feature representation of the source domain and the target domain;

[0160] A feature output unit 703 is used to input the source domain feature representation and the target domain feature representation into the generator, and obtain training features and test features respectively after passing through the first fully connected layer, the first ELU layer, the second fully connected layer, the second ELU layer, the third fully connected layer, and the third ELU layer in the generator in sequence;

[0161] A classification prediction unit 704 is used to input the training features and the corresponding training labels into a classifier, and perform classification prediction through a fourth fully connected layer in the classifier;

[0162] The discriminant prediction unit 705 is used to input the training features and the test features into the discriminator, and perform discriminant prediction through the fifth fully connected layer, the RELU layer and the sixth fully connected layer in the discriminator in sequence.

[0163] In one embodiment, the first prediction unit 602 includes:

[0164] A representation extraction unit, used to respectively extract a high-resolution feature representation, a medium-resolution feature representation and a low-resolution feature representation of the EEG data of a target user;

[0165] A representation input unit, used to sequentially input the high-resolution feature representation into the first generator and the first classifier to obtain a high-resolution label; sequentially input the medium-resolution feature representation into the second generator and the second classifier to obtain a medium-resolution label; sequentially input the low-resolution feature representation into the third generator and the third classifier to obtain a low-resolution label;

[0166] The voting prediction unit is used to vote for the high-resolution label, the medium-resolution label and the low-resolution label, and use the voting result as the individual emotion prediction label data.

[0167] In one embodiment, the multimodal emotion data prediction device 600 based on EEG data further includes:

[0168] The first function construction unit is used to construct the domain adversarial training objective function E of the domain adaptive neural network according to the following formula:

[0169]

[0170] In the formula, and denote the source domain and the target domain respectively, x l is the EEG data with emotion labels, z l for In the unlabeled EEG data, θ, σ and μ are all parameters;

[0171] The second function construction unit is used to construct the binarized cross-entropy loss function of the discriminator according to the following formula:

[0172]

[0173] In the formula, r θ and d μ Represent the generator and discriminator respectively;

[0174] The third function construction unit is used to construct the loss function of the classifier according to the following formula:

[0175]

[0176] In the formula, is the classification loss of the source domain.

[0177] In one embodiment, if Figure 8 As shown, the feature extraction unit 603 includes:

[0178] A frame visual extraction unit 801 is used to extract all frame visual information from the preset audio-visual content, and input each frame visual information into the VGG16 network;

[0179] A feature map extraction unit 802 is used to extract a feature map of each frame of visual information using each convolutional layer in the VGG16 network, and calculate a corresponding average feature map under the feature map of each convolutional layer;

[0180] A key frame extraction unit 803 is used to extract key frame features of each convolution layer using an adaptive method based on the average feature map of each convolution layer;

[0181] A first splicing unit 804 is used to splice the key frame features corresponding to the last two convolutional layers into the deep visual features;

[0182] The second splicing unit 805 is used to divide the auditory information in the preset audio-visual content into multiple auditory segments without overlapping, calculate the average feature map corresponding to each of the auditory segments using each convolution layer in the VGGish network, and splice the average feature maps corresponding to the last two convolution layers into the deep auditory feature;

[0183] The feature fusion unit 806 is used to fuse the deep visual feature and the deep auditory feature into the deep visual and auditory fusion feature.

[0184] In one embodiment, the second prediction unit 604 includes:

[0185] The hypergraph construction unit is used to set the audiovisual content segments corresponding to the deep visual features, deep auditory features and deep audiovisual fusion features as vertices of the hypergraph, and calculate the similarity between any two vertices according to the following formula, and then construct the hypergraph based on this:

[0186]

[0187] In the formula, and For any two vertices, N M is the feature dimension;

[0188] A cluster segmentation unit, used for segmenting the hypergraph into a plurality of clusters corresponding to the emotional states by a spectral hypergraph segmentation method;

[0189] The optimal segmentation unit is used to normalize the clusters, optimally segment the normalized clusters by a real value optimization method, and use the optimal segmentation results as the latent emotion prediction label data.

[0190] In one embodiment, the tag fusion unit 605 includes:

[0191] The weighting and fusion unit is used to weight and fuse the individual emotion prediction label data and the latent emotion prediction label data according to the following formula:

[0192]

[0193] In the formula, Predict label data for individual emotions, is the potential emotion prediction label data, w EEG and w MUL are the weights of individual emotion prediction label data and latent emotion prediction label data in the fusion process, This is the final multimodal fusion emotion prediction result.

[0194] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, which will not be repeated here.

[0195] The embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed, the steps provided in the above embodiment can be implemented. The storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0196] The embodiment of the present invention also provides a computer device, which may include a memory and a processor, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps provided in the above embodiment may be implemented. Of course, the computer device may also include various network interfaces, power supplies and other components.

[0197] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

[0198] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.

Claims

1. A multimodal emotion data prediction method based on EEG data, It is characterized in that include: Extracting differential entropy features of EEG data for training for different sub-bands at different resolutions, and constructing a domain adaptive neural network based on the differential entropy features; Based on the domain adaptive neural network, predictive voting is performed on the EEG data of the target user to obtain individual emotion prediction label data; Extracting deep visual features and deep auditory features from preset audio-visual content through a deep convolutional network model, and fusing the deep visual features and deep auditory features into deep audio-visual fusion features; Constructing a hypergraph based on the deep visual features, deep auditory features and deep audio-visual fusion features, and obtaining latent emotion prediction label data corresponding to the deep visual features, deep auditory features and deep audio-visual fusion features through hypergraph segmentation; The individual emotion prediction label data and the latent emotion prediction label data are weighted and fused, and the fused result is used as the emotion data prediction result.

2. The multimodal emotion data prediction method based on EEG data according to claim 1, It is characterized in that The step of extracting differential entropy features of EEG data for training from different sub-bands at different resolutions, and constructing a domain adaptive neural network based on the differential entropy features, includes: Obtain EEG data with emotion labels in the training set and set it as the source domain; and obtain EEG data without emotion labels in the test set and set it as the target domain; Obtaining a source domain feature representation and a target domain feature representation of the source domain and the target domain respectively; The source domain feature representation and the target domain feature representation are input into the generator, and the training features and the test features are obtained after passing through the first fully connected layer, the first ELU layer, the second fully connected layer, the second ELU layer, the third fully connected layer, and the third ELU layer in the generator in sequence; Inputting the training features and corresponding training labels into a classifier, and performing classification prediction through a fourth fully connected layer in the classifier; The training features and the test features are input into the discriminator, and discriminant prediction is performed in sequence through the fifth fully connected layer, the RELU layer and the sixth fully connected layer in the discriminator.

3. The multimodal emotion data prediction method based on EEG data according to claim 1, It is characterized in that The predictive voting on the EEG data of the target user based on the domain adaptive neural network to obtain individual emotion prediction label data includes: Extract high-resolution feature representation, medium-resolution feature representation and low-resolution feature representation of the target user's EEG data respectively; The high-resolution feature representation is sequentially input into the first generator and the first classifier to obtain a high-resolution label; the medium-resolution feature representation is sequentially input into the second generator and the second classifier to obtain a medium-resolution label; the low-resolution feature representation is sequentially input into the third generator and the third classifier to obtain a low-resolution label; Voting is performed on the high-resolution labels, medium-resolution labels, and low-resolution labels, and the voting results are used as individual emotion prediction label data.

4. The multimodal emotion data prediction method based on EEG data according to claim 2, It is characterized in that Also includes: The domain adversarial training objective function E of the domain adaptive neural network is constructed according to the following formula: In the formula, and denote the source domain and the target domain respectively, x l is the EEG data with emotion labels, z l for In the unlabeled EEG data, θ, σ and μ are all parameters; The binary cross-entropy loss function of the discriminator is constructed according to the following formula: In the formula, r θ and d μ Represent the generator and discriminator respectively; The loss function of the classifier is constructed as follows: In the formula, is the classification loss of the source domain.

5. The multimodal emotion data prediction method based on EEG data according to claim 1, It is characterized in that The extracting of deep visual features and deep auditory features from preset audio-visual content by using a deep convolutional network model, and fusing the deep visual features and deep auditory features into deep audio-visual fusion features, includes: Extracting all frames of visual information from the preset audio-visual content, and inputting each frame of visual information into the VGG16 network; Utilizing each convolutional layer in the VGG16 network to extract a feature map of each frame of visual information, and calculating a corresponding average feature map under the feature map of each convolutional layer; Based on the average feature map of each convolutional layer, the key frame features of each convolutional layer are extracted using an adaptive method; Concatenating the key frame features corresponding to the last two convolutional layers into the deep visual features; The auditory information in the preset audio-visual content is divided into multiple auditory segments without overlap, each convolution layer in the VGGish network is used to calculate the average feature map corresponding to each of the auditory segments, and the average feature maps corresponding to the last two convolution layers are spliced ​​into the deep auditory feature; The deep visual feature and the deep auditory feature are fused into the deep audio-visual fusion feature.

6. The multimodal emotion data prediction method based on EEG data according to claim 1, It is characterized in that The step of constructing a hypergraph based on the deep visual features, the deep auditory features, and the deep audiovisual fusion features, and obtaining latent emotion prediction label data corresponding to the deep visual features, the deep auditory features, and the deep audiovisual fusion features through hypergraph segmentation includes: The audiovisual content segments corresponding to the deep visual features, deep auditory features, and deep audiovisual fusion features are set as vertices of the hypergraph, and the similarity between any two vertices is calculated according to the following formula, and then the hypergraph is constructed based on this: In the formula, and For any two vertices, N M is the feature dimension; Segmenting the hypergraph into a plurality of clusters corresponding to the emotional states by a spectral hypergraph segmentation method; The clusters are normalized, the normalized clusters are optimally segmented using a real value optimization method, and the optimal segmentation results are used as the latent emotion prediction label data.

7. The multimodal emotion data prediction method based on EEG data according to claim 1, It is characterized in that The step of weighting and fusing the individual emotion prediction label data and the latent emotion prediction label data, and using the fused result as the emotion prediction result, includes: The individual emotion prediction label data and the latent emotion prediction label data are weighted and fused according to the following formula: In the formula, Predict label data for individual emotions, is the potential emotion prediction label data, w EEG and w MUL are the weights of individual emotion prediction label data and latent emotion prediction label data in the fusion process, This is the final multimodal fusion emotion prediction result.

8. A multimodal emotion data prediction device based on EEG data, It is characterized in that include: A network construction unit, used for extracting differential entropy features of EEG data for training for different sub-bands at different resolutions, and constructing a domain adaptive neural network based on the differential entropy features; A first prediction unit is used to perform prediction voting on the EEG data of the target user based on the domain adaptive neural network to obtain individual emotion prediction label data; A feature extraction unit, configured to extract deep visual features and deep auditory features from preset audio-visual content through a deep convolutional network model, and fuse the deep visual features and deep auditory features into deep audio-visual fusion features; A second prediction unit is used to construct a hypergraph based on the deep visual features, the deep auditory features and the deep audio-visual fusion features, and obtain latent emotion prediction label data corresponding to the deep visual features, the deep auditory features and the deep audio-visual fusion features through hypergraph segmentation; The label fusion unit is used to assign weights to the individual emotion prediction label data and the latent emotion prediction label data and fuse them, and use the fused result as the emotion data prediction result.

9. A computer device, It is characterized in that It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the multimodal emotion data prediction method based on EEG data as described in any one of claims 1 to 7.

10. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multimodal emotion data prediction method based on EEG data as described in any one of claims 1 to 7 is implemented.