A multimodal emotion recognition method based on hypergraph hierarchical contrastive learning

By introducing hypergraph structure and multi-level contrast learning in multi-modal emotion recognition, the problems of insufficient utilization of modal connections and poor robustness in traditional technology are solved, and more accurate and efficient emotion recognition is achieved.

CN119739990BActive Publication Date: 2025-05-13湖南工商大学
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510253260.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-05-13
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

Traditional multimodal emotion recognition technology fails to make full use of the intrinsic connections and timing consistency between each mode, resulting in insufficient information fusion, insufficient cross-modal feature relationship modeling capabilities, and poor model robustness in actual scenarios such as noise, occlusion or modal loss.

Method used

Using a method based on hypergraph hierarchical comparison learning, by obtaining multimodal data and constructing a model interconnection graph, aggregating it into superpoints to generate hyperedges, performing multi-level comparison learning, calculating joint losses, enhancing the fusion of modal features, and improving the generalization ability and robustness of emotional recognition.

Benefits of technology

Effectively explore the potential correlation and timing consistency between multimodal data, improve the accuracy and robustness of emotion recognition, and enhance the ability to adapt to noise and modal heterogeneity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119739990B_ABST
    Figure CN119739990B_ABST
Patent Text Reader

Abstract

The present application relates to a multimodal emotion recognition method based on hypergraph hierarchical contrast learning, which extracts features from speech data, text data, and image data respectively; constructs nodes based on time series and based on the extracted speech modal features, text modal features, and image modal features, and constructs a modal interconnection graph with the similarity between two nodes as the weight of the edge; aggregates the modal interconnection graph of each time step into different superpoints, determines the generation of hyperedges with the semantic similarity between two superpoints, and constructs a first hypergraph; performs member masking and node masking on the first hypergraph to obtain a second hypergraph and a third hypergraph respectively; performs multi-level contrast learning based on the first hypergraph, the second hypergraph, and the third hypergraph, and calculates the joint loss; enhances the fusion features of the three modal features to obtain enhanced features, and jointly optimizes the enhanced features based on the joint loss to obtain a multimodal emotion representation, and predicts the emotion recognition result based on the multimodal emotion representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of multimodal emotion recognition, and in particular to a multimodal emotion recognition method based on hypergraph hierarchical contrast learning. Background Art

[0002] In the field of multimodal emotion recognition, the development of technology is moving towards a more precise and intelligent direction. As an important part of human-computer interaction, emotion recognition is widely used in social media analysis, online education, customer service systems, and medical emotion detection. Multimodal emotion recognition usually requires processing and analyzing data from multiple modalities such as speech, text, and images. These data provide emotional information in different dimensions. Speech data can capture emotional clues such as intonation and speaking speed, text data provides rich semantic information, and image data reflects intuitive emotional features such as facial expressions and postures.

[0003] However, traditional multimodal emotion recognition technologies still have some limitations. These technologies usually fail to fully utilize the intrinsic connections and temporal consistency between the modalities when processing multimodal data, resulting in insufficient information fusion and insufficient cross-modal feature relationship modeling capabilities. In addition, data of different modalities are heterogeneous, and there are significant differences in the feature spaces of speech, text, and images, which makes it particularly difficult to effectively model multimodal data under a unified framework. At the same time, existing methods have cross-model robustness when faced with actual scenarios such as noise, occlusion, or modality loss, making it difficult to meet complex application requirements. Summary of the invention

[0004] Based on this, it is necessary to provide a multimodal emotion recognition method based on hypergraph hierarchical contrast learning, which includes:

[0005] S1: Acquire multimodal data including speech, text, and image distributed in time, and perform feature extraction on speech data, text data, and image data respectively;

[0006] S2: Based on the time series, nodes are constructed based on the extracted speech modal features, text modal features, and image modal features, and the similarity between two nodes is used as the weight of the edge to construct a modal interconnection graph;

[0007] S3: Aggregate the modal interconnection graph of each time step into different superpoints, determine the generation of hyperedges based on the semantic similarity between two superpoints, and construct the first hypergraph;

[0008] S4: performing member masking and node masking on the first hypergraph to obtain a second hypergraph and a third hypergraph respectively; performing multi-level contrastive learning based on the first hypergraph, the second hypergraph, and the third hypergraph, and calculating a joint loss;

[0009] S5: Enhance the fusion features of the three modal features to obtain enhanced features, jointly optimize the enhanced features based on the joint loss to obtain a multimodal emotion representation, and predict the emotion recognition result based on the multimodal emotion representation.

[0010] Beneficial effects: This method introduces a hypergraph structure to model the relationship between multimodal data to explore the potential correlation and temporal consistency across modalities, and adopts a multi-level contrastive learning strategy to improve the generalization and robustness of emotion recognition. This new method can not only make full use of the complementarity of multimodal data, but also enhance the adaptability to noise and modal heterogeneity, providing technical support for more accurate and efficient emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0012] Figure 1 This is a flowchart of a multimodal emotion recognition method based on hypergraph hierarchical contrast learning in an embodiment of the present application. DETAILED DESCRIPTION

[0013] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present application, so the present application is not limited by the specific embodiments disclosed below.

[0014] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0015] like Figure 1 As shown, this embodiment provides a multimodal emotion recognition method based on hypergraph hierarchical contrast learning, the method comprising:

[0016] S1: Acquire multimodal data including speech, text, and image distributed in time, and perform feature extraction on the speech data, text data, and image data respectively.

[0017] Specifically, multimodal data is represented as: ,in, Indicates the length of the time series; in this embodiment, the value range of the time series length is 10-50 time steps.

[0018] The voice data is recorded as , which represents the speech signal at the tth time step, is the sampling frequency (typical value is 16kHz), is the signal length; the extraction of speech modal features includes:

[0019] Extract MFCC features, spectrogram features, pitch features, and volume features of speech data distributed over time.

[0020] Indicates The text sequence of time steps is composed of the character sequence {w1, w2, ..., wn}. The extraction of text modality features includes:

[0021] For text data distributed by time, the WordPiece algorithm is used for word segmentation, numerical normalization and special character processing. The processed text is input into the pre-trained BERT model to encode the text modal features.

[0022] , indicating the image data of time steps, where , are the image height and width (uniformly scaled to 224×224), 3 means the number of RGB channels is 3; the image modality feature extraction includes:

[0023] For image data distributed by time, they are uniformly scaled to 224×224 size and then standardized; the standardization calculation formula is:

[0024] ;

[0025] in, Represents pixels in the normalized image data; Represents pixels in the scaled image data; represents the mean value of pixels in the image data, Represents the standard deviation of pixels in the image data;

[0026] The standardized image data is augmented using random cropping and horizontal flipping techniques, and the augmented image data is input into the ResNet pre-trained model to extract the features of the pool5 layer output in the ResNet pre-trained model, which are recorded as image modality features. Among them, adaptive mean pooling adjusts the dimension of the image modality features.

[0027] Through the above-mentioned feature extraction process, this embodiment extracts multi-dimensional, high-quality feature representations from the three modalities of speech, text, and image, laying a solid foundation for subsequent correlation modeling between modalities.

[0028] S2: Based on the time series and the extracted speech modal features, text modal features, and image modal features, nodes are constructed, and the similarity between two nodes is used as the weight of the edge to construct a modal interconnection graph.

[0029] Specifically, the process of constructing the modal interconnection graph includes:

[0030] At time step t In the method, the speech modality feature, the text modality feature, and the image modality feature are respectively subjected to two layers of perceptrons to obtain the speech modality feature, the text modality feature, and the image modality feature with dimension adjustment; the calculation formula is:

[0031] ;

[0032] in, Represents the time step after dimension adjustment t The k-th modal feature at ; Represents the time step t The k-th modality feature at the time, including speech modality feature / text modality feature / image modality feature; Representing the two-layer perceptron of the corresponding modality, its forward propagation process is:

[0033] ;

[0034] in, Represents a fully connected layer with an output dimension of 512; m represents layer normalization, which is used to stabilize the training process; ReLU is the activation function (Rectified Linear Unit), which is used to introduce nonlinearity; Indicates that the dropout probability of the Dropout layer is 0.1, which is used to prevent overfitting.

[0035] The dimensionally adjusted speech modality features, text modality features, and image modality features are respectively used as nodes to construct a time step t The set of nodes at the time; time stept The node set at this time is recorded as: , Represents the time step after dimension adjustment t The speech modality characteristics of Represents the time step after dimension adjustment t The text modality characteristics of the time; Represents the time step after dimension adjustment t The image modality features at the time.

[0036] Calculate time steps t The similarity between any two nodes in the node set at the time is calculated as:

[0037] ;

[0038] in, Indicates i The node and j The similarity between nodes; Represents the time step t The first node in the node set i nodes; Represents the time step t The first node in the node set j nodes; represents scaled dot product attention; represents the first projection matrix; represents the second projection matrix; T represents transpose.

[0039] When the similarity between two nodes is greater than the adaptive threshold, an edge is generated, otherwise no edge is generated and the time step is constructed. t The edge set at the time; the calculation formula of the adaptive threshold is: ,in, , are the mean and standard deviation of similarity, is a tuning parameter (default is 2).

[0040] The weight of an edge is the similarity between the corresponding two nodes;

[0041] Based on the similarity measure between two nodes, the adjacency matrix at time step t is constructed. The calculation formula of the elements in the adjacency matrix is:

[0042] ;

[0043] in, Represents the time step t In the adjacency matrix i Line j Elements of a column; Indicates iThe node and j The edges between nodes; Represents the time step t The edge set of time; represents the sine function;

[0044] Based on time step t The node set, edge set, and adjacency matrix at the time step t Modal interconnection diagram at time step t The modal interconnection diagram is recorded as: ,in, Represents the time step t The node set when ; Represents the time step t The edge set when , that is, the edges connecting nodes of different modes; Represents the time step t The adjacency matrix is ​​used to describe the connection relationship between nodes.

[0045] By constructing a modal interconnection graph, this embodiment converts multimodal features at different times into structured node and edge forms, and uses an adaptive threshold similarity calculation method to dynamically generate connection relationships between modalities, thereby forming a graph structure representation with rich semantic information.

[0046] S3: Aggregate the modal interconnection graph of each time step into different superpoints, determine the generation of hyperedges based on the semantic similarity between two superpoints, and construct the first hypergraph.

[0047] Specifically, the construction process of the first hypergraph includes:

[0048] Step 1: Set the time step t The modal interconnection graph at the time of is passed through the graph convolutional network to obtain the hidden states of all the nodes in it;

[0049] Furthermore, the process of obtaining the hidden states of all nodes includes:

[0050] In an attention head, the feature of the node is updated as follows:

[0051] ;

[0052] ;

[0053] in, Indicated in d Graph Convolutional Network with Attention Heads l Layer i Feature representation of nodes; Indicates d Graph Convolutional Network with Attention Headsl The attention coefficient of the layer; represents the ELU activation function; Represents the graph convolutional network l The weight matrix of the layer; Indicated in d Graph Convolutional Network with Attention Heads l- 1st floor j Feature representation of neighbor nodes; represents the softmax function; represents the LeakyReLU activation function; T represents transposition; Represents a splicing operation; Indicates i The set of neighbor nodes of a node; Represents a learnable vector used to calculate attention weights;

[0054] Aggregate the updated node features in multiple attention heads. The aggregation formula is:

[0055] ;

[0056] in, Indicates that in the graph convolutional network l Layer i feature representation of nodes; K represents the number of attention heads (K=8); represents the softmax activation function; Indicates d The attention coefficient of the attention head; Indicates d Graph Convolutional Network with Attention Heads l The weight matrix of the layer; Indicates that in the graph convolutional network l- 1st floor i Feature representation of nodes;

[0057] The updated node features are residually connected and layer normalized, and the calculation formula is:

[0058] ;

[0059] in, Indicates that in the graph convolutional network l Layer i The hidden state of a node, Represents the layer normalization operation.

[0060] Step 2: Aggregate the hidden states of all nodes through the global attention pooling operation to obtain the time step t The aggregation result when ; the calculation formula is:

[0061] ;

[0062] in, Represents the time step t Aggregation results when Indicates that in the graph convolutional network l The hidden state of the first node in the layer; Indicates that in the graph convolutional network l Layer n The hidden state of each node; represents the global attention pooling operation, which is used to aggregate the hidden states of all nodes, which can be decomposed into:

[0063] ;

[0064] ;

[0065] in, Indicates i The pooling score of each node; Represents the sigmoid activation function; Represents the weight matrix of the pooling layer; Represents the bias vector of the pooling layer.

[0066] Step 3: At time step t Perform position encoding and get the time step t Positional encoding ; The calculation formula is:

[0067] ;

[0068] ;

[0069] in, Indicates that at time step t When the position code is i Dimensions; Indicates positions; c represents the total dimension of the position encoding; Represents a scaling factor used to adjust the frequency of position encoding.

[0070] Step 4: Set the time step t The aggregation result at time t is connected with the position encoding by residual connection, and the time step t The modal interconnection graph at this time is aggregated into a super point; the calculation formula is:

[0071] ;

[0072] in, Represents the time step t The modal interconnection graph at that time is aggregated into a super point.

[0073] Step 5: Repeat steps 1-4 until the modal interconnection graphs of all time steps are traversed, several superpoints are aggregated, and a superpoint set is constructed.

[0074] Step 6: Extract the text features from any two superpoints and convert them into the input format of the CLIP model, input them into the CLIP model, obtain the embedding vectors of the two text features, calculate the cosine similarity between the two embedding vectors to obtain the semantic similarity between the two superpoints;

[0075] The text feature extraction formula is:

[0076] ;

[0077] ;

[0078] in, Indicates i The text features of the super points; Indicates j The text features of the super points; Represents a text convolutional neural network model; Indicates i A super point; Indicates j A super point; Represents the text in a superpoint;

[0079] The semantic similarity calculation formula is:

[0080] ;

[0081] in, Indicates semantic similarity; Indicates i The embedding vector of each superpoint; Indicates j The embedding vector of each superpoint; represents the L2 norm.

[0082] Step 7: When the semantic similarity between two hyperpoints is greater than or equal to the preset threshold, a hyperedge is generated, otherwise no hyperedge is generated and a hyperedge set is constructed.

[0083] Step 8: The degree, closeness centrality, and betweenness centrality of the super-point are passed through a multi-layer perceptron to obtain the importance score of the super-point. The importance scores of the super-points connected by the hyper-edges are aggregated to calculate the weight of the hyper-edge.

[0084] The importance score calculation formula of the super point is:

[0085] ;

[0086] in, Indicates i The importance score of each super point; represents a multi-layer perceptron; Indicates i The degree of the superpoint, that is, i The number of hyperedges connecting the hyperpoints; Indicates i The closeness centrality of the superpoint is used to measure the i The average distance from a hyperpoint to all other hyperpoints in the hypergraph; Indicates i The betweenness centrality of the superpoint is used to measure the i The importance of a hyperpoint in all shortest paths in the hypergraph.

[0087] The calculation formula of the closeness centrality of the super point is:

[0088] ;

[0089] in, Indicates v The closeness centrality of superpoints; N represents the number of superpoints; V represents the set of superpoints; Indicates v The super point and u The shortest path length between super points (can be calculated using Euclidean distance).

[0090] The calculation formula of the betweenness centrality of a superpoint is:

[0091] ;

[0092] in, Indicates v Betweenness centrality of superpoints; Indicates s Super point to t The total number of shortest paths between superpoints; Indicates s Super point to t The shortest path between the superpoints passes through v The number of paths to each superpoint.

[0093] The weight calculation formula of the hyperedge is:

[0094] ;

[0095] in, Represents a hyperedge The weight of Represents a hyperedge The number of connected superpoints.

[0096] Step 9: Construct the first hypergraph based on the superpoint set and the hyperedge set.

[0097] Through hypergraph construction, this embodiment aggregates the modal interconnection graph into hyperpoints, generates hyperedges based on semantic similarity, and finally constructs a complete hypergraph structure.

[0098] S4: Perform member masking and node masking on the first hypergraph to obtain a second hypergraph and a third hypergraph respectively; perform multi-level comparative learning based on the first hypergraph, the second hypergraph, and the third hypergraph, and calculate the joint loss.

[0099] Specifically, the construction process of the second hypergraph includes:

[0100] Perform member masking and node masking on the first hypergraph to obtain a second hypergraph and a third hypergraph respectively;

[0101] Define the member mask matrix, expressed as:

[0102] ;

[0103] ;

[0104] in, Indicates k The member mask matrix of each modality feature; Indicates k The member mask probability of each modal feature (the value range of each modality is [0.1, 0.3, 0.2]); Indicates k The basis member mask probability of each modal feature; represents the attenuation coefficient (the value is 0.1); t Represents the time step t ; Indicates the temperature parameter (the value is 100); Indicates the maximum member masking probability (value is 0.5); Indicates the probability of masking according to members Generates a Bernoulli distributed random variable.

[0105] The member mask feature is constructed based on each modal feature and the member mask matrix. The calculation formula is:

[0106] ;

[0107] in, Indicates kMember mask features of modal features; Indicates k modal features, including speech modal features / text modal features / image modal features; represents the Hadamard product; Represents the second multi-layer perceptron, whose forward propagation process is:

[0108] ;

[0109] in, Represents a fully connected layer, with both input and output dimensions being ; BN stands for Batch Normalization, which is used to stabilize the training process.

[0110] The member mask features are compensated based on the features of each modality, and the calculation formula is:

[0111] ;

[0112] ;

[0113] in, Indicates the first k Member mask features of modal features; represents the compensation coefficient; Indicates k Member mask features of modal features; Indicates k modal features, including speech modal features / text modal features / image modal features; Represents the sigmoid activation function; represents the gating function; Represents a splicing operation;

[0114] Based on the member mask features after feature compensation, a block-wise masking strategy is used to perform member masking on each modal feature in the corresponding super point to obtain the second hypergraph.

[0115] The construction process of the third hypergraph includes:

[0116] Step 1: Sample the hyperpoints in the first hypergraph based on Beta distribution, and use DropEdge technology to trim the hyperedges in the first hypergraph to obtain a subgraph;

[0117] Beta distribution: , =2, =5;

[0118] DropEdge Technology: , Indicates u A super point, Indicates v A degree of super point; Represents the maximum degree among all hyperpoints in the hypergraph.

[0119] Step 2: Based on the adjacency matrix of the first hypergraph and the adjacency matrix of the subgraph, regularize the subgraph; the regularization formula is:

[0120] ;

[0121] in, represents the regularized subgraph; represents the square of the LF norm; represents the adjacency matrix of the first hypergraph; Represents the adjacency matrix of the subgraph; represents the regularization coefficient (the value is 0.01); represents the trace of the matrix; L represents the Laplacian matrix of the subgraph; H represents the hypergraph incidence matrix; T represents the transpose.

[0122] Step 3: The regularized subgraph passes through the graph convolutional network. In the graph convolutional network, the hidden states of neighboring super-points are aggregated through the aggregation function to obtain the information transfer result of the super-point. The calculation formula is:

[0123] ;

[0124] ;

[0125] ;

[0126] in, Indicates that in the graph convolutional network l Layer v The information transmission result of super points, represents the attention mechanism; Represents the transformer function; Indicates that in the graph convolutional network l- 1st floor u The hidden states of the neighboring superpoints; Indicates v The set of neighboring superpoints of a superpoint; ReLU activation function. represents the first weight matrix; represents the second weight matrix; represents the first bias vector; represents the second bias vector; Indicates u The attention weight of each superpoint; q represents the query vector, i.e. ; represents the hyperbolic tangent function; Indicates hidden state The weight of .

[0127] Step 4: Based on the graph convolutional network l Layer v The information transfer results of super points and the l- 1st floor v The hidden state of the superpoint is updated in the graph convolutional network through the gated recurrent unit. l Layer v The hidden state of a superpoint is calculated as:

[0128] ;

[0129] in, Indicates that in the graph convolutional network l Layer v The hidden state of a superpoint; represents a gated recurrent unit; Indicates that in the graph convolutional network l Layer v The information transmission result of each super point; Indicates that in the graph convolutional network l- 1st floor v The hidden state of a superpoint.

[0130] Step 5: Add the 0th and 1st layers of the graph convolutional network l Layer v The hidden states of the super nodes are skipped and connected to obtain v The aggregate hidden state of super points is calculated as:

[0131] ;

[0132] in, Indicates v The aggregate hidden state of super points; Indicates that in the 0th layer of the graph convolutional network v The hidden state of a superpoint; represents the learning parameters.

[0133] Step 6: Repeat steps 3-5, traverse all the superpoints in the subgraph, and obtain all the aggregated hidden states in the subgraph; construct a third hypergraph based on all the aggregated hidden states in the subgraph and the pruned hyperedges.

[0134] Furthermore, multi-level contrastive learning includes:

[0135] In the second hypergraph and the third hypergraph, two nodes of the same modality are used as member-level positive sample pairs, and two nodes of different modalities are used as member-level negative sample pairs;

[0136] Member-level contrastive learning is performed by calculating the member loss based on member-level positive sample pairs and member-level negative sample pairs. The member loss is to maximize the similarity of member-level positive sample pairs and minimize the similarity of member-level negative sample pairs.

[0137] Member loss The expression is:

[0138] ;

[0139] in, is the member-level positive sample pair set, is a set of member-level negative sample pairs; , , Respectively represent i , j , k The projection representation of nodes; represents cosine similarity; is the temperature parameter that controls the scale of contrastive learning.

[0140] In the second hypergraph and the third hypergraph, the two superpoints connected by the hyperedge are used as node-level positive sample pairs, and the two superpoints connected by the non-superpoints are used as node-level negative sample pairs;

[0141] Node-level contrastive learning is performed by calculating the node loss based on node-level positive sample pairs and node-level negative sample pairs. The node loss is to maximize the similarity of node-level positive sample pairs and minimize the similarity of node-level negative sample pairs.

[0142] Node loss The expression is:

[0143] ;

[0144] Among them, V is the super point set; is a set of node-level negative sample pairs; , , Super point v , node-level positive samples , node-level negative samples The feature representation of is the temperature parameter.

[0145] Performing data enhancement on the first hypergraph by using a random walk subgraph method to obtain an enhanced hypergraph; using the first hypergraph and the enhanced hypergraph as a graph-level positive sample pair, and using the first hypergraph and the second hypergraph / the third hypergraph as a graph-level negative sample pair;

[0146] Graph-level comparative learning is performed by calculating the hypergraph loss based on graph-level positive sample pairs and graph-level negative sample pairs. The hypergraph loss is to maximize the similarity of graph-level positive sample pairs and minimize the similarity of graph-level negative sample pairs.

[0147] Hypergraph Loss The expression is:

[0148] ;

[0149] in, represents a hypergraph set, represents the set of negative sample hypergraphs; , , Hypergraph , image-level positive samples , image-level negative samples Representation.

[0150] Furthermore, the member loss and node loss in the second hypergraph / third hypergraph are weighted summed with the hypergraph loss to obtain the intra-level contrast loss. , the calculation formula is:

[0151] ;

[0152] in, , , They are the first balance factor, the second balance factor, and the third balance factor, which are used to adjust the weights of different losses.

[0153] The first interaction loss is constructed by maximizing the similarity between a node and its corresponding superpoint and minimizing the similarity between a node and other non-corresponding superpoints. ; The calculation formula is:

[0154] ;

[0155] in, Indicates the number of nodes; Indicates i Feature representation of nodes; Representation Node i The corresponding super point j The feature representation of Represents a negative sample node.

[0156] Through the loss function , the consistency between node features and super-point features can be learned, further enhancing the relationship modeling between multi-modalities.

[0157] The second interaction loss is constructed by maximizing the similarity between a hyperpoint and its corresponding hypergraph and minimizing the similarity between the node and other non-corresponding hypergraphs. ; The calculation formula is:

[0158] ;

[0159] Where N represents the number of super points; Indicates i Feature representation of super points; Indicates super point i Hypergraph g The global representation of Indicates super point i Non-belonging hypergraph h The global representation of .

[0160] Loss Function It is used to model the relationship between the hyperpoint and the global representation, ensuring the feature consistency between the local (hyperpoint) and the whole (hypergraph), thereby enhancing the robustness of emotion recognition.

[0161] The third interaction loss is constructed by maximizing the similarity between a node and its corresponding hypergraph and minimizing the similarity between a node and other non-corresponding hypergraphs. ;

[0162] ;

[0163] in, Representation Node i Hypergraph g' The global representation of Representation Node i Non-belonging hypergraph h' The global representation of .

[0164] Through the loss function ,The nodes directly interact with the hypergraph, which can better align local and global information across levels and improve the model's ability to capture emotional information.

[0165] The first interaction loss, the second interaction loss, and the third interaction loss are weighted summed to get the total interaction loss. , the calculation formula is:

[0166] ;

[0167] in, , , They are the fourth balance factor, the fifth balance factor, and the sixth balance factor, which are used to adjust the weights of different losses.

[0168] The contrast loss within the layer is weighted summed with the total interaction loss, and a regularization term is added to obtain the joint loss , the calculation formula is:

[0169] ;

[0170] in, , , Respectively represent the first weight hyperparameter, the second weight hyperparameter, and the third weight hyperparameter; Represents a regularization term, which is used to control the complexity of the model.

[0171] Loss Function It can fully model multimodal sentiment information at different levels and scales, and take advantage of the hypergraph structure to improve the robustness and accuracy of recognition.

[0172] S5: Enhance the fusion features of the three modal features to obtain enhanced features, jointly optimize the enhanced features based on the joint loss to obtain a multimodal emotion representation, and predict the emotion recognition result based on the multimodal emotion representation.

[0173] Specifically, the enhanced features obtained by enhancing the fusion features of the three modal features include:

[0174] Step 1: Perform modal feature internal fusion on speech modal features through temporal attention weighting, local and global feature aggregation, and channel feature recalibration to obtain speech fusion features;

[0175] For text modality features, the hidden states of multiple layers in the BERT model are fused to obtain text fusion features;

[0176] The multi-scale feature extraction method is used to fuse the image modality features through the spatial attention mechanism and the channel attention mechanism to obtain the image fusion feature;

[0177] Step 2: Project the speech fusion features, text fusion features, and image fusion features into a common space;

[0178] Step 3: In the common space, the tensors of speech fusion features, text fusion features, and image fusion features are fused and compressed by Tucker decomposition;

[0179] Step 4: After tensor decomposition, perform weighted fusion on the speech fusion features, text fusion features, and image fusion features to obtain preliminary fusion features; calculation formula:

[0180] ;

[0181] in, represents the preliminary fusion features; Indicates k The fusion features of the modalities; Indicates that it is used for modal k Multilayer Perceptron (MLP); Representing modality k The weight of .

[0182] Step 5: Perform residual connection on the preliminary fusion feature and the vector formed by concatenating the speech fusion feature, text fusion feature, and image fusion feature to obtain the fusion feature; the calculation formula is:

[0183] ;

[0184] in, Indicates fusion features; represents the weight matrix; Represented by speech fusion features , text fusion features , Image fusion features The concatenated vector; Represents a concatenation operation.

[0185] Step 6: Pass the fused features through the Transformer encoder to obtain the enhanced features.

[0186] After tensor fusion and compression in step 3, a more compact and richer feature representation is obtained, which contains the comprehensive information of speech, text and image modalities; in step 4, the fused features of each modality are further processed by multi-layer perceptrons and assigned different weights to emphasize the importance of each modality in the final fused features. The features after tensor fusion can be regarded as a high-level representation that integrates multimodal information. Through Tucker decomposition, not only the key modal interaction information is retained, but also the redundancy is reduced, making the subsequent weighted fusion more efficient and meaningful. After such processing, each multi-layer perceptron in step 4 can focus more on extracting the unique information of the corresponding modality, and the weight corresponding to the modality ensures that the contribution of each modality to the final decision is appropriate.

[0187] Furthermore, the step of jointly optimizing the enhanced features based on the joint loss to obtain a multimodal emotion representation includes:

[0188] Step 1: Input the enhanced features into the emotion representation generation module to generate an emotion representation vector;

[0189] Step 2: Based on the joint loss Evaluating the difference between the sentiment representation vector and its corresponding true label;

[0190] Step 3: Minimize the joint loss , calculate the gradient through back-propagation, and use an optimizer (e.g., Adam optimizer) to update the parameters of the emotion representation generation module;

[0191] Step 4: Input the enhanced features into the optimized emotion representation generation module, and output the multimodal emotion representation.

[0192] In this embodiment, the emotion representation generation module adopts a Transformer encoder, whose self-attention mechanism can effectively integrate features from different modalities, capture global dependencies, and generate high-quality emotion representation vectors. The Transformer encoder used in this embodiment includes: an input layer for inputting enhanced features into the Transformer encoder; an encoding layer for integrating multimodal information in the enhanced features through several layers of self-attention and feedforward networks; and an output layer for extracting the final hidden state of the encoding layer and using it as the emotion representation vector.

[0193] Furthermore, predicting the emotion recognition result based on the multimodal emotion representation includes:

[0194] Step 1: Input the multimodal emotion representation into the fully connected layer and output the emotion classification probability; the calculation formula is:

[0195] ;

[0196] Among them, P represents the probability of sentiment classification; Represents the first weight matrix of the fully connected layer; Represents the second weight matrix of the fully connected layer; Represents the first bias vector of the fully connected layer; Represents the second bias vector of the fully connected layer; Representing multimodal sentiment representation;

[0197] Step 2: The multimodal sentiment representation is passed through a softmax function containing different modal weights, and the importance scores of different modalities are output respectively; the calculation formula is:

[0198] ;

[0199] in, Representing modalityk Importance score; Representing modality k The weight of ; T represents transpose.

[0200] Step 3: The emotion label corresponding to the maximum emotion classification probability is used as the final emotion label; the calculation formula is:

[0201] ;

[0202] in, Indicates the final sentiment label;

[0203] Step 4: The final emotion label and the importance score of each modality are output as the emotion recognition result.

[0204] In this embodiment, the emotion tags include: anger, surprise, frustration, happiness, fear, and sadness.

[0205] The multimodal emotion recognition method based on hypergraph hierarchical contrast learning provided in this embodiment introduces a hypergraph structure to model the relationship between multimodal data to explore the potential correlation and temporal consistency across modalities, and adopts a multi-level contrast learning strategy to improve the generalization and robustness of emotion recognition. This new method can not only make full use of the complementarity of multimodal data, but also enhance the adaptability to noise and modal heterogeneity, providing technical support for more accurate and efficient emotion recognition. This innovative technology is expected to promote the development of the field of multimodal emotion computing and inject new impetus into the advancement of human-computer interaction technology.

[0206] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0207] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the patent application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent application shall be subject to the attached claims.

Claims

1. A multimodal emotion recognition method based on hypergraph hierarchical contrast learning, characterized in that: include: S1: Acquire multimodal data including speech, text, and image distributed in time, and perform feature extraction on speech data, text data, and image data respectively; S2: Based on the time series, nodes are constructed based on the extracted speech modal features, text modal features, and image modal features, and the similarity between two nodes is used as the weight of the edge to construct a modal interconnection graph; The construction process of the modal interconnection graph includes: At time step t , the speech modality feature, the text modality feature, and the image modality feature are respectively passed through two layers of perceptrons to obtain the speech modality feature, the text modality feature, and the image modality feature with dimension adjustment; The dimensionally adjusted speech modality features, text modality features, and image modality features are respectively used as nodes to construct a time step t The node set at the time; Calculate time steps t The similarity between any two nodes in the node set at that time; When the similarity between two nodes is greater than the adaptive threshold, an edge is generated, otherwise no edge is generated and the time step is constructed. t The edge set of time; The weight of an edge is the similarity between the corresponding two nodes; Based on the similarity measure between two nodes, the adjacency matrix at time step t is constructed, and the calculation formula is: ; in, Represents the time step t In the adjacency matrix i Line j Elements of a column; Represents the time step t The first node in the node set i nodes; Represents the time step t The first node in the node set j nodes; Indicates i The node and j The edges between nodes; Represents the time step t The edge set of time; represents the sine function; Based on time step t The node set, edge set, and adjacency matrix at the time step t Modal interconnection diagram when ; S3: Aggregate the modal interconnection graph of each time step into different superpoints, determine the generation of hyperedges based on the semantic similarity between two superpoints, and construct the first hypergraph; S4: performing member masking and node masking on the first hypergraph to obtain a second hypergraph and a third hypergraph respectively; performing multi-level contrastive learning based on the first hypergraph, the second hypergraph, and the third hypergraph, and calculating a joint loss; S5: Enhance the fusion features of the three modal features to obtain enhanced features, jointly optimize the enhanced features based on the joint loss to obtain a multimodal emotion representation, and predict the emotion recognition result based on the multimodal emotion representation.

2. The multimodal emotion recognition method based on hypergraph hierarchical contrast learning according to claim 1, characterized in that: The extraction of speech modal features includes: Extract MFCC features, spectrogram features, pitch features, and volume features of speech data distributed over time; The extraction of text modality features includes: For text data distributed by time, the WordPiece algorithm is used for word segmentation, numerical normalization and special character processing. The processed text is input into the pre-trained BERT model to encode text modal features. The extraction of image modality features includes: For image data distributed by time, they are uniformly scaled to a size of 224×224 and then standardized. The standardized image data are augmented using random cropping and horizontal flipping techniques. The augmented image data are input into the ResNet pre-trained model, and the features output by the pool5 layer in the ResNet pre-trained model are extracted. The features are recorded as image modality features.

3. The multimodal emotion recognition method based on hypergraph hierarchical contrast learning according to claim 1, characterized in that: In S2, the time step is calculated t The similarity between any two nodes in the node set at the time is calculated as: ; in, Indicates i The node and j The similarity between nodes; Represents the time step t The first node in the node set i nodes; Represents the time step t The first node in the node set j nodes; represents scaled dot product attention; represents the first projection matrix; represents the second projection matrix; T represents transpose.

4. The multimodal emotion recognition method based on hypergraph hierarchical contrast learning according to claim 1, characterized in that: In S3, the construction process of the first hypergraph includes: Step 1: Set the time step t The modal interconnection graph at the time of is passed through the graph convolutional network to obtain the hidden states of all the nodes in it; Step 2: Aggregate the hidden states of all nodes through the global attention pooling operation to obtain the time step t Aggregation results when Step 3: At time step t Perform position encoding and get the time step t Position coding when Step 4: Set the time step t The aggregation result at time t is connected with the position encoding by residual connection, and the time step t The modal interconnection graph when is aggregated into a super point; Step 5: Repeat steps 1-4 until the modal interconnection graphs of all time steps are traversed, several superpoints are aggregated, and a superpoint set is constructed; Step 6: Extract the text features from any two superpoints, convert them into the input format of the CLIP model, input them into the CLIP model, obtain the embedding vectors of the two text features, calculate the cosine similarity between the two embedding vectors to obtain the semantic similarity between the two superpoints; Step 7: When the semantic similarity between two hyperpoints is greater than or equal to the preset threshold, a hyperedge is generated, otherwise no hyperedge is generated and a hyperedge set is constructed; Step 8: The degree, closeness centrality, and betweenness centrality of the super-point are passed through a multi-layer perceptron to obtain the importance score of the super-point. The importance scores of the super-points connected by the hyper-edges are aggregated to calculate the weight of the hyper-edge. Step 9: Construct the first hypergraph based on the superpoint set and the hyperedge set.

5. The multimodal emotion recognition method based on hypergraph hierarchical contrast learning according to claim 1, characterized in that: The construction process of the second hypergraph includes: Perform member masking and node masking on the first hypergraph to obtain a second hypergraph and a third hypergraph respectively; Define the member mask matrix, expressed as: ; ; in, Indicates k The member mask matrix of each modality feature; Indicates k Member mask probability of each modal feature; Indicates k The basis member mask probability of each modal feature; represents the attenuation coefficient; t Represents the time step t ; represents the temperature parameter; represents the maximum member mask probability; Indicates the probability of masking according to members Generate Bernoulli distributed random variables; The member mask feature is constructed based on each modal feature and the member mask matrix. The calculation formula is: ; in, Indicates k Member mask features of modal features; Indicates k modal features, including speech modal features / text modal features / image modal features; represents the second multi-layer perceptron; represents the Hadamard product; The member mask features are compensated based on the features of each modality, and the calculation formula is: ; ; in, Indicates the first k Member mask features of modal features; represents the compensation coefficient; Represents the sigmoid activation function; represents the gating function; Represents a splicing operation; Based on the member mask features after feature compensation, a block-wise masking strategy is used to perform member masking on each modal feature in the corresponding super point to obtain the second hypergraph.

6. The multimodal emotion recognition method based on hypergraph hierarchical contrast learning according to claim 1, characterized in that: The construction process of the third hypergraph includes: Step 1: Sample the hyperpoints in the first hypergraph based on Beta distribution, and use DropEdge technology to trim the hyperedges in the first hypergraph to obtain a subgraph; Step 2: Regularize the subgraph based on the adjacency matrix of the first hypergraph and the adjacency matrix of the subgraph; Step 3: The regularized subgraph passes through the graph convolutional network. In the graph convolutional network, the hidden states of neighboring super-points are aggregated through the aggregation function to obtain the information transfer result of the super-point. The calculation formula is: ; in, Indicates that in the graph convolutional network l Layer v The information transmission result of super points, represents the attention mechanism; Represents the transformer function; Indicates that in the graph convolutional network l- 1st floor u The hidden states of neighbor superpoints; Indicates v The set of neighboring superpoints of a superpoint; Step 4: Based on the graph convolutional network l Layer v The information transfer results of super points and the l- 1st floor v The hidden state of the superpoint is updated in the graph convolutional network through the gated recurrent unit. l Layer v The hidden state of a superpoint; Step 5: Add the 0th and 1st layers of the graph convolutional network l Layer v The hidden states of the super nodes are skipped and connected to obtain v The aggregate hidden state of super points; Step 6: Repeat steps 3-5, traverse all the superpoints in the subgraph, and obtain all the aggregated hidden states in the subgraph; construct a third hypergraph based on all the aggregated hidden states in the subgraph and the pruned hyperedges.

7. The multimodal emotion recognition method based on hypergraph hierarchical contrast learning according to claim 1, characterized in that: Multi-level contrastive learning includes: In the second hypergraph and the third hypergraph, two nodes of the same modality are used as member-level positive sample pairs, and two nodes of different modalities are used as member-level negative sample pairs; Member-level contrastive learning is performed by calculating the member loss based on the member-level positive sample pairs and the member-level negative sample pairs. The member loss is to maximize the similarity of the member-level positive sample pairs and minimize the similarity of the member-level negative sample pairs. In the second hypergraph and the third hypergraph, the two superpoints connected by the hyperedge are used as node-level positive sample pairs, and the two superpoints connected by the non-superpoints are used as node-level negative sample pairs; Node-level contrastive learning is performed by calculating node loss based on node-level positive sample pairs and node-level negative sample pairs. The node loss is to maximize the similarity of node-level positive sample pairs and minimize the similarity of node-level negative sample pairs. Performing data enhancement on the first hypergraph by using a random walk subgraph method to obtain an enhanced hypergraph; using the first hypergraph and the enhanced hypergraph as a graph-level positive sample pair, and using the first hypergraph and the second hypergraph / the third hypergraph as a graph-level negative sample pair; Graph-level comparative learning is performed by calculating the hypergraph loss based on graph-level positive sample pairs and graph-level negative sample pairs. The hypergraph loss is to maximize the similarity of graph-level positive sample pairs and minimize the similarity of graph-level negative sample pairs.

8. The multimodal emotion recognition method based on hypergraph hierarchical contrast learning according to claim 7, characterized in that: Performing a weighted summation of the member loss and the node loss in the second hypergraph / the third hypergraph and the hypergraph loss to obtain a contrast loss within the hierarchy; The first interaction loss is constructed by maximizing the similarity between a node and its corresponding superpoint and minimizing the similarity between the node and other non-corresponding superpoints; A second interaction loss is constructed by maximizing the similarity between a hyperpoint and its corresponding hypergraph and minimizing the similarity between the node and other non-corresponding hypergraphs; The third interaction loss is constructed by maximizing the similarity between a node and its corresponding hypergraph and minimizing the similarity between a node and other non-corresponding hypergraphs; The first interaction loss, the second interaction loss, and the third interaction loss are weighted summed to obtain the total interaction loss; The contrast loss within the layer is weightedly summed with the total interaction loss, and a regularization term is added to obtain the joint loss.

9. The multimodal emotion recognition method based on hypergraph hierarchical contrast learning according to claim 1, characterized in that: The enhanced features obtained by enhancing the fusion features of the three modal features include: Step 1: Perform modal feature internal fusion on speech modal features through temporal attention weighting, local and global feature aggregation, and channel feature recalibration to obtain speech fusion features; For text modality features, the hidden states of multiple layers in the BERT model are fused to obtain text fusion features; The multi-scale feature extraction method is used to fuse the image modality features through the spatial attention mechanism and the channel attention mechanism to obtain the image fusion feature; Step 2: Project the speech fusion features, text fusion features, and image fusion features into a common space; Step 3: In the common space, the tensors of speech fusion features, text fusion features, and image fusion features are fused and compressed by Tucker decomposition; Step 4: After tensor decomposition, the speech fusion features, text fusion features and image fusion features are weighted and fused to obtain preliminary fusion features; Step 5: Perform residual connection on the preliminary fusion feature and the vector formed by concatenating the speech fusion feature, text fusion feature, and image fusion feature to obtain the fusion feature; Step 6: Pass the fused features through the Transformer encoder to obtain the enhanced features.

10. The multimodal emotion recognition method based on hypergraph hierarchical contrast learning according to claim 1, characterized in that: The predicting of emotion recognition results based on the multimodal emotion representation comprises: Step 1: Input the multimodal emotion representation into the fully connected layer and output the emotion classification probability; Step 2: The multimodal sentiment representation is passed through a softmax function containing different modal weights, and the importance scores of different modalities are output respectively; Step 3: The emotion label corresponding to the maximum emotion classification probability is used as the final emotion label; Step 4: The final emotion label and the importance score of each modality are output as the emotion recognition result.

Citation Information

Patent Citations

  • Multivariable time series data anomaly detection method and system based on dynamic graph learning and long and short term convolution

    CN117251731A

  • Cloud side-end collaborative sparse traffic flow prediction privacy protection method and device

    CN119066712A