A classroom emotion recognition method based on multi-view graph neural network
Visual, voice, and text modal diagrams are constructed through multi-view diagram neural network, and the graph attention network is used for encoding and fusion, which solves the problem that teachers find it difficult to identify students' emotions in real time, and achieves high-precision and robust classroom sentiment analysis.
Patent Information
- Application Number
- CN202510235345.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-02-28
AI Technical Summary
In large-scale classrooms or online education scenarios, it is difficult for teachers to pay attention to each student's learning status and emotional response in real time. The existing technology is mostly limited to EEG signals and online dialogue data, and it is impossible to effectively use multimodal data for student sentiment analysis.
A multi-view diagram neural network is adopted to predict students' classroom emotions by constructing visual, speech, and text modal graphs, using graph attention networks to perform GAT encoding, and perform multi-modal adaptive fusion, combining multi-task total loss function training.
It realizes rapid and accurate identification of students' classroom emotions, solves the problems of modal isolation, noise sensitivity and insufficient long-range dependence modeling, improves the accuracy and robustness of sentiment analysis, and reduces manual observation and data organization.
Smart Images

Figure CN120147929B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning and emotion recognition technology, and in particular to a classroom emotion recognition method based on a multi-view graph neural network. Background Art
[0002] In modern education, the state of students in the classroom and the interaction between teachers and students are crucial factors influencing teaching quality. However, in large classrooms or online learning scenarios, it is often difficult for teachers to monitor each student's learning status and emotional reactions in real time. Currently, many studies are using artificial intelligence technologies, particularly graph neural networks (GNNs), to analyze students' emotional states in the classroom. However, these studies are often limited to EEG signals and online conversation data.
[0003] Therefore, in the classroom environment, how to seamlessly analyze students' emotions through existing multimodal data such as video surveillance is an urgent problem to be solved. Summary of the Invention
[0004] In order to overcome the defects in the above-mentioned prior art, the present invention provides a classroom emotion recognition method based on multi-perspective graph neural network, which uses graph neural network to construct multi-perspective features and combines multi-perspective feature fusion technology to more comprehensively and accurately identify students' classroom emotions.
[0005] To achieve the above object, the present invention adopts the following technical solutions, including:
[0006] A classroom emotion recognition method based on a multi-view graph neural network. The training method of the multi-view graph neural network is as follows:
[0007] S1, divide the student's classroom video into several video segments, and extract visual features, voice features, and text features from each video segment; each video segment is marked with classroom emotions;
[0008] S2: Based on visual features, speech features, and text features, a visual modal graph, a speech modal graph, and a text modal graph are constructed respectively. In each modal graph, the video segment is used as a node, and the visual features, speech features, and text features corresponding to each video segment are used as node features of the corresponding modal graph.
[0009] S3, use the graph attention network to perform GAT encoding on each modal graph respectively, update the node features in each modal graph respectively, and obtain the updated modal graphs;
[0010] S4, performing multimodal adaptive fusion based on the updated modal graphs to obtain multimodal adaptive fusion features;
[0011] S5, predicting students’ classroom emotions using multimodal adaptive fusion features;
[0012] S6, establish a multi-task total loss function, and train the multi-view graph neural network with the goal of minimizing the multi-task total loss function; the trained multi-view graph neural network is used to predict classroom emotions based on students' classroom videos.
[0013] Preferably, in step S2, the specific construction of the visual modality graph is as follows:
[0014] Calculate the cosine similarity of the visual features between two nodes as the visual similarity of the two nodes; set the threshold τ video , if the visual similarity between two nodes is greater than the threshold τ video , then an edge is built between the two nodes; otherwise, no edge is built between the two nodes;
[0015] The normalized visual similarity is used as the edge weight of the visual modality graph.
[0016] Preferably, in step S2, the speech modality graph is constructed as follows:
[0017] Calculate the DTW distance of speech features between two nodes and set the threshold τ audio ;
[0018]
[0019] Among them, Edge a(i,j) Indicates the edge construction between nodes i and j in the speech modality graph. If Edge a(i,j) =1, then an edge is built between nodes i and j. If Edge A(i,j) =0, no edge is built between nodes i and j; is the DTW distance of speech features between nodes i and j; max(·) is the maximum value function; are the speech features of nodes i and j respectively;
[0020] The normalized DTW distance is used as the edge weight of the speech modality graph.
[0021] Preferably, in step S2, the text modal diagram is constructed as follows:
[0022] Compute the semantic information increment and information flow coverage of the node:
[0023]
[0024] Among them, u i 、u j are the text segments corresponding to nodes i and j respectively, are the text features of nodes i and j respectively; SII(u i) is the semantic information increment of node i, k represents the k nodes before node i, that is, in the text segment u i The previous k text segments; IFC(u i ) is the information flow coverage of node i, m represents the m nodes after node i, that is, in text segment u i The following m text segments; w(u i ,u j ) is the text segment u i 、u j The semantic similarity between i ,u j ) is the text segment u i 、u j the distance between them;
[0025] The semantic information increment and information flow coverage of the node are weighted and calculated to obtain the comprehensive semantic information flow score of the node:
[0026] SIFS(u i )=α·SII(u i )+β·IFC(u i )
[0027] Among them, α and β are weight coefficients; SIFS(u i ) is the comprehensive score of the semantic information flow of node i;
[0028] Build the text enhancement feature of the node:
[0029]
[0030] Among them, Enhanced(u i ) is the text enhancement feature of node i;
[0031] Calculate the cosine similarity of the text enhancement features between two nodes as the text similarity and set the threshold τ text ;
[0032]
[0033] Among them, Edge t (i,j) represents the edge construction between nodes i and j in the text modal graph. t (i,j)=1, then an edge is built between nodes i and j. If Edge t (i, j) = 0, then no edge is built between nodes i and j; D is the maximum text distance; Sim(u i ,u j ) is the cosine similarity of the text enhancement features between nodes i and j, that is, the text similarity between nodes i and j;
[0034] The normalized text similarity is used as the edge weight of the text modality graph.
[0035] Preferably, in step S3, the GAT encoding is specifically as follows:
[0036] For node i and its neighbor nodes in the modality graph, calculate the attention coefficient between node i and each neighbor node:
[0037]
[0038] Among them, α ij is the attention coefficient between node i and its neighbor node j; W is the linear transformation matrix; a is the attention vector; ∥ represents vector concatenation; h i is the node feature of node i; is the set of neighbor nodes of node i; LeakyReLU is the activation function;
[0039] By using the attention coefficient between node i and each neighbor node, the features of each neighbor node are weighted summed to update the feature representation of node i:
[0040]
[0041] Where σ(·) is the activation function; is the updated node feature of node i;
[0042] Preferably, in step S4, the multimodal adaptive fusion is specifically as follows:
[0043] S41, perform nonlinear transformation and normalization on the node features after each modal update:
[0044]
[0045] Among them, the updated node features in the visual modal graph, speech modal graph, and text modal graph are W' v 、W' a 、W' t are the mapping matrices for vision, speech, and text respectively; Norm(·) is the normalization function; σ(·) is the activation function; They are respectively the normalized visual features, speech features, and text features;
[0046] S42 introduces a transformer-based fusion mechanism. For each attention head in the transformer, the corresponding query, key, and value matrices are constructed:
[0047]
[0048] in, are the linear mapping matrices for generating queries, keys, and values for the h-th attention head; Q h , K h 、V h are the query, key, and value matrices of the h-th attention head, respectively; h = 1, 2, …, H;
[0049] For each attention head, the cross attention output is calculated as follows:
[0050]
[0051] Where: d represents the feature dimension of each attention head; O h is the cross attention output of the h-th attention head;
[0052] Concatenate all attention heads to obtain multimodal fusion features:
[0053]
[0054] in, is the multimodal fusion feature of node i; Concat(·) is the concatenation function; W O is the linear mapping matrix of the multi-head attention output;
[0055] S43, introduces a multimodal adaptive gating mechanism to respectively normalize the visual features, speech features, and text features Generate the corresponding adaptive gating weights:
[0056]
[0057] in, is the linear mapping matrix used to generate adaptive gating weights; is the bias term; softmax(·) is the activation function; g v 、g a 、g t They are Adaptive gating weights of
[0058] Using adaptive gating weight g v 、g a 、g t right Perform weighted summation to obtain adaptive fusion features:
[0059]
[0060] in, is the adaptive fusion feature of node i;
[0061] S44, based on the multimodal fusion features and the adaptive fusion features, obtain the multimodal adaptive fusion features:
[0062]
[0063] in, is the multimodal adaptive fusion feature of node i.
[0064] Preferably, in step S5, the multimodal adaptive fusion features are used as input and passed through a two-layer fully connected network to output the classroom emotion prediction result.
[0065] Preferably, in step S6, a multi-task total loss function is constructed The details are as follows:
[0066] S61, using the multimodal adaptive fusion features of node i Perform emotion classification and obtain the probability of each emotion category of node i as p i,k , the true label of the node is y i,k , k=1,2,...,K, there are K emotion categories, calculate the multimodal classification loss
[0067]
[0068] Among them, α k is the category balance factor; γ is the modulation factor.
[0069] S62, based on the visual features of node i Voice Features Text features And the multimodal adaptive fusion features of node i Calculating contrastive loss
[0070]
[0071] Where s(·,·) represents the cosine similarity function; τ is a parameter;
[0072] S63, use the unimodal features of node i to perform emotion classification, and obtain the probability of each emotion category of node i in visual, speech, and text modes: Calculate unimodal classification loss
[0073]
[0074] Among them, mod∈{v,a,t} represents the visual, speech, and text modalities respectively;
[0075] S64, multi-task total loss function for:
[0076]
[0077] Among them, λ1, λ2, and λ3 are weight hyperparameters.
[0078] Preferably, in step S1, the I3D model is used to extract visual features from the visual information of the video segment, the Wav2Vec2.0 model is used to extract voice features from the voice information of the video segment, and the BERT Large encoding is used to extract text features from the text information of the video segment; the visual information includes facial expressions, the voice information includes voice intonation, and the text information includes text content.
[0079] The present invention also provides a computer program product, characterized in that it includes a computer program / instruction, which, when executed by a processor, implements the above-mentioned classroom emotion recognition method based on a multi-view graph neural network.
[0080] The advantages of the present invention are:
[0081] (1) This invention innovatively uses graph neural networks to construct multi-perspective features, integrate students' visual, voice, and text features, and quickly and accurately identify students' classroom emotions.
[0082] (2) The present invention adopts Graph Attention Networks (GAT) technology to collect students' multi-dimensional classroom performance data based on continuous video contactlessly, and combines it with multi-view feature fusion technology to obtain a more comprehensive and accurate classroom emotion recognition method based on multi-view graph neural network.
[0083] (3) The present invention solves the problems of modal isolation, noise sensitivity, and insufficient long-range dependency modeling in traditional classroom emotion recognition through multimodal graph structure modeling, cross-modal attention fusion, and multi-task joint optimization. It achieves high-precision and robust emotion analysis in complex classroom scenarios, providing a reliable technical foundation for intelligent education.
[0084] (4) The present invention constructs visual, speech, and text modal graphs respectively, and uses similarity metrics (cosine similarity between visual features, DTW distance between speech features, and cosine similarity between text enhancement features) to define the relationship between nodes, thereby solving the problem of data heterogeneity between different modalities.
[0085] (5) Visual modal graphs capture changes in facial expressions, speech modal graphs analyze temporal differences in intonation, and text modal graphs construct semantic coherence. The three cover the physical features (expression), temporal features (speech), and semantic features (text) of emotional expression, forming a multi-dimensional emotional representation and improving the accuracy and reliability of students' classroom emotion recognition.
[0086] (6) Semantic information increment and information flow coverage are introduced into the text modal graph to generate text enhancement features. The semantic evolution of the text (such as the emotional fluctuations caused by students' questions) is captured through information increment (forward k nodes) and coverage (backward m nodes). The text enhancement features integrate historical and future text segment information to solve the limitation of traditional text models that only focus on the current segment.
[0087] (7) The present invention can improve the efficiency of student classroom emotion assessment and reduce a large amount of manual observation and data collation work. The present invention can avoid the sampling errors caused by traditional methods and the experimental errors of directly collecting data from non-real environments. The multi-view graph neural network constructed by the present invention can dynamically obtain the results of student classroom emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 Flowchart of the training method for multi-view graph neural networks.
[0089] Figure 2 This is a comparison chart of the classification effects of the method of the present invention and other methods. DETAILED DESCRIPTION
[0090] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0091] A classroom emotion recognition method based on multi-view graph neural network uses multi-view graph neural network to predict students' classroom emotions based on classroom videos. Figure 1 As shown, the training method of the multi-view graph neural network is as follows:
[0092] S1, acquires classroom video data based on the camera, performs data preprocessing on the classroom video data, and detects faces in the classroom video based on the mature face detection algorithm in the existing technology, and then identifies the identity of the students through face recognition technology, and extracts the classroom video of each student. The student's classroom video is divided into several video segments. In this embodiment, the video is divided according to the student's speaking content, and each video segment corresponds to visual information (facial expression), voice information (voice intonation), and text information (speech content). Visual features, voice features, and text features are extracted from each video segment; each video segment is marked with classroom emotions. Since the present invention is a related study on classroom emotions, facial expressions, voice intonation, speech content, etc. may be important indicators reflecting students' classroom emotions. Therefore, these indicator data are collected from continuous video frames through deep learning technology.
[0093] S2: Based on the visual features, speech features, and text features, a visual modal graph, a speech modal graph, and a text modal graph are constructed. Each modal graph uses video segments as nodes, and the visual, speech, and text features corresponding to each video segment serve as the node features of the corresponding modal graph.
[0094] S21, Construction of visual modality graph.
[0095] The visual information includes students’ facial expressions, and a visual modal graph based on visual features is constructed.
[0096] Each video segment is used as a node of the visual modality graph, and the corresponding visual features are the node features of the visual modality graph. In this embodiment, the visual features are extracted from the visual information of the video segment using the I3D model.
[0097] Calculate the cosine similarity of the visual features between two nodes as the visual similarity of the two nodes; set the threshold τ video , if the visual similarity between two nodes is greater than the threshold τ video , then an edge is built between the two nodes; otherwise, no edge is built between the two nodes. video The choice of can be optimized through experiments to ensure that only strong sentiment similarity relationships are retained on the edges. The edge construction of the visual modality graph is as follows:
[0098]
[0099] Among them, Edge v(i,j) Indicates the edge construction between nodes i and j in the visual modality graph. If Edge v(i,j) =1, then an edge is built between nodes i and j. If Edge v(i,j) =0, no edge is built between nodes i and j; are the visual features of nodes i and j respectively, is the cosine similarity of the visual features between nodes i and j, that is, the visual similarity.
[0100] The edge weights of the visual modal graph are assigned according to the visual similarity values to represent the similarity of the students' emotions. In this embodiment, the standardized visual similarity is used as the edge weight of the visual modal graph.
[0101] S22, Construction of speech modality map.
[0102] Voice information can reflect the intensity and volatility of students' emotions and construct a voice modal graph based on voice features.
[0103] Each video segment is used as a node of the speech modality graph, and the corresponding speech feature is used as the node feature of the visual modality graph. In this embodiment, the Wav2Vec 2.0 model is used to extract speech features from the speech information of the video segment.
[0104] Calculate the DTW distance of the speech features between two nodes and use the dynamic time warping algorithm (DTW algorithm) to measure the similarity of time series, which can capture the similarity of potential intonation and emotional changes in speech. Set the threshold τ audio The edge construction of the speech modal graph is as follows:
[0105]
[0106] Among them, Edge a(i,j) Indicates the edge construction between nodes i and j in the speech modality graph. If Edge a(i,j) =1, then an edge is built between nodes i and j. If Edge A(i,j) =0, no edge is built between nodes i and j; is the DTW distance of speech features between nodes i and j; max(·) is the maximum value function; are the speech features of nodes i and j respectively.
[0107] The edge weight of the speech modal graph is assigned according to the DTW distance value of the speech feature. In this embodiment, the standardized DTW distance is used as the edge weight of the speech modal graph.
[0108] By building a graph structure based on cosine similarity and DTW distance, the visual modality graph captures visual emotion similarity, while the speech modality graph focuses on the similarity of speech emotion changes. This dual structure can more comprehensively depict the potential connections between students' emotions, providing rich graph information for subsequent emotion classification.
[0109] S23, Construction of Text Modal Diagram
[0110] Each video segment is used as a node of the text modality graph, and the corresponding text features are the node features of the text modality graph. In this embodiment, BERT Large encoding is used to extract text features from the text information (text segment) of the video segment.
[0111] Compute the semantic information increment and information flow coverage of the node:
[0112]
[0113] Among them, u i 、u j are the text segments corresponding to nodes i and j respectively, are the text features of nodes i and j respectively; SII(u i ) is the semantic information increment of node i, k represents the k nodes before node i, that is, in the text segment u i The previous k text segments; IFC(u i ) is the information flow coverage of node i, m represents the m nodes after node i, that is, in text segment u i The following m text segments; w(u i ,u j ) is the text segment u i 、u j The semantic similarity between i ,u j ) is the text segment u i 、u j The distance between (text segments u i 、u j relative position or time difference between them).
[0114] Semantic Information Increment SII(u i ) is used to measure u i The amount of newly introduced information, information flow coverage IFC (u i ) is used to measure u i The breadth of spread among text segments.
[0115] The semantic information increment and information flow coverage of the node are weighted and calculated to obtain the comprehensive semantic information flow score of the node:
[0116] SIFS(u i )=α·SII(u i )+β·IFC(u i )
[0117] Among them, α and β are weight coefficients used to balance the contribution of semantic information increment and information flow coverage. i ) is the comprehensive score of the semantic information flow of node i.
[0118] Build the text enhancement feature of the node:
[0119]
[0120] Among them, Enhanced(u i ) is the text enhancement feature of node i, For the text segment u i The feature vector obtained after BERT Large encoding is the text feature.
[0121] Calculate the cosine similarity of the text enhancement features between two nodes as the text similarity and set the threshold τ text ;
[0122]
[0123] Among them, Edge t (i,j) represents the edge construction between nodes i and j in the text modal graph. t (i,j)=1, then an edge is built between nodes i and j. If Edge t (i, j) = 0, then no edge is built between nodes i and j; D is the maximum text distance set; Sim(u i ,u j ) is the cosine similarity of the text enhancement features between nodes i and j, that is, the text similarity between nodes i and j.
[0124] Sim(u i ,u j )=cosine(Enhanced(u i ),Enhanced(u j ))
[0125] Where cosine(·) is the cosine function.
[0126] The edge weights of the text modal graph are assigned according to the text similarity values. In this embodiment, the normalized text similarity is used as the edge weights of the text modal graph.
[0127] S3, use the graph attention network (GAT) to perform GAT encoding on each modal graph respectively, update the node features in each modal graph respectively, and obtain the updated modal graphs.
[0128] The GAT encoding is as follows:
[0129] For node i and its neighbor nodes in the modality graph, calculate the attention coefficient between node i and each neighbor node:
[0130]
[0131] Among them, α ij is the attention coefficient between node i and its neighbor node j; W is the linear transformation matrix, which is used to make the input node feature h i Adapt attention calculation, so the linear transformation matrix W is used to transform h i Mapping is performed to obtain the mapped feature Wh i ; a is the attention vector; ∥ represents vector concatenation; h i is the node feature of node i; is the set of neighbor nodes of node i; LeakyReLU is the activation function.
[0132] By using the attention coefficient between node i and each neighbor node, the features of each neighbor node are weighted summed to update the feature representation of node i:
[0133]
[0134] Where σ(·) is the activation function; is the updated node feature of node i.
[0135] After GAT encoding of each modal graph, the updated node features in the visual modal graph, speech modal graph, and text modal graph are obtained respectively.
[0136] S4, performing multimodal adaptive fusion based on the updated modal graphs to obtain multimodal adaptive fusion features.
[0137] The specific process of step S4 is as follows:
[0138] In S41, in order to fully integrate the features of the three modalities, the node features of each modality after GAT encoding are unified into the common feature space, and the updated node features of each modality are subjected to nonlinear transformation and normalization respectively:
[0139]
[0140] Among them, the updated node features in the visual modal graph, speech modal graph, and text modal graph are W' v 、W' a 、W' t are the mapping matrices for vision, speech, and text, respectively, used for feature dimension alignment; Norm(·) is the normalization function; σ(·) is the activation function; They are respectively the normalized visual features, speech features, and text features.
[0141] S42, in order to deeply capture the relationship between the three modal features of vision, speech, and text, this paper introduces a transformer-based fusion mechanism to improve the ability to share information between the modalities. For each attention head in the transformer, a query, key, and value matrix is constructed:
[0142]
[0143] in, are the linear mapping matrices for generating queries, keys, and values for the h-th attention head; Q h , K h 、V h are the query, key, and value matrices of the h-th attention head, respectively; h = 1, 2, …, H.
[0144] For each attention head, the cross attention output is calculated as follows:
[0145]
[0146] Where: d represents the feature dimension of each attention head; O h is the cross attention output of the h-th attention head, which fully integrates the information between the three modalities.
[0147] Concatenate all attention heads to obtain multimodal fusion features:
[0148]
[0149] in, is the multimodal fusion feature of node i; Concat(·) is the concatenation function, which concatenates the outputs of each attention head along the feature dimension; W O is the linear mapping matrix of the multi-head attention output.
[0150] S43, in order to adaptively assign weights according to the different contributions of each modality in emotion recognition, the present invention introduces a multimodal adaptive gating mechanism. Generate the corresponding adaptive gating weights:
[0151]
[0152] in, is the linear mapping matrix used to generate adaptive gating weights; is the bias term; softmax(·) is the activation function to ensure that the weights of each modality meet the normalization requirements; g v 、g a 、g t They are Adaptive gating weight, g v +g a +g t =1.
[0153] Using adaptive gating weight g v 、g a 、g t right Perform weighted summation to obtain adaptive fusion features:
[0154]
[0155] in, is the adaptive fusion feature of node i.
[0156] S44, based on the multimodal fusion features and the adaptive fusion features, obtain the multimodal adaptive fusion features:
[0157]
[0158] in, is the multimodal adaptive fusion feature of node i, which fully integrates the three modal information of vision, speech and text.
[0159] S5, predicting students’ classroom emotions using multimodal adaptive fusion features.
[0160] Adaptively fuse the multimodal features of the node As input, it passes through a two-layer fully connected network (MLP) and outputs the classroom sentiment prediction result.
[0161] S6, establish loss function The multi-view graph neural network is trained with the goal of minimizing the loss function; the trained multi-view graph neural network is used to predict classroom emotions based on students' classroom videos.
[0162] To further improve the model performance, this paper adopts a multi-task training strategy, combines the fusion of multimodal emotion classification, unimodal emotion classification and contrastive learning, optimizes the multi-view graph neural network, and constructs a loss function The details are as follows:
[0163] S61, classification loss for multimodal sentiment classification (multimodal classification loss )
[0164] Utilize the multimodal adaptive fusion features of node i Perform emotion classification and obtain the probability of each emotion category of node i as p i,k , the true label of the node is y i,k(one-hot vector), k = 1, 2, ..., K, there are K emotion categories, calculate the multimodal classification loss
[0165]
[0166] Among them, α k is the category balance factor; γ is the modulation factor used to reduce the loss weight of easy-to-classify samples.
[0167] S62, Contrastive Loss for Contrastive Learning
[0168] To enhance the multimodal adaptive fusion features The alignment effect between each unimodal feature defines contrastive learning. The unimodal feature specifically includes the visual features of node i. Voice Features Text features The goal is to maximize the multimodal adaptive fusion features With each single modal feature The cosine similarity between them is used to calculate the contrast loss.
[0169]
[0170] where s(·,·) represents the cosine similarity function; τ is a parameter; and mod∈{v,a,t} represents the visual, speech, and text modalities, respectively.
[0171] S63, classification loss for unimodal sentiment classification (unimodal classification loss )
[0172] In addition to the multimodal adaptive fusion features, each single-modal feature is also classified into emotion categories to enhance the discrimination ability. Assume that each single-modal feature is passed through the corresponding fully connected network (MLP) to obtain the corresponding prediction probability. The probability of each emotion category of node i in the visual, speech, and text modalities is Calculate unimodal classification loss
[0173]
[0174] Among them, mod∈{v,a,t} represents the visual, speech, and text modalities respectively;
[0175] S64, multi-task total loss function for:
[0176]
[0177] Among them, λ1, λ2, and λ3 are weight hyperparameters used to adjust the relative importance of different loss terms.
[0178] This embodiment collects class videos of multiple classes in actual classrooms, with a total of 308 students, and obtains the classroom emotion results of each student through the above-mentioned algorithm method, and then compares them with the annotations of their class teachers. The verification standard is the comprehensive score of multiple subject teachers and class teachers (6 types of classroom expressions: happy, sad, calm, surprised, angry, contempt) (reliability>0.9). The Pearson correlation coefficient is significant (p<0.001), which shows that the classroom expression recognition results of the present invention are highly consistent with the teacher's evaluation, which effectively verifies the model and method of the present invention. The research results show that the present method is highly consistent with the teacher's evaluation, and exhibits stable recognition capabilities in all 6 emotion categories, which has significant advantages over traditional classroom expression recognition methods.
[0179] The model's key innovations lie in its use of a multimodal feature fusion method to integrate information from three modalities: video, audio, and text; its effective modeling and capturing of emotional connections between student groups through a graph structure; and its strong adaptability to real-world classroom scenarios. This method can dynamically identify students' emotional states in the classroom, making it suitable for large classrooms or online education scenarios, and providing effective technical support for teachers to monitor students' emotional states in real time.
[0180] Furthermore, the data of the present invention was used to verify other common classroom expression recognition methods, and it was found that the prediction results of other common classroom expression recognition methods could not achieve significant correlation. This may be because classroom expression recognition is a regular classroom scene, which requires more comprehensive information about students' emotional expressions, and also considers group factors. However, the calculation method of the present invention includes the fusion of multiple features, which is more in line with the expression of students' classroom emotion recognition in actual classroom scenes. Therefore, the multi-perspective graph neural network classroom emotion recognition model constructed by the present invention will be able to dynamically obtain students' classroom emotion recognition results. The comparison of the classification effects of the present invention and other technologies is as follows. Figure 2 As shown in the figure, MVGNN represents a method based on multi-view graph neural network (i.e., this method), CNN-LSTM represents a method based on convolutional neural network-long short-term memory network, Transformer represents a method based on attention mechanism converter, GCN represents a method based on graph convolutional neural network, GAT represents a method based on graph attention network, and Traditional ML represents a traditional machine learning method (SVM+RF).
[0181] Table 1 below shows a performance comparison of the proposed MVGNN method and other comparative methods on the classroom emotion recognition task. The data demonstrates that MVGNN significantly outperforms other methods across all evaluation metrics, achieving an improvement of approximately 17 percentage points over traditional methods and a further improvement of approximately 7 percentage points over the closest CNN-LSTM method.
[0182] Table 1
[0183] method Accuracy Recall F1 value MVGNN 92.5% 91.8% 92.1% CNN-LSTM 85.2% 84.5% 84.8% Transformer 83.7% 82.9% 83.3% GCN 81.4% 80.8% 81.1% TraditionalML 75.3% 74.8% 75.0%
[0184] The effectiveness of multimodal fusion in this method is analyzed, as shown in Table 2 below, which shows the performance of different modal combinations. The results show that trimodal fusion achieves the best results, achieving an accuracy of 92.5%. Recognition accuracy gradually decreases as the number of modalities decreases. Among single modalities, video performs best, followed by audio. Multimodal fusion significantly improves recognition performance. These experimental results fully verify the effectiveness of this method and the necessity of multimodal fusion.
[0185] Modal Combination Accuracy Trimodal fusion (video + audio + text) 92.5% Dual-modal fusion (video + audio) 87.3% Video only 82.1% Audio only 79.5% Text only 77.8% The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A classroom emotion recognition method based on multi-view graph neural network, characterized by: The training method of the multi-view graph neural network is as follows: S1, divide the student's classroom video into several video segments, and extract visual features, voice features, and text features from each video segment; each video segment is marked with classroom emotions; S2: Based on visual features, speech features, and text features, a visual modal graph, a speech modal graph, and a text modal graph are constructed respectively. In each modal graph, the video segment is used as a node, and the visual features, speech features, and text features corresponding to each video segment are used as node features of the corresponding modal graph. S3, use the graph attention network to perform GAT encoding on each modal graph respectively, update the node features in each modal graph respectively, and obtain the updated modal graphs; S4, using the updated modal graphs to perform multimodal adaptive fusion based on transformer to obtain multimodal adaptive fusion features; S5, predicting students’ classroom emotions using multimodal adaptive fusion features; S6, establish a multi-task total loss function, and train the multi-view graph neural network with the goal of minimizing the multi-task total loss function; the trained multi-view graph neural network is used to predict classroom emotions based on students' classroom videos.
2. A classroom emotion recognition method based on multi-view graph neural network according to claim 1, characterized in that: In step S2, the specific construction of the visual modality graph is as follows: Calculate the cosine similarity of the visual features between two nodes as the visual similarity of the two nodes; set the threshold τ video , if the visual similarity between two nodes is greater than the threshold τ video , then an edge is built between the two nodes; otherwise, no edge is built between the two nodes.
3. The classroom emotion recognition method based on multi-view graph neural network according to claim 1 is characterized in that: In step S2, the speech modal graph is constructed as follows: Calculate the DTW distance of speech features between two nodes and set the threshold τ audio ; Among them, Edge a(i,j) Indicates the edge construction between nodes i and j in the speech modality graph. If Edge a(i,j) =1, then an edge is built between nodes i and j. If Edge A(i,j) =0, no edge is built between nodes i and j; is the DTW distance of speech features between nodes i and j; max(·) is the maximum value function; are the speech features of nodes i and j respectively.
4. The classroom emotion recognition method based on multi-view graph neural network according to claim 1 is characterized in that: In step S2, the construction of the text modal graph is specifically as follows: Compute the semantic information increment and information flow coverage of the node: Among them, u i 、u j are the text segments corresponding to nodes i and j respectively, are the text features of nodes i and j respectively; SII(u i ) is the semantic information increment of node i, k represents the k nodes before node i; IFC(u i ) is the information flow coverage of node i, m represents the m nodes after node i; w(u i ,u j ) is the text segment u i 、u j The semantic similarity between i ,u j ) is the text segment u i 、u j the distance between them; The semantic information increment and information flow coverage of the node are weighted and calculated to obtain the comprehensive semantic information flow score of the node: SIFS(u i )=α·SII(u i )+β·IFC(u i ) Among them, α and β are weight coefficients; SIFS(u i ) is the comprehensive score of the semantic information flow of node i; Build the text enhancement feature of the node: Among them, Enhanced(u i ) is the text enhancement feature of node i; Calculate the cosine similarity of the text enhancement features between two nodes as the text similarity and set the threshold τ text ; Among them, Edge t (i,j) represents the edge construction between nodes i and j in the text modal graph. t (i,j)=1, then an edge is built between nodes i and j. If Edge t (i, j) = 0, then no edge is built between nodes i and j; D is the maximum text distance; Sim(u i ,u j ) is the cosine similarity of the text enhancement features between nodes i and j, that is, the text similarity between nodes i and j.
5. The classroom emotion recognition method based on multi-view graph neural network according to claim 1 is characterized in that: In step S3, the GAT encoding is specifically as follows: For node i and its neighbor nodes in the modality graph, calculate the attention coefficient between node i and each neighbor node: Among them, α ij is the attention coefficient between node i and its neighbor node j; W is the linear transformation matrix; a is the attention vector; || represents vector concatenation; h i is the node feature of node i; is the set of neighbor nodes of node i; LeakyReLU is the activation function; By using the attention coefficient between node i and each neighbor node, the features of each neighbor node are weighted summed to update the feature representation of node i: Where σ(·) is the activation function; is the updated node feature of node i.
6. The classroom emotion recognition method based on multi-view graph neural network according to claim 1 is characterized in that: In step S4, multimodal adaptive fusion is specifically as follows: S41, perform nonlinear transformation and normalization on the node features after each modal update: Among them, the updated node features in the visual modal graph, speech modal graph, and text modal graph are W' v 、W' a 、W' t are the mapping matrices for vision, speech, and text respectively; Norm(·) is the normalization function; σ(·) is the activation function; They are respectively the normalized visual features, speech features, and text features; S42 introduces a transformer-based fusion mechanism. For each attention head in the transformer, the corresponding query, key, and value matrices are constructed: in, are the linear mapping matrices for generating queries, keys, and values for the h-th attention head; Q h , K h 、V h are the query, key, and value matrices of the h-th attention head, respectively; h = 1, 2, …, H; For each attention head, the cross attention output is calculated as follows: Where: d represents the feature dimension of each attention head; O h is the cross attention output of the h-th attention head; Concatenate all attention heads to obtain multimodal fusion features: in, is the multimodal fusion feature of node i; Concat(·) is the concatenation function; W O is the linear mapping matrix of the multi-head attention output; S43, introduces a multimodal adaptive gating mechanism to respectively normalize the visual features, speech features, and text features Generate the corresponding adaptive gating weights: in, is the linear mapping matrix used to generate adaptive gating weights; is the bias term; softmax(·) is the activation function; g v 、g a 、g t They are Adaptive gating weights of Using adaptive gating weight g v 、g a 、g t right Perform weighted summation to obtain adaptive fusion features: in, is the adaptive fusion feature of node i; S44, based on the multimodal fusion features and the adaptive fusion features, obtain the multimodal adaptive fusion features: in, is the multimodal adaptive fusion feature of node i.
7. The classroom emotion recognition method based on multi-view graph neural network according to claim 1 is characterized in that: In step S5, the multimodal adaptive fusion features are used as input and passed through a two-layer fully connected network to output the classroom emotion prediction results.
8. The classroom emotion recognition method based on multi-view graph neural network according to claim 1 is characterized in that: In step S6, construct the multi-task total loss function The details are as follows: S61, using the multimodal adaptive fusion features of node i Perform emotion classification and obtain the probability of each emotion category of node i as p i,k , the true label of the node is y i,k , k=1,2,...,K, there are K emotion categories, calculate the multimodal classification loss Among them, α k is the category balance factor; γ is the modulation factor; S62, based on the visual features of node i Voice Features Text features And the multimodal adaptive fusion features of node i Calculating contrastive loss Where s(·,·) represents the cosine similarity function; τ is a parameter; S63, use the unimodal features of node i to perform emotion classification, and obtain the probability of each emotion category of node i in visual, speech, and text modes: Calculate unimodal classification loss Among them, mod∈{v,a,t} represents the visual, speech, and text modalities respectively; S64, multi-task total loss function for: Among them, λ1, λ2, and λ3 are weight hyperparameters.
9. The classroom emotion recognition method based on multi-view graph neural network according to claim 1 is characterized in that: In step S1, the visual features of the visual information of the video segment are extracted using the I3D model, the speech features of the speech information of the video segment are extracted using the Wav2Vec 2.0 model, and the text features of the text information of the video segment are extracted using the BERT Large codec; Visual information includes facial expressions, voice information includes voice intonation, and text information includes text content.
10. A computer program product, characterized in that It includes a computer program / instruction, which, when executed by a processor, implements a classroom emotion recognition method based on a multi-view graph neural network as described in any one of claims 1-9.
Citation Information
Patent Citations
Image joint text sentiment analysis method based on modal fusion graph convolutional network
CN117540023A
Emotion analysis method and system based on multi-modal fusion
CN119272224A