Student classroom participation degree analysis method based on double-flow hypergraph convolutional network
By analyzing students' classroom participation through a two-stream hypergraph convolutional network, the problem that existing models fail to consider the impact of students' emotions and behavioral interactions is solved, achieving higher prediction accuracy.
Patent Information
- Application Number
- CN202510692371.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
AI Technical Summary
Existing student classroom engagement analysis models fail to fully consider the impact of emotional and behavioral interactions among students, resulting in insufficient prediction accuracy of student engagement.
A method based on a two-stream hypergraph convolutional network is adopted to extract the multidimensional features of image data through a multi-feature encoder, and the multi-dimensional propagation module and multi-frequency propagation module are used to model the contagious effect of student participation, which is then combined with a participation classifier for prediction.
The accuracy of predicting student classroom engagement is improved, especially in considering the correlation and dependence of multidimensional characteristics and frequency information, which enhances the capture and understanding of the contagious effect of student participation.
Smart Images

Figure CN120599347A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and in particular to a method for analyzing student classroom participation based on a dual-stream hypergraph convolutional network. Background Art
[0002] Student engagement is a key factor influencing academic success and learning outcomes. It is a multidimensional construct encompassing behavioral, affective, and cognitive aspects that together provide a holistic view of student engagement in the learning process. With the increasing popularity of online learning platforms and digital classrooms, accurately predicting and understanding student engagement is crucial for improving instructional strategies and providing personalized interventions.
[0003] Given the complexity of engagement, traditional methods typically assess student engagement through surveys and teacher observations. Subsequently, researchers have adopted a range of machine learning and data mining techniques to develop predictive models using different data sources (such as facial expressions, body language, eye movements, and physiological signals). However, most existing methods focus on analyzing a single dimension, such as emotion or behavior, and fail to fully capture the multidimensional characteristics of engagement and the interactions between these dimensions. This limitation often leads to a one-sided understanding of student engagement and fails to fully reflect the complex changes in students' behavior during the learning process.
[0004] The impact of classroom engagement contagion on individual engagement remains underexplored. In recent years, much research on student engagement has focused on assessing engagement based on individual student factors. Several studies have shown that a student's engagement can be positively predicted by their peers' engagement, confirming the existence of social contagion effects within student groups. Emotions and behaviors can spread within a classroom, significantly influencing individual learning engagement and motivation. However, existing student engagement prediction models often ignore this phenomenon and fail to consider the impact of emotional and behavioral interactions between students in the classroom. Furthermore, current models struggle to capture the complex interactions between emotional and behavioral dimensions. Therefore, effectively leveraging student classroom images to model engagement contagion between students is a key challenge to improving the accuracy of engagement prediction. Summary of the Invention
[0005] The purpose of this invention is to provide a method for analyzing student classroom engagement based on a two-stream hypergraph convolutional network. This method is used to address the technical problem that existing student classroom engagement analysis models fail to consider the impact of emotional and behavioral interactions between students in the classroom.
[0006] A method for analyzing student classroom participation based on a two-stream hypergraph convolutional network. The specific steps are as follows:
[0007] S1: collects image data of students during class;
[0008] S2: Build an engagement model based on a two-stream hypergraph convolutional network. The engagement model includes a multi-feature encoder, a multi-element propagation module, a multi-frequency propagation module, and an engagement classifier.
[0009] S3: Extract multi-dimensional features of image data through a multi-feature encoder;
[0010] S4: Input the multidimensional features into the multivariate propagation module and the multi-frequency propagation module respectively, model the contagious effect of student participation through the multivariate propagation module, and capture the multi-frequency information from the extracted features through the multi-frequency propagation module;
[0011] S5: The engagement classifier fuses the outputs of the input multi-propagation module and the multi-frequency propagation module to predict the student engagement.
[0012] Optionally, the multidimensional features in step S1 include three unimodal features: visual attention features, body behavior features, and emotion features.
[0013] Optionally, the multi-propagation module is a hypergraph convolutional neural network, and the multi-frequency propagation module is a multi-frequency graph convolutional neural network.
[0014] Optionally, the specific steps in step S3 are:
[0015] S3.1: Process the images of N students taken at the same time as a group, extract the multi-dimensional features of the N students, and obtain multiple unimodal features for each student;
[0016] S3.2: Input multiple single-modal features into multiple multi-layer perceptrons to obtain the multi-dimensional feature encoding of each student, and add the student embedding S to the modal encoding of a single student. i .
[0017] Optionally, the specific steps of modeling the contagion effect of student participation through the multivariate communication module in step S4 are:
[0018] S4.1.1: Construct a classroom role sequence with N students participating as a node set Hyperedge set Hypergraph of node weights and hyperedge weights, where: the set of nodes A single node corresponds to a single modal feature of a single student, and the hyperedge set In the case of a single hyperedge encoding multimodal or group-related effects among students;
[0019] S4.1.2: Perform node convolution, update hyperedge features by aggregating node features, and then propagate hyperedge information to nodes through hyperedge convolution;
[0020] S4.1.3: Repeat step S4.1.2. After L iterations, the output of the last iteration is used as the output of the multi-propagation module.
[0021] Optionally, in step S4.1.2, the node weight is adjusted by calculating the attention score of the node vertex and its associated hyperedge.
[0022] Optionally, the specific steps of capturing multi-frequency information from the extracted features by the multi-frequency propagation module in step S4 are:
[0023] S4.2.1: Construct an undirected graph in parallel with the multi-propagation module, the undirected graph including a node set and hyperedge sets Node Set A single node in the CNN corresponds to a single modal feature of a single student;
[0024] S4.2.2: Connect a single node to all nodes representing the same dimension as other students, as well as nodes of other dimensions of the same student, to obtain the adjacency matrix, and normalize it to obtain the Laplacian matrix of the graph;
[0025] S4.2.3: Obtain multi-frequency features using low-pass and high-pass filters, where the high-pass filter is equal to the Laplacian of the normalized graph, and the low-pass filter is equal to the difference between two identity matrices and the Laplacian matrix of the normalized graph.
[0026] S4.2.4: Combine the low-frequency and high-frequency information obtained by the low-pass filter and the high-pass filter through adaptive weighting;
[0027] S4.2.5: Each student node aggregates the low-frequency and high-frequency information of neighboring nodes through iterative propagation, and after K layers of message passing, the output of the multi-frequency propagation module is obtained.
[0028] Optionally, the specific method in step S5 is:
[0029] Aggregate the outputs from the multi-frequency propagation module and the multivariate propagation model to construct the student's classroom participation μ i , the classroom participation μ i The input is normalized into the softmax layer to obtain the probability distribution of class participation categories and the student's predicted participation label.
[0030] Optionally, the engagement classifier in step S5 is optimized using L2 regularized categorical cross entropy loss.
[0031] Due to the adoption of the above technical solution, the present invention has the following advantages:
[0032] 1. This application uses parallel hypergraphs and graph convolutional networks to explore high-order and complex relationships between various student features, and fully utilizes the correlation and dependency of features between different frequency information to improve the accuracy of analysis and prediction.
[0033] 2. This application uses a hypergraph structure to encode student engagement contagion and captures both differences and commonalities in emotions and behaviors through multi-frequency signals. To address individual differences in engagement propagation, this application introduces hypergraph attention to adjust propagation weights. The model achieves high accuracy across various engagement prediction tasks.
[0034] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings of the present invention are described below.
[0036] Figure 1 This is a flow chart of the student classroom participation analysis method based on the dual-stream hypergraph convolutional network of the present invention.
[0037] Figure 2 Schematic diagram of the structure of the hypergraph convolutional network of the present invention.
[0038] Figure 3 This is the node-edge-node transformation graph in the hypergraph convolutional network of the present invention.
[0039] Figure 4 Schematic diagram of the structure of the multi-frequency graph convolutional neural network of the present invention.
[0040] Figure 5 This is the loss matrix diagram of the student engagement classification task of the present invention. DETAILED DESCRIPTION
[0041] The present invention will be further described below with reference to the accompanying drawings and examples.
[0042] Example:
[0043] A method for analyzing student classroom participation based on a two-stream hypergraph convolutional network. The specific steps are as follows:
[0044] S1: collects image data of students during class;
[0045] S2: Build an engagement model based on a two-stream hypergraph convolutional network. The engagement model includes a multi-feature encoder, a multi-element propagation module, a multi-frequency propagation module, and an engagement classifier.
[0046] In this example, the multi-feature encoder includes a visual attention feature extractor, a limb behavior feature extractor, and an emotion feature extractor, wherein: the emotion feature extractor can adopt a VGG-Face model or a ResNet-50 model or a more complex deep network composed of the above models; the visual attention feature extractor can adopt an OpenFace model or an OpenCV model; the limb behavior feature extractor can adopt an HRNet model; the feature extraction of the above models are all prior arts and this application will not go into details.
[0047] In this example, if Figure 1 、 Figure 2 and Figure 3 As shown, the multi-propagation module is a hypergraph convolutional neural network HGCN, and the multi-frequency propagation module is a multi-frequency graph convolutional neural network MFGCN.
[0048] S3: The multi-dimensional features of the image data are extracted through the multi-feature encoder. The multi-dimensional features are extracted by the visual attention feature extractor, the body behavior feature extractor and the emotion feature extractor to obtain three unimodal features: visual attention feature, body behavior feature and emotion feature. The specific steps are as follows:
[0049] S3.1: Take the images X of N students taken at the same time as N} are processed as a group, where: x i is the image of the nth student, i∈N; the corresponding label is represented as Y={y1,y2,…,y N}. The student image sequence is represented as {(S1,x1),(S2,x2),…(S N ,x N )}, this sequence contains the student S i Multidimensional features. Including emotional features in the emotional dimension Visual attention characteristics and upper body limb behavior characteristics There are three unimodal features. Unique hot vector s i Represent each student and build a lookup table for students in the same classroom to calculate the student embedding for each student
[0050]
[0051] S3.2: After passing these unimodal features through their respective feature encoders, the three unimodal features are input into three multilayer perceptrons W1, W2, and W3 respectively to obtain the multidimensional feature encoding of each student:
[0052]
[0053] in, is the bias parameter of the multilayer perceptron. In order to consider the impact of students’ personality characteristics on their participation, the student embedding S is added to the modal encoding of each student. i :
[0054]
[0055] S4: Input the multidimensional features into the multi-dimensional propagation module and the multi-frequency propagation module respectively, model the contagious effect of student participation through the multi-dimensional propagation module, and capture multi-frequency information from the extracted features through the multi-frequency propagation module; the specific steps are:
[0056] S4.1: Model the contagion effect of student participation through the multivariate transmission module:
[0057] S4.1.1: Construct a classroom role sequence with N students participating as a node set Hyperedge set Hypergraph with node weight γ and hyperedge weight ω Where: Node Set A single node v corresponds to a single modal feature of a single student, and the above encoding Initialized as node embeddings of the hypergraph in: Hyperedge set A single hyperedge e in encodes multimodal or group-related influences between students, where:
[0058] In this embodiment, the hypergraph can be represented as an incidence matrix The non-zero term H ve =1 indicates that hyperedge e is associated with node v, otherwise H ve =0:
[0059]
[0060] In this embodiment, the student class participation is mainly determined by the multimodal characteristics of the students and the group influence between students, and there are potential multivariate relationships. Therefore, a multimodal hyperedge and a group influence hyperedge are constructed for each node. Specifically: Figure 2 As shown, each node First, all nodes with the same modality as other students in the class There are group influence hyperedges between these nodes. In addition, each node Other modalities connected to this student These nodes have multimodal hyperedges. In this way, the constructed hypergraph captures multivariate and group influence information beyond pairwise representation.
[0061] In this embodiment, in order to avoid model complexity, random initialization weight values are used. Specifically, the hyperedge weights are defined in the hypergraph. The hyperedge e is assigned an edge weight ω(e), and each node in this hyperedge is assigned a node weight γ e (v).
[0062] S4.1.2: Perform node convolution, update hyperedge features by aggregating node features, and then propagate hyperedge information to nodes through hyperedge convolution;
[0063] In this embodiment, through the hypergraph convolution operation, information propagation can be performed on the graph structure and multi-element embedding can be fused. Specifically, the main purpose of defining the convolution operator in the hypergraph is to measure the transition probability between two vertices, through which the embedding (or features) of each vertex can be propagated in the graph neural network. First, node convolution is performed to update the hyperedge embedding by aggregating node features, and then the hyperedge information is propagated to the node through hyperedge convolution. According to the hypergraph structure, a hyperedge convolution layer f(Q, W, P) is constructed according to the following formula:
[0064]
[0065] in, represents the hypergraph signal of the lth layer, Q(0)=Q, σ is the nonlinear activation function, is the weight matrix of the hyperedge, and are the node degree matrix and the hyperedge degree matrix respectively. is the weight matrix between the (l)th layer and the (l+1)th layer. By symmetric normalization of the above formula, we finally get the following formula:
[0066]
[0067] Because Q (l+1) Q (l) Differentiable, the hypergraph convolution is performed in this way and optimized by gradient descent. In this way, Figure 3 As shown in Figure 3, the multi-propagation module utilizes the properties of high-order correlations between data and gradually refines high-order multimodal relationships by performing node-edge-node transformations.
[0068] S4.1.3: Repeat step S4.1.2 for L iterations, and use the output of the last iteration as the output of the multi-propagation module:
[0069]
[0070] As an embodiment of the present application, in step S4.1.2, the node weight is adjusted by calculating the attention score of the node vertex and its associated hyperedge.
[0071] In this embodiment, hypergraph convolution has an innate attention mechanism. Specifically: As can be seen from the above formula, the transition probability between vertices is not binary, which means that the information propagated between vertices is assigned different importance. However, this attention mechanism is not learnable or trainable after the association matrix H is determined. The purpose of introducing hypergraph attention is to dynamically learn the propagation weights of information between nodes, ensuring that the model can capture the individual differences of students' more fine-grained emotional reactions and behavioral patterns in the process of propagation influence. Calculate v i Vertex and its associated hyperedge e j Attention score:
[0072]
[0073] Among them, σ(·) is a nonlinear activation function, It is v i Neighborhood set of , sim(·) is a similarity function used to calculate the pairwise similarity between two vertices, defined as:
[0074] sim(v i ,e j )=a T [v i ][e j ]
[0075] Where [·][·] represents a connection and a is a weight vector that outputs a scalar similarity value. In this way, the initial association matrix of the hypergraph is enriched to represent the attention score as the node weight γ of the edge dependency. e (v), as a dynamic weighted incidence matrix:
[0076]
[0077] Intuitively, γ e (v) The contribution of node v to hyperedge e is measured, thereby strengthening the subtle multimodal dependencies and group influence of student features.
[0078] S4.2: Capturing multi-frequency information from the extracted features through the multi-frequency propagation module:
[0079] S4.2.1: Construct an undirected graph in parallel with the multi-propagation module The undirected graph Include Node Set and hyperedge sets Node Set A single node in the CNN corresponds to a single modal feature of a single student;
[0080] In this embodiment, the node set With Hypergraph The node sets are the same, respectively represented as {f i e ,f i a ,f i u} feature dimension.
[0081] S4.2.2: Connect a single node to all nodes representing the same dimension as other students, as well as nodes of other dimensions of the same student, to obtain the adjacency matrix, and normalize it to obtain the Laplacian matrix of the graph;
[0082] In this embodiment, Different from relying on hyperedges to encode complex group dependencies, Based on the pairwise connections, direct frequency-aware feature propagation is achieved. Specifically, for each node f i x , connect it to all nodes representing the same dimension as other students {f j x |j∈[1,N],j≠i}, and other dimension nodes of the same student {f i z |z∈{e,a,u},z≠x}. The constructed graph like Figure 4 As shown, the adjacency matrix of the graph is The normalized graph Laplacian matrix can be expressed as:
[0083]
[0084] in, is the diagonal matrix, and I is the identity matrix. The integration of the data into the engagement prediction model enhances the model's predictive ability.
[0085] S4.2.3: Use a low-pass filter and high-pass filter Get multi-frequency features, where: high-pass filter Equal to the normalized Laplace operator L of the graph, low-pass filter It is equal to the difference between the two identity matrices I and the normalized Laplace matrix L of the graph:
[0086]
[0087] In this embodiment, the high-pass filter Emphasize high-frequency information within the feature. According to Fourier transform theory, the eigenvectors of the normalized Laplacian matrix are used as the basis of the Fourier transform of the graph. and The filtering operation can be interpreted as a signal And the convolution operation of the corresponding convolution kernel *c:
[0088]
[0089] S4.2.4: Low-pass filter by adaptive weighting and high-pass filter The obtained low-frequency and high-frequency information are combined;
[0090] In this embodiment, after obtaining multi-frequency features, different frequency components are adaptively integrated to perform student participation representation learning. Figure 2 As shown, low-frequency and high-frequency information are combined through adaptive weighting:
[0091]
[0092] Among them, F (k) is the feature representation of the k layer, and W l and W h It is a learnable weight matrix that adaptively balances low-frequency and high-frequency components.
[0093] Based on the information propagation between different nodes, the above formula can be re-expressed as:
[0094]
[0095] in, Represents the adjacent nodes of node i. and Represents the contribution of node j to the low-frequency and high-frequency information of node i, satisfying the constraint
[0096] S4.2.5: Each student node aggregates the low-frequency and high-frequency information of neighboring nodes through iterative propagation. After K layers of message passing, the output of the multi-frequency propagation module is obtained:
[0097]
[0098] S5: The engagement classifier fuses the outputs of the input multi-propagation module and the multi-frequency propagation module to predict the student engagement μ i .
[0099] In this embodiment, the student participation representation μ is constructed i :
[0100]
[0101] Among them, μ i As a comprehensive feature representation, it captures the multi-frequency interactions related to participation in infection and participation prediction. Then, μ i Input the softmax layer for normalization and obtain the participation status:
[0102]
[0103] Among them, W4 is a trainable weight matrix, represents the probability distribution of engagement categories, It is student S i The predicted engagement labels for .
[0104] In this example, the engagement classifier is optimized using the L2 regularized categorical cross entropy loss:
[0105]
[0106] Among them, Num is the number of classes, s(i) represents the number of students, and y i,j They represent the probability distribution of the predicted results and the true label respectively, λ is the weight of the L2 regularization term, and θ represents the trainable parameters in the model.
[0107] S6: Simulation Validation: Verify the effectiveness of the proposed student engagement prediction method. When there is a certain class imbalance in the data, turn to the confusion matrix to confirm the accuracy of the results. Figure 5 Shows the confusion matrix of this application method in two-class and three-class classification tasks.
[0108] S6.1: Experimental setup:
[0109] S6.1.1: Data Preparation: To evaluate the effectiveness of this approach, we use the RoomReader dataset as a benchmark. This dataset measures both off-task and on-task engagement and includes multimodal data from student-tutor interactions, collected over 30 sessions involving 118 participants in an online setting. This dataset provides continuous annotations of engagement, with engagement labels ranging from [2, -2].
[0110] 23 usable videos were filtered out from 30 conference videos and participation annotations, and students from the same class were merged together. To ensure data accuracy, invalid classroom segments were removed from the beginning and end of each video. Frames were extracted from all students' videos and combined with the annotated participation values. Existing literature studies show that the participation labels in the RoomReader dataset are highly imbalanced: approximately 80.2% of the samples are labeled in the range (1,2], approximately 18.3% are labeled in (0,1], 1.3% are labeled in (-1,0], and 0.2% are labeled in [-2, -1].
[0111] Both binary and ternary classification tasks were performed simultaneously. For the ternary classification task, 948 low-engagement samples, 1,629 medium-engagement samples, and 2,431 high-engagement samples were randomly sampled from the ranges [-2, -1), [-1, 1], and (1, 2], respectively. This process generated 1,252 sets of input data, which were based on the number of students in each class. For the binary classification task, keeping the same 1,252 sets of input data, 948 low-engagement samples were classified as "disengaged", and 1,629 medium-engagement samples and 2,431 high-engagement samples were classified as "engaged". To mitigate the potential impact of class imbalance, class weights were applied to the categorical cross-entropy loss function. Subsequently, multidimensional features were extracted for each set of classroom data and encoded using the multi-feature encoding strategy described above.
[0112] S6.1.2: All experiments were implemented using PyTorch in Python 3.8.20 on an NVIDIA 3090 GPU. Standard evaluation protocols were followed, using accuracy, F1-score, and AUC as performance metrics. During training, the Adam optimizer was used with a learning rate of 1×10 -5 To alleviate the overfitting problem, we also used the dropout technique with a ratio of p = 0.5. We tested the hypergraph convolution layer L and graph convolution layer K in the range of 1 to 6, and reported the results with the best performance. The model hyperparameters are shown in Table 1.
[0113] Table 1 Model hyperparameters
[0114] Optimizer Batch <![CDATA[D h ]]> L K Dropout <![CDATA[Adam optimizer (1×10 -5 )]]> 16 512 3 3 0.5
[0115] S6.2 Baseline comparison: To provide a comprehensive performance evaluation, the proposed method is compared with several existing engagement prediction models. Seven baseline models are selected: TCCT-NET model, EnsModel model, ConvLSTM model, EG-NET model, ED-MTT model, Bootstrap model and Haar-MGL model. The above models are all existing models and have been used for engagement prediction in existing literature. This application does not go into details about the specific structure of the above models. These baselines are re-implemented using the RoomReader3 dataset, following the specified data format. The results of the comparison experiment on the RoomReader dataset are shown in Table 2.
[0116] Table 2 Comparative experiment results on the RoomReader dataset
[0117] method Binary classification accuracy (%) Three-category accuracy (%) Number of features TCCT-NET model 73.99±1.43 60.48±1.18 single EnsModel 75.30±3.50 69.53±3.54 single ConvLSTM model 76.50±1.85 74.53±1.50 single EG-NET model 77.38±1.27 72.76±2.24 single ED-MTT model 73.80±3.35 71.20±1.35 single Bootstrap Model 82.42±2.15 75.42±2.43 Multiple Haar-MGL model 90.18±1.34 - Multiple DS-HGCN (this application) 93.23±0.79 80.13±1.24 Multiple
[0118] In this example, the results in Table 2 were obtained using an 80% / 20% training / testing split, with the mean and standard deviation calculated for five independent trials. Notably, methods using multidimensional features consistently outperform those relying on single-dimensional features. Furthermore, by accounting for student engagement contagion, this method achieves higher accuracy. While Haar-MGL achieves significant results through multimodal graph learning, it lacks interpretability in extracting and analyzing effective features. Unlike Haar-MGL, this method leverages more readily available image data to more effectively analyze student engagement across multiple dimensions.
[0119] S6.3: Ablation experiments:
[0120] In order to better evaluate the effectiveness of the model components, an ablation study is conducted on the key elements of DS-HGCN, and the results are shown in Table 3. Specifically, the following factors are considered in the ablation experiment:
[0121] S6.3.1: Impact of Multivariate Information: We removed the hypergraph component and performed binary and ternary classification based solely on the multi-frequency representation, referred to as Variant 1 in Table 3. The results show a 4.08% decrease in accuracy for the binary classification task and a 5.67% decrease for the ternary classification task, with corresponding decreases in F1-score and AUC. This highlights the effectiveness of introducing student engagement contagion, which successfully captures the complex relationship between engagement and multidimensional features.
[0122] S6.3.2: Impact of Multi-Frequency Information: By removing the multi-frequency component and relying solely on the multivariate module (variable 2 in Table 3), the results show that the accuracy of the two-class classification task decreased by 1.89% and the accuracy of the three-class classification task decreased by 4.18%. Experimental results show that multi-frequency information has the least effect on distinguishing between non-engaged and engaged tasks, but has a significant impact on the three-class classification task. This emphasizes the importance of multi-frequency feature information for predicting engagement.
[0123] S6.3.3: Impact of Multi-Feature Information: The graph structure of DS-HGCN incorporates features from three dimensions: emotion, visual attention features, and upper body behavior. We conducted experiments by sequentially removing features from each dimension to predict student engagement using the remaining two dimensions. The results are shown in Table 3 for variables 3 through 5. We found that the absence of any single feature significantly degraded the model's accuracy. In particular, removing the emotion feature led to a sharp drop in performance. These results are consistent with the fundamental definition of student engagement and highlight the importance of leveraging multi-dimensional features.
[0124] S6.3.3: Effect of Hypergraph Attention: We removed the hypergraph attention module from both binary and ternary classification tasks, referred to as variant 6, and showed a significant drop in performance on these tasks. This suggests that dynamically learning the association matrix can significantly enhance training and lead to better results.
[0125] Table 3 summarizes the ablation results of DS-HGCN, which show that all modified models have lower accuracy, F1 score and AUC than the full model. These components enhance the performance of the model individually and collectively, confirming the contribution of each module.
[0126] Table 3 DS-HGCN ablation experiment results
[0127]
[0128] In summary, this application uses parallel hypergraphs and graph convolutional networks to explore high-order and complex relationships between various student features and fully exploits the correlations and dependencies between features at different frequencies. Extensive experiments were conducted on a real-world educational dataset, and the model achieved an accuracy of 94.02% for binary classification and 81.37% for ternary classification. Therefore, the method described in this application improves the accuracy of predictive analysis of student engagement in class.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A method for analyzing student classroom participation based on a two-stream hypergraph convolutional network, characterized by: The specific steps are: S1: collects image data of students during class; S2: Build an engagement model based on a two-stream hypergraph convolutional network. The engagement model includes a multi-feature encoder, a multi-element propagation module, a multi-frequency propagation module, and an engagement classifier. S3: Extract multi-dimensional features of image data through a multi-feature encoder; S4: Input the multidimensional features into the multivariate propagation module and the multi-frequency propagation module respectively, model the contagious effect of student participation through the multivariate propagation module, and capture multi-frequency information from the extracted features through the multi-frequency propagation module; S5: The engagement classifier fuses the outputs of the input multi-propagation module and the multi-frequency propagation module to predict the student engagement.
2. The method for analyzing student classroom participation based on a two-stream hypergraph convolutional network according to claim 1 is characterized in that: The multidimensional features in step S1 include three unimodal features: visual attention features, body behavior features, and emotional features.
3. The method for analyzing student classroom participation based on a two-stream hypergraph convolutional network according to claim 1 is characterized in that: The multi-propagation module is a hypergraph convolutional neural network, and the multi-frequency propagation module is a multi-frequency graph convolutional neural network.
4. A method for analyzing student classroom participation based on a two-stream hypergraph convolutional network according to claim 1 or 2, characterized in that: The specific steps in step S3 are: S3.1: Process the images of N students taken at the same time as a group, extract the multi-dimensional features of the N students, and obtain multiple unimodal features for each student; S3.2: Input multiple single-modal features into multiple multi-layer perceptrons to obtain the multi-dimensional feature encoding of each student, and add the student embedding S to the modal encoding of a single student. i .
5. The method for analyzing student classroom participation based on a two-stream hypergraph convolutional network according to claim 3 is characterized in that: The specific steps for modeling the contagion effect of student participation through the multivariate communication module in step S4 are: S4.1.1: Construct a classroom role sequence with N students participating as a node set Hyperedge set Hypergraph of node weights and hyperedge weights, where: the set of nodes A single node corresponds to a single modal feature of a single student, and the hyperedge set In the case of a single hyperedge encoding multimodal or group-related effects among students; S4.1.2: Perform node convolution, update hyperedge features by aggregating node features, and then propagate hyperedge information to nodes through hyperedge convolution; S4.1.3: Repeat step S4.1.
2. After L iterations, the output of the last iteration is used as the output of the multi-propagation module.
6. The method for analyzing student classroom participation based on a two-stream hypergraph convolutional network according to claim 5 is characterized in that: In step S4.1.2, the node weight is adjusted by calculating the attention score of the node vertex and its associated hyperedge.
7. The method for analyzing student classroom participation based on a two-stream hypergraph convolutional network according to claim 3 is characterized in that: The specific steps of capturing multi-frequency information from the extracted features by the multi-frequency propagation module in step S4 are as follows: S4.2.1: Construct an undirected graph in parallel with the multi-propagation module, the undirected graph including a node set and hyperedge sets Node Set A single node in the CNN corresponds to a single modal feature of a single student; S4.2.2: Connect a single node to all nodes representing the same dimension as other students, as well as nodes of other dimensions of the same student, to obtain the adjacency matrix, and normalize it to obtain the Laplacian matrix of the graph; S4.2.3: Obtain multi-frequency features using a low-pass filter and a high-pass filter, where the high-pass filter is equal to the Laplacian of the normalized graph, and the low-pass filter is equal to the difference between two identity matrices and the Laplacian matrix of the normalized graph. S4.2.4: Combine the low-frequency and high-frequency information obtained by the low-pass filter and the high-pass filter through adaptive weighting; S4.2.5: Each student node aggregates the low-frequency and high-frequency information of neighboring nodes through iterative propagation, and after K layers of message passing, the output of the multi-frequency propagation module is obtained.
8. The method for analyzing student classroom participation based on a two-stream hypergraph convolutional network according to claim 3 is characterized in that: The specific method in step S5 is: Aggregate the outputs from the multi-frequency propagation module and the multivariate propagation model to construct the student's classroom participation μ i , the classroom participation μ i The input is normalized into the softmax layer to obtain the probability distribution of class participation categories and the student's predicted participation label.
9. The method for analyzing student classroom participation based on a two-stream hypergraph convolutional network according to claim 8 is characterized in that: In step S5, the engagement classifier is optimized using the L2 regularized categorical cross entropy loss.