Score prediction method and system based on adaptive graph neural network
Through the adaptive graph neural network method, additional features are generated using decision trees and conditional variational autoencoder, and combined with graph convolutional neural network training score prediction model, the problem of low accuracy of grade prediction in the existing technology is solved, and higher prediction accuracy and generalization ability are achieved.
Patent Information
- Application Number
- CN202510455124.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-19
AI Technical Summary
In the prior art, adaptive graph neural networks have not been applied to achievement prediction, resulting in low accuracy of achievement prediction.
Adaptive graph neural network method is used to filter feature data through decision trees to form a graph, use conditional variational autoencoder to generate additional features, and train it in combination with graph convolutional neural network to establish a score prediction model.
It improves the accuracy, accuracy, recall and F1 value of student achievement prediction, and has strong generalization ability of model and is suitable for online learning environment.
Smart Images

Figure CN120509513A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of educational technology, and in particular to a performance prediction method and system based on an adaptive graph neural network. Background Art
[0002] Online learning is a novel teaching method that is not affected by time or location, making it popular among teachers and students. However, compared with traditional offline learning, the educational quality of online learning is less than satisfactory. Student performance prediction is a key task in the field of educational data mining. Previous research generally falls into three categories: the first category involves correlation analysis between learner characteristics and academic performance; the second category involves student performance prediction based on traditional machine learning methods; and the third category involves student performance prediction based on deep learning methods.
[0003] The first approach involves analyzing and exploring the correlation between various learner characteristics and learning outcomes, specifically, academic performance. The impact of learner characteristics on academic performance is a crucial factor in understanding learning outcomes. Sunar, Bonafini, Hussain, Yang, Summers, and others have studied the correlation between learner characteristics and academic performance. However, these studies primarily explore the various characteristics that influence student academic performance and fail to provide timely intervention for students at risk of failure.
[0004] In recent years, many researchers have focused on predicting student grades using various machine learning methods (such as logistic regression, decision trees, naive Bayes, and support vector machines). Huang, Brooks, Chen Zijian, Lacave, Francis, Ban Wenjing, and others have conducted research using various machine learning methods. Traditional machine learning methods require extracting useful features from the data and combining them with different classifiers to predict student grades. While these studies have achieved some progress, this feature extraction method limits predictive performance.
[0005] Deep learning technology automatically learns student performance characteristics, replacing traditional feature engineering techniques. Recently, using deep learning methods to predict student academic performance has attracted widespread research attention. Existing literature shows that deep learning methods outperform traditional methods in predicting student performance. Tomasevic, Hassan, Waheed, Mubarak, and others have studied the effectiveness of performance prediction using various deep learning techniques, achieving promising results. However, research on the application of graph neural networks in education is rare. Summary of the Invention
[0006] The embodiments of the present application provide a performance prediction method and system based on an adaptive graph neural network, which is used to solve the problem in the prior art that the adaptive graph neural network is not applied to performance prediction, resulting in low accuracy of performance prediction.
[0007] On the one hand, embodiments of the present application provide a performance prediction method based on an adaptive graph neural network, including:
[0008] Obtain training data, which includes feature data and score labels of multiple students;
[0009] Input the feature data into the decision tree, use the decision tree to filter the feature data, use the students with retained feature data as nodes, and connect edges between adjacent nodes to form a graph;
[0010] Taking each node in the graph as the center, a conditional variational autoencoder is used to generate additional features for each central node, and the additional features are combined with the feature data of the central node to form enhanced features;
[0011] Extracting the adjacency matrix of the graph with enhanced features, and establishing a graph convolutional neural network based on the adjacency matrix. The graph convolutional neural network includes a first layer and subsequent graph convolution layers. The graph convolution layers other than the first layer are represented by the sum of the first layer and the previous layer of the graph convolution layer through adaptive parameters. The adaptive parameters are determined by adaptive residuals.
[0012] The feature data and enhanced features are input into the first layer. The parameters of the graph convolutional neural network are adjusted based on the difference between the classification results of the graph convolutional neural network and the grade labels to complete the training of the graph convolutional neural network and obtain the grade prediction model.
[0013] The student's feature data to be predicted is input into the performance prediction model to obtain the performance prediction results.
[0014] On the other hand, the embodiment of the present application also provides a performance prediction system based on an adaptive graph neural network, including:
[0015] The data acquisition module is used to obtain training data, which includes feature data and score labels of multiple students;
[0016] A graph building module is used to input feature data into a decision tree, filter the feature data using the decision tree, use students with retained feature data as nodes, and connect edges between adjacent nodes to form a graph;
[0017] The feature enhancement module is used to generate additional features for each central node using a conditional variational autoencoder, with each node in the graph as the center. The additional features are combined with the feature data of the central node to form enhanced features.
[0018] A network building module is used to extract the adjacency matrix of the graph with enhanced features and build a graph convolutional neural network based on the adjacency matrix. The graph convolutional neural network includes a first layer and subsequent graph convolution layers. The graph convolution layers other than the first layer are represented by the sum of the first layer and the previous layer of the graph convolution layer through adaptive parameters. The adaptive parameters are determined by adaptive residuals.
[0019] The model training module is used to input feature data and enhanced features into the first layer, and adjust the parameters of the graph convolutional neural network based on the difference between the classification results of the graph convolutional neural network and the grade labels to complete the training of the graph convolutional neural network and obtain the grade prediction model;
[0020] The grade prediction module is used to input the student's feature data to be predicted into the grade prediction model to obtain the grade prediction result.
[0021] The adaptive graph neural network-based grade prediction method and system in this application first uses semantic mining to extract the correlation between student feature data and grades. Second, it performs feature enhancement on the graph, improving the performance of automatic grade prediction. Experimental results show that compared with existing mainstream student grade prediction methods, this method achieves better accuracy, precision, recall, and F1 scores on academic data from xAPI (learning behavior data) and OULA (Open University Learning Analytics Data), and demonstrates strong model generalization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] Figure 1 Flowchart of the performance prediction method based on adaptive graph neural network provided in an embodiment of the present application. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0025] Figure 1Flowchart of a performance prediction method based on an adaptive graph neural network provided in an embodiment of the present application. This embodiment of the present application provides a performance prediction method based on an adaptive graph neural network, including:
[0026] S100, obtaining training data, where the training data includes feature data and score labels of multiple students.
[0027] For example, the training data used in the embodiments of the present application can use an xAPI dataset or an OULA dataset. Regardless of which dataset is used, the training data is in CSV (comma-separated value) format. Each row of these training data represents the data of a student, where the first few columns in a row of data are feature data, and the last column is the score label.
[0028] S110, inputting the feature data into a decision tree, using the decision tree to screen the feature data, taking the students with the retained feature data as nodes, connecting edges between adjacent nodes to form a graph.
[0029] For example, there may be some potential connections between the feature data of different students. For example, students with good grades may have some similar behavioral characteristics. These potential connections help to construct edges in the graph, which is very similar to the homogeneity principle in GCN (graph convolutional neural network). Therefore, this application uses data mining methods to model the potential connections between feature data, and then uses these established edges to further predict.
[0030] Decision trees can perform correlation analysis on feature data very well, aggregate feature data with strong similarity, and distinguish feature data with weak similarity. In the process of constructing the graph, this application uses a decision tree to analyze the input feature data to evaluate the importance of each feature data. The decision tree calculates the effect of each feature data in the splitting process and selects the feature data that is most important for score prediction based on the Gini coefficient indicator. Based on the size of the Gini coefficient, the way to connect the student nodes is determined. Different decision tree depths, Gini coefficient thresholds, and numbers of connections are used for different data sets. For the xAPI dataset, the depth of the decision tree is set to 3. The Gini coefficient of each student leaf node is first calculated in the decision tree, and then random connections are made between the student leaf nodes with a Gini coefficient less than 0.4, and then the graph is constructed. The number of connections is set to 5. For OULA data, the depth of the decision tree is 4. Similarly, random connections are made between leaf nodes with a Gini coefficient less than 0.4, and the number of connections is set to 5.
[0031] This application constructs a graph based on the Gini coefficient to maximize the amount of correlated feature data in the training data. This allows relatively highly correlated feature data to be connected as nodes without introducing excessive noise edges. Each student leaf node has 5 edges connected to it. Graph convolutional neural networks are more prone to oversmoothing for dense graphs with many edges, as the number of message passing operations increases with the number of edges, causing some deep information in the input data to be lost as the number of model layers increases.
[0032] S120, taking each node in the graph as the center, using a conditional variational autoencoder to generate additional features for each center node, and combining the additional features with the feature data of the center node to form enhanced features.
[0033] For example, the task of predicting student grades belongs to the node classification in the graph neural network. Since the training data may have problems such as noise and insufficient data volume, it is particularly important to use data enhancement. In particular, the sample size of grade data in the xAPI dataset is small, and the graph model has less information to learn during the learning process, which will lead to the inability to fully explore the correlation between feature data and grades. Data enhancement can learn more information related to student grade prediction in addition to the original data. The feature enhancement method can be embedded in the node classification. Therefore, this application uses node feature enhancement to solve the above problems.
[0034] Feature enhancement is to enhance the representation learning ability of the model by adding additional features. During the enhancement process, each node in the graph is regarded as a central node, and the feature data of its first-order neighbor nodes are used to generate new features of the central node for further enhancement. Specifically, this application adopts a conditional variational autoencoder (CVAE) to regard all nodes in the graph as central nodes, and the first-order neighbor nodes of the central node obey the conditional distribution conditional on the central node. The conditional variational autoencoder includes an encoder and a decoder. The feature data of the central node and the feature data of the neighbor nodes are input into the encoder to generate corresponding hidden variables and maximize the lower bound of evidence. The hidden variables and the feature data of the central node are input into the decoder to generate additional features.
[0035] This application will use conditional variational autoencoders to generate additional features, and input relevant features beyond the existing student feature data into the model as additional information, further improving the learnability of the model and reducing overfitting caused by the small amount of academic performance data.
[0036] S130, extracting the adjacency matrix of the graph with enhanced features, and establishing a graph convolutional neural network based on the adjacency matrix, wherein the graph convolutional neural network includes a first layer and subsequent graph convolution layers, and graph convolution layers other than the first layer are represented by adding adaptive parameters of the first layer and the previous layer of the graph convolution layer, and the adaptive parameters are determined using adaptive residuals.
[0037] For example, in order to increase the representation learning ability of the model and enable the model to more deeply explore the correlation between students' feature data and grades, adaptive initial residual (AIR) is introduced on the basis of the conventional GCN network architecture.
[0038] AIR decomposes each layer of graph convolution in GCN into a combination of P and T operations. P operations represent message propagation between neighboring nodes, and T operations apply nonlinear transformations to node representations, allowing the model to capture the distribution of training data. Larger P operations allow node representations to derive from multi-order neighborhood information, while larger T operations empower the model with stronger expressive power. Excessive P operations can lead to oversmoothing, while excessive T operations can cause performance degradation due to performance saturation. Model degradation is the primary cause of early performance degradation, and the use of adaptive residuals can prevent this, enabling better predictions of online student performance.
[0039] In the first layer of the GCN network, the feature data in the training data is combined with the generated additional features to form enhanced data, which is input to the first layer of weight sharing. The output of this first layer is as follows:
[0040]
[0041] Among them, H (0) represents the output feature of the first layer after the graph convolution operation, σ1 represents the nonlinear activation function, represents the normalized adjacency matrix, represents the degree matrix of the graph, represents the adjacency matrix of the graph, X represents the feature data, Represents the enhanced feature data, W represents the trainable weight matrix, and || represents splicing.
[0042] In the other graph convolution layers after the first layer, a learnable weight parameter α is used to make the representation of the lth graph convolution layer the sum of the first layer and the l-1th graph convolution layer through adaptive parameter matching. Each graph convolution layer is represented as follows:
[0043]
[0044]
[0045] in, represents the learnable parameters of the lth layer, σ2 represents the Sigmoid function, represents the feature representation of the i-th node in the l-1 layer, represents the feature representation of the i-th node in the input layer (i.e., layer 0), | represents the concatenation operation, and H (l) and H (l-1) Represent the output features of the lth layer and the k-1th layer respectively, σ3 represents the ReLu activation function, α l-1 represents the learnable parameters of the l-1th layer, and ⊙ represents element-wise multiplication.
[0046] S140: Input the feature data and enhanced features into the first layer, and adjust the parameters of the graph convolutional neural network according to the difference between the classification results of the graph convolutional neural network and the score labels to complete the training of the graph convolutional neural network and obtain the score prediction model.
[0047] S150, inputting the student's feature data to be predicted into the performance prediction model to obtain the performance prediction result.
[0048] The present application also provides a performance prediction system based on an adaptive graph neural network, the system comprising:
[0049] The data acquisition module is used to obtain training data, which includes feature data and score labels of multiple students;
[0050] A graph building module is used to input feature data into a decision tree, filter the feature data using the decision tree, use students with retained feature data as nodes, and connect edges between adjacent nodes to form a graph;
[0051] The feature enhancement module is used to generate additional features for each central node using a conditional variational autoencoder, with each node in the graph as the center. The additional features are combined with the feature data of the central node to form enhanced features.
[0052] A network building module is used to extract the adjacency matrix of the graph with enhanced features and build a graph convolutional neural network based on the adjacency matrix. The graph convolutional neural network includes a first layer and subsequent graph convolution layers. The graph convolution layers other than the first layer are represented by the sum of the first layer and the previous layer of the graph convolution layer through adaptive parameters. The adaptive parameters are determined by adaptive residuals.
[0053] The model training module is used to input feature data and enhanced features into the first layer, and adjust the parameters of the graph convolutional neural network based on the difference between the classification results of the graph convolutional neural network and the grade labels to complete the training of the graph convolutional neural network and obtain the grade prediction model;
[0054] The grade prediction module is used to input the student's feature data to be predicted into the grade prediction model to obtain the grade prediction result.
[0055] Experimental Description
[0056] In this example, experiments were conducted using the xAPI and OULA academic performance datasets. The academic performance data used in the experiments was balanced. The hyperparameters of the proposed model were: 3 convolutional layers in the graph convolutional neural network model, and 2 sampling of the data. The model was optimized using the Adam optimizer with a learning rate of 0.01. The number of iterations was set to 100, and the dropout setting was set to 0.5.
[0057] 1. Experimental results
[0058] To demonstrate the effectiveness of the proposed model, the performance of the comparison method and the proposed ALGCN model are verified on the academic achievement dataset.
[0059] On the xAPI academic performance data, the ALGCN model outperformed all comparison methods. The ALGCN model achieved an accuracy of 93.91%. The nested ensemble classifier, bagging+2 algorithms, CatBoost, and hybrid approaches achieved accuracy scores of 79.17%, 80.83%, 92.27%, and 93%, respectively. Compared to these methods, the ALGCN model achieved accuracy improvements of 14.74%, 13.08%, 1.64%, and 0.91%, respectively. This is primarily due to the GCN's representational learning capabilities, which effectively learn the inherent correlations between students in online learning, thus improving student performance prediction. Neural network methods were also compared on the xAPI data. The ANN achieved an accuracy score of 78.1%. These results demonstrate that the proposed ALGCN model outperforms neural network methods on all accuracy metrics. This is due to the inclusion of adaptive residuals and local data augmentation in the ALGCN model, which further improve the results. Experimental results demonstrate the effectiveness of the ALGCN model in predicting student grades.
[0060] The ALGCN model achieved an accuracy of 94.00% on the OULA academic performance data. This demonstrates that the ALGCN model outperforms existing traditional machine learning and deep learning methods. These results demonstrate the suitability of the proposed ALGCN model for modeling student performance prediction in online learning.
[0061] The performance of the ALGCN model is compared with the Naive Bayes method used by Azizah et al. The Naive Bayes method achieved an accuracy of 67.66%. Compared to the Naive Bayes method, the ALGCN model achieved a 26.34% improvement in accuracy. This demonstrates the powerful representation learning capabilities of the adaptive graph neural network approach for student performance prediction, effectively modeling the relationship between student characteristics and performance. The experimental results of the ALGCN model are further compared with those of the LSTM, Multilayer Perceptron, and Hybrid 2D-CNN methods. The LSTM, Multilayer Perceptron, and Hybrid 2D-CNN methods achieved accuracy of 80.40, 78.2, and 88, respectively. Compared to the LSTM method, the ALGCN model achieved a 13.6% improvement in accuracy. Experimental results demonstrate that data augmentation in the ALGCN model can learn more information useful for student performance prediction. Compared to the Multilayer Perceptron method, the ALGCN model achieved a 15.8% improvement in accuracy. Compared to the Hybrid2D-CNN approach, the ALGCN model achieved a 6% improvement in accuracy. Observing these experimental results, we found that the student score prediction model based on graph neural networks performed better than the non-graph neural network-based model. These experimental results demonstrate that the graph neural network-based student score prediction model is more suitable for modeling student score prediction tasks in online learning. The ALGCN model has strong representational learning capabilities and can deeply explore the semantic relationship between student characteristics and scores.
[0062] The ALGCN model in this experiment achieved optimal performance on both xAPI and OULA academic data, indicating that the proposed ALGCN model is effective in learning the diverse features of different academic data and can accurately identify students who are academically successful and those who are academically unsuccessful. This demonstrates the model's strong generalization ability and verifies the universality of this application method.
[0063] 2. Ablation studies
[0064] To further verify the effectiveness of the model, an ablation study experiment of the ALGCN model will be conducted. The main model outperforms the performance of other sub-modules, although the sub-modules may have good performance on other tasks. The ALGCN main model and its sub-modules (i.e., without feature enhancement and without adaptive residual) are applied to two academic performance datasets, xAPI and OULA, respectively. Table 1 shows the experimental results of the ablation study. The results in the "ALGCN w / o local" row refer to the experimental results of the ALGCN model without feature enhancement. The results in the "ALGCN w / o AIR" row refer to the experimental results of the ALGCN model without adaptive residual.
[0065] Table 1 Ablation performance analysis of the ALGCN model for score prediction
[0066]
[0067] From Table 1 we can see that:
[0068] 1) Under the same settings, the proposed ALGCN model outperforms all submodules when the student performance data volume is large or small. The ALGCN model achieves superior performance on the smaller xAPI academic performance data because data augmentation allows the proposed model to learn more information useful for student performance prediction. The addition of adaptive residuals further enhances the model's learning ability and reduces overfitting. The performance improvement is particularly pronounced on the larger OULA academic performance data.
[0069] 2) As expected, the simplified model without feature augmentation showed a decrease in performance on the academic performance prediction task. This decrease in performance suggests that feature augmentation enables the model to learn more useful representational information from the data itself, which is useful for predicting student performance. This demonstrates the importance of the feature augmentation module.
[0070] 3) As expected, the simplified model without the adaptive residual also showed decreased performance on the academic performance prediction task. This suggests that the adaptive residual can enhance the model's representational learning capabilities, enabling deeper exploration of the mutual information between student characteristics and grades, and better modeling the relationship between these characteristics and grades. This validates the effectiveness of the adaptive residual module.
[0071] 4) When the data size is small, the impact of adaptive residuals on experimental results is smaller than that of feature augmentation, while the opposite is true when the data size is large. When the amount of student performance data is small, feature augmentation plays a more significant role because the model can learn additional representational information that is useful for predicting student performance. When the data size is large, the model's ability to learn representational information becomes more important because the dataset itself can provide relatively more information useful for predicting student performance, and the addition of adaptive residuals can significantly improve prediction results.
[0072] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0073] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. The performance prediction method based on adaptive graph neural network is characterized by: include: Acquire training data, wherein the training data includes feature data and score labels of multiple students; Inputting the feature data into a decision tree, screening the feature data using the decision tree, using students who retain the feature data as nodes, and connecting edges between adjacent nodes to form a graph; Taking each node in the graph as the center, using a conditional variational autoencoder to generate additional features for each center node, and combining the additional features with the feature data of the center node to form enhanced features; Extracting an adjacency matrix of the graph with the enhanced features, and establishing a graph convolutional neural network based on the adjacency matrix, wherein the graph convolutional neural network includes a first layer and subsequent graph convolution layers, wherein the graph convolution layers other than the first layer are represented by adding adaptive parameters of the first layer and a previous layer of the graph convolution layer, and the adaptive parameters are determined using an adaptive residual; Inputting the feature data and the enhanced features into the first layer, and adjusting the parameters of the graph convolutional neural network based on the difference between the classification result of the graph convolutional neural network and the grade label to complete the training of the graph convolutional neural network and obtain a grade prediction model; The student's feature data to be predicted is input into the performance prediction model to obtain the performance prediction result.
2. The performance prediction method based on adaptive graph neural network according to claim 1, characterized in that: The decision tree predicts the performance of the input feature data and connects the bottom leaf nodes of the decision tree.
3. The performance prediction method based on adaptive graph neural network according to claim 2, characterized in that: Before connecting edges, the Gini coefficients of the leaf nodes are calculated first, and then edges are randomly connected between the leaf nodes whose Gini coefficients are less than a threshold.
4. The performance prediction method based on adaptive graph neural network according to claim 3 is characterized in that: The depth of the decision tree, the threshold, and the number of edges are all set according to the dataset to which the training data belongs.
5. The performance prediction method based on adaptive graph neural network according to claim 1, characterized in that: The conditional variational autoencoder includes an encoder and a decoder. The feature data of the central node and the feature data of the neighboring nodes are input into the encoder to generate corresponding latent variables. The latent variables and the feature data of the central node are input into the decoder to generate the additional features.
6. The performance prediction method based on adaptive graph neural network according to claim 5, characterized in that: The neighbor nodes obey a conditional distribution with the central node as a condition.
7. A system using the performance prediction method based on an adaptive graph neural network according to any one of claims 1 to 6, characterized in that: include: A data acquisition module is used to acquire training data, wherein the training data includes feature data and score labels of multiple students; A graph building module, configured to input the feature data into a decision tree, screen the feature data using the decision tree, use students who retain the feature data as nodes, and connect edges between adjacent nodes to form a graph; a feature enhancement module, configured to use a conditional variational autoencoder to generate additional features for each central node with each node in the graph as the center, and combine the additional features with the feature data of the central node to form enhanced features; A network establishment module, configured to extract an adjacency matrix of the graph having the enhanced features, and establish a graph convolutional neural network based on the adjacency matrix, wherein the graph convolutional neural network includes a first layer and subsequent graph convolution layers, wherein the graph convolution layers other than the first layer are represented by adding adaptive parameters of the first layer and the layer before the graph convolution layer, and the adaptive parameters are determined using an adaptive residual; a model training module, configured to input the feature data and the enhanced features into the first layer, and adjust the parameters of the graph convolutional neural network based on the difference between the classification results of the graph convolutional neural network and the grade labels, so as to complete the training of the graph convolutional neural network and obtain a grade prediction model; The performance prediction module is used to input the student's feature data to be predicted into the performance prediction model to obtain the performance prediction result.