Knowledge tracking and test question answering prediction method based on heterogeneous graph attention network
By adopting heterogeneous graph attention network in educational data mining, combining node embedding and graph attention mechanisms, the problem of insufficient accuracy in the existing technology when dealing with heterogeneous graph data in educational scenarios is solved, and more efficient prediction effects and flexible analysis capabilities are achieved.
Patent Information
- Application Number
- CN202510398036.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-04-01
AI Technical Summary
When processing heterogeneous graph data in educational scenarios, it is difficult for the existing technology to effectively capture the complex relationships of different types of nodes and edges, resulting in insufficient accuracy in predicting the accuracy of students' answering questions, and the model structure is single, making it difficult to adapt to the analysis needs in different scenarios.
The knowledge tracking and test answer prediction method based on heterogeneous graph attention network is used, and different types of nodes are mapped into low-dimensional vector representations through the node embedding layer, and the graph attention layer is used for multi-layer attention mechanism processing, combined with the GCN layer fusion graph structure information, and finally the accuracy prediction is made through the output layer.
It realizes in-depth semantic mining of educational heterogeneous graph data, improves prediction accuracy, and has flexibility to adapt to the analytical needs of individual students and all students, and meets the needs of diverse educational scenarios.
Smart Images

Figure CN119917815A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of educational data mining, and specifically to a knowledge tracking and test answer prediction method based on a heterogeneous graph attention network. Background Art
[0002] In the field of educational data mining, early research focused on simple statistical analysis and association rule mining, which made it difficult to deeply mine complex graph structure relationships. Graph neural networks have developed rapidly recently, but existing heterogeneous graph models mostly use a unified node and edge processing method, which does not fully consider the differences in node types (students, questions, knowledge, timestamps) and multiple edge types (such as student answers, knowledge associations, etc.) in educational scenarios; Disadvantages of existing technologies: Traditional graph neural network models often cannot effectively capture the complex relationships between different types of nodes and edges when processing heterogeneous graph data, especially in heterogeneous graphs composed of students, questions, knowledge, and timestamps in the field of education. This leads to insufficient accuracy in tasks such as predicting the accuracy of students' answers to questions; at the same time, the model has a single structure and is difficult to adapt to the analysis needs of individual students and all students in different scenarios, and its scalability and flexibility are poor. Summary of the invention
[0003] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a knowledge tracking and test answer prediction method based on a heterogeneous graph attention network, including: The node embedding layer is used to map different types of nodes into low-dimensional vector representations, including: The student node embedding layer is used to map the discrete numbers of students into low-dimensional vector representations with a dimension of student_embed_dim; The question node embedding layer is used to represent the attribute characteristics of the question and map the discrete number of the question to a vector with dimension question_embed_dim; The knowledge node embedding layer converts the discrete number of knowledge into a vector with dimension knowledge_embed_dim; The timestamp node embedding layer converts the timestamp information into a vector of dimension timestamp_embed_dim; The edge embedding layer embeds different types of edges in the heterogeneous graph, with the dimension edge_embed_dim; Graph attention layer, including: Linear transformation layer, which converts the dimension of input features into a form suitable for multi-head attention calculation; Type-specific learnable parameter tensors used to compute attention coefficients between nodes connected by different types of edges; LeakyReLU activation function, used to enhance the expressiveness of the model; Type-specific GAT layer, a ModuleList consisting of multiple GraphAttentionLayers, is used to perform multi-layer attention mechanism processing on node features; The GCN layer uses the GCNConv layer as a feature enhancement module to fuse graph structure information to update node features; The output layer, including the linear layer, maps the high-dimensional features to a single prediction value, which is the prediction value of the accuracy of the student's answer to the question.
[0004] Furthermore, it also includes a graph convolution layer for aggregating the information of all students, whose input and output dimensions are both student embedding dimensions student_embed_dim. By constructing a graph structure relationship between students and performing graph convolution operations, the feature representation of each student can be integrated with the relevant information of other students.
[0005] Furthermore, it also includes parameters for constructing student association relationships, which are used to represent knowledge similarity weights and problem similarity weights. When constructing the adjacency matrix between students, the adjacency matrix constructed based on knowledge point associations and problem associations is integrated.
[0006] Furthermore, the forward propagation process of the model includes: S1, data preprocessing and embedding layer operation, processing the graph data of a single student or all students according to the mode selection, extracting the characteristics of each type of node and the information of the edge, and converting the discrete number into a vector representation through the corresponding embedding layer; S2, graph attention layer processing, which passes the embedded representation through multiple graph attention layers in sequence for feature extraction and update; S3, aggregate all student information, construct the associated adjacency matrix of students on knowledge points and questions, and perform graph convolution operations to obtain the aggregated student feature representation; S4, GCN layer and output layer operation, after aggregating all students’ information, the embedded representation is enhanced through the GCNConv layer, and the node pair embedding representation for prediction is constructed and input into the output layer to obtain the prediction accuracy value.
[0007] Furthermore, the process of building the student association relationship includes: S11, build_knowledge_adjacency_matrix method, used to calculate similarity based on knowledge embedding in student embedding, and determine adjacency relationship based on threshold setting to obtain adjacency matrix based on knowledge embedding similarity; S12, build_question_adjacency_matrix method, used to calculate similarity based on question embeddings in student embeddings, and determine adjacency relationships based on threshold settings to obtain an adjacency matrix based on question embedding similarity; S13, compute_similarity method, is used to calculate the cosine similarity between the input embedding vectors and return the similarity matrix.
[0008] The beneficial effects of the present invention are: Mode flexibility: The analysis mode can be switched flexibly between single-student and all-student modes. The single-student mode provides fine-grained support for personalized learning diagnosis, while the all-student mode helps in educational policy formulation and overall teaching evaluation, meeting the needs of diverse educational scenarios, which are difficult to be taken into account with existing technologies.
[0009] Enhanced feature expression: Multi-layer GAT and GCN work together to make node feature expressions richer and effectively mine the deep semantics of educational heterogeneous graphs, which is better than most single-structure models. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a flowchart of the knowledge tracking and test answer prediction method based on heterogeneous graph attention network; Figure 2 Schematic diagram of the implementation of the knowledge tracking and test answer prediction method based on heterogeneous graph attention network. DETAILED DESCRIPTION
[0011] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings, but the protection scope of the present invention is not limited to the following.
[0012] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.
[0013] like Figure 1 As shown in the figure, the knowledge tracking and test answer prediction method based on heterogeneous graph attention network includes: The node embedding layer is used to map different types of nodes into low-dimensional vector representations, including: The student node embedding layer is used to map the discrete numbers of students into low-dimensional vector representations with a dimension of student_embed_dim; The question node embedding layer is used to represent the attribute characteristics of the question and map the discrete number of the question to a vector with dimension question_embed_dim; The knowledge node embedding layer converts the discrete number of knowledge into a vector with dimension knowledge_embed_dim; The timestamp node embedding layer converts the timestamp information into a vector of dimension timestamp_embed_dim; The edge embedding layer embeds different types of edges in the heterogeneous graph, with the dimension edge_embed_dim; Graph attention layer, including: Linear transformation layer, which converts the dimension of input features into a form suitable for multi-head attention calculation; Type-specific learnable parameter tensors used to compute attention coefficients between nodes connected by different types of edges; LeakyReLU activation function, used to enhance the expressiveness of the model; Type-specific GAT layer, a ModuleList consisting of multiple GraphAttentionLayers, is used to perform multi-layer attention mechanism processing on node features; The GCN layer uses the GCNConv layer as a feature enhancement module to fuse graph structure information to update node features; The output layer, including the linear layer, maps the high-dimensional features to a single prediction value, which is the prediction value of the accuracy of the student's answer to the question.
[0014] It also includes a graph convolution layer for aggregating the information of all students. Its input and output dimensions are both student embedding dimensions student_embed_dim. By constructing a graph structure relationship between students and performing graph convolution operations, the feature representation of each student can be integrated with the relevant information of other students.
[0015] It also includes parameters for constructing student association relationships, which are used to represent knowledge similarity weights and problem similarity weights. When constructing the adjacency matrix between students, the adjacency matrix constructed based on knowledge point associations and problem associations is integrated.
[0016] The forward propagation process of the model includes: S1, data preprocessing and embedding layer operation, processing the graph data of a single student or all students according to the mode selection, extracting the characteristics of each type of node and the information of the edge, and converting the discrete number into a vector representation through the corresponding embedding layer; S2, graph attention layer processing, which passes the embedded representation through multiple graph attention layers in sequence for feature extraction and update; S3, aggregate all student information, construct the associated adjacency matrix of students on knowledge points and questions, and perform graph convolution operations to obtain the aggregated student feature representation; S4, GCN layer and output layer operation, after aggregating all students’ information, the embedded representation is enhanced through the GCNConv layer, and the node pair embedding representation for prediction is constructed and input into the output layer to obtain the prediction accuracy value.
[0017] The process of building a student association relationship includes: S11, build_knowledge_adjacency_matrix method, used to calculate similarity based on knowledge embedding in student embedding, and determine adjacency relationship based on threshold setting to obtain adjacency matrix based on knowledge embedding similarity; S12, build_question_adjacency_matrix method, used to calculate similarity based on question embeddings in student embeddings, and determine adjacency relationships based on threshold settings to obtain an adjacency matrix based on question embedding similarity; S13, compute_similarity method, is used to calculate the cosine similarity between the input embedding vectors and return the similarity matrix.
[0018] Specifically, Figure 2 As shown, the present invention includes a node embedding layer: student_embed: Student node embedding layer, which maps the discrete number of students into a low-dimensional vector representation with the dimension student_embed_dim.
[0019] question_embed: question node embedding layer, which maps the discrete number of questions into a vector with the dimension of question_embed_dim. It is used to represent the attribute features of questions.
[0020] knowledge_embed: knowledge node embedding layer, which converts the discrete number of knowledge into a vector with the dimension of knowledge_embed_dim.
[0021] timestamp_embed: Timestamp node embedding layer, which converts timestamp information into a vector with a dimension of timestamp_embed_dim.
[0022] edge_embed: edge embedding layer, which embeds different types of edges in heterogeneous graphs, with the dimension edge_embed_dim.
[0023] GraphAttentionLayer: Core components and calculation process: W: Linear transformation layer, which converts the dimension of input feature x from in_channels to num_heads*out_channels for multi-head attention calculation. Its function is to perform preliminary feature transformation on the input feature to adapt it to the requirements of the multi-head attention mechanism and provide suitable feature dimensions for subsequent attention calculation and message passing.
[0024] type_specific_a: type-specific learnable parameter tensor with dimension (edge_types, num_heads, 2*out_channels). Specific attention parameters are selected according to the edge type to calculate the attention coefficients between nodes connected by different types of edges. This parameter enables the model to learn different attention modes for different types of edge relationships, thereby processing the diverse edge information in heterogeneous graphs in a more detailed manner.
[0025] leaky_relu: LeakyReLU activation function, which is used to introduce nonlinear factors, enhance the expressiveness of the model, and activate the intermediate results during the attention calculation process.
[0026] Forward propagation process: First, use the add_self_loops function to add self-loop edges to the input edge_index to ensure that the node can consider its own feature information during the message passing process, making the graph structure information more complete and rich.
[0027] Transform the input feature x, get h through the W layer, and adjust the shape for subsequent multi-head attention calculation. This step transforms and reshapes the original node features to prepare for the parallel calculation of the multi-head attention mechanism. Get the node index row and col of the edge connection according to edge_index, and select the corresponding type of attention parameter from type_specific_a according to the edge type edge_type. Concatenate the node feature h according to the edge connection to get a_input, and integrate the performance_factor based on the performance of all students to consider the information related to the rules of all students, so that the attention calculation is more targeted and effective. The attention score attention_scores is obtained through a series of tensor operations, leaky_relu activation and softmax operations, and dropout is applied for random inactivation to prevent overfitting.
[0028] Finally, by calling the propagate method inherited from the MessagePassing base class, the features of the neighbor nodes are aggregated according to the attention scores, the message passing process is completed, and finally the processed and reshaped output features are returned.
[0029] Type specific GAT layers: ModuleList, which consists of multiple GraphAttentionLayers, is used to perform multi-layer attention mechanism processing on node features. By stacking multiple graph attention layers, the model can gradually extract higher-level and more abstract feature representations, and deeply explore the complex relationships and potential patterns between nodes in heterogeneous graphs. Each layer performs further attention calculations and feature updates based on the output features of the previous layer, so that node features can integrate more levels of graph structure information and neighbor node information, enhancing the model's feature extraction and expression capabilities for heterogeneous graph data.
[0030] GCN layer: As a feature enhancement module, the GCNConv layer further uses graph convolution operations to fuse graph structure information to update node features based on the features processed by the graph attention layer.
[0031] Output layer: The linear layer output_layer has an input dimension of total_embed_dim*num_heads*2 and an output dimension of 1. The function of this layer is to perform the final linear transformation on the node features processed by the previous multi-layer graph neural network, and map the high-dimensional features to a single prediction value, that is, the prediction value of the correctness of the student's answer to the question. By adjusting the weight and bias parameters, the output layer can output a prediction probability value between 0 and 1 based on the input feature information, indicating the possibility that the student's answer to the question is correct, thereby achieving the prediction goal of the model.
[0032] Graph convolutional layer for aggregating all student information: The graph convolution layer is specially designed to aggregate the relevant information of all students. The input and output dimensions are both student embedding dimensions student_embed_dim. Its purpose is to fully consider the performance rules of all students on the same knowledge points and problems in the model, and to enable the feature representation of each student to integrate the relevant information of other students by constructing the graph structure relationship between students and performing graph convolution operations.
[0033] Parameters for constructing student association relationships: Learnable parameter matrices are used to represent knowledge similarity weights and problem similarity weights. When constructing the adjacency matrix between students, these weight parameters are used to fuse the adjacency matrix constructed based on knowledge point associations and problem associations. By adjusting the weights, the model can flexibly control the importance of different aspects of association in constructing relationships between students, thereby more accurately capturing the complex association patterns between students based on knowledge and problems, making the process of aggregating all student information more adaptable and effective.
[0034] Model forward propagation process Data preprocessing and embedding layer operations: According to the different values of is_single_student_mode, the graph data of a single student or all students is processed respectively. The corresponding features of each type of node such as students, questions, knowledge, timestamps, and edge information are extracted from the input data dictionary data. Then, the discrete numbers of each type of node are converted into vector representations through the corresponding embedding layer, and these vectors are spliced along the last dimension to obtain the combined embedding representation combined_embeds.
[0035] Figure attention layer processing: Combined_embeds is passed through multiple graph attention layers (gat_layer) for feature extraction and update. In each graph attention layer, feature transformation, attention calculation and message passing are performed according to the calculation process of the graph attention layer, so that the node features gradually integrate more graph structure-related information and neighbor node information, thereby extracting more representative and discriminative feature representations. Through the stacking of multiple layers of graph attention mechanism.
[0036] Aggregate all student information: Call the aggregate_all_students_info method to aggregate the information of all students. This method first extracts the embedding part student_embeds of the student node from combined_embeds, and then builds the adjacency matrix based on the association between students in knowledge points and questions. Specifically, the adjacency matrix based on knowledge embedding similarity is obtained through the build_knowledge_adjacency_matrix method, and the adjacency matrix based on question embedding similarity is obtained through the build_question_adjacency_matrix method, and the two adjacency matrices are fused according to knowledge_similarity_weight and question_similarity_weight to obtain the final adjacency matrix adj_matrix between students. Then, the adjacency matrix is converted into the edge index format edge_index, and the student embedding student_embeds and edge index edge_index are passed to the student_agg_gcn graph convolution layer for graph convolution operation, so that the information between students is propagated and aggregated in the constructed graph structure, and the aggregated student feature representation aggregated_embed is obtained. Finally, the aggregated features are expanded to the same shape as the input combined_embeds and added to the original combined_embeds to incorporate the features into the information related to all students. This process fully utilizes the performance rules of all students on the same knowledge points and questions by constructing a reasonable student association graph structure and performing graph convolution operations, providing more comprehensive information reference for predicting the correct answer rate of a single student, and enhancing the model's prediction ability and generalization performance.
[0037] GCN layer and output layer operations: After aggregating all the student information, combined_embeds is enhanced by the GCNConv layer, and the node features are further updated by graph convolution operations, so that the node features can integrate a wider range of graph structure information. Then, the node pair embedding representation pair_embeds for prediction is constructed according to the prediction task requirements. In the single student mode, the features of the student node and the question node are spliced together in order through specific index operations; in the all-student mode, the corresponding node pairs are generated according to the number of all students and all questions. Finally, pair_embeds is input to the output layer output_layer, and the predicted accuracy value is obtained and returned through linear transformation. The output layer maps the high-dimensional features processed by the previous multi-layer graph neural network to a single prediction value, realizing the conversion from feature extraction to the final prediction result, and completing the model's prediction task of the accuracy of students' answers to questions.
[0038] Methods for building adjacency matrices build_knowledge_adjacency_matrix method: First, the corresponding knowledge embeddings knowledge_embeds are extracted from the student embeddings student_embeds. The knowledge embeddings are in a fixed interval of the student embedding features (implemented by the extract_knowledge_embeds method). This interval needs to be adjusted and determined according to the feature arrangement in the actual data.
[0039] Then, the similarity between knowledge embeddings is calculated, and the cosine similarity is calculated using the compute_similarity method to obtain the similarity matrix similarity_matrix between knowledge embeddings.
[0040] Finally, the adjacency relationship is determined by setting a threshold based on the similarity, and the position corresponding to the element that meets the threshold condition is set to an adjacency relationship (value is 1), otherwise it is a non-adjacency relationship (value is 0), and the adjacency matrix adj_matrix based on knowledge embedding similarity is obtained. The adjacency matrix constructed in this way can reflect the association relationship between students based on their knowledge mastery, and provides a graph structure information foundation based on knowledge factors for subsequent graph convolution operations.
[0041] build_question_adjacency_matrix method: Similar to the way of constructing the knowledge adjacency matrix, we first extract the question embeddings question_embeds from the student embeddings, and then calculate the similarity between the question embeddings to obtain the similarity matrix. Then, we determine the adjacency relationship according to different threshold settings to obtain the adjacency matrix based on the similarity of the question embeddings.
[0042] compute_similarity method: Calculate the cosine similarity between the input embedding vectors and return the similarity matrix. The specific implementation process is to traverse all pairs of embedding vectors, for each pair of embedding vectors, use the F.cosine_similarity function to calculate the cosine similarity between them, and fill the result into the corresponding position of the similarity matrix.
[0043] The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not deviate from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.
Claims
1. A knowledge tracking and test answer prediction method based on heterogeneous graph attention network, characterized in that: include: The node embedding layer is used to map different types of nodes into low-dimensional vector representations, including: The student node embedding layer is used to map the discrete numbers of students into low-dimensional vector representations with a dimension of student_embed_dim; The question node embedding layer is used to represent the attribute characteristics of the question and map the discrete number of the question to a vector with dimension question_embed_dim; The knowledge node embedding layer converts the discrete number of knowledge into a vector with dimension knowledge_embed_dim; The timestamp node embedding layer converts the timestamp information into a vector with a dimension of timestamp_embed_dim; The edge embedding layer embeds different types of edges in the heterogeneous graph, with the dimension edge_embed_dim; Graph attention layer, including: Linear transformation layer, which converts the dimension of input features into a form suitable for multi-head attention calculation; Type-specific learnable parameter tensors used to compute attention coefficients between nodes connected by different types of edges; LeakyReLU activation function, used to enhance the expressiveness of the model; Type-specific GAT layer, a ModuleList consisting of multiple GraphAttentionLayers, is used to perform multi-layer attention mechanism processing on node features; The GCN layer uses the GCNConv layer as a feature enhancement module to fuse graph structure information to update node features; The output layer, including the linear layer, maps the high-dimensional features to a single prediction value, which is the prediction value of the accuracy of the student's answer to the question.
2. The method for knowledge tracking and test answer prediction based on heterogeneous graph attention network according to claim 1 is characterized in that: It also includes a graph convolution layer for aggregating the information of all students. Its input and output dimensions are both student embedding dimensions student_embed_dim. By constructing a graph structure relationship between students and performing graph convolution operations, the feature representation of each student can be integrated with the relevant information of other students.
3. The method for knowledge tracking and test answer prediction based on heterogeneous graph attention network according to claim 2 is characterized in that: It also includes parameters for constructing student association relationships, which are used to represent knowledge similarity weights and problem similarity weights. When constructing the adjacency matrix between students, the adjacency matrix constructed based on knowledge point associations and problem associations is integrated.
4. The method for knowledge tracking and test answer prediction based on heterogeneous graph attention network according to claim 3 is characterized in that: The forward propagation process of the model includes: S1, data preprocessing and embedding layer operation, processing the graph data of a single student or all students according to the mode selection, extracting the characteristics of each type of node and the information of the edge, and converting the discrete number into a vector representation through the corresponding embedding layer; S2, graph attention layer processing, which passes the embedded representation through multiple graph attention layers in sequence for feature extraction and update; S3, aggregate all student information, construct the associated adjacency matrix of students on knowledge points and questions, and perform graph convolution operations to obtain the aggregated student feature representation; S4, GCN layer and output layer operation, after aggregating all students’ information, the embedded representation is enhanced through the GCNConv layer, and the node pair embedding representation for prediction is constructed and input into the output layer to obtain the prediction accuracy value.
5. The method for knowledge tracking and test answer prediction based on heterogeneous graph attention network according to claim 4 is characterized in that: The process of building a student association relationship includes: S11, build_knowledge_adjacency_matrix method, used to calculate similarity based on knowledge embedding in student embedding, and determine adjacency relationship based on threshold setting to obtain adjacency matrix based on knowledge embedding similarity; S12, build_question_adjacency_matrix method, used to calculate similarity based on question embeddings in student embeddings, and determine adjacency relationships based on threshold settings to obtain an adjacency matrix based on question embedding similarity; S13, compute_similarity method, is used to calculate the cosine similarity between the input embedding vectors and return the similarity matrix.
Citation Information
Patent Citations
Intelligent student ability assessment method based on associated skill knowledge
CN116823027A
Knowledge tracking method based on knowledge concept association and historical attention information
CN118747527A
Static heterogeneous network link prediction method and system based on graph attention network
CN119004255A
Method and system for relation learning by multi-hop attention graph neural network
US20220092413A1
Classroom teaching cognitive load measuring system
WO2020010785A1