Knowledge tracking method and model fusing heterogeneous graph neural network and historical performance

By constructing heterogeneous global graphs and adaptive fusion knowledge states, the shortcomings of the existing knowledge tracking model in question and concept representation, global information utilization and long-term dependency modeling are solved, and better prediction performance and knowledge state understanding are achieved.

CN120146161APending Publication Date: 2025-06-13ARMY ENG UNIV OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226074.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing knowledge tracking model has difficulty in diagnosing students' learning status, lacks global information utilization and long-term dependency modeling, resulting in limited model performance.

Method used

The knowledge tracking method of fusion of heterogeneous graph neural network and historical performance is adopted. By constructing a heterogeneous global graph, the graph neural network algorithm is used to enhance the characterization of questions, and the importance weighted state is extracted from the historical knowledge state, adaptively fuse the current and historical knowledge states to predict students' answering performance in the next moment.

Benefits of technology

It improves the prediction performance of the knowledge tracking model, overcomes long-term dependence problems, can more accurately model the potential relationship between questions and knowledge points, and enhances the understanding of students' knowledge state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146161A_ABST
    Figure CN120146161A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge tracking method and model fusing a heterogeneous graph neural network and historical performance, and the method comprises the steps: firstly, jointly constructing a heterogeneous graph structure through the combination of an inclusion relation between questions and concepts, a transfer relation between the questions and a co-occurrence relation between the concepts, and enhancing the question representation through a graph neural network algorithm; secondly, in order to overcome the problem of long-term dependence, starting from the importance of the historical knowledge state of the student, dividing the knowledge state into a current-moment knowledge state and a historical important knowledge state, obtaining the current-moment knowledge state through a long-short-term memory network, and generating the historical important knowledge state based on importance coefficient weighted aggregation; and finally, predicting the answering performance of the student by fusing the two knowledge states and combining the question characterization. Experimental results show that the model QRHKT provided by the invention obtains a better knowledge tracking result on three common data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge tracing, and specifically relates to a knowledge tracing method and model that integrates heterogeneous graph neural networks and historical performance. Background Art

[0002] The integration of intelligent tutoring systems (ITS) in online education has become increasingly important. As a key task in these systems, knowledge tracing (KT) has become an effective means of applying artificial intelligence in the field of educational technology and has received extensive attention in multiple fields. Essentially, knowledge tracing models use students' historical learning data to construct models of their knowledge mastery status and predict students' future learning performance through these models.

[0003] Although existing knowledge tracing models provide a certain degree of accuracy in diagnosing students' learning status, they still face some limitations. A major challenge lies in the representation of questions and concepts. Current datasets mainly only contain information about the inclusion relationships between questions and concepts, lacking other effective information such as concept hierarchical relationships. In the case of a large number of questions, the interaction times between students and each question may be less, which increases the difficulty of learning effective question representations. Manually annotating the hierarchical information of concepts by experts is both time-consuming and laborious. Previous studies have focused on the implicit relationships that may exist between questions and between concepts, but most studies have concentrated on students' individual question-solving information and have not fully utilized global information such as the response records of all students.

[0004] In addition, learning is a gradual process, and students gradually improve their mastery of concepts through continuous practice. There are sequential dependencies in the question sequence, and past question-solving records will have varying degrees of influence on current learning. Some records may have a greater impact, while other records may have a smaller impact. Modeling these sequential dependencies is crucial for improving model performance. As the number of historical question-solving records increases, the difficulty of modeling dependencies also increases. Some studies have tried to solve the long-term dependencies in the question-solving sequence and addressed the problem of long-term dependencies from the perspective of question similarity without distinguishing the importance of the question-solving sequence. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a knowledge tracing method and model that integrates heterogeneous graph neural networks and historical performance in view of the deficiencies of the prior art.

[0006] To solve the above technical problem, the technical solution of the present invention is as follows:

[0007] A knowledge tracing method that integrates heterogeneous graph neural networks and historical performance, comprising the following steps:

[0008] S1. Construct a heterogeneous global graph that contains problem nodes and concept nodes. The inclusion relationships between problem nodes and concept nodes, the co-occurrence relationships between concept nodes, and the transition relationships between problem nodes form the edges of the graph. Then, use the graph neural network algorithm to encode the nodes in the heterogeneous global graph.

[0009] S2. According to the problem representation obtained after passing through the graph neural network, obtain the student's current knowledge state and historical important knowledge state at the current moment. The historical important knowledge state is the knowledge state of the student after answering questions containing higher-level concepts correctly. Then, adaptively learn the weights between the current knowledge state and the historical important knowledge state to fuse these two knowledge states and form a comprehensive knowledge state representation.

[0010] S3. Concatenate the fused knowledge state with the problem embedding at the next moment and input it into a fully connected network to predict the student's answering performance at the next moment.

[0011] Preferably, in step S1, the problem data for constructing the heterogeneous global graph is collected from an online education platform, and then the concept data included in the problems is obtained. Clean the collected problem and concept data to remove duplicate, incorrect, or incomplete data records and unify the data format. The problem data and concept data after cleaning are used as the problem nodes and concept nodes of the heterogeneous global graph respectively, and both have unique identifiers.

[0012] Preferably, in step S1, for each problem node in the graph, according to the concepts included in the problem, create two types of directed inclusion edges in the graph: one is from the problem node to the corresponding concept node, and the other is from the corresponding concept node to the problem node.

[0013] Preferably, in step S1, according to the historical problem-solving records of all students, count the occurrence patterns of different concepts during the problem-solving process of a large number of students. If they often appear successively or alternately in the problem-solving sequence of the same student, or similar situations frequently occur in the problem-solving sequences of different students, then it is determined that there is a co-occurrence relationship between these concepts. And according to the frequency of this co-occurrence relationship, model the co-occurrence relationship between concepts as a directed co-occurrence edge in the heterogeneous global graph.

[0014] Preferably, according to the frequency of co-occurrence relationships, the co-occurrence relationships between concepts are modeled as directed edges of co-occurrence relationships in a heterogeneous global graph, specifically as follows: First, consecutive concept question sequences are merged to obtain a concept co-occurrence sequence. Then, for the concepts in the concept co-occurrence sequence, their adjacent co-occurring concepts are calculated based on all historical concept co-occurrence sequences, obtaining a collection of concept co-occurrence sequences in which the concepts appear. And k concepts with the highest co-occurrence frequency with the concept are selected from it, and the co-occurrence relationships between the concept and these k concepts are modeled as directed edges of co-occurrence relationships in the heterogeneous global graph.

[0015] Preferably, in step S1, according to the historical question-solving records of all students, the correlation coefficient between question A and question B that appears before question A is calculated. If this correlation coefficient is greater than a preset transfer relationship threshold, the transfer relationship between question A and question B is modeled as a directed edge of the transfer relationship in the heterogeneous global graph.

[0016] Preferably, if question B appears multiple times before question A, the question B closest to question A is used to calculate the correlation.

[0017] Preferably, in step S1, a graph neural network is used to encode the nodes in the heterogeneous global graph, specifically including: 1) Initializing the embedding: Using the IDs of questions and concepts to obtain the initial embedding vectors; 2) Message passing: For each node, according to the features of its neighbor nodes and the relationship types of the edges, update the embedding representation of the node through an aggregation operation; 3) Multi-layer propagation: Through multi-layer GNN propagation, accumulate the embeddings of each layer to form the final representations of questions and concepts.

[0018] Preferably, step S2 specifically includes: 1) Concatenating the question representations obtained through the graph neural network with the answer responses to encode the question-answer pairs; 2) Inputting the question-answer pair embeddings into a long short-term memory network to obtain the student's current knowledge state; 3) Using activation functions and learnable weight parameters to calculate the importance coefficients of each historical knowledge state, and weighted aggregate the student's knowledge states at each historical time step based on these importance coefficients to obtain the historical important knowledge states; 4) Adaptively learn the weights between the current knowledge state and the historical important knowledge state, and fuse these two knowledge states to form a comprehensive knowledge state representation.

[0019] A knowledge tracing model that fuses a heterogeneous graph neural network and historical performance, and this knowledge tracing model is constructed and generated by the above knowledge tracing method.

[0020] The beneficial effects of the present invention are:

[0021] The present invention provides a knowledge tracing method and model that integrates heterogeneous graph neural networks and historical performance. First, by combining the inclusion relationship between questions and concepts, the transfer relationship between questions, and the co-occurrence relationship between concepts, a heterogeneous global graph is jointly constructed, and the question representation is enhanced through graph neural network algorithms. Then, to overcome the problem of long-term dependence, from the perspective of the importance of historical knowledge states, the knowledge state of students is divided into the knowledge state at the current moment and the historically important knowledge state. The knowledge state at the current moment is obtained through a long short-term memory network, and the historically important knowledge state is generated by weighted aggregation based on importance coefficients. An adaptive fusion module is used to fuse the knowledge state at the current moment and the historically important knowledge state. Finally, the next answering performance of students is predicted by combining the knowledge state and the question representation.

[0022] On three public datasets, the comparative experimental results show that the QRHKT proposed by the present invention achieves better prediction performance. The ablation experiment proves the necessity of the heterogeneous global graph question representation module and the three relationships it contains, as well as the historically important state perception module, and their impact on the overall performance; the clustering analysis results show that QRHKT has a positive impact on the representation of question nodes; the visualization results show that QRHKT can model the accurate and reasonable knowledge state for learners. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is an architecture diagram of the knowledge tracing model QRHKT based on a heterogeneous global graph neural network;

[0024] Figure 2 It is a construction diagram of the inclusion relationship between questions and concepts and the corresponding directed edges;

[0025] Figure 3 It is a construction diagram of the co-occurrence relationship between concepts and the corresponding directed edges;

[0026] Figure 4 It is a construction diagram of the transfer relationship between questions and the corresponding directed edges;

[0027] Figure 5 It is the difference graph of the relevance φ i,j value of different questions;

[0028] Figure 6 It is a visual display after clustering of ten concepts and associated questions in the ASSISTchall data after t-SNE dimensionality reduction;

[0029] Figure 7 It is a visual case of the knowledge state of learners tracked by the model QRHKT. DETAILED DESCRIPTION OF THE INVENTION

[0030] To facilitate the understanding of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that the embodiments are only for helping to understand the present invention and should not be regarded as specific limitations on the present invention.

[0031] Aiming at the two limitations of the existing knowledge tracing model, one is the lack of utilization of global information in the representation of questions, and the other is the failure to overcome the long-term dependence on the student's question-solving sequence. The present invention proposes a knowledge tracing method integrating heterogeneous graph neural network and historical performance, including the following steps:

[0032] S1. Construct a heterogeneous global graph, which contains question nodes and concept nodes, and the inclusion relationship between question nodes and concept nodes, the co-occurrence relationship between concept nodes, and the transition relationship between question nodes form the edges of the graph; then use the graph neural network algorithm to encode the nodes in the heterogeneous global graph, and learn the embedded representations of questions and concepts through message passing and aggregation.

[0033] S2. According to the question representation obtained through the graph neural network, obtain the student's current knowledge state and historical important knowledge state, where the historical important knowledge state is the knowledge state of the student after answering questions containing higher-level concepts; then adaptively learn the weights between the current knowledge state and the historical important knowledge state to fuse these two knowledge states and form a comprehensive knowledge state representation.

[0034] S3. Concatenate the fused knowledge state with the question embedding at the next moment and input it into a fully connected network to predict the student's answering performance at the next moment.

[0035] In step S1, the question data for constructing the heterogeneous global graph is collected from an online education platform, and then the concept data included in the questions is obtained; the collected question and concept data are cleaned to remove duplicate, incorrect or incomplete data records and unify the data format; the cleaned question data and concept data are used as the question nodes and concept nodes of the heterogeneous global graph respectively, and both have unique identifiers.

[0036] In step S1, for each question node in the graph, according to the concepts included in the question, create two types of inclusion relationship directed edges in the graph: one is from the question node to the corresponding concept node, and the other is from the corresponding concept node to the question node.

[0037] In step S1, based on the historical problem-solving records of all students, the occurrence patterns of different concepts in the problem-solving processes of a large number of students are statistically analyzed. If they often appear successively or alternately in the problem-solving sequences of the same student, or similar situations frequently occur in the problem-solving sequences of different students, it is determined that there is a co-occurrence relationship between these concepts. And the strength of the correlation between concepts can be measured according to the frequency of this co-occurrence relationship. The co-occurrence relationship between concepts is modeled as a directed edge of the co-occurrence relationship in the heterogeneous global graph.

[0038] According to the frequency of the co-occurrence relationship, the co-occurrence relationship between concepts is modeled as a directed edge of the co-occurrence relationship in the heterogeneous global graph. Specifically: First, the continuous concept problem-solving sequences are merged to obtain the concept co-occurrence sequence S. Then, for the concept c in the concept co-occurrence sequence S i , calculate the concepts that co-occur adjacent to it according to all historical concept co-occurrence sequences, and obtain the set of concept co-occurrence sequences in which the concept c i appears, and select the k concepts with the highest co-occurrence frequency with the concept c i from it. Model the co-occurrence relationship between the concept c i and these k concepts as a directed edge of the co-occurrence relationship in the heterogeneous global graph.

[0039] In step S1, based on the historical problem-solving records of all students, calculate the correlation coefficient between the question q i and the questions q i that appear before the question q j . If this correlation coefficient is greater than the preset transfer relationship threshold, model the transfer relationship between the question q i and the question q j as a directed edge of the transfer relationship in the heterogeneous global graph. And if the question q i appears multiple times before the question q j , then use the question q i closest to the question q j to calculate the correlation.

[0040] In step S1, use a graph neural network to encode the nodes in the heterogeneous global graph, specifically including: 1) Initialize the embedding: Use the IDs of the questions and concepts to obtain the initial embedding vectors; 2) Message passing: For each node, update the embedding representation of the node through an aggregation operation according to the features of its neighbor nodes and the type of edge relationship; 3) Multilayer propagation: Through multilayer GNN propagation, accumulate the embeddings of each layer to form the final representations of the questions and concepts.

[0041] Step S2 specifically includes: 1) Concatenating the question representation obtained through the graph neural network with the answer response to encode the question-answer pair; 2) Embedding the question-answer pair into the input long short-term memory network to obtain the student's current knowledge state; 3) Using an activation function and learnable weight parameters to calculate the importance coefficients of each historical knowledge state, and aggregating the knowledge states of the student at each historical time step weighted based on these importance coefficients to obtain the historical important knowledge state; 4) Adaptively learning the weights between the current knowledge state and the historical important knowledge state, and fusing these two knowledge states to form a comprehensive knowledge state representation.

[0042] The present invention also provides a knowledge tracking model QRHKT that fuses heterogeneous graph neural networks and historical performance, which is constructed and generated by the above knowledge tracking method.

[0043] As Figure 1 shown, the knowledge tracking model QRHKT includes a heterogeneous global graph question representation module, a historical important knowledge state perception module, and a prediction module; the heterogeneous global graph question representation module is used to construct a heterogeneous global graph through the inclusion relationship between questions and concepts, the co-occurrence relationship between concepts, and the transfer relationship between questions, and then enhance the question representation through the graph neural network algorithm; the historical important knowledge state perception module is used to obtain the student's current knowledge state and historical important knowledge state according to the question representation obtained by the heterogeneous global graph question representation module, and then adaptively learn the weights between the current knowledge state and the historical important knowledge state, and fuse these two knowledge states; the prediction module is used to predict the student's answering performance at the next moment according to the fused knowledge state obtained by the adaptive knowledge state fusion module and the question representation at the next moment.

[0044] The task of the knowledge tracking model QRHKT is to process and analyze this sequence of data by establishing a model based on the student's historical question-solving records, and mine potential information from the data for answering prediction at the next moment.

[0045] A student's question-solving record can be expressed as:

[0046] χ = ((q 1 , r 1 ), (q 2 , r 2 ), (q 3 , r 3 ),..., (q T , r T ))

[0047] In the formula, q i represents the question, and r i is a binary variable of 0 or 1, representing the student's answering result, that is, whether the student answered the question q correctlyi (0 means incorrect, 1 means correct).

[0048] Given a question record χ, knowledge tracking needs to predict whether the student can answer q correctly T+1 , that is, calculate the probability:

[0049] P(r T+1 =1|q T+1 ,X)

[0050] The heterogeneous global graph question representation module utilizes the global information of students' question-answering sequences. Based on the sequence information of all students, it calculates the inclusion relationship between questions and concepts, the co-occurrence relationship between concepts, and the transfer relationship between questions. It then constructs a heterogeneous global graph based on these three relationships. It then uses a graph neural network algorithm to propagate and aggregate messages and learn the features of the nodes in the graph.

[0051] Heterogeneous Global Graph It is expressed as:

[0052]

[0053] In the formula, Representing a heterogeneous global graph The title node is included and concept nodes The node set, q i represents the i-th question, c i represents the i-th concept, is the total number of questions involved in the process of answering questions by all students, is the total number of concepts involved in the question; ε represents the heterogeneous global graph The set of edge relationships in , namely, the inclusion relationship between questions and concepts, the co-occurrence relationship between concepts, and the transfer relationship between questions.

[0054] The construction of topic-concept inclusion relationship: Figure 2 As shown, the heterogeneous global graph There are two types of nodes in the graph, namely question nodes and concept nodes. A question contains one or more concepts, and a concept also has multiple questions. Therefore, the relationship between questions and concepts is a bipartite graph structure. In addition, questions and the concepts they contain are closely related, and questions containing the same concept should be closer. Therefore, the relationship between questions and concepts is modeled as a heterogeneous global graph There are two types of inclusion relationship directed edges in (c j ,q i ,r include ) and (q i ,c j ,r include_by ), which indicates that question q i Contains concept cj 。

[0055] Construction of co-occurrence relationships between concepts: When students do exercises, they often do multiple exercises on one concept in a row and then move on to exercises on another concept. As shown in the construction process of the co-occurrence relationship between the concepts and the corresponding directed edges, from the exercise level, the student has done exercises q Figure 3 →q 1 →q 2 →q 3 →q 4 →q 5 →q 6 →q 7 →q 8 →q 9 These 9 exercises; from the concept level, the student has done 1 exercise with the concept of c 1 →1 exercise with the concept of c 2 →1 exercise with the concept of c 3 →2 exercises with the concept of c 2 →3 exercises with the concept of c 3 →1 exercise with the concept of c 4 。Since the learning process is hierarchical, students need to start learning from low-level concepts and gradually master and transition to more advanced concepts. For example, for the exercises corresponding to the two concepts c 2 and c 3 , if there is an overlap in the exercise sequence, it can be considered that such concepts have a strong correlation. Generally speaking, frequent co-occurrence behaviors in different exercise sequences can show strong concept correlations. Therefore, construct the relationships between concepts based on the co-occurrence information in the historical exercise records of all students.

[0056] Specifically, first merge the consecutive concept exercise sequences to obtain the concept co-occurrence sequence S, such as the {c Figure 3 , c 1 , c 2 , c 3 , c 2 , c 2 , c 3 , c 3 , c 3 , c 4} in 1 is merged to obtain {c 2 , c 3 , c 2 , c 3 , c 4}. Next, for the concept c i in the concept co-occurrence sequence S, calculate its adjacent co-occurring concepts based on all historical concept co-occurrence sequences. The concepts c i and c jThe co-occurrence frequency f between c (c j ,c i ) is calculated as follows:

[0057]

[0058] In the formula, N(c j ) is the set of concept co-occurrence sequences in which the concept c i appears.

[0059] In addition, to avoid introducing too much noise, only the k concepts with the highest co-occurrence frequency with the concept c i are finally selected. That is, if f c (c j ,c i ) is the top-k co-occurrence frequency of the concept node c i , then the co-occurrence relationship between the concepts c i and c j is modeled as a directed edge of the co-occurrence relationship in the heterogeneous global graph : (c j ,c i ,r similar ).

[0060] Construction of the transfer relationship between questions: In addition to the relationships between questions and concepts, and between concepts and concepts, there is also a correlation between questions. For example, there are many transfer relationships between questions. When a student completes question q 1 , he is likely to be able to correctly answer question q 2 . To represent this potential relationship between questions, the Phi correlation coefficient is used to measure the correlation between two questions. Specifically, the transfer relationship table between questions as shown in Table 1 is constructed.

[0061] Table 1 Transfer relationship table between questions

[0062]

[0063] Based on the question-solving records of all students, count the number of pairs of (incorrect, incorrect), (incorrect, correct), (correct, incorrect), and (correct, correct) between two questions q j and q i , and in the question sequence, q j appears before q i . If q i appears multiple times before q j , only consider the most recent q j . Finally, calculate its correlation using the following formula:

[0064]

[0065] where φ i,j ranges from [-1, 1], representing how much the performance of a student on question q j affects their performance on question q i . If φ i,j > 0, it indicates a positive correlation between these two questions; otherwise, it indicates a negative correlation.

[0066] Similarly, to avoid introducing too much noise, a transfer frequency threshold θ is set. When φ i,j > θ, the transfer relationship between question q i and q j is modeled as a directed edge of the transfer relationship in the heterogeneous global graph : (q j , q i , r transfer ).

[0067] Question representation aggregation: Use a graph neural network to encode the nodes in the heterogeneous global graph . In the heterogeneous global graph , there are two types of edges connecting neighbors for question nodes, namely r include and r transfer ; there are also two types of edges connecting neighbors for concept nodes, namely r include_by and r similar . In the GNN layer, by aggregating the features of neighbor nodes, messages are passed along the edges to update the nodes. First, obtain the initial embeddings and of questions and concepts using the IDs of questions and concepts, and and represent the node embeddings of questions and concepts after propagation in layer l, where is the embedding vector of question q i , and is the embedding vector of concept c i . Specifically, for question q i , let represent its neighbor nodes with edge type r x , and the aggregation formula is as follows:

[0068]

[0069] where || is the concatenation operation, is the weight matrix, is the bias vector, f(·) is the activation function, and the relu activation function is selected here, is or

[0070] Then, different messages propagated by different types of edges are accumulated, and the topic node representation is updated as shown in the following formula:

[0071]

[0072] In the formula, accum(·) can be operations such as summation, concatenation, etc. Here, the method of taking the mean value is adopted.

[0073] The aggregation operation of the concept node is similar to the topic node aggregation operation described above. The updated concept node representation is as follows:

[0074]

[0075] After propagation in the L-th layer, the embeddings of each layer are combined to form the final representations of topics and concepts:

[0076]

[0077] In the formula, α l ≥0 represents the importance of the output of the l-th layer of the graph neural network. In the present invention, α l is set to 1 / (L + 1).

[0078] Historical important knowledge state perception module: Certain historical knowledge states of students will always have a relatively important impact on subsequent predictions. After answering questions containing higher-level concepts, students can often answer questions with lower-level concepts. In the present invention, the knowledge state after answering questions containing higher-level concepts is regarded as the historical important knowledge state. Based on this idea, an adaptive knowledge state fusion module is designed. Specifically, the knowledge state of students is divided into the knowledge state at the current moment and the historical important knowledge state. The knowledge state at the current moment is obtained through LSTM; the historical important knowledge state is obtained by weighted aggregation of the knowledge states of students at each historical time step according to the importance weights. Then, the knowledge state at the current moment and the historical important knowledge state are adaptively combined and applied to the prediction at the next moment.

[0079] Interactive embedding: After obtaining the representation of the question through a multi-layer graph neural network, in order to obtain the performance of the student's answer, it is necessary to concatenate the question representation and the answer response. In order to emphasize the difference between correct and wrong answers, the following method is used to encode the question-answer pair:

[0080]

[0081] In the formula, is the concatenation operation, x t is the embedding vector of the question q t , r t is the answer result of the student, which is related to xt All-zero or all-one vectors in the same dimension.

[0082] Knowledge state at the current moment: Embed the question-answer pair into the input LSTM and calculate the knowledge state of the student at each time step:

[0083] h t = LSTM(a t , h t-1 )

[0084] Historical important knowledge state: At time t, the knowledge state h at the current moment is obtained t and a series of historical knowledge states {h 1 , h 2 , h 3 ,..., h t-1} before time t. Considering that each historical knowledge state will have different degrees of influence on the current prediction, first calculate the importance coefficient for each historical knowledge state:

[0085] α i = Z T tanh(W T h i + b)

[0086]

[0087] In the formula, tanh is the tanh activation function, W is the weight matrix, b is the bias vector, Z is the learnable weight parameter, and α i is the importance coefficient of the historical knowledge state h i . Next, obtain the historical important knowledge state h' through weighted summation t-1 :

[0088]

[0089] Finally, through the adaptive knowledge state fusion module, adaptively learn the weights between the knowledge state at the current moment and the historical important knowledge state, and determine the amount of information retained by different knowledge states based on these weights:

[0090]

[0091] In the formula, σ is the Sigmoid activation function, ⊙ is the Hadamard product, is the weight matrix, b z1 , b z2 is the bias vector, is the fused knowledge state.

[0092] Prediction module: In this module, the fused knowledge state and the question embedding x at the next moment t+1 They are concatenated and input into a fully-connected network for prediction:

[0093]

[0094] In the formula, is the weight matrix, and b out is the bias vector, is the probability that the predicted student answers the question correctly.

[0095] Finally, the model is trained through the cross-entropy loss function:

[0096]

[0097] Experimental verification:

[0098] This experiment aims to answer the following questions: RQ1: Compared with existing sequential models and graph-based knowledge tracing models, how is the performance prediction effect of the proposed QRHKT of the present invention for learners? RQ2: How do the various modules in QRHKT affect the performance of the model? RQ3: How to determine that the heterogeneous global graph question representation module achieves the optimal effect? RQ4: How does the heterogeneous global graph question representation module realize the modeling of the potential relationship between questions and knowledge points?

[0099] Dataset: The present invention uses three benchmark datasets to evaluate the prediction performance of the QRHKT model, as shown in Table 2.

[0100] Table 2 Datasets

[0101]

[0102]

[0103] The ASSIST2009 dataset comes from the ASSISTments online education platform, which records the data of students using the ASSISTments platform for learning during 2009 - 2010. ASSISTChall is a dataset with the richest descriptive information among all ASSISTments public datasets, collected from the public data mining competition held in 2017. This dataset contains a total of 942,816 learning records and has the longest average learning sequence length. The EdNet-KT1 dataset is a large-scale hierarchical dataset of various student activities collected by Santa. It has collected 131,417,236 interactions of 784,309 students, becoming the largest public IES dataset released so far. In addition, Santa provided a total of 13,169 questions, annotated with 293 concepts, and the present invention uses the KT1 version.

[0104] Experimental Setup: Before the experiment, data records with empty problem ID and kc fields in the dataset were deleted. In terms of model training, 5-fold cross-validation was adopted to find the optimal parameters. For each fold, 20% of the response records were used as the test dataset, 20% as the validation set, and 60% as the training set. The embedding dimensions of both the problem nodes and the concept nodes were 64, the dimension of the hidden vector in the LSTM was 64, the number of layers of the graph neural network was 2, all learnable parameters were optimized by Adam, the Batchsize was set to 64, and the Learning rate was set to 0.005, which decreased with the number of iterations. For student exercise sequences with inconsistent lengths, the length of each input to the sequence model was fixed at 50 exercises. All experiments were conducted on a 64-bit Ubuntu 20.04.5 LTS server equipped with an Intel(R) Xeon(R) Gold 6238R CPU @ 2.30GHz and 32GB NVIDIA A100 GPU. All code implementations were executed using the PyTorch framework.

[0105] Baseline: To evaluate the effectiveness of QRHKT, it was compared with several different baseline methods. To ensure the fairness of the comparison, the same training, validation, and test sets as QRHKT were used. All these methods were adjusted to ensure the best performance was obtained from the implementation.

[0106] Graph-Structure-Free Models: DKT was the first deep learning-based knowledge tracing model, which represented the student's knowledge state through the hidden vector in the RNN; DKT+F incorporated three features, namely the time interval since the student's last learning, the time interval at the same time point of the last learning, and the historical number of times of learning the same concept, on the basis of DKT to model the student's forgetting behavior; DKVMN introduced a dynamic key-value memory network, which stored the static embedding of the concept and the student's mastery level on each concept through a static matrix and a dynamic matrix respectively; EKT considered the relevance of each problem to all concepts and used the semantic information of the exercises as the input; SAKT used a self-attention mechanism to identify the interactions related to the given concept from the student's past exercise interactions and predicted their knowledge mastery based on the performance on the relevant interactions; AKT used a new monotonic attention mechanism to link the performance of the student on the problem to be predicted with the historical exercise performance.

[0107] Graph Structure-based Models: GKT learns the relationships between knowledge concepts and models knowledge tracing as a time series node-level classification problem, and uses GCN for training; GIKT uses the graph convolutional network GCN to aggregate item embeddings and concept embeddings; SGKT uses an association graph to model the relationships between item concepts and uses a gated graph neural network to obtain the student knowledge state from the student's answering process. To ensure the fairness of the experiment, the item attribute features such as difficulty level, type, and average response time were not utilized in the test.

[0108] Table 3 Network Structures Used by the Comparative Models

[0109]

[0110] To answer RQ1, a comparative experiment was conducted, and the comparative experiment results shown in Table 4 were obtained.

[0111] Table 4 Comparative Experiment Results AUC (%)

[0112]

[0113] As shown in Table 4, first of all, QRHKT achieved better results than other baseline models on the three datasets of ASSIST2009, ASSISTchall, and EdNet-KT1. The AUC values were 2.93%, 4.84%, and 9.56% higher than the best-performing models in all three datasets respectively. The results indicate that QRHKT can enhance the prediction performance by learning the item representations and fusing the student historical states. Secondly, ERHKT showed a greater advantage in ASSISTchall because ASSISTchall contains richer practice answering records, from which QRHKT can build a denser heterogeneous graph structure and can better learn the item representations, resulting in better performance. The graph structure-based GKT, GIKT, and SGKT also performed better in ASSISTchall than in ASSIST09 and EDNET, which also shows the effectiveness of the graph structure models in learning item representations. The AUC value of the ERH knowledge tracing model was higher than that of other graph structure-based models, which may be due to the role of the historical important knowledge state perception module using the LSTM structure. Since the model structures are inconsistent, this conclusion needs to be verified in ablation experiments. Finally, it can be seen that the improvement of the model on the ASSIST2009 dataset is not obvious. This is because in the ASSIST2009 dataset, the number of items is large, the ratio of items to concepts is too high, and the complexity of the knowledge structure is higher than that of other datasets.

[0114] To answer RQ2, an ablation experiment was conducted: The QRHKT model learns question representations through the heterogeneous global graph question representation module and fuses the student's historical and current states for prediction through the historical important knowledge state perception module. To verify the effectiveness of the two modules, an ablation experiment was carried out to verify the performance of QRHKT with different components. The results of the ablation experiment are shown in Table 5.

[0115] The heterogeneous global graph question representation module includes three relationships: the inclusion relationship between questions and concepts, the co-occurrence relationship between concepts, and the transfer relationship between questions. To verify the necessity of these three relationships, the experimental group was further divided. The specific experimental groups are as follows: QRHKT-RCR removes the co-occurrence relationship between concepts; QRHKT-RQR removes the transfer relationship between questions; QRHKT-RQC only uses the inclusion relationship between questions and concepts; QRHKT-RGG removes the heterogeneous global graph question representation module; QRHKT-RHP removes the historical important knowledge state perception module and only uses the knowledge state at the current moment for prediction; QRHKT-RGGHP removes both the heterogeneous global graph module and the historical important knowledge state perception module.

[0116] Table 5 Results of the ablation experiment AUC (%)

[0117]

[0118] It can be found from the results that: 1) After removing the heterogeneous global graph title representation module in QRHKT-RGG, its AUC decreased by 5.89%, 3.41%, and 4.49% on ASSIST09, ASSISTchall, and EdNet-KT1 respectively, indicating that the graph structure constructed by the present invention using global information can effectively learn the title representation and significantly improve the model prediction performance; the performance of QRHKT-RHP is higher than that of QRHKT-RGGHP, which can also prove this conclusion. 2) When QRHKT-RCR removes the co-occurrence relationship between concepts, the performance of the model decreases to a certain extent, indicating that the co-occurrence relationship between concepts can improve the model performance. The performance of QRHKT-RQR is higher than that of QRHKT-RQC, which can also prove this conclusion. 3) When QRHKT-RQR removes the transfer relationship between questions, the performance of the model decreases to a certain extent, indicating that the transfer relationship between questions can improve the model performance. The performance of QRHKT-RCR is higher than that of QRHKT-RQC, which can also prove this conclusion. 4) The performance of QRHKT-RQC is higher than that of QRHKT-RGG, which can prove that the inclusion relationship between questions and concepts can improve the model performance. 5) After removing the historical important knowledge state perception module in QRHKT-RHP, its AUC decreased by 3.56%, 5.07%, and 4.83% on ASSIST09, ASSISTchall, and EdNet-KT1 respectively, indicating that the role of adding historical important knowledge state to the model is significant.

[0119] To answer RQ3, a hyperparameter θ analysis was conducted: When constructing the transfer relationship between questions into the heterogeneous global graph, in order to avoid introducing too much noise, a suitable θ value needs to be set to filter out some noisy data. Therefore, multiple groups of experiments were set up to find the appropriate θ. As shown in Table 6, the best θ value for the ASSIST09, ASSISTchall, and EdNet-KT1 datasets is 0.4. This result indicates that when the θ value is small, noisy data will be introduced, affecting the model performance, while when the θ value is large, too little information will be introduced, making it impossible to effectively improve the model performance.

[0120] Table 6 Analysis results of hyperparameter θ

[0121]

[0122] To visually show the role of the hyperparameter θ, taking the ASSISTchall dataset as an example, 8 different exercises were randomly selected for analysis. The numbers of the exercises and the corresponding concepts are shown in Table 7. It can be clearly seen from Figure 5 the different exercise correlations φ i,jThe difference in values indicates that the interactive representation module is effective in representing different exercise nodes after training; when the hyperparameter θ is set to 0.4, for example, the correlation φ i,j value between Exercise 1206 and Exercise 1207 is 0.42. It can be observed in Table 7 that Exercise 1206 and Exercise 1207 correspond to the same concept of perimeter, indicating that when θ takes the value of 0.4, the heterogeneous graph neural network discovers the deep relationship between nodes and concepts.

[0123] Table 7 Numbers of 8 questions selected from ASSISTchall and the corresponding concepts

[0124]

[0125] To answer RQ4, clustering analysis is carried out: mainly to verify whether the heterogeneous global graph question representation module has a positive impact on the learning of question representation. Since GIKT and SGKT, like the model of the present invention, are knowledge tracing models based on heterogeneous graphs and all three discuss question representation, GIKT and SGKT are selected as comparison models for the clustering experiment. In the experiment, the concept numbers are used as the labels of question nodes, the K-Means algorithm is used to cluster the question nodes, the number of clusters is set to the number of concept nodes in each dataset, and the Normalized Mutual Information (NMI) and Adjusted Rand Index (ARI) are used as clustering evaluation indicators. To avoid the contingency of experimental results, 10 repeated experiments are carried out and the experimental mean is calculated. Table 8 shows the experimental results of clustering.

[0126]

[0127] Where MI is the mutual information value between variables and H is the entropy value of variables.

[0128] The Rand Index RI (RandIndex) is:

[0129]

[0130] The Adjusted Rand Index ARI is:

[0131]

[0132] Table 8 Clustering results of question nodes

[0133]

[0134] As can be seen from Table 8, QRHKT has better performance in clustering tasks. Its normalized mutual information (NMI) is 0.38% - 6.21% higher than that of other models, and its adjusted rand index (ARI) is 0.8% - 14.94% higher than that of other models. Compared with other baseline models, after considering the association between students' questions and the ability coefficients of students' answers, the QRHKT model has better discrimination ability of question nodes in terms of concepts. Therefore, adding the association between students and questions and students' answering ability has a positive impact on the classification and grading of questions.

[0135] To more clearly show the clustering effect of question nodes, ten concepts in the ASSISTchall dataset and the questions associated with these concepts are selected for visual display, as Figure 6 shown. The feature vectors of question nodes learned by the three models (a) GIKT, (b) SGKT, and (c) QRHKT are respectively reduced in dimension by the t-SNE algorithm and projected into a two-dimensional space for visual comparison. For each question node representation vector, different colors are used according to the concept number it belongs to. As can be seen from Figure 6 it, compared with GIKT and SGKT, QRHKT can better map questions of different concepts to different regions, and the boundaries between questions of different concept categories are clearer. In addition, in the region where questions of the same concept are clustered, the overlap degree of question nodes is not high, and the discrimination degree between questions is better retained on the basis of clustering.

[0136] Visualization: To achieve the goal of personalized learner modeling, we studied the effectiveness of QRHKT in tracking knowledge states in terms of accuracy and rationality. Figure 7 shows a visual case of tracking knowledge states from the same learning sequence, which is a student interaction segment obtained from the ASSISTChall dataset. From Figure 7 it, some important findings are obtained, which can help build a personalized profile of the learner.

[0137] Figure 7 is a visual case of the learner's knowledge state tracked by QRHKT, using the learning sequence of the ASSISTChall dataset, in which the learner answered 20 questions about 6 concepts. In (a), the concepts included in each question and the student's answering results are shown at the top of the heat map, and the 6 concepts from A to F are shown on the left. In (b), there are 4 radar charts, which are the knowledge states of the student in 6 concepts after answering at the 5th, 10th, 15th, and 20th moments respectively.

[0138] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A knowledge tracking method integrating heterogeneous graph neural networks and historical representations, characterized in that: The following steps are involved: S1, construct a heterogeneous global graph, which contains topic nodes and concept nodes, and the inclusion relationship between topic nodes and concept nodes, the co-occurrence relationship between concept nodes, and the transfer relationship between topic nodes constitute the edges of the graph; then use the graph neural network algorithm to encode the nodes in the heterogeneous global graph; S2, based on the question representation obtained through the graph neural network, obtain the student's current knowledge state and historical important knowledge state, where the historical important knowledge state is the student's knowledge state after answering questions containing higher-level concepts; then adaptively learn the weight between the current knowledge state and the historical important knowledge state to merge the two knowledge states and form a comprehensive knowledge state representation; S3, concatenates the fused knowledge state and the question embedding of the next moment and inputs it into the fully connected network to predict the student’s answering performance at the next moment.

2. The knowledge tracking method integrating heterogeneous graph neural network and historical representation according to claim 1 is characterized in that: In step S1, the topic data for constructing the heterogeneous global graph is collected from the online education platform, and then the concept data contained in the topic is obtained; the collected topic and concept data are cleaned to remove duplicate, erroneous or incomplete data records, and the data format is unified; the cleaned topic data and concept data are respectively used as topic nodes and concept nodes of the heterogeneous global graph, and both have unique identifiers.

3. The knowledge tracking method integrating heterogeneous graph neural network and historical representation according to claim 1 is characterized in that: In step S1, for each topic node in the graph, two types of directed edges of inclusion relationships are created in the graph according to the concepts contained in the topic: one is from the topic node to the corresponding concept node, and the other is from the corresponding concept node to the topic node.

4. The knowledge tracking method integrating heterogeneous graph neural network and historical representation according to claim 1 is characterized in that: In step S1, based on the historical problem-solving records of all students, the occurrence patterns of different concepts in the problem-solving process of students are counted. If they often appear successively or alternately in the problem-solving sequence of the same student, or similar situations also frequently appear in the problem-solving sequences of different students, then it is determined that there is a co-occurrence relationship between these concepts, and based on the frequency of this co-occurrence relationship, the co-occurrence relationship between concepts is modeled as a co-occurrence relationship directed edge in a heterogeneous global graph.

5. The knowledge tracking method integrating heterogeneous graph neural network and historical representation according to claim 4 is characterized in that: According to the frequency of co-occurrence relationship, the co-occurrence relationship between concepts is modeled as a co-occurrence directed edge in the heterogeneous global graph. Specifically, the continuous concept question sequences are first merged to obtain the concept co-occurrence sequence. Then, for the concepts in the concept co-occurrence sequence, the adjacent co-occurring concepts are calculated according to all historical concept co-occurrence sequences to obtain the collection of concept co-occurrence sequences in which the concepts appear. The k concepts with the highest co-occurrence frequency with the concepts are selected from them, and the co-occurrence relationship between the concept and these k concepts is modeled as a co-occurrence directed edge in the heterogeneous global graph.

6. The knowledge tracking method integrating heterogeneous graph neural network and historical representation according to claim 1, characterized in that: In step S1, the correlation coefficient between question A and question B that appears before question A is calculated based on the historical question-answering records of all students. If the correlation coefficient is greater than a preset transfer relationship threshold, the transfer relationship between question A and question B is modeled as a transfer relationship directed edge in a heterogeneous global graph.

7. The knowledge tracking method integrating heterogeneous graph neural network and historical representation according to claim 6 is characterized in that: If question B appears multiple times before question A, the relevance is calculated using question B that is closest to question A.

8. The knowledge tracking method integrating heterogeneous graph neural network and historical representation according to claim 1, characterized in that: In step S1, a graph neural network is used to encode nodes in a heterogeneous global graph, specifically including: 1) initializing embedding: using the IDs of the topic and concept to obtain the initial embedding vector; 2) message passing: for each node, the embedding representation of the node is updated through aggregation operations according to the characteristics of its neighboring nodes and the relationship type of the edges; 3) multi-layer propagation: through multi-layer GNN propagation, the embedding of each layer is accumulated to form the final representation of the topic and concept.

9. The knowledge tracking method integrating heterogeneous graph neural network and historical representation according to claim 1, characterized in that: The step S2 specifically includes: 1) concatenating the question representation obtained through the graph neural network with the answer response to encode the question-answer pair; 2) embedding the question-answer pair into the long short-term memory network to obtain the student's current knowledge state; 3) using activation functions and learnable weight parameters to calculate the importance coefficient of each historical knowledge state, and based on these importance coefficients, weightedly aggregate the student's knowledge state at each historical time step to obtain historical important knowledge states; 4) adaptively learning the weight between the current knowledge state and the historical important knowledge state, and fusing the two knowledge states to form a comprehensive knowledge state representation.

10. A knowledge tracking model integrating heterogeneous graph neural networks and historical representations, characterized in that: The knowledge tracking model is constructed and generated by the knowledge tracking method as described in any one of claims 1-9.