Knowledge tracking method based on question representation and student answering ability interaction

By constructing a knowledge point-question heterogeneous graph and deep residual cross-border network, combined with a time convolutional network, the shortcomings of existing knowledge tracking methods in multi-knowledge point processing and prediction accuracy are solved, and more efficient student knowledge state modeling and prediction are achieved.

CN120124722APending Publication Date: 2025-06-10SHANGHAI UNIV OF ENG SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510109678.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing knowledge tracking methods are difficult to effectively deal with multi-knowledge points, have low prediction accuracy, and have failed to make full use of the intrinsic relationship between the questions and knowledge points.

Method used

Construct knowledge points-the heterogeneous graph of the questions, use graph neural network to learn the interaction relationship between nodes, extract node characteristics as the characterization of the question; combine students' answering ability characteristics and the characterization of the question, capture higher-order feature interactions through deep residual cross networks; use time convolution network to track students' cognitive status and predict the probability of correct answers.

Benefits of technology

It improves the comprehensiveness of the question characterization and the dynamic modeling ability of students' knowledge state, and significantly improves the accuracy of predicting changes in students' knowledge state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124722A_ABST
    Figure CN120124722A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge tracking method based on question representation and student answering ability feature interaction, and belongs to the technical field of knowledge tracking. Comprising the steps of obtaining a historical answer record data set; according to the answering result, using a sliding window to calculate the answering accuracy, and combining the number of times of answering attempts to construct the answering ability characteristics of the student; taking the knowledge points and the questions as nodes of a heterogeneous graph, and constructing a question-knowledge point heterogeneous graph; generating topic representation from the topic-knowledge point heterogeneous graph by using a heterogeneous graph neural network; interactive fusion is carried out on the answering ability features and the question representations of the students through a deep residual network to generate high-order features, and the high-order features are input into a time convolutional network to obtain time sequence changes of knowledge states; and constructing a prediction network, and predicting the future answering accuracy through the prediction network based on the knowledge state and the high-order features. According to the invention, the answering ability of the student is combined with question characterization, so that the understanding of the relationship between the student and the question is improved, and the prediction precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge tracing, and particularly relates to a knowledge tracing method based on the interaction between question representation and student answering ability characteristics. Background Art

[0002] Knowledge Tracing is a data-driven learning subject modeling technology, aiming to establish a tracing model based on the historical behavior data of learners to achieve dynamic assessment of their knowledge states and intelligent prediction of learning performance. In recent years, Knowledge Tracing has attracted the attention of researchers at home and abroad, and they have proposed many related methods. These methods are roughly divided into two categories, methods based on traditional machine learning and methods based on deep learning. Among the methods based on traditional machine learning, Bayesian Knowledge Tracing (BKT) based on the Hidden Markov Model is one of the most representative methods. It models learning sequences based on Bayesian networks to evaluate students' knowledge levels. However, BKT is difficult to handle the problem of multiple knowledge points and has the problem of low prediction accuracy. The Deep Knowledge Tracing model (DKT) first applies the Recurrent Neural Network (RNN) to knowledge tracing. DKT is based on the sequence modeling of RNN and uses a large number of neurons to represent potential knowledge states. Compared with the methods based on traditional machine learning, DKT has greatly improved in prediction performance. However, there are problems such as gradient explosion and lack of learning process characteristics, and it only considers the historical answering sequences of students without making full use of the internal relationship between questions and knowledge points. The Pre-trained Question Embeddings based Knowledge Tracing model (PEBG) improves knowledge tracing by constructing a bipartite graph of question-knowledge point relationships and learning pre-trained embeddings of questions. The Dual Graph Ensemble Learning Method for Knowledge Tracing (DGEKT) establishes a bipartite graph structure of student learning interactions to capture heterogeneous question-knowledge point associations and interaction transfers. These methods use graph structures to reveal the dependence relationships between questions, but they fail to comprehensively and comprehensively consider the question difficulty and the connections between questions; and many knowledge tracing research methods mainly measure students' knowledge states from the overall learning cycle without dynamically tracking and updating students' knowledge mastery. Summary of the Invention

[0003] Aiming at the defects existing in the prior art, the present invention constructs a knowledge point-question heterogeneous graph with the title and knowledge points as vertices, uses a graph neural network to learn the interaction relationship between nodes, extracts node features as question representations; then takes the features representing the student's answering ability and the question representations as the inputs of a deep residual cross network to capture high-order feature interactions, and finally uses a temporal convolutional network to track the student's cognitive state and predict the probability that the student will correctly answer the next question, achieving a better prediction effect compared with the current mainstream knowledge tracking methods.

[0004] To achieve the above object, the present invention provides a knowledge tracking method based on the interaction between question representations and student answering ability features, including the following steps:

[0005] (1) Obtain a historical answering record data set and perform data cleaning, extract the required feature set, including: students, questions, knowledge points, answering results, and the number of answering attempts; and generate a time series of answering records according to the students.

[0006] (2) According to the answering results in the time series of the student's answering records, use a sliding window to calculate the student's answering accuracy rate within a period of time, and combine the number of attempts of the student on the corresponding questions to construct the student's answering ability features.

[0007] (3) Assign unique indexes to the questions and knowledge points respectively, and construct a question-knowledge point correspondence matrix, a question similarity matrix, and a knowledge point similarity matrix.

[0008] (4) Take the questions and knowledge points as the nodes of the heterogeneous graph respectively, and construct the edge connection relationship between the questions and knowledge points through the question-knowledge point correspondence matrix, the question similarity matrix, and the knowledge point similarity matrix to form a question-knowledge point heterogeneous graph.

[0009] (5) Use the heterogeneous graph neural network HGNN to learn the interaction relationship between nodes from the question-knowledge point heterogeneous graph and generate question representations.

[0010] (6) Through a deep residual network, perform interactive fusion on the student's answering ability features and the question representations to generate high-order features.

[0011] (7) Input the high-order features into the temporal convolutional network TCN for modeling to obtain the temporal changes of the knowledge state.

[0012] (8) Construct a prediction network, and predict the future answering accuracy rate based on the knowledge state and high-order features through the prediction network.

[0013] Further, the step (1) is specifically:

[0014] (1.1) Clean the data in the historical answer record dataset, filter out the missing data in the historical answer record dataset, and extract the required features:

[0015]

[0016] Among them: u represents the student, q represents the question, s represents the knowledge point, r represents the answer result, and a represents the number of answer attempts;

[0017] (1.2) For the extracted required features, count the answer records according to the students, generate the time series of the students' answer records, and filter out the students' answer records with fewer interaction records; the answer record is expressed as:

[0018] X = {x 1 , x 2 , …, x t}

[0019] x t = (q t , s t , r t, a t )

[0020] Among them: q t represents the question answered at time t, s t represents the knowledge point to which q t belongs, r t represents the answer result at time t, and a t represents the number of answer attempts by the student on question q t .

[0021] Further, the specific steps for each student in step (2) are:

[0022] (2.1) Use the sliding window algorithm to calculate the answer correct rate of the student at time t:

[0023]

[0024] Among them: W is the window size, and r i represents the answer result at time i;

[0025] (2.2) Obtain the answer ability score AbilityScore(t) of the student at time t according to the answer correct rate and the number of answer attempts of the question at time t:

[0026]

[0027] Among them: a t is the number of answer attempts by the student on the corresponding question at time t;

[0028] (2.3) Aggregate the answering ability scores of the student at all times to form the knowledge state feature vector A of the student, which serves as the answering ability feature of the student.

[0029] Further, the specific steps of step (3) are as follows:

[0030] (3.1) Assign a unique index to each of the said questions and knowledge points, and represent the relationship between the questions and knowledge points by constructing a binary matrix P to obtain the question-knowledge point correspondence matrix:

[0031]

[0032] Where: i represents the question index, and j represents the knowledge point index;

[0033] (3.2) Use the intersection over union formula to calculate the similarity between questions and the similarity between knowledge points;

[0034]

[0035] Where: A q is the similarity matrix between questions, represents the set of knowledge points corresponding to the question q with index i i , represents the set of knowledge points corresponding to the question q with index j j ; A s is the similarity matrix between knowledge points, represents the set of questions included in the knowledge point s with index i i , represents the set of questions included in the knowledge point s with index j j .

[0036] Further, the specific steps of step (4) are as follows:

[0037] (4.1) Calculate the question vertex feature matrix as the question nodes in the heterogeneous graph Where: N q represents the number of questions, X q contains 3 feature dimensions X q [i] = [C q [i], T q [i], R q [i]], C q [i] represents the number of correct answers to the question with index i, T q [i] represents the total number of answers to the question with index i, R q [i] represents the average correct rate of the question with index i;

[0038]

[0039] T q [i] > 0

[0040] Among them: T represents the complete time period for all students to answer questions, q t represents the question answered at time t, r t represents the answer result at time t, q i represents the question with index i;

[0041] (4.2) Calculate the knowledge point vertex feature matrix as the knowledge point nodes in the heterogeneous graph Among them: N s represents the number of knowledge points, X s includes 3 feature dimensions X s [j] = [C s [j], T s [j], R s [j]], C s [j] represents the number of correct answers for the knowledge point with index j, T s [j] represents the total number of times the knowledge point with index j is answered, R s [j] represents the average correct rate of the knowledge point with index j;

[0042]

[0043] T s [j] > 0

[0044] Among them: T represents the complete time period for all students to answer questions, r t represents the answer result at time t, s t represents the knowledge point corresponding to the question answered at time step t, s j represents the knowledge point with index j;

[0045] (4.3) Extract the associated edge E from the question-knowledge point relationship matrix q-s , if the question q with index i i involves the knowledge point s with index j j , then set it to 1, otherwise set it to 0;

[0046] (4.4) Extract the edge set E from the question similarity matrix q-q , and the edge weight is determined by the similarity in the question similarity matrix;

[0047] (4.5) Extract the edge set E from the knowledge point similarity matrix s-s , and the edge weight is determined by the similarity in the knowledge point similarity matrix.

[0048] Furthermore, the heterogeneous graph neural network HGNN consists of an encoder and a decoder. The encoder processes the question-knowledge point heterogeneous graph data through multi-layer convolution operations to learn node embeddings; the decoder maps the node embeddings output by the encoder to reconstruct the weights of the edges in the heterogeneous graph. Specifically:

[0049] (5.1) For node type v, the initial input is:

[0050]

[0051] Extract node features through multi-layer convolution. The formula for message passing between nodes in each layer is:

[0052]

[0053] Among them: represents the feature vector of node v at the l-th layer, R is the set of edge types in the graph, N r (v) represents the set of nodes connected to node v through edge type r, is the weight matrix of edge type r at the l-th layer, and σ is the activation function;

[0054] For each edge type r, the formula for the convolution operation is:

[0055] Conv r (h v ,h u )=α uv ·(W·h u )

[0056]

[0057] Among them: α uv is the attention coefficient;

[0058] After L layers of convolution, the final embedding representation of the node is obtained:

[0059] (5.2) The decoder reconstructs the edge weights based on the node embedding Z v :

[0060]

[0061] Among them: is the reconstructed edge weight, z i and z i are the embedding vectors of nodes i and j respectively; W is the linear transformation matrix, and σ is the activation function;

[0062] (5.3) Use the sum of the positive and negative sample loss functions to train the heterogeneous graph neural network and minimize the reconstruction error;

[0063]

[0064] Among them: ε represents the set of actually existing edges, ε neg represents the set of edges generated by negative sampling;

[0065] (5.4) Update the parameters of the heterogeneous graph neural network by backpropagating the gradient according to the loss function L, and use the trained heterogeneous graph neural network to generate the final node embedding as the question representation.

[0066] Furthermore, the specific content of the step (6) is as follows:

[0067] (6.1) Combine the question representation with the answering ability feature to form a unified input feature vector, which is used as the input of the first layer of the deep residual cross network:

[0068] X = Concat(Z q , A)

[0069] Among them: Z q represents the question representation, and A represents the answering ability feature;

[0070] (6.2) For the input feature of the l-th layer, perform feature mapping through a fully connected layer;

[0071]

[0072] Among them: is the weight matrix, is the bias term, is the feature after being mapped by the fully connected layer;

[0073] (6.3) Use the multi-head attention mechanism to capture the correlation between features;

[0074]

[0075] Among them: Q, K, V are the query, key, and value matrices, and d k is the feature dimension scaling factor;

[0076] (6.4) Fuse the attention mechanism and the input feature through a residual connection, and perform layer normalization and activation function;

[0077]

[0078] Among them: is the output of the attention mechanism;

[0079] (6.5) After being processed by L layers, output the high-order feature X DRCN as the cognitive fusion feature X that can represent the relationship between the student's answering ability and the questionDRCN = X (L) 。

[0080] Further, step (7) is specifically as follows:

[0081] (7.1) Concatenate the high-order features with the student's answer results as the input of the Temporal Convolutional Network (TCN):

[0082] X TCN = Concat(X DRCN , Y)

[0083] where: X DRCN is the high-order feature, Y is the student's answer result, and N TCN is the input of the first layer of the dilated convolutional layer of the TCN;

[0084] (7.2) Input the concatenated features into the dilated convolutional layer of the TCN;

[0085]

[0086] where: k represents the size of the convolutional kernel, d represents the dilation coefficient and increases layer by layer as the network depth changes, and W convl [i] is the weight of the convolutional kernel of the l-th layer; represents the output of the (l - 1)-th layer and is also the input of the l-th layer;

[0087] (7.3) Process using residual connection and activation function as the output of the l-th layer:

[0088]

[0089] (7.4) After L layers of convolutional processing, take as the final knowledge state representation output by the TCN

[0090] Further, the prediction network consists of a fully connected layer, an output layer, and an activation function; specifically:

[0091] (8.1) Concatenate the knowledge state representation with the high-order features as the input of the prediction network;

[0092] X input = [Z tcn , X DRCN

[0093] where: Z tcn represents the knowledge state representation, and X DRCN represents the high-order feature;

[0094] ​(8.2) The input of the prediction network generates hidden layer features through a fully connected layer, and is processed through an output layer and a Sigmoid activation function, and the result is mapped to the interval [0, 1] as the predicted value output;

[0095] H hidden = σ (W fc X input + b fc )

[0096]

[0097] where: W fc is the weight matrix of the fully connected layer, b fc is the bias vector of the fully connected layer, σ is the activation function, W out is the weight matrix of the output layer, b out is the bias vector of the output layer, is the predicted correct rate of the student's answer;

[0098] (8.3) Use the cross-entropy loss function to train the prediction network to minimize the difference between the predicted value and the correct answer a t+1 :

[0099]

[0100] where: P t+1 is the probability that the student answers correctly at the next time step, that is, the predicted value

[0101] Furthermore, the predicted value is converted into a binary value through a condition. If is less than 0.5, it is considered that the student's answer result is wrong. If is greater than or equal to 0.5, it is considered that the student's answer result is correct.

[0102] Advantages of the present invention:

[0103] The present invention uses a heterogeneous graph neural network to learn question representations, comprehensively considering the difficulty of questions, the coverage breadth of knowledge points, the question-knowledge point correspondence, the similarity between questions, etc., thus improving the comprehensiveness of question representations. The sliding window technology is used to dynamically calculate the answering ability of students and combine it with the question representation, and input it into the deep residual cross network. It can effectively distinguish the changes in the knowledge states of different students, and through high-order feature interactions, improve the understanding of the relationship between students and questions, thereby improving the prediction accuracy. On the basis of the above-mentioned feature interactions, a temporal convolutional network is used to track the knowledge states of students. The TCN can capture long-term and short-term dependencies in the time series while avoiding the problem of gradient explosion. Brief Description of the Drawings

[0104] Figure 1 This is a framework diagram of the knowledge tracing method based on the interaction between question representation and student answering ability characteristics in an embodiment of the present invention.

[0105] Figure 2 This is a structural schematic diagram of the deep residual cross network in an embodiment of the present invention.

[0106] Figure 3 This is a structural diagram of the temporal convolutional network in an embodiment of the present invention. Detailed Description of the Invention

[0107] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0108] As Figure 1 shown, the present invention provides a knowledge tracing method based on the interaction between question representation and student answering ability characteristics, including the following steps:

[0109] S101. Obtain the historical answering record dataset and perform data cleaning, extract the required feature set, including: students, questions, knowledge points, answering results, and answering attempt times; and generate the time series of answering records according to the students.

[0110] The datasets used in the present invention are the ASSISTment2017 and Junyi Academy2015 datasets.

[0111] (1) Clean the data of the historical answering record dataset, filter the missing data in the historical answering record dataset, and extract the required features:

[0112]

[0113] Among them: u represents students, q represents questions, s represents knowledge points, r represents answering results, and a represents answering attempt times.

[0114] (2) For the extracted required features, count the answering records according to the students, generate the time series of the students' answering records, and filter out the students' answering records with less interaction records; the answering records are represented as:

[0115] X = {x 1 , x 2 , …, x t}

[0116] x t = (q t , s t , r t, a t )

[0117] Among them: qt Denote the question answered at time t as s t Denote q t as the knowledge point to which it belongs, r t Denote the answering result at time t as a t Denote that the student is on question q t The number of answering attempts on the question.

[0118] S102. According to the answering results in the time series of the student's answering records, use a sliding window to calculate the answering accuracy rate of the student within a period of time, and combine the number of attempts of the student on the corresponding questions to construct the answering ability characteristics of the student.

[0119] (1) Use the sliding window algorithm to calculate the answering accuracy rate of the student at time t:

[0120]

[0121] where: W is the window size, r i Denotes the answering result at time i;

[0122] (2) Obtain the answering ability score AbilityScore(t) of the student at time t according to the answering accuracy rate and the number of answering attempts on the question at time t:

[0123]

[0124] where: a t Is the number of attempts of the student on the corresponding question at time t.

[0125] (3) Aggregate the answering ability scores of the student at all times to form the knowledge state feature vector A of the student, which is used as the answering ability characteristic of the student.

[0126] S103. Assign unique indexes to the questions and knowledge points respectively, and construct a question-knowledge point correspondence matrix, a question similarity matrix, and a knowledge point similarity matrix.

[0127] (1) Assign unique indexes to each question and knowledge point respectively, and obtain the question-knowledge point correspondence matrix by constructing a binary matrix P to represent the relationship between questions and knowledge points.

[0128]

[0129] where: i represents the question index, j represents the knowledge point index, q i Is the question with index i, s j Is the knowledge point with index j.

[0130] (2) Use the intersection-over-union formula to calculate the similarity between questions and the similarity between knowledge points.

[0131]

[0132] Among them: A q is the similarity matrix between questions, indicating the set of knowledge points corresponding to the question q with index i, i and indicating the set of knowledge points corresponding to the question q with index j; A j is the similarity matrix between knowledge points, s indicating the set of questions included in the knowledge point s with index i, and i indicating the set of questions included in the knowledge point s with index j. j

[0133]

[0133] S104. Use the knowledge points and questions as nodes of a heterogeneous graph respectively, and construct the edge connection relationship between knowledge points and questions through the question-knowledge point correspondence matrix, the question similarity matrix, and the knowledge point similarity matrix to form a question-knowledge point heterogeneous graph.

[0134] (1) Calculate the question vertex feature matrix as the question nodes in the heterogeneous graph Among them: N q represents the number of questions, X q contains 3 feature dimensions X q [i] = [C q [i], T q [i], R q [i]], C q [i] represents the number of correct answers to the question with index i, T q [i] represents the total number of times the question with index i has been answered, and R q [i] represents the average correct rate of the question with index i.

[0135]

[0136] T q [i] > 0

[0137] Among them: T represents the complete time period of all students' answers, q t represents the question answered at time t, r t represents the answer result at time t, and q i represents the question with index i.

[0138] (2) Calculate the knowledge point vertex feature matrix as the knowledge point nodes in the heterogeneous graph Among them: N sIndicates the number of knowledge points, N s Includes 3 feature dimensions X s [j] = [C s [j], T s [j], R s [j]], C s [j] represents the number of correct answers for the knowledge point with index j, T s [j] represents the total number of responses for the knowledge point with index j, R s [j] represents the average correct rate for the knowledge point with index j.

[0139]

[0140] T s [j] > 0

[0141] Where: T represents the complete time period of all students' responses, r t represents the response result at time t, s t represents the knowledge point corresponding to the question answered at time step t, s j represents the knowledge point with index j.

[0142] (3) Extract the associated edges E from the question-knowledge point relationship matrix q-s , if the question q with index i i involves the knowledge point s with index j j , then set it to 1, otherwise set it to 0.

[0143] (4) Extract the edge set E from the question similarity matrix q-q , and the edge weights are determined by the similarity in the question similarity matrix.

[0144] (5) Extract the edge set E from the knowledge point similarity matrix s-s , and the edge weights are determined by the similarity in the knowledge point similarity matrix.

[0145] S105. Use the heterogeneous graph neural network HGNN to learn the complex interaction relationships between nodes from the question-knowledge point heterogeneous graph and generate question representations.

[0146] The heterogeneous graph neural network HGNN includes an encoder and a decoder. The encoder processes the question-knowledge point heterogeneous graph data through multiple convolutional operations to learn node embeddings; the decoder maps the node embeddings output by the encoder to reconstruct the weights of the edges in the heterogeneous graph.

[0147] Specifically:

[0148] (1) For the node type v, the initial input of the encoder is:

[0149]

[0150] Extract node features through multi-layer convolution. The message passing formula between nodes in each layer is as follows:

[0151]

[0152] Where: represents the feature vector of node v in the l-th layer, R is the set of edge types in the graph, and N r (v) represents the set of nodes connected to node v through edge type r, is the weight matrix of edge type r in the l-th layer, and σ is the activation function.

[0153] For each edge type r, the formula for the convolution operation is:

[0154] Conv r (h v ,h u ) = α uv ·(W·h u )

[0155]

[0156] Where: α uv is the attention coefficient.

[0157] After L layers of convolution, the embedded representation of the final node is obtained:

[0158]

[0159] (2) The decoder reconstructs the edge weights based on the node embedding Z v :

[0160]

[0161] Where: is the edge reconstruction weight, z i and z j are the embedding vectors of nodes i and j respectively; W is the linear transformation matrix, and σ is the activation function.

[0162] (3) Train the network using the loss function, optimize the model through the positive and negative sample loss functions, and minimize the reconstruction error. For the set of actually existing edges ε and the set of edges ε neg generated by negative sampling, calculate the positive example loss and negative example loss respectively, and finally combine the positive and negative sample losses to obtain the total loss function. The calculation formula is as follows:

[0163]

[0164] (4) Update the model parameters by backpropagating the gradient according to the loss function L, and use the trained heterogeneous graph neural network to generate the final node embedding as the question representation.

[0165] S106. Interact and fuse the answering ability characteristics of the student and the question representation through a deep residual network to generate high-order features.

[0166] (1) Combine the question representation and the answering ability characteristics to form a unified input feature vector, which is used as the input of the first layer of the deep residual cross network:

[0167] X = Concat(Z q , A)

[0168] where: Z q represents the question representation, and A represents the answering ability characteristics.

[0169] (2) For the input features of the l-th layer, perform feature mapping through a fully connected layer.

[0170]

[0171] where: is the weight matrix, is the bias term, is the feature after being mapped by the fully connected layer.

[0172] (3) Use the multi-head attention mechanism to capture the correlation between features.

[0173]

[0174] where: Q, K, V are the query, key, and value matrices, and d k is the feature dimension scaling factor.

[0175] (4) Fuse the attention mechanism and the input features through a residual connection, and perform layer normalization and activation function.

[0176]

[0177] where: is the output of the attention mechanism.

[0178] (5) After L layers of processing, output the high-order feature X DRCN as the cognitive fusion feature X that can represent the relationship between the student's answering ability and the question DRCN = X (L) .

[0179] S107. Input the high-order features into the temporal convolutional network TCN for modeling the temporal changes of the knowledge state.

[0180] (1) Concatenate the high-order features with the student's answer results and use them as the input of the Temporal Convolutional Network (TCN):

[0181] X TCN = Concat(X DRCN , Y)

[0182] where: X DRCN is the high-order feature, Y is the student's answer result, and X TCN is the input of the first layer of the dilated convolutional layer of the TCN.

[0183] (2) Input the concatenated features into the dilated convolutional layer of the TCN.

[0184]

[0185] where: k represents the size of the convolutional kernel, d represents the dilation coefficient and increases layer by layer as the network depth changes, W convl [i] is the weight of the convolutional kernel of the l-th layer, represents the output of the (l - 1)-th layer and is also the input of the l-th layer.

[0186] (3) Process using residual connection and activation function as the output of the l-th layer:

[0187]

[0188] (4) After L layers of convolutional processing, is used as the final knowledge state representation of the TCN

[0189] S108. Construct a prediction network to predict the future answer correct rate based on the knowledge state and high-order features through the prediction network.

[0190] The prediction network consists of a fully connected layer, an output layer, and an activation function. Specifically:

[0191] (1) Concatenate the knowledge state representation with the high-order features as the input of the prediction network.

[0192] X input = [Z tcn , X DRCN

[0193] where: Z tcn represents the knowledge state representation, and X DRCN represents the high-order feature.

[0194] (2) The input of the prediction network generates hidden layer features through the fully connected layer and is processed through the output layer and the Sigmoid activation function, and the result is mapped to the [0, 1] interval as the predicted value output.​

[0195] H hidden = σ(W fc X input + b fc )

[0196]

[0197] where: W fc is the weight matrix of the fully connected layer, b fc is the bias vector of the fully connected layer, σ is the activation function, W out is the weight matrix of the output layer, b out is the bias vector of the output layer, is the predicted correct rate of the student's answer.

[0198] (3) Train the prediction network using the cross - entropy loss function to minimize the difference between the predicted value t+1 and the correct answer:

[0199]

[0200] where: P t+1 is the probability that the student answers correctly at the next time step, that is, the predicted value

[0201] Predict the correct rate of answering questions using the trained prediction network, and convert the predicted value into a binary value through a condition. If is less than 0.5, it is considered that the student's answer result is wrong. If is greater than or equal to 0.5, it is considered that the student's answer result is correct.

[0202] In the embodiments of the present invention, AUC (Area under ROC curve), ACC (Accuracy), and F1 score are used as evaluation metrics to evaluate the performance of the model. AUC refers to the local area under the ROC (Receiver operator characteristic curve), and its value ranges from 0.5 to 1; ACC refers to the percentage of correct prediction results in all results; the F1 score takes into account both the precision and recall of the classification model, and its value ranges from 0 to 1. The higher the values of the three metrics, the better the prediction performance of the model. Among them, AUC is the most commonly used evaluation metric in knowledge tracing tasks. 80% of the data in the dataset is divided into the training set, and 20% is divided into the test set. The embodiments of the present invention use the PyTorch framework to build the model and use the Adam optimizer to update the parameters. The maximum number of training epochs is set to 200, the batch_size is set to 32, the dimension of the hidden layer is set to 256, and the learning rate and Dropout are set to 0.001 and 0.5 respectively to prevent overfitting.

[0203] The present invention conducts comparative experiments and ablation experiments on the knowledge tracing method based on the interaction between item representation and student answering ability to verify that this method is superior to existing knowledge tracing methods and the role of key modules in this method. The results of the comparative experiments are shown in Table 1.

[0204] Table 1

[0205]

[0206] As can be seen from Table 1, the IQAKT model proposed by the present invention is superior to other knowledge tracing methods in terms of the metrics of the selected dataset. This shows that by introducing a deep residual cross network for high-order interaction between item features and answering ability, the ability to model students' knowledge states can be significantly improved. Specifically, compared with traditional sequence modeling methods (such as DKT and DKVMN), the model of the present invention can more effectively capture the complex dependencies between items and dynamically update the knowledge states of students, ensuring the adaptability and interpretability of the model in complex learning scenarios.

[0207] In addition, compared with DKT_F that introduces the forgetting factor and SAKT based on the attention mechanism, the model of the present invention also achieves significant advantages in performance. This indicates that the in-depth mining of question features and the dynamic interaction of students' answering abilities are the keys to improving the model performance. Although PEBG and DGEKT enhance the capture of the relationship between questions and knowledge points through the graph structure, they fail to effectively combine the high-order interaction of students' answering abilities and question features, while the model proposed by the present invention has significant advantages in this regard. On the Junyi Academy2015 dataset, the AUC index of the model of the present invention performs the best. Although the ACC is slightly lower than that of DGEKT, this further proves the overall performance improvement ability of the IQAKT model in predicting students' knowledge states. On the Assistment2017 dataset, the model of the present invention also performs excellently, indicating that it can effectively capture the dynamic knowledge changes of students when processing long sequence data. Generally speaking, the research results show that the method of the present invention has significant advantages in comprehensively mining multi-dimensional features such as question features, answering abilities, and learning processes, providing a new direction for further optimizing the knowledge tracing model in the future.

[0208] The ablation experiment results of the method proposed by the present invention on the ASSISTment2017 and Junyi datasets are shown in Table 2.

[0209] Table 2

[0210]

[0211] As can be seen from Table 2, each module of the IQAKT model of the present invention plays an important role in improving the model performance. The IQAKT-RC without pre-trained weights decreases by 2.45% and 1.75% on the two datasets respectively, indicating that the question representation can effectively capture the difficulty of the questions and the similarity between questions; the performance of IQAKT-RD without the deep residual cross network decreases by 2.67% and 2.65% respectively, indicating that this module can integrate the question and answering ability features and alleviate the problem of gradient disappearance; the performance of IQAKT-RA without the answering ability feature decreases by 2.47% and 2.68%, indicating that the answering ability feature is crucial for measuring students' knowledge states. In addition, the performance of IQAKT-RCD without question representation and the deep residual cross network decreases by 4.49% and 2.78%, further proving the necessity of these two modules in capturing feature interactions. The performance of IQAKT-RDA without the deep residual cross network and the answering ability feature decreases by 2.63% and 2.67%. Losing the information of students' answering behaviors will cause a significant decrease in the model performance. The performance of IQAKT-RCDA without all modules decreases the most (5.48% and 2.79%), indicating that these modules have a significant synergistic effect in improving the model performance.

[0212] In summary, the method of the present invention comprehensively considers the relationship between questions and knowledge points, the similarity between questions, and the dynamic changes in students' answering abilities. The experimental results verify the rationality and effectiveness of the modules proposed by the present invention. Compared with other inventions, the present invention has been improved in various measurement indicators and has good reference significance in practical applications.

[0213] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the principles and spirit of the present invention shall be included within the protection scope of the present invention.

Claims

1. A temporal convolutional knowledge tracking method based on the interaction between question representation and student answering ability characteristics, characterized in that: The steps include: (1) Obtain the historical answer record data set and perform data cleaning to extract the required feature set, including: students, questions, knowledge points, answer results, and number of answer attempts; and According to the time series of students' answer records; (2) Based on the answer results in the time series of students' answer records, a sliding window is used to calculate the correct answer rate of students over a period of time, and the students' answering ability characteristics are constructed based on the number of attempts they made on the corresponding questions; (3) assigning unique indexes to the questions and knowledge points respectively, and constructing a question-knowledge point correspondence matrix, a question similarity matrix, and a knowledge point similarity matrix; (4) The questions and knowledge points are respectively used as nodes of a heterogeneous graph, and edge connection relationships between questions and knowledge points are constructed through the question-knowledge point correspondence matrix, the question similarity matrix, and the knowledge point similarity matrix to form a question-knowledge point heterogeneous graph; (5) Using a heterogeneous graph neural network HGNN to learn the interaction relationship between nodes from the question-knowledge point heterogeneous graph to generate a question representation; (6) interactively fusing the student’s answering ability characteristics and question representations through a deep residual network to generate high-order features; (7) Inputting the high-order features into a temporal convolutional network (TCN) to model the temporal changes of the knowledge state; (8) Construct a prediction network to predict the correctness of future answers based on the knowledge state and high-order features.

2. The temporal convolution knowledge tracking method based on the interaction between question representation and student answering ability characteristics according to claim 1 is characterized in that: The step (1) is specifically: (1.1) Clean the historical answer record data set, filter out missing data in the historical answer record data set, and extract required features: Among them: u represents the student, q represents the question, s represents the knowledge point, r represents the answer result, and a represents the number of attempts to answer the question; (1.2) Based on the extracted required features, the time series of students’ answer records are generated according to the students’ statistical answer records, and the answer records of students with fewer interaction records are filtered out; The answer record is represented as: X={x1,x2,...,x t } x t =(q t ,s t ,r t, a t ) Where: q t represents the question to be answered at time t, s t Indicates q t The knowledge point belongs to, r t represents the answer result at time t, a t Indicates that students are in q t The number of answer attempts on the question.

3. The temporal convolution knowledge tracking method based on the interaction between question representation and student answering ability characteristics according to claim 1 is characterized in that: The specific steps for each student in step (2) are: (2.1) Use the sliding window algorithm to calculate the correct rate of students’ answers at time t: Where: W is the window size, r i Indicates the answer result at time i; (2.2) According to the correct answer rate and the number of answer attempts at time t, the student's answering ability score AbilityScore(t) at time t is obtained: Among them: a t is the number of attempts made by the student to answer the corresponding question at time t; (2.3) Summarize the student’s answering ability scores at all times to form the student’s knowledge state feature vector A, which serves as the student’s answering ability feature.

4. The temporal convolution knowledge tracking method based on the interaction between question representation and student answering ability characteristics according to claim 1 is characterized in that: The step (3) is specifically: (3.1) Assign a unique index to each of the questions and knowledge points, and construct a binary matrix P to represent the relationship between the questions and knowledge points, thus obtaining the question-knowledge point correspondence matrix: Among them: i represents the topic index, j represents the knowledge point index; (3.2) Use the intersection-and-union ratio formula to calculate the similarity between questions and the similarity between knowledge points; Among them: A q is the similarity matrix between topics, Represents the question q with index i i The corresponding knowledge point set, Represents the question q with index j j The corresponding knowledge point set; A s is the similarity matrix between knowledge points, Represents the knowledge point s with index i i The set of topics included, Represents the knowledge point s with index j j The set of topics included.

5. The temporal convolution knowledge tracking method based on the interaction between question representation and student answering ability characteristics according to claim 1 is characterized in that: The step (4) is specifically: (4.1) Calculate the question vertex feature matrix as the question node in the heterogeneous graph Where: N q Indicates the number of questions, X q Contains 3 feature dimensions X q [i]=[C q [i], T q [i], R q [i]],C q [i] represents the number of correct answers to the question with index i, T q [i] represents the total number of answers to the question with index i, R q [i] represents the average correct rate of the question with index i; Where: T represents the complete time period for all students to answer, q t represents the question to be answered at time t, r t represents the answer result at time t, q i Represents the question with index i; (4.2) Calculate the knowledge point vertex feature matrix as the knowledge point node in the heterogeneous graph Where: N s represents the number of knowledge points, X s Includes 3 feature dimensions X s [j]=[C s [j],T s [j],R s [j]], G s [j] represents the number of correct answers to the knowledge point with index j, T s [j] represents the total number of answers to the knowledge point with index j, R s [j] represents the average accuracy of index j; Where: T represents the complete time period for all students to answer, r t represents the answer result at time t, s t represents the knowledge point corresponding to the question answered at time step t, s j Represents the knowledge point with index j; (4.3) Extract the associated edge E from the question-knowledge point relationship matrix q-s , if the question q with index i i Involving knowledge point s with index j j , then it is set to 1, otherwise it is set to 0; (4.4) Extract edge set E from the question similarity matrix q-q ,The edge weight is determined by the similarity in the question similarity matrix; (4.5) Extract edge set E from knowledge point similarity matrix s-s ,The edge weight is determined by the similarity in the knowledge point similarity matrix.

6. The temporal convolution knowledge tracking method based on the interaction between question representation and student answering ability characteristics according to claim 1 is characterized in that: The heterogeneous graph neural network HGNN includes two parts: an encoder and a decoder. The encoder processes the question-knowledge point heterogeneous graph data through multi-layer convolution operations to learn node embedding; the decoder maps the node embedding output by the encoder and reconstructs the weights of the edges in the heterogeneous graph; specifically: (5.1) Encoder For node type v, the initial input is: Node features are extracted through multi-layer convolution, and the message passing formula between nodes in each layer is: in: represents the feature vector of node v in the lth layer, R is the set of edge types in the graph, N r (v) represents the set of nodes connected to node v through edge type r, is the weight matrix of edge type r in layer l, and σ is the activation function; For each edge type r, the formula for the convolution operation is: Conv r (h v ,h u )=α uv ·(W·h u ) Where: α uv is the attention coefficient; After L layers of convolution, the final node embedding representation is obtained: (5.2) The decoder is based on the node embedding Z v Reconstruct edge weights: in: Reconstruct the weight for the edge, z i 、z j are the embedding vectors of nodes i and j respectively; W is the linear transformation matrix, and σ is the activation function; (5.3) using the sum of positive and negative sample loss functions to train the heterogeneous graph neural network to minimize the reconstruction error; Among them: ε represents the actual edge set, ε neg represents the set of edges generated by negative sampling (5.4) Back-propagating the gradient according to the loss function L updates the parameters of the heterogeneous graph neural network, and uses the trained heterogeneous graph neural network to generate the final node embedding as the question representation.

7. The temporal convolution knowledge tracking method based on the interaction between question representation and student answering ability characteristics according to claim 1 is characterized in that: The step (6) is specifically as follows: (6.1) Combine the question representation with the answer ability feature to form a unified input feature vector as the input of the first layer of the deep residual cross network: X=Concat(Z q ,A) Where: Z q represents the question representation, and A represents the answering ability characteristic; (6.2) For the input features of the lth layer, feature mapping is performed through the fully connected layer; in: is the weight matrix, is the bias term, It is the feature after mapping by the fully connected layer; (6.3) Use multi-head attention mechanism to capture the correlation between features; Where: Q, K, V are query, key and value matrices, d k is the feature dimension scaling factor; (6.4) The attention mechanism and input features are fused through residual connections, and layer normalization and activation functions are performed; in: is the output of the attention mechanism; (6.5) After L layers of processing, the output high-order feature X DRCN As a cognitive integration feature that can express the relationship between students' answering ability and questions, DRCN =X (L) .

8. The temporal convolution knowledge tracking method based on the interaction between question representation and student answering ability characteristics according to claim 1 is characterized in that: The step (7) is specifically: (7.1) The high-order features are concatenated with the student’s answer results as the input of the temporal convolutional network (TCN): X TCN =Concat(X DRCN ,Y); Where: X DRCN is the high-level feature, Y is the student's answer, X TCN It is the first layer input of the dilated convolutional layer of TCN; (7.2) Input the concatenated features into the dilated convolutional layer of TCN; Among them: k represents the size of the convolution kernel, d represents the expansion coefficient and increases layer by layer with the change of the number of network layers, W convl [i] is the weight of the convolution kernel in layer l; Represents the output of the l-1th layer and is also the input of the lth layer; (7.3) is processed using residual connection and activation function as the output of the lth layer: (7.4) After L layers of convolution processing, The final knowledge state representation as the output of TCN 9. The temporal convolution knowledge tracking method based on the interaction between question representation and student answering ability characteristics according to claim 1 is characterized in that: The prediction network consists of a fully connected layer, an output layer and an activation function; specifically: (8.1) concatenating the knowledge state representation with the high-order features as input to the prediction network; X input =[Z tcn ,X DRCN ] Where: Z tcn represents the knowledge state, X DRCN Represents high-level features; (8.2) The input of the prediction network generates hidden layer features through the fully connected layer, and is processed through the output layer and the Sigmoid activation function, and the result is mapped to the [0,1] interval as the prediction value output; H hidden =σ(W fc X input +b fc ) Where: W fc is the weight matrix of the fully connected layer, b fc is the bias vector of the fully connected layer, σ is the activation function, and W out is the weight matrix of the output layer, b out is the bias vector of the output layer, is the predicted correctness of students’ answers; (8.3) The prediction network is trained using the cross entropy loss function to minimize the predicted value Between the correct answer t+1 The differences: Where: P t+1 The probability that the student answers the question correctly in the next time step, that is, the predicted value 10. The temporal convolution knowledge tracking method based on the interaction between question representation and student answering ability characteristics according to claim 9 is characterized by: The predicted value Transformed into binary value by condition, if If the answer is less than 0.5, the student's answer is considered wrong. If it is greater than or equal to 0.5, the student’s answer is considered correct.

Citation Information

Cited By

  • Cognitive analysis system for attention ability change and topic characteristics

    CN121599807A