A Knowledge Tracing Method Combining Difficulty Features and Temporal Correlation Features

By combining the knowledge tracking method of problem difficulty and time correlation characteristics, a model is constructed to consider students' individual differences and learning behavior, which improves the prediction accuracy and learning efficiency analysis of knowledge tracking, and solves the problem of failure to effectively consider students' individual differences in the existing technology.

CN119862275BActive Publication Date: 2025-07-29GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411933203.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-07-29
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

The existing knowledge tracking methods fail to effectively consider the impact of different students' knowledge mastery level, time and time interval on learning efficiency and knowledge acquisition of different questions, resulting in low prediction accuracy.

Method used

Combining the problem difficulty characteristics and time correlation characteristics, by constructing a knowledge tracking model, we calculate the embedded vectors of students' time, time interval, knowledge mastery level and answering questions, personalize the difficulty of questions, calculate the learning efficiency, and update the knowledge mastery level to predict students' future answering situation.

Benefits of technology

The accuracy of the knowledge tracking model predicts students' answers is improved, and the interpretability of learning efficiency calculations and the adaptability of personalized problem difficulty is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862275B_ABST
    Figure CN119862275B_ABST
Patent Text Reader

Abstract

The present invention discloses a knowledge tracing method that belongs to the technical field of knowledge tracing, specifically a knowledge tracing method that combines difficulty features and time correlation features, and includes the following specific steps: screening student data, constructing a student answer dataset, and preprocessing the student data; constructing a knowledge tracing model for predicting student answer situations and outputting the student's knowledge mastery level; training the knowledge tracing model; applying the model to recommend personalized questions for students. The present invention enhances the difficulty feature information embedded in the questions through the correlation weight between the questions and each knowledge point and the question difficulty, and gives personalized question difficulties that match the students' knowledge levels according to the knowledge mastery levels of different students; obtains the students' answer performance through the time taken by the students to do the questions, the time interval between doing the questions, and the answer situations, and then combines the personalized question difficulties to calculate the students' knowledge acquisition degree, enriching the semantic information of the knowledge acquisition degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge tracing, and in particular to a knowledge tracing method combining difficulty features and time correlation features. Background Art

[0002] Knowledge tracing is to model the understanding of knowledge points of students based on their historical answering records, so as to track the change of the mastery state of knowledge points of students (also known as the student knowledge mastery level) during their continuous answering process, and then predict the future answering situation of students. Knowledge tracing is essentially a supervised sequence prediction problem, that is, given the answering situations of students on questions q1, q2... q t to predict their answering result of the next question q t+1 .

[0003] Currently, the mainstream algorithms in the field of knowledge tracing include algorithms based on machine learning, algorithms based on deep learning, and algorithms based on dynamic key-value memory networks, etc. The representative of the knowledge tracing model based on machine learning is the Bayesian Knowledge Tracing (BKT) proposed by Corbett et al., which is applied in the Intelligence Tutoring System (ITS); the Deep Knowledge Tracing (DKT) model first applies the Recurrent Neural Network (RNN) to the field of knowledge tracing and achieves good prediction results. When modeling the learning behavior of students, the DKT model not only uses the answering performance at the current moment, but also uses the answering performance at historical moments to assist in modeling, so as to predict the future answering results of students; the knowledge tracing model based on dynamic key-value memory networks (Dynamic Key-Value Memory Networks for Knowledge Tracing, DKVMN) uses a static matrix and a dynamic matrix to update the knowledge state of students. This external reading and writing method makes the DKVMN model have better prediction results than the DKT model using RNN and LSTM, and its advantages are particularly obvious when dealing with a large amount of data, greatly improving the sequence modeling ability of the model.

[0004] Generally speaking, there is a close connection between the difficulty of a question and the level of students' knowledge mastery it reflects. When answering high-difficulty questions, students with a higher level of knowledge mastery can provide more accurate answers than those with a lower level. Simple questions can enhance students' learning enthusiasm, but the learning efficiency will be reduced; while more challenging questions tend to improve students' learning efficiency. In addition, from the perspective of question setting, the difficulty of questions is also of great significance. Questions that are too simple or too difficult cannot distinguish the knowledge states of different students.

[0005] During the learning process of students, the time interval between doing questions for the same knowledge point has a great impact on the students' knowledge state. As time goes by, along with different question-solving experiences, students may forget the knowledge points they learned before, or have a deeper understanding of a certain knowledge point. At the same time, the length of time students spend on doing questions will also affect their learning efficiency. Too short a time for doing questions indicates that the question cannot bring more challenges to students, thus reducing the learning efficiency; too long a time for doing questions may mean that the question difficulty exceeds the students' knowledge mastery level, which will also reduce the learning efficiency.

[0006] Based on the above, the existing knowledge tracking methods have the following problems:

[0007] 1. The existing technology calculates the impact of question difficulty on students' answering questions by setting the question difficulty level, but does not consider that different students have different knowledge mastery levels, and the difficulty of the same question is also different for different students.

[0008] 2. The existing technology updates the students' knowledge mastery level according to the question difficulty and the answering situation, but does not consider the impact of the time students spend on doing questions and the time interval between doing questions on the degree of knowledge acquisition and learning efficiency of students, as well as the impact of the degree of knowledge acquisition and learning efficiency on the knowledge mastery level.

[0009] Therefore, an invention of a knowledge tracking method that combines difficulty features and time correlation features is proposed. By combining difficulty features and time correlation features to analyze and process the students' learning efficiency and knowledge mastery level, and then making predictions, the accuracy of the model's prediction of students' answering situations is improved. Summary of the Invention

[0010] To solve the above technical problems, according to one aspect of the present invention, the following technical solutions are provided:

[0011] A knowledge tracking method that combines difficulty features and time correlation features, which includes the following specific steps:

[0012] S1: Screen student data, construct a student answering question data set, and preprocess the student data;

[0013] S2: Construct a knowledge tracing model for predicting students' answering situations and outputting students' knowledge mastery levels;

[0014] Among them, the specific steps of S2 are as follows:

[0015] S21: Calculate the time used by students to answer questions and the time interval between answering questions, and initialize the embedding vectors of the time used to answer questions, the time interval between answering questions, the knowledge mastery level, and the answering situation;

[0016] S22: Construct question embeddings in combination with question difficulty, and calculate the personalized question difficulty that matches the students' knowledge mastery levels;

[0017] S23: Calculate the degree of knowledge acquisition after students answer questions;

[0018] S24: Calculate the learning efficiency and update the students' knowledge mastery levels;

[0019] S25: Predict the answering situations of students at time t + 1, and determine whether to output the students' knowledge mastery levels;

[0020] S3: Train the knowledge tracing model;

[0021] S4: Apply the model to recommend questions personalized for students.

[0022] As a preferred solution of a knowledge tracing method combining difficulty features and time correlation features according to the present invention, among them: the specific steps of S1 are as follows:

[0023] S11: Screen students' data and classify and organize students' answering data;

[0024] S12: Calculate the difficulty of answering questions and the difficulty of knowledge points, and initialize the embedding matrices of question numbers, knowledge point numbers, and question difficulties.

[0025] As a preferred solution of a knowledge tracing method combining difficulty features and time correlation features according to the present invention, among them: the specific steps of S11 are as follows:

[0026] S111: Obtain students' initial answering data, and each answering record in the answering data includes student numbers, answering situations, start time of answering questions, end time of answering questions, question numbers, knowledge point numbers, knowledge point names, question types, teacher number information;

[0027] S112: Classify and organize the information involved in students' answering data, and after classification and organization, obtain the student list, question list, and knowledge point list as follows:

[0028] Student list S = {s1, s2,..., s I}, s i∈S represents a student, whose value is the student number, and I is the number of students;

[0029] The question list Q = {q1, q2, …, q J}, q j ∈Q represents a question, whose value is the question number; J is the number of questions;

[0030] The knowledge point list K = {k1, k2, …, k N}, k n ∈K represents a knowledge point, whose value is the knowledge point number, and N is the number of knowledge points;

[0031] During the process of doing questions, the answer record sequence of student s i is represented as X i = {(q1, a1), (q2, a2), …, (q T , a T )}, where, (q t , a t ) is the answer record of s i at time t. q t ∈Q is the question answered by s i at time t, and a t is the answer situation of s i to q t at time t. Its value is 0 or 1, 1 means correct answer, 0 means wrong answer. T is the number of questions answered by s i . The start answering time and end answering time of s i answering questions are recorded in the answering time sequence QT i = {(st1, et1), (st2, et2), …, (st T , et T )}, where, (st t , et t ) are the start answering time st i and end answering time et t when s t answers q t .

[0032] As a preferred solution of a knowledge tracking method combining difficulty characteristics and time correlation characteristics according to the present invention, wherein: The specific steps of S12 are as follows:

[0033] S121, calculate the question answering difficulty. The calculation method of the question answering difficulty of any question q j in the question list Q is as follows:

[0034] S1211: Count the number of students who answer q j and calculate the number of students who answer q correctlyj The proportion of the number of students answering q j in the total number of students, and then multiply by the question answering difficulty level C qs , to obtain the question answering difficulty of q j . The calculation formula is as follows:

[0035]

[0036] Among them, qs j is the question answering difficulty of q j , is the number of students answering q j , is the total number of students who answered q j correctly, and C qs is the set question answering difficulty level;

[0037] S1212: Repeat S121 to calculate the question answering difficulties of all questions in Q, and obtain the question answering difficulty list QS = {qs1, qs2,..., qs J}, qs j ∈QS represents the question answering difficulty of q j ;

[0038] S122. Calculate the knowledge point difficulty. The calculation method of the difficulty of any knowledge point k n in the knowledge point list K is as follows:

[0039] S1221: Count the total number of students who answered the questions containing k n , and then calculate the proportion of the number of students who answered the questions containing k n correctly, and multiply by the knowledge point difficulty level C kd , to obtain the difficulty of k n . The calculation formula is as follows:

[0040]

[0041] Among them, kd n is the knowledge point difficulty of k n , is the number of students who answered the questions containing k n , is the number of students who answered the questions containing k n correctly, and C kd is the set knowledge point difficulty level;

[0042] S1222: Repeat S122 to calculate the knowledge point difficulties of all knowledge points in K, and obtain the knowledge point difficulty list KD = {kd1, kd2,..., kd N}, kd n∈KD represents k n ∈K's knowledge point difficulty;

[0043] S123, initialize the embedding matrices of knowledge point numbers and question numbers:

[0044] S1231, initialize the knowledge point number embedding matrix KN: Represent the number of knowledge point k n as a one-hot encoding O(k n )∈R 1×N , set an embedding matrix Map k n into an embedding vector

[0045]

[0046] where d k is the dimension of the knowledge point number embedding vector, calculate the embedding vectors of all knowledge point numbers, and then combine them into a knowledge point number embedding matrix Each d kn -dimensional row vector in KN represents a knowledge point;

[0047] S1232, initialize the question number embedding matrix QN: Represent the number of question q j as a one-hot encoding O(q j )∈R 1 ×J , set an embedding matrix Map q j into an embedding vector

[0048]

[0049] where d q is the dimension of the question number embedding vector, calculate the embedding vectors of all question numbers, and then combine them into a question number embedding matrix Each dq n -dimensional row vector in QN represents a question;

[0050] S124, calculate the question difficulty, initialize the question difficulty embedding matrix QD:

[0051] S1241, calculate the question difficulty: First, construct a matrix QK to describe the association between questions and knowledge points. Each row in QK corresponds to a question, and each column in QK corresponds to a knowledge point. When q j is related to k nWhen associated, the value at the j-th row and n-th column in QK is 1, otherwise it is 0. At the same time, to calculate the degree of association between each knowledge point and the question, q is calculated using QK j The association weights with each knowledge point

[0052]

[0053] where KN is the knowledge point serial number embedding matrix; QK j,: represents the j-th row of QK and is used to describe the association status of q j with each knowledge point; weight represents the degree of association of q j with each knowledge point; then, considering the question difficulty, the difficulty of answering the question and the difficulty of the knowledge points included in the question, through the association weights of q j with each knowledge point weight the knowledge point difficulty KD, and then add the weighted knowledge point difficulty to the difficulty of answering the question to obtain the question difficulty qd of q j , and the calculation formula is as follows: j , the calculation formula is as follows:

[0054]

[0055] where QS j is the j-th bit of the question answering difficulty list QS and represents the question answering difficulty of q j , is the transpose of;

[0056] S1242, initialize the question difficulty embedding matrix QD: Represent the difficulty qd of q j as a one-hot encoding j Set an embedding matrix Set an embedding matrix C qs is the question answering difficulty level set in S1211, map qd j into an embedding vector

[0057]

[0058] where d qd is the dimension of the question difficulty embedding vector, calculate the embedding vectors of all question difficulties, and then merge them into a question difficulty embedding matrix Each d-dimensional row vector in QD represents a question difficulty. qd Each d-dimensional row vector in QD represents a question difficulty.

[0059] As a preferred solution of a knowledge tracing method combining difficulty features and time correlation features according to the present invention, wherein: the specific steps of S21 are as follows:

[0060] S211. Calculate the time used for answering questions and the time interval between answering questions:

[0061] S2111: Obtain the answer record sequence X i of student s i ={(q1,a1),(q2,a2),…,(q T ,a T )} and the answer time sequence QT i ={(st1,et1),(st2,et2),…,(st T ,et T )} from the data sorted out in S11. Subtract the start answering time st i from the end answering time et t when s t answers q t to obtain the time used aq i for s t to answer q t . The calculation formula is as follows:

[0062] aq t =et t -st t

[0063] S2112: Calculate the time used for answering all questions in X i to obtain the time used list AQ i ={aq1,aq2,…,aq i} of s T . Wherein, aq t ∈AQ i indicates the time used for s i to answer q t ;

[0064] S2113: For any question q i answered by s t , obtain the knowledge point k t with the highest degree of association with q according to the association weight t between q t and each knowledge point calculated in S122. Search for k i in the answer record before s t answers q t-1 ,a t-1 )} tThe problem q with the highest degree of association and the smallest interval of the end time of answering questions t The problem q' with the smallest interval of the end time of answering questions t , subtract s i The end time of answering q t et t from the end time of answering q' t et' t , to obtain s i The interval of answering q t bq t . If q' t does not exist, then bq t = 0. The calculation formula is as follows:

[0065]

[0066] S2114: Calculate the intervals of answering questions for all questions in X i to obtain the list of answering time intervals BQ i = {bq1, bq2,..., bq T}, where bq t ∈ EQ i represents the interval of answering q i by s t ;

[0067] S212. Initialize the answering time embedding vector and the answering time interval embedding vector:

[0068] S2121: Represent the answering time aq i of answering q t by s t as a one-hot encoding O(aq t ) ∈ R 1×T , set an embedding matrix to map aq t to an embedding vector d aq is the dimension of the answering time embedding vector:

[0069]

[0070] S2122: Calculate the answering time embedding vectors for all questions in X i and merge them into the answering time embedding matrix

[0071] S2123: Represent the answering time interval bq i of answering q t by s t as a one-hot encoding O(bq t ) ∈ R 1×T , set an embedding matrix bq t Mapping to embedding vector d bq is the dimension of the embedding vector of the time interval between questions:

[0072]

[0073] S2124: Calculate X i The time interval embedding vectors of all questions in the question are merged into the time interval embedding matrix

[0074] S213, initialize the embedding vectors of knowledge mastery level and answer status:

[0075] S2131: Create vector Used to express s i The level of knowledge mastered at time t, d h yes Dimensions;

[0076] S2132: For s i Answer Q t Answers to the questions t , a t Mapping into embedding vectors d a is the dimension of the answer embedding vector, When the element is 1, it means the answer is correct, and when the element is 0, it means the answer is wrong.

[0077] As a preferred solution of the knowledge tracking method combining difficulty characteristics and time correlation characteristics described in the present invention, the specific steps of S22 are as follows:

[0078] S221, build question embedding based on question difficulty: for s i The answer record at time t (q t ,a t ), using the problem difficulty embedding matrix QD to strengthen the problem q t Information representation, combined with q t Question number embedding The association weight between questions and knowledge points and question difficulty embedding Output q through the multi-layer perceptron t Title embedding The calculation is as follows:

[0079]

[0080] in, is a matrix concatenation operation, is the weight matrix, is the bias term; d q is the dimension of the problem sequence embedding vector; N is the number of knowledge points and also the association weight between the problem and the knowledge point of dimension; d qd is the dimension of the problem difficulty embedding vector; d h is the dimension of the knowledge mastery level vector;

[0081] S222, calculate the personalized problem difficulty that matches the student's knowledge mastery level: at time t, input s i in answering q t the knowledge mastery level before to the personalized problem difficulty calculation module, and use q t of the question embedding and calculate and output q t for s i the personalized problem difficulty d sq is the dimension of SQ t The calculation method is as follows:

[0082]

[0083] where, is s i in answering q t the knowledge mastery level before, tanh is a non-linear activation function, and σ is the sigmoid activation function, is the weight matrix, is the bias term, used to calculate the feature representation of the personalized problem difficulty of q t for s i the personalized problem difficulty used to calculate the weight information of the personalized problem difficulty of q t for s i the personalized problem difficulty

[0084] As a preferred solution of the knowledge tracing method combining difficulty features and time correlation features described in the present invention, wherein: the specific steps of S23 are as follows:

[0085] S231: Concatenate the time spent on answering questions embedding vector i when answering q t the time interval between answering questions embedding vector and the answering situation embedding vector used to represent the answering performance of s for s i the answering performance The calculation method is as follows:

[0086]

[0087] S232: Input the difficulty of personalized questions into the knowledge acquisition degree calculation module, and combine with to calculate and output the knowledge acquisition degree s i as follows: The calculation method is as follows:

[0088]

[0089] where is the weight matrix, is the bias term, is used to extract the feature representation of the knowledge acquisition degree of s i , is used to calculate the weight information of the knowledge acquisition degree of s i ; d aq is the dimension of the time spent on answering questions embedding vector; d bq is the dimension of the time interval between questions embedding vector; d a is the dimension of the answering situation embedding vector.

[0090] As a preferred solution of the knowledge tracking method combining difficulty features and time correlation features according to the present invention, wherein: The specific steps of S24 are as follows:

[0091] S241: Design a learning efficiency calculation module to calculate the learning efficiency of student s i based on the answering situation and the time spent on answering questions. When student s i answers question q t , it is considered that the learning efficiency of answering correctly is higher than that of answering wrongly, and the learning efficiency of the time spent on answering questions being moderate is higher than that of being too long or too short. The learning efficiency l t is calculated as follows:

[0092]

[0093] where a is a hyperparameter, σ is the sigmoid activation function, a t is the answering situation of student s answering q t , aq t is the time spent by student s i answering q t , is the average value of the time spent by all students answering q t ;

[0094] S242: Design a knowledge mastery level update module and input the knowledge acquisition degree and the knowledge mastery level to the knowledge mastery level update module, and use l t as the weight, and perform weighted addition on and to output the updated knowledge mastery level The calculation method for updating the knowledge mastery level is as follows:

[0095]

[0096] Among them, is s i the knowledge mastery level before answering question q t , is s i the knowledge mastery level after answering question q t .

[0097] As a preferred solution of a knowledge tracking method combining difficulty characteristics and time correlation characteristics according to the present invention, wherein: the specific steps of S25 are as follows:

[0098] S251: is s i the knowledge mastery level at time t + 1. Combining the question embedding t+1 of q can obtain s i the prediction y t+1 of the answering situation of answering question q t+1 , and the calculation method is as follows:

[0099]

[0100] Among them, y t+1 is the predicted probability of s i answering question q t+1 correct. When y t+1 ≥β, it is considered that s i answers correctly. When y t+1 <β, it is considered that s i answers incorrectly, and β is a manually set threshold;

[0101] S252: If t + 1 ≤ T, jump to S22 to continue processing the answering record (q t+1 , a t+1 ); otherwise, output the knowledge mastery level of s i

[0102] As a preferred solution of a knowledge tracking method combining difficulty characteristics and time correlation characteristics according to the present invention, wherein: the specific steps of S3 are as follows:

[0103] ​S31: Divide the student answer dataset into a training set and a test set according to a certain ratio, initialize the training parameters and model hyperparameters, and input the data of the training set into the knowledge tracing model for training;

[0104] S32: Use the cross-entropy loss function to calculate the loss value L between the actual answer situation a t of the student and the predicted value y t as follows:

[0105]

[0106] where θ represents all trainable parameters and embeddings in the model, and λ θ is the regularization hyperparameter. When both the accuracy rate and the loss value of the model tend to be stable, the training ends. After the training ends, use the test set to test the model to confirm the actual prediction ability of the model;

[0107] The specific steps of the said S4 are as follows:

[0108] S41: After the model training ends, input the answer sequence of a student into the model. After being processed by the model, output the student's knowledge mastery level;

[0109] S42: For the knowledge points that the student has a weak grasp of, select the questions that the student has not answered and contain the knowledge points that the student has a weak grasp of from the question list to form a question sequence to recommend to the student for answering. When the student answers the recommended question sequence, the model simultaneously predicts the student's answer situation and updates the student's knowledge mastery level.

[0110] Compared with the prior art:

[0111] In the present invention, through the correlation weight between the questions and each knowledge point and the question difficulty, the difficulty characteristic information of the question embedding is enhanced, and according to the knowledge mastery levels of different students, personalized question difficulties that conform to the students' knowledge levels are given; the student's answer performance is obtained through the time taken by the student to do the questions, the time interval between doing the questions, and the answer situation, and then the knowledge acquisition degree of the student is calculated in combination with the personalized question difficulty, enriching the semantic information of the knowledge acquisition degree; the learning efficiency of the student is analyzed from the answer situation and the time taken to do the questions, enhancing the interpretability of the learning efficiency calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0112] Figure 1 is a schematic flow chart of the present invention;

[0113] Figure 2 is a schematic diagram of the model structure of the present invention;

[0114] Figure 3 is a schematic diagram of the structure of the personalized question difficulty calculation module of the present invention;

[0115] Figure 4 Schematic diagram of the knowledge acquisition degree calculation module of the present invention;

[0116] Figure 5 Schematic diagram of the learning efficiency calculation module and the knowledge mastery level update module of the present invention. Specific embodiments

[0117] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0118] The present invention provides a knowledge tracking method combining difficulty features and time correlation features. Please refer to Figures 1 - 5 and the specific steps are as follows:

[0119] S1: Screen student data, construct a student answer dataset, and preprocess the student data;

[0120] Among them: The specific steps of S1 are as follows:

[0121] S11: Screen student data and classify and organize student answer data;

[0122] Among them: The specific steps of S11 are as follows:

[0123] S111: Obtain the initial student answer data. Each answer record in the answer data contains student number, answer situation, start time of doing the question, end time of doing the question, question number, knowledge point number, knowledge point name, question type, teacher number information; The specific method is: Obtain the student number, question number, answer situation, knowledge point number, start answer time and end answer time from the initial student answer data, and delete the student answer records with the number of answering students less than 2 and the knowledge point number being empty;

[0124] S112: Classify and organize the information involved in the student answer data. After classification and organization, the following student list, question list and knowledge point list are obtained:

[0125] Student list S = {s1, s2,..., s I}, s i ∈S represents a student, whose value is the student number, and I is the number of students;

[0126] Question list Q = {q1, q2,..., q J}, q j ∈Q represents a question, whose value is the question number; J is the number of questions;

[0127] Knowledge point list K = {k1, k2,..., k N}, k n∈K represents a knowledge point, whose value is the knowledge point serial number, and N is the number of knowledge points;

[0128] During the process of doing exercises, student s i 's answer record sequence is denoted as X i ={(q1,a1),(q2,a2),…,(q T ,a T )}, where (q t ,a t ) is s i 's answer record at time t, q t ∈Q is the question answered by s i at time t, a t is s i 's answer situation regarding q t at time t, whose value is 0 or 1, 1 indicates a correct answer, 0 indicates a wrong answer, T is the number of questions answered by s i , and the start answering time and end answering time of s i answering questions are recorded in the answering time sequence QT i ={(st1,et1),(st2,et2),…,(st T ,et T )}, where (st t ,et t ) is the start answering time st i and end answering time et t of s t answering q t ;

[0129] S12: Calculate the difficulty of question answering and the difficulty of knowledge points, and initialize the embedding matrices of question numbers, knowledge point numbers, and question difficulties;

[0130] Among them: The specific steps of S12 are as follows:

[0131] S121. Calculate the difficulty of question answering. The calculation method of the difficulty of question answering for any question q j in the question list Q is as follows:

[0132] S1211: Count the number of students who answered q j , calculate the proportion of the number of students who answered q j correctly in the total number of students who answered q j , and then multiply it by the question answering difficulty level C qs , and the difficulty of question answering for q j can be obtained. The calculation formula is as follows:

[0133]

[0134] Among them, qs j is the difficulty of answering question q j , is the number of students answering question q j , is the total number of students who answer question q j correctly, and C qs is the set difficulty level of answering questions;

[0135] S1212: Repeat S121 to calculate the difficulty of answering questions for all questions in Q, and obtain the question answering difficulty list QS = {qs1, qs2,..., qs J}, where qs j ∈QS represents the difficulty of answering question q j ;

[0136] S122. Calculate the knowledge point difficulty. The calculation method of the difficulty of any knowledge point k n in the knowledge point list K is as follows:

[0137] S1221: Count the total number of students who have answered questions containing k n , then calculate the proportion of the number of students who have answered questions containing k n correctly, and multiply it by the knowledge point difficulty level C kd to obtain the difficulty of k n . The calculation formula is as follows:

[0138]

[0139] Among them, kd n is the knowledge point difficulty of k n , is the number of students who have answered questions containing k n , is the number of students who have answered questions containing k n correctly, and C kd is the set knowledge point difficulty level;

[0140] S1222: Repeat S122 to calculate the knowledge point difficulty of all knowledge points in K, and obtain the knowledge point difficulty list KD = {kd1, kd2,..., kd N}, where kd n ∈KD represents the knowledge point difficulty of k n ∈K;

[0141] S123. Initialize the embedding matrices of knowledge point numbers and question numbers:

[0142] S1231. Initialize the knowledge point number embedding matrix KN: Represent the number of the knowledge point k n as a one-hot encoding O(kn ) ∈ R 1×N , set an embedding matrix Map k n into an embedding vector

[0143]

[0144] where d k is the dimension of the embedding vector of the knowledge point number. Calculate the embedding vectors of all knowledge point numbers, and then combine them into an embedding matrix of knowledge point numbers Each d kn -dimensional row vector in KN represents a knowledge point;

[0145] S1232, initialize the question number embedding matrix QN: Represent the number of the question q j as a one-hot encoding O(q j ) ∈ R 1 ×J , set an embedding matrix Map q j into an embedding vector

[0146]

[0147] where d q is the dimension of the embedding vector of the question number. Calculate the embedding vectors of all question numbers, and then combine them into an embedding matrix of question numbers Each d qn -dimensional row vector in QN represents a question;

[0148] S124, calculate the question difficulty, initialize the question difficulty embedding matrix QD. The present invention designs a method for calculating the question difficulty. By weighting the difficulty of knowledge points with the correlation weights between the question and each knowledge point, and then combining it with the question answering difficulty, the question difficulty combined with the knowledge point difficulty is obtained. The specific steps are as follows:

[0149] S1241, calculate the question difficulty: First, construct a matrix QK to describe the correlation between the question and the knowledge point. Each row in QK corresponds to a question, and each column in QK corresponds to a knowledge point. When q j is associated with k n , the value at the j-th row and n-th column position in QK is 1, otherwise it is 0; At the same time, to calculate the correlation degree between each knowledge point and the question, use QK to calculate the correlation weight j between q

[0150]

[0151] Among them, KN is the knowledge point serial number embedding matrix; QK j,: represents the j-th row of QK, which is used to describe the correlation status of q j with each knowledge point; the weight represents the correlation degree of q j with each knowledge point; then, considering the problem difficulty, the answering difficulty of the question and the difficulty of the knowledge points included in the question, through the correlation weight of q j with each knowledge point weight the knowledge point difficulty KD, and then add the weighted knowledge point difficulty to the answering difficulty of the question to obtain the question difficulty qd of q j , and the calculation formula is as follows: j , the calculation formula is as follows:

[0152]

[0153] Among them, QS j is the j-th position of the question answering difficulty list QS, which represents the question answering difficulty of q j , is the transpose of;

[0154] S1242, initialize the question difficulty embedding matrix QD: represent the difficulty qd of q j as a one-hot encoding j Set an embedding matrix C C qs is the question answering difficulty level set in S1211, map qd j into an embedding vector

[0155]

[0156] Among them, d qd is the dimension of the question difficulty embedding vector, calculate the embedding vectors of all question difficulties, and then merge them into a question difficulty embedding matrix Each d-dimensional row vector in QD represents a question difficulty; qd

[0157] This step includes but is not limited to the following embodiments:

[0158] Obtain the initial answering data of students, classify and organize the data to obtain the student list S = {s1, s2,..., s I}, the question list Q = {q1, q2,..., q J} and the knowledge point list K = {k1, k2,..., k N}, record the answering record sequence X of student s i ​​i = {(q1, a1), (q2, a2), …, (q T , a T )} and the answering time series QT i = {(st1, et1), (st2, et2), …, (st T , et t )}; Set the difficulty level C of question answering qs = 100, calculate the difficulty of all questions to obtain the question answering difficulty list QS; Set the difficulty level C of knowledge points kd = 100, calculate the difficulty of all knowledge points to obtain the knowledge point difficulty list KD; Initialize the embedding matrices of knowledge point numbers, question numbers, and question difficulties to obtain the knowledge point number embedding matrix Question number embedding matrix and the question difficulty embedding matrix

[0159] S2: Construct a knowledge tracing model for predicting students' answering situations and outputting students' knowledge mastery levels; The present invention constructs a knowledge tracing model that combines difficulty features and time correlation features, and the model structure is as Figure 2 shown;

[0160] Among them, the specific steps of S2 are as follows:

[0161] S21: Calculate the time used by students to do questions and the time interval between doing questions, and initialize the embedding vectors of the time used to do questions, the time interval between doing questions, the knowledge mastery level, and the answering situation; The present invention designs a method for extracting the time correlation features of students' question answering, calculates the time used by students to do questions and the time interval between doing questions, and obtains the time correlation features of students' question answering. These features include the embedding vectors of the time used to do questions, the time interval between doing questions, the knowledge mastery level, and the answering situation, which are used for subsequent operations of the model;

[0162] Among them: The specific steps of S21 are as follows:

[0163] S211, calculate the time used to do questions and the time interval between doing questions:

[0164] S2111: Obtain the answering record sequence X i of student s i = {(q1, a1), (q2, a2), …, (q T , a T )} and the answering time series QT i = {(st1, et1), (st2, et2), …, (st T , et T )} from the data sorted in S11, and let s i answer q tThe end time et of answering questions t Subtract the start time st of answering questions t , to obtain s i Answer q t The time aq taken to answer question q t , and the calculation formula is as follows:

[0165] aq t = et t - st t

[0166] S2112: Calculate the time taken to answer all questions in X i , to obtain the list AQ of the time taken to answer questions s i = {aq1, aq2,..., aq i}, where aq T ∈ AQ t represents the time aq taken to answer question q i s i Answer q t ;

[0167] S2113: For any question q i answered in s t , according to the association weight between q t and each knowledge point calculated in S122 obtain the knowledge point k t with the highest degree of association with q t , and search in the previous answering record of s i when answering q t i.e., {(q1, a1), (q2, a2),..., (q t-1 , a t-1 )} for the question q' t with the highest degree of association with k t and the smallest time interval from the end time of answering q t , subtract the end time et' i of answering q' t from the end time et t of answering q t in s t , to obtain the time interval bq i of answering q t in s t . If q' t does not exist, then bq t = 0, and the calculation formula is as follows:

[0168]

[0169] S2114: Calculate X iThe time intervals for all questions in the question, get the time interval list BQ i ={bq1,bq2,…,bq T}, where bq t ∈BQ i Indicates s i Answer Q t The time interval between questions;

[0170] S212, initialize the embedding vector of the time spent on solving the problem and the embedding vector of the time interval between solving the problem:

[0171] S2121: will s i Answer Q t The time it takes to solve the problem is aq t Expressed as one-hot encoding O(aq t )∈R 1×T , set up an embedding matrix Aq t Mapping to embedding vector d aq is the dimension of the embedding vector when solving the problem:

[0172]

[0173] S2122: Calculate X i The time embedding vectors of all questions in the problem are merged into the time embedding matrix

[0174] S2123: will s i Answer Q t The time interval between questions bq t Expressed as one-hot encoding O(bq t )∈R 1×T , set up an embedding matrix DQ t Mapping to embedding vector d bq is the dimension of the embedding vector of the time interval between questions:

[0175]

[0176] S2124: Calculate X i The time interval embedding vectors of all questions in the question are merged into the time interval embedding matrix

[0177] S213, initialize the embedding vectors of knowledge mastery level and answer status:

[0178] S2131: Create vector Used to express s i The level of knowledge mastered at time t, dh yes Dimensions;

[0179] S2132: For s i Answer Q t Answers to the questions t , a t Mapping into embedding vectors d a is the dimension of the answer embedding vector, When the element is 1, it means the answer is correct, and when the element is 0, it means the answer is wrong;

[0180] S22: Constructing a question embedding based on the difficulty of the question, and calculating the difficulty of the personalized question that matches the student's knowledge mastery level; considering that different students have different knowledge mastery levels, the difficulty of the same question for different students is also different; the present invention constructs a personalized question difficulty calculation module that matches the student's knowledge mastery level, which is used to calculate q t For i the difficulty of personalized questions;

[0181] The specific steps of S22 are as follows:

[0182] S221, build question embedding based on question difficulty: for s i The answer record at time t (q t ,a t ), using the problem difficulty embedding matrix QD to strengthen the problem q t Information representation, combined with q t Question number embedding The association weight between questions and knowledge points and question difficulty embedding Output q through a multi-layer perceptron (MLP) t Title embedding The calculation is as follows:

[0183]

[0184] in, is a matrix concatenation operation, is the weight matrix, is the bias term; d q is the dimension of the question number embedding vector; N is the number of knowledge points, which is also the association weight between the question and the knowledge point Dimension; d qd is the dimension of the question difficulty embedding vector; d h is the dimension of the knowledge mastery level vector;

[0185] S222, calculate the difficulty of personalized questions that match the student's knowledge level: at time t, input si Before answering question q t Previous knowledge mastery level To the personalized question difficulty calculation module, using the question embedding of q t of q and Calculate and output the difficulty of question q t For s i The personalized question difficulty of s d sq is the dimension of SQ t The calculation method is as follows:

[0186]

[0187] Among them, is s i Before answering question q t Previous knowledge mastery level, tanh is a non-linear activation function, σ is the sigmoid activation function, is the weight matrix, is the bias term, used to calculate the feature representation of the personalized question difficulty of q t For s i The feature representation of the personalized question difficulty of s used to calculate the weight information of the personalized question difficulty of q t For s i The weight information of the personalized question difficulty of s

[0188] Among them: The tanh activation function is a commonly used activation function in neural networks, and its calculation formula is:

[0189]

[0190] In deep learning, the tanh activation function is often used in the hidden layer because its output is zero-centered, which helps with data processing and network training;

[0191] The sigmoid activation function is widely used in neural networks, especially in binary classification problems; since its output can be regarded as a probability, it is very useful in the output layer for converting the predictions of the neural network into a probability distribution; in addition, the smoothness of the sigmoid function provides a continuous gradient during the gradient descent optimization process, which is beneficial to model training. The calculation formula is:

[0192]

[0193] S23: Calculate the knowledge acquisition degree after the student answers the question; The present invention constructs a knowledge acquisition degree calculation module for calculating the knowledge acquisition degree after s i After answering question q t The knowledge acquisition degree. Using s iFor q t Calculate the knowledge acquisition degree of the student's answer to the question based on the personalized question difficulty, time taken to answer the question, time interval between answering questions, and answering situation; the structure of the knowledge acquisition degree calculation module is as shown in Figure 4 shown;

[0194] Among them: The specific steps of S23 are as follows:

[0195] S231: Concatenate the time - taken - to - answer question embedding vector i when answering q t , the time - interval - between - answering - questions embedding vector , and the answering - situation embedding vector to represent the answering performance of s . The calculation method is as follows: i The answering performance of s is calculated as follows:

[0196]

[0197] S232: Input the personalized question difficulty into the knowledge acquisition degree calculation module, and combine it with and to calculate and output the knowledge acquisition degree of s i . The calculation method is as follows: The calculation method is as follows:

[0198]

[0199] Among them, is the weight matrix, is the bias term, is used to extract the feature representation of the knowledge acquisition degree of s i , and is used to calculate the weight information of the knowledge acquisition degree of s i ; d aq is the dimension of the time - taken - to - answer - question embedding vector; d bq is the dimension of the time - interval - between - answering - questions embedding vector; d a is the dimension of the answering - situation embedding vector;

[0200] S24: Calculate the learning efficiency and update the student's knowledge mastery level; The present invention designs a learning efficiency calculation module and a knowledge mastery level update module to calculate the learning efficiency and update the student's knowledge mastery level. The structure diagram is as shown in Figure 5 shown;

[0201] Among them: The specific steps of S24 are as follows:

[0202] S241: Design a learning efficiency calculation module. Through the student s iCalculate the learning efficiency based on the answering situation and time spent on answering questions of student s. When student s i is answering question q t it is considered that the learning efficiency of answering correctly is higher than that of answering wrongly, and the learning efficiency of spending an appropriate amount of time on answering questions is higher than that of spending too long or too short time on answering questions. The learning efficiency l t is calculated as follows:

[0203]

[0204] where a is a hyperparameter, σ is the sigmoid activation function, a t is the answering situation of student s answering q t aq t is the time spent by student s i answering q t and is the average value of the time spent by all students answering q t ;

[0205] S242: Design a knowledge mastery level update module. Input the knowledge acquisition degree and the knowledge mastery level to the knowledge mastery level update module. Use l t as the weight, and perform weighted addition on and to output the updated knowledge mastery level The calculation method of the knowledge mastery level update is as follows:

[0206]

[0207] where is the knowledge mastery level before s i answers q t , is the knowledge mastery level after s i answers q t ;

[0208] S25: Predict the answering situation of the student at time t + 1 and determine whether to output the student's knowledge mastery level;

[0209] Among them: The specific steps of S25 are as follows:

[0210] S251: is the knowledge mastery level of s i at time t + 1 (i.e., before answering q t+1 ). Combining the question embedding t+1 of q can obtain the predicted y i of the answering situation of s answering q t+1 t+1 ​, the calculation method is as follows:

[0211]

[0212] Among them, y t+1 is the predicted probability of the correct answer q i of s t+1 When y t+1 ≥β, it is considered that s i answers correctly. When y t+1 <β, it is considered that s i answers incorrectly. β is a manually set threshold;

[0213] S252: If t + 1 ≤ T, jump to S22 to continue processing the answering record (q t+1 , a t+1 ); otherwise, output the knowledge mastery level of s i This step includes but is not limited to the following embodiments:

[0214] Obtain the answering record sequence X

[0215] ={(q1, a1), (q2, a2), …, (q i of student s i , a T , a T )} and the answering time sequence QT i ={(st1, et1), (st2, et2), …, (st T , et T )}, calculate the time used for answering questions and the time interval between answering questions, and obtain the answering time list AQ i ={aq1, aq2, …, aq T} and the answering time interval list BQ i ={bq1, bq2, …, bq T}, respectively construct the answering time embedding vector and the answering time interval embedding vector; create a vector representing the knowledge mastery level of s i at time t. At time t = 1, the knowledge mastery level of s i is Input the answering record (q1, a1) into the knowledge tracing model, and combine the proportion weight of the knowledge points included in the question of q1 and the question difficulty embedding to output the question embedding of q1 through a multi-layer perceptron (MLP) Combine the question embedding i of q1 and the knowledge mastery level of s Calculate the personalized question difficulty of q1 for s i and concatenate s with the time - used embedding vector when answering q1 i the time - interval embedding vector of the time used for doing questions and the answering - situation embedding vector to obtain the answering performance of s Combine the personalized question difficulty of q1 for s i and the answering performance to calculate the knowledge acquisition degree of s i Input the answering situation a of s and the time used for doing questions aq into the learning - efficiency calculation module, set the hyper - parameter a = 0.5, and calculate the learning efficiency l1 when answering q1. Use l1 as the weight to perform weighted addition on the knowledge acquisition degree i and the knowledge - mastery level to obtain the updated knowledge - mastery level i Combine the question embedding of q2 t and t calculate the prediction y2 of the answering situation when answering q2 for s. Set the threshold β = 0.5. When y2≥β, it is considered that s i answers q2 correctly. When y2 < β, it is considered that s answers q2 wrongly; then repeat the above steps to continue processing the answering record (q2, a2) until all the answering records in X are processed, and output the knowledge - mastery level of s Combine the question embedding of q2 and calculate s i the prediction y2 of the answering situation when answering q2. Set the threshold β = 0.5. When y2≥β, it is considered that s i answers q2 correctly. When y2 < β, it is considered that s i answers q2 wrongly; then repeat the above steps to continue processing the answering record (q2, a2) until all the answering records in X i are processed, and output the knowledge - mastery level of s i S3: Train the knowledge - tracking model;

[0216] Among them: The specific steps of S3 are as follows:

[0217] S31: Divide the student answering dataset into a training set and a test set according to a certain proportion, initialize the training parameters and model hyper - parameters, and input the data of the training set into the knowledge - tracking model for training;

[0218] S32: Use the cross - entropy loss function to calculate the loss value L between the student's actual answering situation a

[0219] and the predicted value y t The calculation method is as follows: t

[0220]

[0221] ​Among them, θ represents all trainable parameters and embeddings in the model, and λ θ is a regularization hyperparameter. When both the accuracy rate and loss value of the model tend to be stable, the training ends. After the training ends, the test set is used to test the model to confirm the actual prediction ability of the model;

[0222] This step includes but is not limited to the following embodiments:

[0223] The student answer dataset is divided into a training set and a test set in a ratio of 8:2. Set the training batch size batch_size = 32, the number of training epochs epoch = 200, the training learning rate is 0.002, the weight decay is 0.0001, and the dimension d kn , d qn , d qd , d h , d sq , d a , d aq , d bq of the embedding vector is 128; the data of the training set is input into the knowledge tracing model for training. After each batch of training ends, the loss value and accuracy rate are calculated. When both the accuracy rate and loss value of the model tend to be stable, the training ends; after the training ends, the test set is used to test the model to confirm the actual prediction ability of the model;

[0224] S4: Apply the model to recommend questions for students individually;

[0225] Among them: The specific steps of S4 are as follows:

[0226] S41: After the model training ends, input the answer sequence of a student into the model. After being processed by the model, output the student's knowledge mastery level;

[0227] S42: For the knowledge points that the student has a weak grasp of, select questions from the question list that the student has not answered and contain the knowledge points that the student has a weak grasp of to form a question sequence to recommend to the student for answering. When the student answers the recommended question sequence, the model simultaneously predicts the student's answering situation and updates the student's knowledge mastery level.

[0228] Based on the above, the present invention designs a knowledge tracing model that combines difficulty features and time correlation features. Among them, a personalized question difficulty calculation module for students is designed. By the correlation weight between the question and each knowledge point and the question difficulty, the difficulty feature information of the question embedding is enriched, and according to the knowledge mastery level of different students, a personalized question difficulty that conforms to the student's knowledge mastery level is given; a knowledge acquisition degree calculation module is designed. The student's answering performance is obtained by using the time taken by the student to do the questions, the time interval between doing the questions, and the answering situation, and then combined with the personalized question difficulty to calculate the student's knowledge acquisition degree, enriching the semantic information of the knowledge acquisition degree; a learning efficiency calculation module is designed. After subtracting the time taken by the student to do the questions from the average time taken to do the questions, and then passing it through the learning efficiency calculation function together with the student's answering situation, the learning efficiency of the student is analyzed from the answering situation and the time taken to do the questions, enhancing the interpretability of the learning efficiency calculation.

[0229] Although the present invention has been described above with reference to the embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the embodiments disclosed in the present invention can be combined with each other in any way, and the situations of these combinations are not exhaustively described in this specification only for the sake of saving space and resources. Therefore, the present invention is not limited to the specific embodiments disclosed in the text, but includes all technical solutions falling within the scope of the claims.

Claims

1. A knowledge tracing method combining difficulty features and time correlation features, characterized in that, The specific steps are as follows: S1: Screen student data, construct a student answer dataset, and preprocess the student data; S2: Construct a knowledge tracing model for predicting student answer situations and outputting the student's knowledge mastery level; Among them, the specific steps of S2 are as follows: S21: Calculate the time taken by the student to do the questions and the time interval between questions, and initialize the embedding vectors of the time taken to do the questions, the time interval between questions, the knowledge mastery level, and the answer situation; S22: Construct a question embedding in combination with the question difficulty, and calculate the personalized question difficulty that meets the student's knowledge mastery level; S23: Calculate the degree of knowledge acquisition after the student answers the question; S24: Calculate the learning efficiency and update the student's knowledge mastery level; S25: Predict the student's answer situation at time t+1 and determine whether to output the student's knowledge mastery level; S3: Train the knowledge tracing model; S4: Apply the model to recommend questions personalized for students; The specific steps of S22 are as follows: S221, build question embedding based on question difficulty: for s i The answer record at time t (q t ,a t ), using the problem difficulty embedding matrix QD to strengthen the problem q t Information representation, combined with q t Embedded question number The association weight between questions and knowledge points and question difficulty embedding Output q through the multi-layer perceptron t Title embedding The calculation is as follows: Among them, is the matrix concatenation operation, is the weight matrix, is the bias term; d q is the dimension of the question number embedding vector; N is the number of knowledge points and also the correlation weight between questions and knowledge points of the dimension; d qd is the dimension of the question difficulty embedding vector; d h is the dimension of the knowledge mastery level vector; S222, Calculate the difficulty of personalized questions that match the student's knowledge mastery level: At time t, input s i Before answering q t The knowledge mastery level To the personalized question difficulty calculation module, use q t Of the question embedding And Calculate and output q t For s i The difficulty of personalized questions d sq Is the dimension of SQ t The calculation method is as follows: Among them, is s i Answer q t The prior knowledge mastery level before, tanh is a non-linear activation function, and σ is the sigmoid activation function. is the weight matrix, is the bias term, used to calculate q t for s i The feature representation of the personalized question difficulty of used to calculate q t for s i The weight information of the personalized question difficulty of The specific steps of S23 are as follows: S231: Serialize s i Answer q t Embedding vector of the time taken to answer questions Embedding vector of the time interval between answering questions And embedding vector of the answering situation Used to represent s i Answering performance The calculation method is as follows: S232: Input the personalized question difficulty into the knowledge acquisition degree calculation module, and combine with to calculate and output the knowledge acquisition degree i of s The calculation method is as follows: Among them, is the weight matrix, is the bias term, used to extract the feature representation of the knowledge acquisition degree of s i ; used to calculate the weight information of the knowledge acquisition degree of s i ; d aq is the dimension of the embedding vector of the time used for answering questions; d bq is the dimension of the embedding vector of the time interval between answering questions; d a is the dimension of the embedding vector of the answering situation.

2. The knowledge tracing method combining difficulty features and time correlation features according to claim 1, wherein The specific steps of S1 are as follows: S11: Screen student data and classify and organize the student answer data; S12: Calculate the answer difficulty of the questions and the difficulty of the knowledge points, and initialize the embedding matrices of the question numbers, knowledge point numbers, and question difficulty.

3. The knowledge tracing method combining difficulty characteristics and time correlation characteristics according to claim 2, wherein The specific steps of S11 are as follows: S111: Obtain the initial student answer data, and each answer record in the answer data contains student number, answer situation, start time of doing the questions, end time of doing the questions, question number, knowledge point number, knowledge point name, question type, teacher number information; S112: Classify and organize the information involved in the student answer data, and after classification and organization, obtain the student list, question list, and knowledge point list as follows: Student list S = {s1, s2, …, s I}, s i ∈ S represents a student, whose value is the student serial number, and I is the number of students; Question list Q = {q1, q2, …, q J}, q j ∈Q represents a question, whose value is the question serial number; J is the number of questions; Knowledge point list K = {k1, k2, …, k N}, where k n ∈K represents a knowledge point, whose value is the knowledge point serial number, and N is the number of knowledge points; During the process of doing questions, student s i 's answer record sequence is represented as X i = {(q1, a1), (q2, a2), …, (q T , a T )}, where (q t , a t ) is s i 's answer record at time t, q t ∈ Q is the question that s i answers at time t, a t is s i 's answer situation regarding q t at time t, and its value is 0 or 1, 1 indicates a correct answer, 0 indicates a wrong answer, T is the number of questions that s i answers. Record s i 's start answering time and end answering time in the answering time sequence QT i = {(st1, et1), (st2, et2), …, (st T , et T )}, where (st t , et t ) is s i 's start answering time st t and end answering time et t at time t when answering q t .

4. A knowledge tracing method combining difficulty characteristics and time correlation characteristics according to claim 3, characterized in that The specific steps of S12 are as follows: S121. Calculate the difficulty of answering the question. For any question q in the question list Q j , the calculation method of the difficulty of answering the question is as follows: S1211: Count the number of all students who answered q j , calculate the number of students who answered q j correctly, and calculate the proportion of the number of students who answered q j correctly in the total number of students who answered q qs . Then multiply it by the question answering difficulty level C j to obtain the question answering difficulty of q . The calculation formula is as follows: Among them, qs j is the difficulty of answering question q j , is the number of students answering question q j , is the total number of students who answer question q j correctly, and C qs is the set difficulty level for answering questions. S1212: Repeat S121 to calculate the question answering difficulty of all questions in Q, and obtain the question answering difficulty list QS = {qs1, qs2, …, qs J}, where qs j ∈QS represents the question answering difficulty of q j ; S122, calculate the difficulty of knowledge points. For any knowledge point k in the knowledge point list K n the calculation method of its difficulty is as follows: S1221: Count the total number of students who answered questions containing k n , then calculate the proportion of students who answered questions containing k n correctly, and multiply it by the knowledge point difficulty level C kd to obtain the difficulty of k n . The calculation formula is as follows: Among them, kd n is the knowledge point difficulty of k n , is the number of students who answered the questions containing k n , is the number of students who answered the questions containing k n correctly, and C kd is the set knowledge point difficulty level; S1222: Repeat S122 to calculate the knowledge point difficulty of all knowledge points in K, and obtain the knowledge point difficulty list KD = {kd1, kd2, …, kd N}, where kd n ∈KD represents the knowledge point difficulty of k n ∈K; S123, initialize the embedding matrices of the knowledge point numbers and question numbers: S1231, Initialize the knowledge point serial number embedding matrix KN: Represent the serial number of knowledge point k n as a one-hot encoding O(k n ) ∈ R 1×N , and set an embedding matrix to map k n into an embedding vector Among them, d k is the dimension of the embedding vector of the knowledge point serial number. Calculate the embedding vectors of all knowledge point serial numbers and then merge them into a knowledge point serial number embedding matrix Each d kn -dimensional row vector in KN represents a knowledge point; S1232, Initialize the question number embedding matrix QN: Represent the serial number of question q j as a one-hot encoding O(q j ) ∈ R 1×J , and set an embedding matrix to map q j into an embedding vector Among them, d q is the dimension of the question number embedding vector. Calculate the embedding vectors of all question numbers and then combine them into a question number embedding matrix Each d qn -dimensional row vector in QN represents a question; S124, calculate the question difficulty and initialize the question difficulty embedding matrix QD: S1241, Calculate the problem difficulty: First, construct the matrix QK to describe the association between problems and knowledge points. Each row in QK corresponds to a problem, and each column in QK corresponds to a knowledge point. When q j is associated with k n , the value at the j-th row and n-th column in QK is 1, otherwise it is 0. At the same time, to calculate the degree of association between each knowledge point and the problem, use QK to calculate the association weight between q j and each knowledge point Among them, KN is the knowledge point serial number embedding matrix; QK j,: represents the j-th row of QK, which is used to describe the j correlation status between q and each knowledge point; the weight j represents the degree of correlation between q j and each knowledge point; then, considering the problem difficulty, the difficulty of answering the question and the difficulty of the knowledge points included in the question, through the correlation weight between q j and each knowledge point, the knowledge point difficulty KD is weighted, and then the weighted knowledge point difficulty is added to the difficulty of answering the question to obtain the j problem difficulty qd of q, and the calculation formula is as follows: Among them, QS j is the i-th bit of the question answering difficulty list QS, representing q j 's question answering difficulty, is 's transpose; S1242, Initialize the difficulty embedding matrix QD: For q j whose difficulty level is qd j represented as one-hot encoding Set an embedding matrix C qs as the difficulty level of the answer to the question set in S1211. Map qd j into an embedding vector where q qd is the dimension of the problem difficulty embedding vector. Calculate the embedding vectors of all problem difficulties and then combine them into a problem difficulty embedding matrix Each d in QD qd dimensional row vector represents a problem difficulty.

5. A knowledge tracking method combining difficulty characteristics and time correlation characteristics according to claim 1, characterized in that The specific steps of S21 are as follows: S211, calculate the time taken to do the questions and the time interval between questions: S2111: Obtain the answer record sequence X i of student s i ={(q1,a1),(q2,a2),…,(q T ,a T )} and the answer time sequence QT i ={(st1,et1),(st2,et2),…,(st T ,et T )}. Subtract the start answering time st i from the end answering time et t when s t answers question q t to get the answering time aq i when s t answers question q t . The calculation formula is as follows: aq t = et t - st t S2112: Calculate X i The time taken to answer all questions in i is obtained as s i 's question answering time list AQ T = {aq1, aq2, …, aq t}, where aq i ∈AQ i represents the time taken to answer question q t ; S2113: For s i Answer any question q t , according to q calculated in S122 t The association weight with each knowledge point Get with q t The knowledge point k with the highest degree of correlation t , in s i Answer Q t The answer record before is {(q1,a1),(q2,a2),…,(q t-1 ,a t-1 )}∈X i Search and k t The highest correlation with q t The question with the shortest time interval to finish the question q' t , will s i Answer Q t End of question time et t Subtract answer q' t End of question time et' t , get s i Answer Q t The time interval between questions bq t , if q' t If it does not exist, then bq t =0, the calculation formula is as follows: S2114: Calculate X i For all the questions in i , calculate the time intervals between answering them to obtain the time interval list BQ i ={bq1, bq2, …, bq T}, where bq t ∈BQ i represents the time interval for answering question q i ; t ​ S212, initialize the embedding vectors of the time taken to do the questions and the time interval between questions: S2121: Set s i Answer q t The time taken to answer question q t is represented as a one-hot encoding O(aq t ) ∈ R 1×T , and set an embedding matrix Map aq t to an embedding vector d aq is the dimension of the embedding vector for the time taken to answer questions: S2122: Calculate X i The time-embedded vectors for answering all questions in i are merged into a time-embedded matrix for answering questions S2123: Set s i Answer q t The time interval bq for answering questions t Represented as one-hot encoding O(bq t ) ∈ R 1×t , Set an embedding matrix Map dq t to an embedding vector d bq is the dimension of the embedding vector of the question answering time interval: S2124: Calculate X i The time interval embedding vectors for all questions in i are merged into a time interval embedding matrix S213, initialize the embedding vectors of the knowledge mastery level and the answer situation: S2131: Create a vector used to represent s i 's knowledge mastery level at time t, where d h is the dimension of; S2132: For s i Answer q t Regarding the answering situation a t , map a t to an embedding vector d a is the dimension of the answering situation embedding vector. When the element of is 1, it means the answer is correct; when the element is 0, it means the answer is wrong.

6. A knowledge tracing method combining difficulty characteristics and time correlation characteristics according to claim 5, characterized in that, The specific steps of S24 are as follows: S241: Design a learning efficiency calculation module to calculate the learning efficiency of student s i based on their answering situation and time spent on questions. When student s i answers question q t , it is considered that the learning efficiency of answering correctly is higher than that of answering wrongly, and the learning efficiency of spending a moderate amount of time on questions is higher than that of spending too long or too short a time. The calculation method of learning efficiency l t is as follows: where a is a hyperparameter, σ is the sigmoid activation function, a t is the student's answer q t in terms of the answering situation, aq t is the student s i answering q t the time taken to answer the question, is the average time taken by all students to answer q t to answer the question; S242: Design a knowledge mastery level update module, and input the degree of knowledge acquisition and the knowledge mastery level to the knowledge mastery level update module, and use l t as the weight to and perform weighted addition and output the updated knowledge mastery level The calculation method for updating the knowledge mastery level is as follows: Among them, is s i knowledge mastery level before answering q t is s i knowledge mastery level after answering q t ​​ 7. A knowledge tracing method combining difficulty characteristics and time correlation characteristics according to claim 6, characterized in that The specific steps of S25 are as follows: S251: For s i The knowledge mastery level at time t+1, combined with q t+1 The question embedding of We can obtain s i Answer q t+1 The prediction y of the answering situation t+1 The calculation method is as follows: Among them, y t+1 is the predicted s i answer q t+1 correct probability. When y t+1 ≥β, it is considered that s i answers correctly. When y t+1 <β, it is considered that s i answers incorrectly. β is a manually set threshold; S252: If t + 1 ≤ T, jump to S22 to continue processing the answer record (q t+1 ,a t+1 ); otherwise, output the knowledge mastery level of s i ​ 8. A knowledge tracing method combining difficulty characteristics and time correlation characteristics according to claim 1, characterized in that, The specific steps of S3 are as follows: S31: Divide the student answer dataset into a training set and a test set according to a certain proportion, initialize the training parameters and model hyperparameters, and input the data of the training set into the knowledge tracing model for training; S32: Calculate the actual answering situation a of the student using the cross-entropy loss function t and the predicted value y t The loss value L between them is calculated as follows: where θ represents all trainable parameters and embeddings in the model, and λ θ is a regularization hyperparameter. When both the accuracy rate and loss value of the model tend to be stable, the training ends. After the training ends, the test set is used to test the model to confirm the actual prediction ability of the model; The specific steps of S4 are as follows: S41: After the model training is completed, input the answer sequence of a student into the model, and after being processed by the model, output the student's knowledge mastery level; S42: For the knowledge points that the student has a weak grasp of, select the questions that the student has not answered and contain the knowledge points that the student has a weak grasp of from the question list to form a question sequence for the student to recommend to answer. When the student answers the recommended question sequence, the model simultaneously predicts the student's answer situation and updates the student's knowledge mastery level.

Citation Information

Patent Citations

  • Comprehensive knowledge tracking method for enhancing difficulty of test questions

    CN115795015A

  • Deep knowledge tracing with transformers

    US20210390873A1