A Knowledge Tracing Method Based on Meta-Path and Attention Mechanism

Through methods based on metapaths and attention mechanisms, path instances are generated and embedded and fusion are solved, and the problem of failing to effectively utilize metapaths and higher-order semantic information in the prior art is improved, and the accuracy and efficiency of knowledge tracking are improved.

CN116821497BActive Publication Date: 2025-07-29NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310780118.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-07-29
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

The existing knowledge tracking methods fail to effectively utilize the metapathic information and advanced semantic information between students, exercises and knowledge points, resulting in poor evaluation of students' knowledge proficiency.

Method used

Using a method based on metapathic path and attention mechanism, path examples are generated through a random walk algorithm, convolutional neural network and attention mechanism are used to embed students, exercises and knowledge points, and high-order semantic information is extracted in combination with long and short-term memory networks to construct fusion embedding of knowledge points, and evaluate students' knowledge proficiency.

Benefits of technology

It improves the evaluation effect of students' knowledge proficiency, solves the problem that metapathic paths and higher-order semantic information are not used in existing methods, and improves the accuracy and efficiency of knowledge tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116821497B_ABST
    Figure CN116821497B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of knowledge tracing, and proposes a knowledge tracing method based on meta-path and attention mechanism. After the collected student learning behavior data is processed, according to the defined meta-path, the random walk method is used to generate path instances, and a convolutional neural network is constructed to obtain the semantic information of the path instances. The embeddings of students and exercises are obtained by using the Embedding method. The embedding of knowledge points is obtained by combining Word2Vec and long short-term memory network, and then the fused embedding of knowledge points is obtained according to the self-attention mechanism. The final embeddings of students and exercises are obtained by combining the attention mechanism and the semantic information of path instances respectively. The obtained embeddings are concatenated with the historical problem-solving sequences covered by the dataset and then input into the long short-term memory network to obtain the knowledge point proficiency. Based on the traditional technology, the present invention strengthens the evaluation of knowledge point proficiency by using the semantic information embedding of the meta-path among students, exercises and knowledge points and the fused embedding of knowledge points, and improves the effect of knowledge tracing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of knowledge tracing, and particularly relates to a knowledge tracing method based on meta-path and attention mechanism. Background Art

[0002] With the deep integration of the Internet and the education industry, education characterized by intelligence, digitization, and personalization has given rise to a large number of new education models such as mobile learning and intelligent learning. While the drawbacks of the traditional education model have been effectively alleviated, online learning technology has also tended to mature. However, in the existing teaching system, there is a lack of in-depth mining and analysis of the cognitive structures of students and groups, teaching evaluation data, and effective research on the recommendation methods of exercises and learning paths. Therefore, in order to enhance the ability of autonomous learning, improve the knowledge state level of learners themselves, and improve teachers' overall understanding of students and related systems, knowledge tracing has emerged as the times require.

[0003] Knowledge tracing is a technology that models the knowledge mastery of students based on the student behavior sequence according to the past answering situations of students, so as to predict the degree of students' knowledge mastery. At the same time, it is also the key point of the adaptive education system. Whether it is the planning of the learning path or the recommendation of personalized exercises, the first step is to accurately predict the proficiency of students in knowledge. Knowledge tracing originated from the cognitive diagnosis task in the computer field. Initially, its purpose was to distinguish the cognitive diagnosis task in the psychological aspect. However, with the development of computer technology, knowledge tracing has evolved into an indispensable research direction in the field of educational big data and is quite different from the original task of cognitive diagnosis. First of all, different from cognitive diagnosis, knowledge tracing models all consider time series, and most of them are sequence models, such as those proposed based on sequence models such as long short-term memory network (LSTM), memory-augmented network, hidden Markov model, etc. Even though a small number of knowledge tracing models are non-sequence models, they also consider time series factors, especially by extracting time series features to consider time series.

[0004] According to existing technologies, knowledge tracing methods can be broadly categorized into three main categories: factor analysis-based knowledge tracing using feature engineering to extract artificial features, such as Item Response Theory (IRT), Performance Factors Analysis (PFA), Knowledge Tracing Machines (KTM), and the Difficulty, Ability, Skill, and Student Skill Practice History (DASH) model from Lindsey (DAS3H). Knowledge tracing based on probabilistic graphical models, such as the Bayesian Knowledge Tracing Model (BKT), Knowledgeable Prompt-tuning (KPT), and the fuzzy cognitive diagnosis framework (Fuzzy CDF), and deep learning-based knowledge tracing, such as the Exercise-Enhanced Recurrent Neural Network (EERNN), EKT, the Self-Attentive Model for Knowledge Tracing (SAKT), Relation-Aware Self-Attention for Knowledge Tracing (RKT), and Context-Aware Attentive Knowledge Tracing (AKT). Currently, all three models are used in industry, each with its own advantages and disadvantages. Factor analysis-based models effectively utilize the knowledge point information contained in exercises through forms such as the Q matrix, improving model performance while also intuitively reflecting users' mastery of the knowledge points. However, these models lack scalability and struggle to utilize other static information, such as the interval between exercises and the exercise text, wasting a significant amount of semantic information. Probabilistic graph-based models mostly incorporate educational theory and are highly interpretable. They perform well when the probability graphs are constructed appropriately, but perform poorly when they are not. These models rely on experts' understanding of the educational environment. Deep learning-based models, which emerged with the advent of heterogeneous data, achieve excellent results by effectively processing heterogeneous data. However, the uninterpretability of deep learning technology limits its application in knowledge tracking. Current deep learning-based knowledge tracking applications also neglect the impact of the explicit information of meta-paths between heterogeneous data and the high-level information of knowledge points on predicting students' proficiency in knowledge points, which impacts model effectiveness. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a knowledge tracing method based on meta - paths and attention mechanism, including:

[0006] Step 1: Collect students' learning behavior data, define invalid data and clean it, organize the cleaned data into a data set, define meta - path P and generate path instances p according to the meta - path using the random walk algorithm;

[0007] Step 1.1: Collect students' learning behavior data, including three types of entities: students U, exercises I, and knowledge points KC;

[0008] Step 1.2: Define data with the number of submissions (the number of times a student does the same exercise) less than α times, the correct rate (the correct rate of a student doing the same exercise) lower than γ, and the number of records (the number of times an exercise is practiced) less than λ as invalid data. Clean the invalid data in the students' learning behavior data and organize the cleaned data into a data set;

[0009] Step 1.3: Define meta - paths P containing different semantic information, including: UIUI and UIKCI. The meta - path UIUI means that students who have done the same exercise i a have also done other exercises i b , and the meta - path UIKCI means that students who have done exercise i a have also done other exercises i of the same knowledge point b , and generate path instances p based on the data set according to the meta - path P using the random walk algorithm;

[0010] Step 2: Construct a 1×1 convolutional neural network CNN, embed the path instances p, and then use the max - pooling operation to obtain the embedding l of the meta - path P ;

[0011] Step 2.1: Use the path instances of students' learning behavior entities as nodes, and use the Embedding method to obtain node embeddings. Then connect the node embeddings to obtain the embedding matrix X formed by connecting the node embeddings p , X p ∈R L×d , where R is the set of real numbers, L is the length of the path instance, and d is the embedding dimension of the entity;

[0012] Step 2.2: Construct a 1×1 convolutional neural network model CNN composed of a single convolutional layer and a single pooling layer. Input the embedding matrix X formed by connecting the node embeddings p into the convolutional neural network to obtain the embedding matrix t p , and the formula is as follows:

[0013] t p =CNN(X p ,β) (1)

[0014] where t p is the embedding matrix of the path instance p, CNN(·) is the convolutional neural network, and X p is the embedding matrix formed by connecting node embeddings, and β represents the number of input channels, output channels, number of convolutional kernels, stride, and padding in the convolutional neural network;

[0015] Step 2.3: Randomly select the embeddings of s path instances in the path instance embedding matrix t p and use the max-pooling operation to derive the embedding l of the meta-path P from the embeddings of the s path instances P ;

[0016]

[0017] where l P is the embedding of the meta-path P, max-pooling(·) is the max-pooling operation, are the embeddings of the selected s path instances;

[0018] Step 3: Use the numpy library to perform one-hot representation operations on the students and exercises in the dataset, and use the Embedding method in the torch framework to convert the one-hot representations of the students and exercises into low-dimensional embeddings to obtain the student embedding X u and the question embedding Y i ;

[0019] Step 4: Use the attention mechanism to obtain the meta-path P attention weight θ representing the interaction of different students U, exercises I, and knowledge points KC u,i,P , and obtain the embedding l of the semantic information of the path instance u→i ;

[0020] Step 4.1: Use the Linear function of the fully connected layer in the torch framework to input the dimensional information of the student, exercise, and meta-path embeddings into the Xavier initialization function to obtain the weight matrix and bias vector of the fully connected layer;

[0021] Step 4.2: Adopt a two-layer architecture to obtain the meta-path attention weights of the interaction of different students U, exercises I, and knowledge points KC on the meta-paths UIUI and UIKCI respectively. The formula is as follows:

[0022]

[0023]

[0024] In Equation (3), is the meta-path attention weight of the interaction between students and exercises in the first layer, and W u(1) is the weight matrix for the first - layer students, \(W\) i (1) is the weight matrix for the first - layer exercises, \(W\) P (1) is the weight matrix for the first - layer meta - path embedding \(l\) P is the bias vector for the first - layer, and \(f(\cdot)\) is the ReLU function; (1) In Equation (4),

[0025] is the meta - path attention weight for the interaction between the second - layer students and exercises, \(W\) (2) is the weight matrix for the second - layer, \(b\) (2) is the bias vector for the second - layer;

[0026] Step 4.3: Use the softmax function to normalize the meta - path attention weight for the interaction between the second - layer students and exercises, and obtain the meta - path attention weights \(\theta\) of UIUI and UIKCI respectively. The formula is as follows: u,i,P

[0027]

[0028] In the formula, \(\theta\) u,i,P is the meta - path attention weight, \(M\) u→i is the set of meta - paths, and \(\exp()\) is the exponential function with the natural constant \(e\) as the base;

[0029] Step 4.4: Perform a dot - product on the embedding \(l\) P of the meta - path \(P\) obtained in Step 2.3 and the meta - path attention weight \(\theta\) u,i,P to obtain the embedding \(l\) u→i of the path - instance semantic information. The formula is as follows:

[0030]

[0031] Step 5: Use the embedding \(l\) u→i of the path - instance semantic information to obtain the attention weight vectors of students and exercises respectively, and obtain the final embeddings of students and exercises by element - wise multiplication and

[0032] Step 5.1: Use the Linear function in the torch framework to input the dimension information of the student embedding, exercise embedding, and path - instance semantic information embedding into the Xavier initialization function to obtain the weight matrix and bias vector of the fully - connected layer specific to the student and exercise embeddings;

[0033] Step 5.2: Use the Linear function to obtain the attention weight vectors of students and exercises. The formula is as follows:​​

[0034] θ u = f(W′ u X u + W′ u→i l u→i + b′ u ) (7)

[0035] θ i = f(W′ i Y i + W′ u→i l u→i + b′ i ) (8)

[0036] In Equation (7), θ u is the attention weight vector of the student, W′ u is the weight matrix of the fully connected layer specific to the student, W′ u→i is the weight matrix of the path instance semantic information embedding, b′ u is the bias vector specific to the student;

[0037] In Equation (8), θ i is the attention weight vector of the exercise, W′ i is the weight matrix of the fully connected layer specific to the exercise, b′ i is the bias vector specific to the exercise;

[0038] Step 5.3: Multiply the student and exercise attention weight vectors element-wise with the student embedding and exercise embedding to obtain the final embeddings of the student and the exercise and

[0039]

[0040]

[0041] Step 6: Use the word2vec algorithm and the long short-term memory network to embed the knowledge point KC to obtain the embedding vector kc of the knowledge point. Use the method of traversing the data to construct the exercise knowledge point matrix, and then use the self-attention mechanism to obtain the fused embedding vector of all the knowledge points contained in the same exercise

[0042] Step 6.1: Convert the pre-trained word vector file based on the glove model into a word2vec word vector file. Retrieve the word vectors corresponding to the knowledge points in the dataset from the word2vec word vector file, and record the retrieved word vectors to construct the word vector matrix v, v ∈ R d×h , where h is the number of knowledge points;

[0043] Step 6.2: Input the word vector matrix v into the long short-term memory network, and output the embedded vector kc of the knowledge point, where kc ∈ R h×1 ;

[0044] Step 6.3: Use the method of traversing data to count all the knowledge points included in each exercise in the dataset, and construct the exercise knowledge point matrix Q, where Q ∈ R e×g , e is the number of all knowledge points included in the same exercise, and g is the number of exercises;

[0045] Step 6.4: Use the self-attention mechanism based on the knowledge point embedded vector kc and the exercise knowledge point matrix Q to obtain the fused embedded vector of all the knowledge points included in the same exercise Obtain the high-order semantic information between knowledge points. The formula is as follows:

[0046]

[0047] In the formula, is the fused embedded vector of all the knowledge points included in the same exercise, SA(·) is the self-attention mechanism, attention pooling() is the attention pooling operation, and q u,i is the student embedding X u and the exercise embedding Y i concatenated to obtain the query matrix;

[0048] Step 7: Combine the path instance semantic information embedding l u→i , the knowledge point fusion embedding the final embedding of the student and the exercise and together with the historical problem-solving sequence S covered by the dataset, to obtain the interaction information between the student, the exercise, the knowledge point, and the path instance semantic information Substitute it into the long short-term memory network to obtain the relevant knowledge proficiency of the student

[0049] Step 7.1: Organize all the exercises done by the same student in the dataset into the historical problem-solving sequence S, and combine the path instance semantic information embedding l u→i , the knowledge point fusion embedding the final embedding of the student and the exercise and the historical problem-solving sequence S to obtain the interaction information between the student, the exercise, the knowledge point, and the path instance semantic information The combination formula is as follows:

[0050]

[0051] Step 7.2: The interaction information Input into the long short-term memory network (LSTM) to obtain the proficiency of the student's knowledge points The formula is as follows:

[0052]

[0053]

[0054] In Equation (13), h t+1 is the hidden state at time t + 1, and c t+1 is the cell state at time t + 1, h t is the hidden state at time t, and c t is the cell state at time t, and α t+1 represents the input gate, forget gate, and output gate parameters;

[0055] In Equation (14), W r is the initial weight matrix automatically generated by the RELU function, and b r is the initial bias vector automatically generated by the RELU function.

[0056] The beneficial effects of the present invention are:

[0057] The present invention proposes a knowledge tracing method based on meta-path and attention mechanism, which uses meta-path and attention mechanism to evaluate the proficiency of students' knowledge points.

[0058] First, a knowledge tracing method based on meta-path and attention mechanism uses meta-path, combines rich path instances among students, exercises, and knowledge points, improves the embeddings of students and exercises, and solves the problem that the meta-path information between nodes is not utilized when embedding students and exercises.

[0059] Second, a knowledge tracing method based on meta-path and attention mechanism constructs a knowledge point fusion embedding, extracts the high-order semantic information between knowledge points, and uses it for the generation of interaction information, solves the problem that the existing knowledge tracing methods lack the utilization of high-order semantic information between knowledge points, and improves the evaluation effect of students' knowledge proficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is the general flow chart of the present invention.

[0061] Figure 2 is the schematic diagram of the knowledge tracing process in the present invention.

[0062] Figure 3 is the schematic diagram of the meta-path in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0063] In order to make the advantages and technical solutions of the present invention more clear and explicit, the present invention is further described below with reference to the accompanying drawings and specific implementation examples.

[0064] This paper proposes a knowledge tracking method based on meta-path and attention mechanism. The knowledge tracking process is as follows: Figure 2 As shown, the overall flow chart of the present invention is as follows Figure 1 The specific steps are as follows:

[0065] Step 1: Collect student learning behavior data, define invalid data and perform data cleaning, organize the cleaned data into a data set, define the meta-path P and use the random walk algorithm to generate path instances p based on the meta-path;

[0066] Step 1.1: Collect student learning behavior data, including three types of entities: student U, exercise I, and knowledge point KC;

[0067] This embodiment uses the public data of student learning behavior history records compiled by Carnegie Learning, and selects data from August 30, 2005 to November 30, 2006 as standby data. The standby data includes information such as students, exercises, knowledge points, submission time, and correctness.

[0068] Step 1.2: Define data with fewer than α submission times (set to 30 in this invention), a correct rate lower than γ (set to 10% in this invention), and fewer than λ records (set to 30 in this invention) as invalid data. Clean the invalid data in the student learning behavior data and organize the cleaned data into a data set.

[0069] After cleaning, the data set for use in this embodiment includes: 135 students, 2720 exercises, 76 knowledge points, 12347 records, and an accuracy rate of 0.73.

[0070] Because defining a long metapath may introduce noisy semantics, the present invention defines two short metapaths of at most four steps.

[0071] Step 1.3: Define a meta-path P containing different semantic information, such as Figure 3 As shown, including: UIUI and UIKCI, the meta-path UIUI indicates that the same exercise i has been done a Students also did other exercises b , meta-path UIKCI means that the student has done exercise i a I also did other exercises on the same knowledge point. b , according to the meta-path P, a random walk algorithm is used to generate a path instance p based on the data set;

[0072] The generated path instance is used to aggregate information. The path instance is essentially a sequence of entity nodes. In order to embed such a sequence of nodes into a low-dimensional vector, the present invention uses the common convolutional neural network CNN to process the path instance sequence.

[0073] Step 2: Construct a 1×1 convolutional neural network CNN, embed the path instance p, and then use the max-pooling operation to obtain the embedding l of the meta-path P ;

[0074] Step 2.1: Use the student learning behavior entity of the path instance as a node, and use the Embedding method to obtain the node embedding. Then connect the node embeddings to obtain the embedding matrix X formed by connecting the node embeddings p , X p ∈R L×d , where R is the set of real numbers, L is the length of the path instance, and d is the embedding dimension of the entity;

[0075] Step 2.2: Construct a convolutional neural network model CNN composed of a single convolutional layer and a single pooling layer, and input the embedding matrix X formed by connecting the node embeddings p into the convolutional neural network to obtain the embedding matrix t p , and the formula is as follows:

[0076] t p =CNN(X p ,β) (1)

[0077] In the formula, t p is the embedding matrix of the path instance p, CNN(·) is the convolutional neural network, X p is the embedding matrix formed by connecting the node embeddings, and β is the number of input channels, output channels, number of convolutional kernels, stride, and padding in the convolutional neural network;

[0078] Step 2.3: Randomly select the embeddings of s path instances in the path instance embedding matrix t p , and use the max-pooling operation to derive the embedding l of the meta-path P from the embeddings of the s path instances P ;

[0079]

[0080] In the formula, l P is the embedding of the meta-path P, max-pooling(·) is the max-pooling operation, and {t p} s p=1 are the embeddings of the selected s path instances;

[0081] The present invention selects the embeddings of s path instances, aiming to capture important dimensional features from multiple path instances.

[0082] Step 3: Use the numpy library to obtain one-hot representations of students and exercises in the dataset, and use the Embedding method in the torch framework to convert the one-hot representations of students and exercises into low-dimensional embeddings, obtaining student embeddings X u and problem embeddings Y i ;

[0083] The present invention sets an embedding layer to convert the one-hot representations of students and exercises into low-dimensional embeddings. Formally, given a student-exercise pair <u, i>, let k u ∈R |u|×1 and j i ∈R |i|×1 represent the one-hot representations of students and exercises. The embedding layer corresponds to two parameter matrices K∈R |U|×d and J∈R |I|×d , which store the latent factors of students and exercises respectively. d is the dimensionality size of the student and exercise embeddings, and |U| and |I| are the total numbers of students and exercises respectively. The embedding operation is implemented as follows:

[0084] X u =K T *k u (15)

[0085] Y i =J T *j i (16)

[0086] Step 4: Use the attention mechanism to obtain the meta-path P attention weights θ representing the interactions of different students U, exercises I, and knowledge points KC u,i,P , obtaining the embedding l of the path instance semantic information u→i ;

[0087] Step 4.1: Use the Linear function of the fully connected layer in the torch framework to input the dimensionality information of the student, exercise, and meta-path embeddings into the Xavier initialization function to obtain the weight matrix and bias vector of the fully connected layer;

[0088] Step 4.2: Adopt a two-layer architecture to obtain the meta-path attention weights of the interactions of different students U, exercises I, and knowledge points KC on the meta-paths UIUI and UIKCI respectively. The formula is as follows:

[0089]

[0090]

[0091] In formula (3), is the meta-path attention weight for the first-layer student-exercise interaction, and W u (1) is the weight matrix of the first-layer student embedding, and W i (1) is the weight matrix of the first-layer exercise embedding, and W P (1) is the weight matrix of the first-layer meta-path embedding l P , b (1) is the bias vector of the first layer, and f(·) is the ReLU function;

[0092] In formula (4), is the meta-path attention weight for the second-layer student-exercise interaction, and W (2) is the weight matrix of the second layer, and b (2) is the bias vector of the second layer;

[0093] Step 4.3: Use the softmax function to normalize the meta-path attention weight of the second-layer student-exercise interaction, and obtain the meta-path attention weights θ u,i,P of UIUI and UIKCI respectively. The formula is as follows:

[0094]

[0095] In the formula, θ u,i,P is the meta-path attention weight, M u→i is the set of meta-paths, and exp() is the exponential function with the natural constant e as the base;

[0096] Step 4.4: Perform a dot product on the embedding l P of the meta-path P obtained in Step 2.3 and the meta-path attention weight θ u,i,P to obtain the embedding l u→i of the path instance semantic information. The formula is as follows:

[0097]

[0098] Step 5: Use the embedding l u→i of the path instance semantic information to respectively obtain the attention weight vectors of the students and exercises, and obtain the final embeddings and

[0099] by element-wise multiplication. Step 5.1: Use the Linear function of the fully connected layer in the torch framework to input the dimensional information of the student embedding, exercise embedding, and path instance semantic information embedding into the Xavier initialization function to obtain the weight matrix and bias vector of the fully connected layer specific to the student and exercise embeddings.

[0100] The weight matrix and bias vector specific to the meta-path UIUI or the meta-path UIKCI can be obtained using the Xavier initialization function.

[0101] Step 5.2: Use the Linear function to obtain the attention weight vectors of the student and the exercise, with the formula as follows:

[0102] θ u = f(W' u X u + W' u→i l u→i + b' u ) (7)

[0103] θ i = f(W' i Y i + W' u→i l u→i + b' i ) (8)

[0104] In Equation (7), θ u is the attention weight vector of the student, W' u is the weight matrix of the fully connected layer specific to the student, W' u→i is the weight matrix of the semantic information embedding of the path instance, and b' u is the bias vector specific to the student;

[0105] In Equation (8), θ i is the attention weight vector of the exercise, W' i is the weight matrix of the fully connected layer specific to the exercise, and b' i is the bias vector specific to the exercise;

[0106] Step 5.3: Multiply the student and exercise attention weight vectors element-wise with the student embedding and exercise embedding to obtain the final embeddings of the student and the exercise and

[0107]

[0108]

[0109] Step 6: Use the word2vec algorithm and the long short-term memory network to embed the knowledge point KC to obtain the embedding vector kc of the knowledge point. Use the method of traversing the data to construct the exercise-knowledge point matrix, and then use the self-attention mechanism to obtain the fused embedding vector of all the knowledge points included in the same exercise

[0110] Step 6.1: Convert the word vector file pre-trained based on the glove model into a word2vec word vector file, retrieve the word vector corresponding to the knowledge point in the dataset in the word2vec word vector file, record the retrieved word vectors and construct the word vector matrix v, v∈R d×h , h is the number of knowledge points;

[0111] Step 6.2: Input the word vector matrix v into the long short-term memory network and output the embedding vector kc of the knowledge point, kc∈R h×1 , the formula is as follows:

[0112]

[0113] In the formula, c, h, are related parameters;

[0114] An exercise may contain multiple knowledge points. In order to effectively utilize the potential connection information between related knowledge points, the present invention adopts the self-attention mechanism Self Attention (SA).

[0115] Step 6.3: Use the method of traversing data to count all the knowledge points contained in each exercise in the data set and construct the exercise knowledge point matrix Q, Q∈R e×g , e is the number of all knowledge points contained in the same exercise, g is the number of exercises;

[0116] Step 6.4: Based on the knowledge point embedding vector kc and the exercise knowledge point matrix Q, use the self-attention mechanism to obtain the fused embedding vector of all knowledge points contained in the same exercise The high-level semantic information between knowledge points is obtained, and the formula is as follows:

[0117]

[0118] In the formula, is the fusion embedding vector of all knowledge points contained in the same exercise, SA(·) is the self-attention mechanism, attention pooling() is the attention pooling operation, q u,i Embed X for students u and the problem embedding Y i The concatenated query matrix;

[0119] Step 7: Embed path instance semantic information into l u→i , knowledge point fusion embedding Final embedding of students and exercises and Together with the historical question sequence S covered by the dataset, we can obtain the interactive information between students, questions, knowledge points and path instance semantic information Substitute it into the long short-term memory network to obtain the relevant knowledge proficiency of the student

[0120] Step 7.1: Organize all the exercises done by the same student in the dataset into a historical exercise sequence S, and embed the path instance semantic information l u→i and the knowledge point fusion embedding The final embedding of the student and the exercise and Combine with the historical exercise sequence S to obtain the interaction information among the student, exercise, knowledge point, and path instance semantic information The combination formula is as follows

[0121]

[0122] Step 7.2: Input the interaction information into the long short-term memory network LSTM to obtain the knowledge proficiency of the student The formula is as follows

[0123]

[0124]

[0125] In formula (13), h t+1 is the hidden state at time t + 1, c t+1 is the cell state at time t + 1, h t is the hidden state at time t, c t is the cell state at time t, and α t+1 represents the input gate, forget gate, and output gate parameters

[0126] In formula (14), W r is the initial weight matrix automatically generated by the RELU function, and b r is the initial bias vector automatically generated by the RELU function

[0127] Step 8: Establish an experimental environment and use the dataset to verify the effectiveness of a knowledge tracing method based on meta-path and attention mechanism

[0128] Establish the experimental environment as shown in Table 1. In this embodiment, the effectiveness of the method is evaluated by two indicators, AUC (the area under the ROC curve) and MAE (mean absolute error). The experimental results show that a knowledge tracing method based on meta-path and attention mechanism (MAKT) is superior to other methods on the dataset adopted in the present invention. The comparison results are shown in Table 2

[0129] Table 1 Experimental environment of the present invention

[0130]

[0131] In the knowledge tracing method based on factor analysis, the KTM method performs the best. This method introduces a decomposition machine, which can more conveniently integrate various auxiliary static information, and can also effectively alleviate the problem that it is difficult to train the model due to the excessive sparsity of cross features. In addition, the innovation of separately counting the number of correct and incorrect knowledge points related to specific exercises proposed by this method is also effective. Although it is relatively rare, it undoubtedly improves the effect to a certain extent.

[0132] Knowledge tracing methods based on deep learning (AKT, DKT) are superior to knowledge tracing methods based on factor analysis (AFM, PFA, and KTM) in most cases. It is worth noting that the DKT method has good results. This method uses LSTM, so DKT combines the historical problem-solving situations of students with the recent problem-solving situations to determine the current knowledge proficiency. The design of the forgetting gate also conforms to the realistic feature of the Ebbinghaus forgetting curve from the perspective of educational psychology. At the same time, this means the superiority of deep neural networks in capturing the interaction relationship between users and exercises for tracing.

[0133] The present invention also considers a variant method of MAKT (MAKT-M) and finds that the variant method is inferior to the complete method in overall performance. The experimental results show that the co-attention mechanism can better utilize the meta-path-based neighbor information to comprehensively improve the representation of nodes. First, specific interactions determine the importance of each meta-path. Second, the meta-path provides important explicit information for the interaction between students and exercises, which has a potential impact on the learned representations of users, exercises, and knowledge points. Ignoring this impact may prevent further performance improvement. In addition, although MAKT-M achieves competitive performance compared to the baseline, it is still significantly worse than the complete MAKT.

[0134] Table 2 Comparison of the experimental results of the present invention and the prior art

[0135] model AUC MAE AFM 0.660 0.341 PFA 0.665 0.347 KTM 0.665 0.344 AKT 0.687 0.342 DKT 0.690 0.336 MAKT-M 0.716 0.337 MAKT 0.729 0.339

Claims

1. A knowledge tracing method based on meta-path and attention mechanism, characterized in that Including: Step 1: Collect students' learning behavior data, define invalid data and perform data cleaning on it, organize the cleaned data into a dataset, define a meta-path P, and generate path instances p using a random walk algorithm based on the meta-path; Step 2: Construct a 1×1 convolutional neural network CNN, embed the path instance p, and then use the max pooling operation to obtain the embedding l of the meta-path P ; Step 3: Use the numpy library to perform one-hot representation operations on the students and exercises in the dataset, and use the Embedding method in the torch framework to convert the one-hot representations of the students and exercises into low-dimensional embeddings, obtaining student embeddings X u and question embeddings Y i ; Step 4: Use the attention mechanism to obtain the attention weight θ of the meta-path P representing the interaction of different students U, exercises I, and knowledge points KC u,i,P , and obtain the embedding l of the semantic information of the path instance u→i ; Step 5: Use the embedding l of the path instance semantic information u→i Respectively obtain the attention weight vectors of the students and exercises, and obtain the final embeddings of the students and exercises by element-wise multiplication and Step 6: Embed the knowledge point KC using the word2vec algorithm and the long short-term memory network to obtain the embedded vector kc of the knowledge point. Use the method of traversing the data to construct the exercise knowledge point matrix, and then use the self-attention mechanism to obtain the fused embedded vector of all the knowledge points included in the same exercise Step 7: Embed the path instance semantic information into l u→i , knowledge point fusion embedding The final embedding of the student and the exercise and Together with the historical problem-solving sequence S covered by the dataset, the interaction information between the student, the exercise, the knowledge point, and the path instance semantic information is obtained Substitute it into the long short-term memory network to obtain the relevant knowledge proficiency of the student 2. The knowledge tracking method based on meta-path and attention mechanism according to claim 1, wherein Step 1 includes: Step 1.1: Collect students' learning behavior data, including three types of entities: students U, exercises I, and knowledge points KC; Step 1.2: Define data with the number of submissions less than α times, the correct rate lower than γ, and the number of records less than λ as invalid data, perform data cleaning on the invalid data in the students' learning behavior data, and use the linear interpolation algorithm to fill in the missing values for the missing human activity data in the organized dataset; Step 1.3: Define meta-paths P containing different semantic information, including: UIUI and UIKCI. The meta-path UIUI indicates that students who have done the same exercise i a have also done other exercises i b , and the meta-path UIKCI indicates that students have done exercise i a and have also done other exercises i on the same knowledge point b . According to the meta-path P, use the random walk algorithm to generate path instances p based on the dataset.

3. A knowledge tracking method based on meta-path and attention mechanism according to claim 1, characterized in that Step 2 includes: Step 2.1: Take the path instance student learning behavior entity as a node, use the Embedding method to obtain the node embedding, and then connect the node embeddings to obtain the embedding matrix X formed by connecting the node embeddings p , X p ∈R L×d , where R is the set of real numbers, L is the length of the path instance, and d is the embedding dimension of the entity; Step 2.2: Construct a 1×1 convolutional neural network model CNN consisting of a single convolutional layer and a single pooling layer, and input the embedding matrix X formed by connecting nodes p into the convolutional neural network to obtain the embedding matrix t p , and the formula is as follows: t p = CNN(X p ,β) where t p is the embedding matrix of the path instance p, CNN(·) is the convolutional neural network, and X p is the embedding matrix formed by connecting node embeddings, and β is the number of input channels, output channels, number of convolutional kernels, stride, and padding in the convolutional neural network; Step 2.3: Randomly select the embeddings of s path instances in the path instance embedding matrix t p and derive the embedding l of the meta-path P from the embeddings of the s path instances using the max-pooling operation P ; where l P is the embedding of the meta-path P, and max-pooling(·) is the max pooling operation, is the embedding of the selected s path instances.

4. A knowledge tracking method based on meta-path and attention mechanism according to claim 1, characterized in that Step 4 includes: Step 4.1: Use the Linear function in the torch framework to input the dimensional information of the embeddings of students, exercises, and meta-paths into the Xavier initialization function to obtain the weight matrix and bias vector of the fully connected layer; Step 4.2: Adopt a two-layer architecture to obtain the meta-path attention weights for the interactions of different students U, exercises I, and knowledge points KC on the meta-paths UIUI and UIKCI respectively. The formula is as follows: In the formula, is the meta-path attention weight for the first-layer student-exercise interaction, W u (1) is the weight matrix of the first-layer students, W i (1) is the weight matrix of the first-layer exercises, W P (1) is the weight matrix of the first-layer meta-path embedding l P and b (1) is the bias vector of the first layer, and f(·) is the ReLU function; Wherein, is the meta-path attention weight for the interaction between students and exercises at the second layer, W (2) is the weight matrix at the second layer, and b (2) is the bias vector at the second layer; Step 4.3: Use the softmax function to normalize the meta-path attention weights of the second-layer student-exercise interaction to obtain the meta-path attention weights θ of UIUI and UIKCI respectively u,i,P , and the formula is as follows: where θ u,i,P is the meta-path attention weight, M u→i is the meta-path set, and exp() is the exponential function with the natural constant e as the base; Step 4.4: Dot product the embedding \(l\) of the meta-path \(P\) P and the meta-path attention weight \(\theta\) u,i,P to obtain the embedding \(l\) of the path instance semantic information u→i , and the formula is as follows:

5. A knowledge tracing method based on meta-path and attention mechanism according to claim 1, characterized in that Step 5 includes: Step 5.1: Use the Linear function in the torch framework to input the dimensional information of the student embedding, exercise embedding, and path instance semantic information embedding into the Xavier initialization function to obtain the weight matrix and bias vector of the fully connected layer specific to the student and exercise embeddings; Step 5.2: Use the Linear function to obtain the attention weight vectors of students and exercises. The formula is as follows: θ u = f(W′ u X u + W′ u→i l u→i + b′ u ) θ i = f(W′ i Y i + W′ u→i l u→i + b′ i ) where, θ u is the attention weight vector of the student, W′ u is the fully connected layer weight matrix specific to the student, W′ u→i is the weight matrix for path instance semantic information embedding, b′ u is the bias vector specific to the student; where θ i is the attention weight vector of the exercise, W′ i is the fully connected layer weight matrix specific to the exercise, and b′ i is the bias vector specific to the exercise; Step 5.3: Element-wise multiply the student and exercise attention weight vectors with the student embedding and exercise embedding to obtain the final embeddings of the student and the exercise and 6. The knowledge tracking method based on meta-path and attention mechanism according to claim 1, characterized in that Step 6 includes: Step 6.1: Convert the pre-trained word vector file based on the glove model into a word2vec word vector file, retrieve the word vectors corresponding to the knowledge points in the dataset in the word2vec word vector file, and record the retrieved word vectors to construct a word vector matrix v, where v ∈ R d×h , and h is the number of knowledge points; Step 6.2: Input the word vector matrix v into the long short-term memory network, and output the embedded vector kc of the knowledge point, where kc ∈ R h ×1 ; Step 6.3: Use the method of traversing data to count all the knowledge points included in each exercise in the dataset, and construct an exercise knowledge point matrix Q, Q ∈ R e×g , where e is the number of all knowledge points included in the same exercise, and g is the number of exercises; Step 6.4: Based on the knowledge point embedding vector kc and the exercise knowledge point matrix Q, use the self-attention mechanism to obtain the fused embedding vector of all the knowledge points contained in the same exercise The high-order semantic information between knowledge points is obtained, and the formula is as follows: In the formula, is the fusion embedding vector of all knowledge points included in the same exercise. SA(·) is the self-attention mechanism, attention pooling() is the attention pooling operation, and q u,i is the student embedding X u and the exercise embedding Y i The query matrix obtained by concatenation.

7. A knowledge tracking method based on meta-path and attention mechanism according to claim 1, characterized in that, Step 7 includes: Step 7.1: Organize all the exercises done by the same student in the dataset into a historical exercise sequence S, and embed the path instance semantic information l u→i , the knowledge point fusion embedding The final embedding of the student and the exercise and Combine with the historical exercise sequence S to obtain the interaction information between the student, the exercise, the knowledge point, and the path instance semantic information The combination formula is as follows: Step 7.2: Input the interaction information into the long short-term memory network (LSTM) to obtain the proficiency of the student's knowledge points The formula is as follows: where h t+1 is the hidden state at time t + 1, c t+1 is the cell state at time t + 1, h t is the hidden state at time t, c t is the cell state at time t, and α t+1 represents the input gate, forget gate, and output gate parameters; Where, W r is the initialization weight matrix automatically generated by the RELU function, and b r is the initialization bias vector automatically generated by the RELU function.

Citation Information

Patent Citations

  • Knowledge tracking method based on heterogeneous graph learning and fusion learning participation state

    CN113947262A

  • Semantic representation model-based text classification method and apparatus, and computer device

    WO2021051503A1