A personalized knowledge tracking method based on knowledge graph recommendation
By introducing personalized recommendation methods based on knowledge graphs and hierarchical structure update mechanisms into the knowledge tracking model, the shortcomings in personalization and interpretation of the deep learning knowledge tracking model are solved, and more efficient knowledge tracking and model performance improvement are achieved.
Patent Information
- Application Number
- CN202211462249.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-11-22
AI Technical Summary
The existing deep learning knowledge tracking model performs poorly in personalization and fails to effectively utilize features such as topics, disciplines, and pre-order exercises, resulting in insufficient models in interpretability and data sparse issues.
Using a personalized knowledge tracking method based on knowledge graph recommendation, we use user-question interaction matrix and question knowledge graph to obtain dense feature vectors using BERT embedding and free text embedding, and use knowledge graph recommendation method to capture higher-order structures and personalized preferences, and introduce a hierarchical structure update mechanism in long-term and short-term memory networks.
It improves the personalized tracking ability, interpretability and versatility of the model, alleviates the problem of data sparseness, and improves the model performance through a more reasonable knowledge state update mechanism.
Smart Images

Figure CN115827968B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of online learning and educational data mining, and in particular to a knowledge tracing method based on knowledge graph recommendation. Background Art
[0002] The knowledge tracing task was proposed by Corbett et al. in 1995. Since the rapid development of the Internet and the widespread use of online education, knowledge tracing has attracted more and more researchers. It is very similar to cognitive diagnosis. The ultimate goal of both is to evaluate and predict learners' learning ability. The difference is that cognitive diagnosis considers cognitive factors, and most of them consider the cognitive factor of knowledge points. The limitations of cognitive diagnosis are obvious: the amount of data cannot be too large. Whether the number of knowledge points is large or the types of cognitive factors are many, cognitive diagnosis is difficult to calculate. However, knowledge tracing performs better in the big data environment.
[0003] The current knowledge tracing models can be generally divided into two categories: traditional methods and deep learning methods. And the traditional methods can be further subdivided into knowledge tracing methods based on probability models and knowledge tracing methods based on logistic regression. The deep learning methods can be further subdivided into knowledge tracing methods without graph structures and knowledge tracing methods with graph structures.
[0004] For traditional methods, a typical method based on probability models - Bayesian Knowledge Tracing (BKT) uses the Hidden Markov Model (HMM) to model the knowledge state of learners. Another traditional method, the knowledge tracing method based on logistic regression such as Performance Factors Analysis (PFA), modifies the Learning Factors Analysis (LFA) to make it suitable for adaptive learning.
[0005] For deep learning methods, Deep Knowledge Tracing (DKT) was the first model to introduce deep learning methods in the field of knowledge tracing. The deep learning method it used, Long Short-Term Memory (LSTM), brought about a huge improvement in performance compared to traditional methods. Subsequent derivative methods can be divided into two categories according to whether they have a graph structure. Deep learning methods without a graph structure, such as Dynamic Key-Value Memory Networks for Knowledge Tracing (DKVMN) and A Self-Attentive model for Knowledge Tracing (SAKT), all perform well in terms of performance. However, due to the low interpretability of neural networks, knowledge tracing models with a graph structure have emerged. For example, the Graph-based Interaction Model for knowledge (GIKT) constructs a bipartite graph of questions and knowledge points, and then uses a Graph Convolutional Network (GCN) to extract high-order information of questions and knowledge points for input into the knowledge tracing model. Graph-based Knowledge Tracing (Modeling Student Proficiency Using Graph Neural Network, GKT) constructs a graph of all the knowledge points involved in the questions according to the precedence relationship, and then uses a Graph Neural Network (GNN) to process and obtain the prediction results. Knowledge structure-enhanced graph representation learning model for attentive knowledge tracing (KSGKT) considers both aspects: the knowledge point structure graph and the question-knowledge point bipartite graph, and uses the knowledge point structure graph to enhance the question-knowledge point bipartite graph to obtain more information, improve interpretability, and alleviate the data sparsity problem. The above deep learning methods with a graph structure prove that for the knowledge tracing task, the addition of a graph structure not only performs excellently in terms of performance, but also improves the interpretability of the model and alleviates the data sparsity problem compared to traditional methods.
[0006] However, the existing problem is that deep learning methods consistently exclude user features and do not focus on characterizing personalization. Knowledge tracing models with graph structures are more difficult to capture personalization due to their irregularity, and only consider two aspects: questions and knowledge points, without considering features such as topics, disciplines, and prior exercises. Our method incorporates user-question interaction information to capture personalized features and constructs a knowledge graph by considering all important attributes of questions, solving the problem that knowledge tracing models cannot model personalization, and further improving the interpretability of the model and alleviating the data sparsity problem. Additionally, a common problem with current knowledge tracing models is that information such as questions in the dataset is represented using numbers instead of free text with semantic information. The disadvantages of this approach include not only the loss of semantic information but also the difficulty in obtaining complete information when looking for relationships between questions and other attributes such as question knowledge points. Another benefit of using free text is that it can help address the general problem of cross-disciplinary models. Even a model trained for the mathematics discipline can be used for the physics discipline because the knowledge of most disciplines can be decomposed into questions, knowledge points, topics, etc., only the specific content is different.
[0007] Another deficiency of deep learning methods is that: the update mechanism of the knowledge state h t (i.e., the hidden state in the network) updates all state information based on the latest input, ignoring the hierarchical structure information in the knowledge state. For example, if a learner has currently mastered the four arithmetic operations, equations, and plane geometry, and the current input is a mixed operation of addition, subtraction, multiplication, and division, then obviously the level of plane geometry is higher than the current input and should be retained, while the state of the four arithmetic operations should be replaced with the current input. Some parts of the equations also involve the knowledge of the current input, so they should be updated by a certain mechanism.
[0008] Therefore, our invention well fills the gap that deep knowledge tracing methods cannot achieve personalization, improves the graph structure and the knowledge state update mechanism, and further enhances the interpretability, generality of the model and alleviates the data sparsity problem. Summary of the Invention
[0009] The object of the present invention is to address the problems and unresolved tasks existing in the existing deep learning knowledge tracing models and the graph structures used, and propose a personalized knowledge tracing method based on knowledge graph recommendation.
[0010] To achieve the above object, the present invention adopts the following technical solutions:
[0011] A personalized knowledge tracing method based on knowledge graph recommendation, the steps are as follows:
[0012] In the first step, obtain publicly available practice data;
[0013] Step 2: Construct the user-question interaction matrix and the question knowledge graph, i.e., the dataset required by the model;
[0014] Construct the learner-question interaction matrix and the question knowledge graph from the obtained public practice data as the experimental dataset. The specific operations are as follows:
[0015] In the learner-question interaction matrix, for each learner, if there is a positive performance (i.e., answering correctly) on a certain question, it is 1; otherwise, it is 0. The question knowledge graph consists of many triples (head entity, relation, tail entity). The head entity is the question, and the head entity + relation locates the tail entity. The tail entity is the set of all attribute entities. The attribute entity in the following text is the tail entity. For example, (mixed operation_1 + related topic -> four arithmetic operations), the tail entity "four arithmetic operations" is the attribute entity of the question "mixed operation_1".
[0016] Step 3: Obtain the dense feature vectors of the text representation in the question knowledge graph using BERT embedding and free text embedding;
[0017] For the text representation in the question knowledge graph, use BERT embedding. BERT embedding is used as a feature extractor to adjust the text attributes of entities. Free text embedding directly embeds the text of a single word and uses mean pooling for the text of multiple words. Finally, connect the two embedding representations together as the final feature representation, which is the dense feature vector input in Step 4;
[0018] Step 4: Use the knowledge graph recommendation method for the dense feature vectors obtained in Step 3 to obtain the high-order structure and neighborhood information of each entity, and distinguish the importance of each relationship and attribute to the user;
[0019] The purpose of using knowledge graph recommendation is to capture the high-order structure and neighborhood information of the nodes in the knowledge graph while also capturing the personalized preferences of learners, which are ignored in previous knowledge tracing models.
[0020] Consider the candidate pair of user u and question entity q. Let N(q) denote the set of entities directly connected to the question entity q (i.e., the set of attribute entities of the question), and r q,e denote the relationship between the question entity q and the attribute entity e. Use the inner product function g: to calculate the score between the user and (relationship + attribute entity) to obtain the different degrees of influence of each attribute entity and each relationship on each user. The calculation method is as follows:
[0021]
[0022] where u ∈ R d and e ∈ R d and r ∈ R dThey are the feature representations of user u, attribute entity e, and relationship r respectively, where d represents the dimension. Generally speaking It characterizes the importance of relationship r + attribute entity e to user u. For example, among the attribute entities with the same relationship of "topic", a user may have more potential interest and positive performance in the practice of "four arithmetic operations" in his historical activities, indicating that the user has a good grasp of the relevant knowledge of "four arithmetic operations". If the question to be predicted also belongs to "four arithmetic operations", then he has a high probability of getting it right. Another user may be more proficient in "plane geometry" and have more positive performance and potential interest in the knowledge in this area.
[0023] Another benefit of obtaining personalized preferences is that if the platform wants to design an exercise recommendation function, then this part can be used for an adjustment: when recommending new knowledge, mix in the questions that the learner is interested in to avoid the learner losing interest in practicing.
[0024] To characterize the topological proximity structure of question entity q, we calculate the linear combination of the neighborhood of question entity q
[0025]
[0026] is the normalized user and (relationship + attribute entity) score, that is, the normalized The calculation method is as follows:
[0027]
[0028] e represents the attribute entity, and exp(·) represents the exponential function with the natural constant as the base. When calculating the neighborhood representation of the entity, the user and (relationship + attribute entity) score acts as a personalized filter because we aggregate the neighborhoods with deviations from these user-specific scores.
[0029] In the actual knowledge graph, the number of entities in N(q) may vary greatly for different questions because some questions only involve one knowledge point, while some questions involve multiple knowledge points. Questions with richer knowledge points also have more attributes such as "topic" and "previous practice". To ensure the stability and effectiveness of each batch of calculations, we use a fixed number of neighbors for each question. We calculate the neighborhood representation of question entity q as where S q is defined as follows:
[0030]
[0031] N(q) is the set of entities directly connected to entity q, and K is a configurable constant. We will find the optimal value of K in the subsequent experiments. is the set of secondary entities after fusing K neighbors. We call S q is the single-layer receptive field of entity q, because the final representation of q is determined by its h-layer receptive field.
[0032] Through the neighborhood aggregation operation, the local neighboring structure is successfully captured and stored in each entity. The neighbors are weighted by relying on the connection relationship and the scores of specific users, which not only represents the semantic information of the knowledge graph but also represents the user's personalized interest in the relationship. The neighborhood definition of a given entity can also be extended hierarchically to multi-hop to simulate high-order entity dependencies and capture the user's potential long-distance interests.
[0033] The last step of knowledge graph recommendation is to aggregate the topic entity representation q and its neighborhood representation into a dense vector.
[0034] Step 5: Input the information obtained in the fourth step and the performance of the joint learner on the corresponding topic into a long short-term memory network with a hierarchical structure, namely the ordered neuron long short-term memory network, to model and update the learner's knowledge state;
[0035] For each topic, which corresponds to a specific time point t, we concatenate the topic representation processed in step (3) and its corresponding right or wrong answer a t , and apply a non-linear transformation to obtain the learning activity representation at the current time point t:
[0036]
[0037] RELU represents the non-linear transformation, W and b represent the weight matrix and bias respectively, and concat(·) represents concatenation.
[0038] To capture the long-term dependencies of the entire historical activities and simulate the formation of different knowledge states during the learner's learning process, we use an improved LSTM - ON-LSTM (Ordered Neurons: Integrating Tree Structures into Recurrent Neural Networks, ON-LSTM) to jointly model the input learning activity representation. ON-LSTM integrates the hierarchical structure of the tree into the LSTM, enabling the improved network to learn the hierarchical structure information unsupervised, which has a huge impact on tasks such as knowledge discovery and translation. The definition is as follows:
[0039] First is the standard input gate i t , output gate o t and forget gate f t :
[0040] i t= σ(W i X t + U i h t-1 + b i )
[0041] o t = σ(W o X t + U o h t-1 + b o )
[0042] f t = σ(W f X t + U f h t-1 + b f )
[0043]
[0044] Calculate the level of the historical knowledge state h t-1 :
[0045]
[0046]
[0047]
[0048] Calculate the level of the current input x t :
[0049]
[0050]
[0051]
[0052] Calculate the level l of the knowledge state h t-1 and the level l of the current input x his : t The intersection ω now : t :
[0053]
[0054] When performing state update: The intersecting part of the two is updated using LSTM, and the part higher than max(l now , l his ) is retained, and the part lower than min(l now , l his ) is directly replaced with the current input, that is
[0055]
[0056] Finally, update the knowledge state to obtain the latest knowledge state h t :
[0057]
[0058] Among them, W represents the weight matrix of the current input X t U represents the weight matrix of the current knowledge state h t-1 b represents the bias, and σ(·) refers to the sigmoid function. Use the hyperbolic tangent function tanh(·) to calculate the candidate value for subsequent determination of the forgotten content. represents the intersection operation.
[0059] Step 6: Select appropriate information to interact with the current input and knowledge state to obtain the prediction result;
[0060] For the current input and the current knowledge state h t-1 , we select two question entities q1, q2 with the same knowledge points as and the two attribute entities e1, e2 with the highest scores for interaction. π is the score of the two, R is the ranking of the scores, and k indicates selecting the top k.
[0061]
[0062]
[0063] π is the score of the two, R is the ranking of the scores, and k indicates selecting the top k.
[0064] and simulate that the learner practices the question at time point t Use the following method to obtain the prediction result p t :
[0065] α i,j = Softmax(W T concat(f i , f j ) + b)
[0066]
[0067] Among them, f = <·> represents the interaction, and W and b represent the weight and bias respectively. represents aggregation neighborhood attribute e j embedding, and g represents the inner product function.
[0068] During the training process, gradient descent is used to minimize the cross-entropy loss between the predicted value p t and the true value a t :
[0069]
[0070] Advantages of the present invention: The present invention uses the knowledge graph recommendation idea to assist in solving the problem that the knowledge tracing model with a heterogeneous graph structure can achieve personalized tracing, and the free text and its encoding used also assist in solving the general problem of cross-science models. The constructed knowledge graph and the user question interaction information matrix further alleviate the data sparsity problem, and the more reasonable hierarchical knowledge state update mechanism adopted enables the model to further obtain better performance. Brief Description of the Drawings
[0071] Figure 1 is the overall flowchart of the present invention.
[0072] Figure 2 is the overall framework diagram of the present invention
[0073] Figure 3 is a schematic diagram of BERT embedding and free text embedding.
[0074] Figure 4 is a schematic diagram of single-node aggregation.
[0075] Figure 5 is a two-layer receptive field diagram for a given entity, where K = 2.
[0076] Figure 6 is a schematic diagram of learning behavior modeling, knowledge state update, and prediction. Detailed Description of the Invention
[0077] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0078] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0079] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. All data acquisition in this embodiment is based on compliance with laws and regulations, platform privacy protection, and user consent, and is a legal application of data.
[0080] An embodiment of the present invention provides a personalized knowledge tracking method based on knowledge graph recommendation, as Figure 1 and 2 shown, including the following steps:
[0081] (1) Obtain a number of learning data of the learner
[0082] Obtain data from the publicly available educational dataset Assistment2009, including learner identity information, exercise information, answering situation, knowledge points involved in the questions, chapters, etc. Prepare to process the data to obtain two data files: one is the scoring file in the knowledge graph recommendation task, and the other is the knowledge graph of the learner. The ratio of the training, validation, and test datasets is 6:2:2, and the statistical information of this dataset is shown in Table 1.
[0083] Table 1 Statistical Information
[0084]
[0085] (2) Obtain the scoring data file
[0086] In the scoring file, according to the performance of each learner, if they successfully answer a question, we consider their performance on this question to be positive, and record this interaction (user, question, 1). If they fail to answer the question successfully, we consider their performance on this question to be negative, and record this interaction (user, question, 0).
[0087] (3) Construct the knowledge graph KG
[0088] In the knowledge graph KG = {H, R, T}, H represents the head entity, T represents the tail entity, and R represents the relationship between H and R. Therefore, we have H + R → T. According to the characteristics of the ASSISTment2009 dataset, we design and select four tuples: (question, type, question type), (question, involved knowledge point, knowledge point), (question, belonging question set, question set), (question, belonging template, template).
[0089] (4) Construct the knowledge graph recommendation layer
[0090] The role of this layer is to effectively capture the correlation between items by mining relevant attributes on the KG. Inspired by the knowledge graph recommendation task, in order to automatically discover the high-order structural information and semantic information of the KG, we first design an embedding layer to process free text. Among them, BERT embedding is used as a pre-trained language model, and free text embedding is used to obtain the dense feature vectors of entities. Sample the neighborhood of each entity in the KG as their receptive field, and then merge the neighborhood information with the bias when calculating the representation of a given entity. Such a design has two advantages: 1) The neighborhood information can be stored in the structure of each entity; 2) The neighborhood is weighted by relying on the connection relationship and the scores of specific learners, which represents both the semantic information of the KG and the personalized interests of learners in relationships. For example, some learners like to do multiple-choice questions, while some learners prefer calculation questions. To avoid the huge space overhead caused by too large a neighborhood, we sample the neighborhood of each entity and use a fixed size N = 4 as its receptive field, which makes the cost within an acceptable range. The neighborhood of a given entity can also be extended to multiple hops to simulate high-order entity dependencies and capture the potential interests of learners.
[0091] For each triple, calculate the score between learner l and relationship r + tail entity e using the inner product to obtain the different degrees of influence of each attribute entity and its relationship on each user:
[0092]
[0093] (6) Characterize the topological proximity structure of the head entity (i.e., the question entity) q
[0094] We calculate the linear combination of its neighborhood:
[0095]
[0096] is the normalized learner and (relationship + tail entity) score, and the calculation method is as follows:
[0097]
[0098] Let \(e\) denote the attribute entity. When calculating the neighborhood representation of an entity, the learner (relation + tail entity) score acts as a personalized filter because we aggregate neighborhoods that are biased towards the learner-specific scores, as Figure 4 shown.
[0099] (7) Neighborhood Aggregation
[0100] We calculate the neighborhood representation of the head entity \(q\) as where \(S\) q is defined as follows:
[0101]
[0102] \(N(q)\) is the set of entities directly connected to entity \(q\), \(K\) is a configurable constant, and we will find the optimal value of \(K\) in subsequent experiments. is the set of sub-entities after fusing \(K\) neighbors. We call \(S\) q the single-layer receptive field of entity \(q\) because the final representation of \(q\) is determined by its \(n\)-layer receptive field. Figure 5 shows the two-layer receptive field diagram of a given entity, where \(K = 2\).
[0103] We tried using three aggregators \(agg: R\) d × \(R\) d → \(R\) d , and a suitable aggregator can be selected according to the characteristics of the dataset during the specific implementation process. The three aggregators are as follows:
[0104] · Sum aggregator: Sum the vectors and then use a non-linear transformation
[0105]
[0106] · Concat aggregator: Concatenate the vectors and then use a non-linear transformation
[0107]
[0108] · Neighbor aggregator: Directly use the neighborhood representation of entity \(q\) to represent entity \(q\)
[0109]
[0110] Since the selected dataset in this example is not large, the Sum aggregator is selected to aggregate entity \(q\) and its neighborhood representation into a dense vector.
[0111] (8) Modeling the Learner's Knowledge State
[0112] Extract the problem representation with high-order neighborhood information from the KG after aggregation In the combined scoring file, the performance a of the learner on this question is represented using a non - linear transformation for the current input:
[0113]
[0114] Input the data into the network for modeling.
[0115] First, calculate the level of the historical knowledge state h t-1 :
[0116]
[0117]
[0118]
[0119]
[0120] and the level of the current input x t :
[0121]
[0122]
[0123]
[0124] Update the knowledge state according to the following rules:
[0125]
[0126]
[0127]
[0128] where represents the intersection.
[0129] In this network, we set the learning rate to 0.001 and the learning decay rate to 0.92 to make the learning process smoother and more in line with the knowledge acquisition process of the human brain.
[0130] (9) Prediction
[0131] For the current input and the current knowledge state h t-1 , we select two question entities q1, q2 with the same knowledge points as and the two attribute entities e1, e2 with the highest scores for interaction.
[0132]
[0133]
[0134] π is the score of the two, R is the ranking of the scores, and k indicates selecting the top k ones.
[0135] and simulates that the learner practices the questions at time point t Obtain the prediction result p using the following method t :
[0136] α i,j = Softmax(W T concat(f i , f j ) + b)
[0137]
[0138] where represents aggregating the neighborhood attribute e j embedding, and g represents the inner product function.
[0139] (10) Optimization method
[0140] During the model training process, the parameters are updated using gradient descent. The specific approach is to minimize the cross-entropy loss between the prediction result p t and the true result a t . The loss function is as follows:
[0141]
[0142] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A personalized knowledge tracking method based on knowledge graph recommendation, characterized in that The steps are as follows: First, obtain publicly available practice data; Second: Construct a user-question interaction matrix and a question knowledge graph, which are the datasets required by the model; Construct a learner-question interaction matrix and a question knowledge graph from the obtained publicly available practice data, which are the experimental datasets. The specific operations are as follows: In the learner-question interaction matrix, for each learner, if there is a positive performance (i.e., answering correctly) on a certain question, it is 1, otherwise it is 0; the question knowledge graph consists of many triples (head entity, relation, tail entity). The head entity is the question, and the head entity + relation locates the tail entity. The tail entity is the whole attribute entity, and the attribute entity in the following text is the tail entity; Third: Use BERT embedding and free text embedding for the text representation in the question knowledge graph to obtain its dense feature vectors; Use BERT embedding for the text representation in the question knowledge graph. BERT embedding is used as a feature extractor to adjust the text attributes of entities; free text embedding directly embeds the text of a single word and uses mean pooling for the text of multiple words; finally, the two embedding representations are concatenated together as the final feature representation, which is the dense feature vector input in the fourth step; Fourth: Use the knowledge graph recommendation method for the dense feature vectors obtained in the third step to obtain the high-order structure and neighborhood information of each entity, and distinguish the importance of each relationship and attribute to the user; Using the knowledge graph recommendation can capture the high-order structure and neighborhood information of the nodes in the knowledge graph while also capturing the personalized preferences of the learners; Consider the candidate pair of user \(u\) and question entity \(q\). Let \(N(q)\) denote the set of entities directly connected to the question entity \(q\), that is, the set of attribute entities of the question; \(r\) q,e represents the relationship between the question entity \(q\) and the attribute entity \(e\). Use the inner product function \(g\): to calculate the score between the user and (relationship + attribute entity), so as to obtain the different degrees of influence of each attribute entity and each relationship on each user. The calculation method is as follows: where \(u\in\mathbb{R}\) d , \(e\in\mathbb{R}\) d and \(r\in\mathbb{R}\) d are the feature representations of user \(u\), attribute entity \(e\) and relationship \(r\) respectively, and \(d\) represents the dimension; characterizes the importance of relationship \(r\) + attribute entity \(e\) to user \(u\); To characterize the topological proximity structure of the problem entity q, calculate the linear combination of the neighborhood of the problem entity q is the normalized user sum (relationship + attribute entity) score, i.e., the normalized The calculation method is as follows: Among them, e represents the attribute entity, and exp(·) represents the exponential function with the natural constant as the base; when calculating the neighborhood representation of the attribute entity, the user and (relation + attribute entity) scores act as personalized filters because they aggregate the neighborhoods with biases for these user-specific scores; To ensure stable and effective computation for each batch, a fixed number of neighbors are used for each problem; the neighborhood representation of problem entity q is calculated as where S q is defined as follows: N(q) is the set of entities directly connected to entity q, and K is a configurable constant for which the optimal value will be found in subsequent experiments. is the secondary entity set after fusing K neighbors; S q is the single-layer receptive field of the topic entity q, and the final representation of q is determined by its h-layer receptive field. Through the neighborhood aggregation operation, the local proximity structure is successfully captured and stored in each entity. The neighbors are weighted by relying on the connection relationship and the scores specific to the user, which represents both the semantic information of the knowledge graph and the user's personalized interest in the relationship; the neighborhood definition of a given entity can also be hierarchically extended to multi-hop to simulate the high-order entity dependencies and capture the user's potential long-distance interests; The last step of knowledge graph recommendation is to aggregate the topic entity representation q and its neighborhood representation into a dense vector; Fifth: Input the information obtained in the fourth step and the learner's performance on the corresponding questions into a long short-term memory network with a hierarchical structure, that is, an ordered neuron long short-term memory network, to model and update the learner's knowledge state; For each question, which corresponds to a specific time point t, the concatenated representation of the questions after the fourth step of processing and its corresponding correct / incorrect value a t , and a non-linear transformation is applied to obtain the representation of the learning activity at the current time point t: Among them, RELU represents a non-linear transformation, W and b represent the weight matrix and bias respectively, and concat(·) represents concatenation; To capture the long-term dependencies of the entire historical activities and simulate the formation of different knowledge states during the learner's learning process, use the improved LSTM, that is, ON-LSTM, to jointly model the input learning activity representation; ON-LSTM integrates the hierarchical structure of the tree into the LSTM, enabling the improved network to learn the hierarchical structure information without supervision. The definition is as follows: First is the standard input gate i t , output gate o t and forget gate f t : i t = σ(W i X t + U i h t-1 + b i ) o t = σ(W o X t + U o h t-1 + b o ) f t = σ(W f X t + U f h t-1 + b f ) Calculate the historical knowledge state h t-1 Level of: Calculate the current input x t at the level of: Calculate the knowledge state h t-1 at level l his and the current input x t at level l now to obtain the intersection ω t : When performing status update: the intersecting part of the two is updated using LSTM, and the part higher than max(l now , l hi ) is retained, and the part lower than min(l now , l his ) is directly replaced with the current input, that is Finally, update the knowledge state to obtain the latest knowledge state h t : Among them, W represents the weight matrix of the current input X t and U represents the weight matrix of the current knowledge state h t-1 b represents the bias, and σ(·) refers to the sigmoid function. The hyperbolic tangent function tanh(·) is used to calculate the candidate value for subsequent determination of the forgotten content. denotes the intersection operation; Sixth: Select appropriate information to interact with the current input and knowledge state to obtain the prediction result; For the current input and the current state of knowledge h t-1 , select and two question entities q1 and q2 with the same knowledge points and the two attribute entities e1 and e2 with the highest scores for interaction; π is the score of the two, R is the ranking of the scores, and k indicates selecting the top k; and simulates the learner practicing the questions at time point t obtain the prediction result p using the following method t : α i,j = Softmax(W T concat(f i , f j ) + b) Among them, f = <·> represents interaction, W and b represent weights and biases respectively, represents aggregation neighborhood attribute e j embedding, g represents an inner product function; Minimize the cross - entropy loss between the predicted value p t and the true value a t during the training process:
2. The personalized knowledge tracking method based on knowledge graph recommendation according to claim 1, wherein In the first step, the data includes the questions done by all learners, whether the answers are correct, and the question attributes. The question attributes include the knowledge points of the questions, the themes they belong to, the disciplines, and the previous exercises. The data is in the form of free text or other forms of data represented by their numbers.
3. The personalized knowledge tracking method based on knowledge graph recommendation according to claim 2, wherein In the user-question interaction matrix, if the user answers the question correctly, there is a positive interaction value of 1 between the two, otherwise it is 0; for the question knowledge graph, all questions and their attributes form triples (question, relationship, attribute). The relationships include the knowledge points of the questions, the themes they belong to, the previous exercises, and the disciplines.
4. The personalized knowledge tracking method based on knowledge graph recommendation according to claim 1, characterized in that Calculate the knowledge level of the currently input question and the level of the current knowledge state, compare the two, update them in sections. The intersecting part of the two is updated using a long short-term memory network; the part below the minimum of the two is directly replaced with the information of the currently input question to replace the old information; the part above the maximum of the two is retained as it is.
5. The personalized knowledge tracking method based on knowledge graph recommendation according to claim 1, wherein Select two knowledge points directly related to the question to be predicted and the two highest-scoring attribute entities to interact with the current knowledge state and the question to be predicted. The interaction result is the prediction result.
Citation Information
Patent Citations
Knowledge tracking method and system integrating long-term and short-term memory and Bayesian network
CN110807469A
Self-adaptive learning method and system based on knowledge model and storage medium
CN110991645A