Explanatable knowledge tracking method for Markov blanket optimization

By introducing Markov blanket optimization method into the knowledge tracking model, the shortcomings of existing models in interpretability and causality processing are solved, and more efficient and accurate knowledge tracking and learning suggestions are achieved.

CN120180037APending Publication Date: 2025-06-20EAST CHINA NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510287945.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing knowledge tracking models have shortcomings in improving interpretability and handling complex causal relationships, making it difficult to provide clear explanations for educators, and computational resources are high and inefficient when processing large-scale data sets.

Method used

The interpretable knowledge tracking method optimized by Markov blanket is adopted. Through prior knowledge constraints, the Markov blanket learning algorithm is used to calculate the mutual information value between nodes, determine the priority order of nodes, and limit the maximum number of parent nodes. Combining the fast greedy equivalent search algorithm and Bayesian information criterion, Markov blanket features are extracted, and a knowledge tracking model is constructed using logistic regression, decision tree and random forest algorithm.

Benefits of technology

It improves the interpretability and prediction accuracy of the knowledge tracking model, can more accurately identify the causal relationships that affect learning results, provide more personalized and accurate learning suggestions, reduces the demand for computing resources and improves efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180037A_ABST
    Figure CN120180037A_ABST
Patent Text Reader

Abstract

The invention discloses an interpretable knowledge tracking method for Markov blanket optimization, and the method employs a Markov blanket technology to extract key causal features of a learning process from education big data, so as to accurately track the knowledge mastering state of students. The method comprises the following steps of: constraining a Markov blanket learning algorithm by utilizing priori knowledge, and reducing search space and calculation cost by calculating mutual information among nodes and determining priorities of the nodes and the maximum number of father nodes; determining a Markov blanket feature subset of the target node by using a fast greedy equivalent search algorithm in combination with a scoring function and a Bayesian information criterion; based on the features, a model instance capable of explaining the knowledge tracking method is constructed, and training and optimization are carried out. The model instance is suitable for various intelligent education systems, has a wide vertical application scene in the field of intelligent education, is helpful to provide more personalized learning experience for learners, and helps teachers to specifically understand knowledge mastering conditions of students.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of intelligent education and artificial intelligence, and particularly relates to an interpretable knowledge tracing model based on Markov blankets and its application in educational data analysis. Background Art

[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] With the development of artificial intelligence technology, existing knowledge tracing technologies, especially deep learning-driven knowledge tracing models, although have made certain progress in predicting students' knowledge mastery status, have significant deficiencies in interpretability. Due to the complex network structure and black box characteristics of traditional deep knowledge tracing models, it is difficult to provide clear explanations for educators, which limits the application of these models in educational scenarios. In addition, when dealing with large-scale datasets, many knowledge tracing models face problems of high computational resource requirements and low efficiency, especially unable to effectively extract causal relationships in the data for interpretation and application. Most existing models rely on correlation analysis and cannot deeply understand and explain the key factors affecting students' learning outcomes. In order to improve the interpretability of knowledge tracing models, exploring effective feature extraction and learning methods has become an urgent problem to be solved.

[0004] In summary, existing knowledge tracing models lack effective solutions in improving interpretability and dealing with complex causal relationships. Summary of the Invention

[0005] Aiming at the deficiencies in the application of knowledge tracing models in the prior art, the purpose of the present invention is to provide an interpretable knowledge tracing method optimized by Markov blankets, so as to realize more effective and reliable personalized learning process analysis and recommendation in the educational field, and effectively improve learning effects and teaching quality.

[0006] The specific technical solution to achieve the purpose of the present invention is as follows:

[0007] An interpretable knowledge tracing method optimized by Markov blankets, the method includes:

[0008] Step S1: Use prior knowledge to constrain the score-based Markov blanket learning algorithm. First, calculate the mutual information values between all nodes, determine the priority order of each node based on mutual information, physical laws, and expert knowledge elements, and set the maximum number of parent nodes for each node according to the priority of the node;

[0009] Step S2: Adopt the fast greedy equivalence search algorithm, combine the scoring function and the Bayesian information criterion, conduct forward search and backward search on the target node, extract the Markov blanket of the target node, identify the corresponding parent nodes, child nodes and spouse nodes of the target node, and use independence test and scoring function to learn the Markov blanket of the target variable;

[0010] Step S3: According to the Markov blanket features obtained from Step S2, select and apply three machine learning algorithms, namely logistic regression, decision tree and random forest, for model construction and training to create a knowledge tracing model;

[0011] Step S4: Based on the knowledge tracing model constructed in Step S3, analyze the learning behavior data of students, predict their knowledge mastery status, and generate suggestions for learning path adjustment to optimize the subsequent learning effect and tracking.

[0012] Furthermore, Step S1 specifically includes:

[0013] Step 1.1: Calculate the mutual information values between all nodes, determine the priority order of each node through the node ranking algorithm based on mutual information, and further optimize the priority arrangement of nodes by combining expert knowledge, physical laws and the generally recognized knowledge dependence relationship in the education field;

[0014] Step 1.2: According to the node priority order, set the maximum number of parent nodes for each node, so as to limit the scoring function of the Markov blanket learning algorithm, reduce the search space and calculation cost, and provide an effective prior knowledge basis for the subsequent Markov blanket learning process;

[0015] The determination of the maximum number of parent nodes is based on the mutual information value between nodes, and the priority order will affect the subsequent search process.

[0016] Furthermore, Step S2 specifically includes:

[0017] Step 2.1: Apply the fast greedy equivalence search algorithm, combine the node priority order and the maximum number of parent nodes determined in Step S1, add a constraint condition to limit the search scope, and determine the Markov blanket set initially containing the target variable;

[0018] Step 2.2: Use the search strategy and scoring function in the global Bayesian network structure learning algorithm to evaluate the relationship between nodes one by one, update the initial graph structure, and score the relevance of candidate nodes to optimize the network structure of the Markov blanket set;

[0019] Step 2.3: Evaluate the direct association structure between the target node and candidate nodes, identify and retain the parent nodes, child nodes and spouse nodes with high scores, eliminate irrelevant or low-correlation nodes, and update the graph structure;

[0020] Step 2.4: Based on the evaluated direct association structure, conduct forward search, gradually add directed edges that meet the scoring requirements, connect the eligible candidate nodes to the target node through edges, and expand the Markov blanket of the target node;

[0021] Step 2.5: Perform backward search. For each directed edge added during the forward search, check its validity one by one, delete the edges with scores lower than the threshold to clean up redundant or unnecessary connections and maintain the simplicity of the graph structure;

[0022] Step 2.6: Based on the backward search, prune the final graph structure, remove redundant edges, ensure that the Markov blanket of the target node only contains the nodes with the highest correlation coefficients, and form the final feature subset.

[0023] Furthermore, Step 2.1 specifically includes:

[0024] Step 2.1.1: Calculate the mutual information values between all nodes, determine the priority order of each node, and set the maximum number of parent nodes for each node according to the mutual information values and the priority order;

[0025] Step 2.1.2: Use the set scoring function to evaluate the relevance between nodes, and find all possible adjacent node pairs of the target node T, where T is the representative node of the target variable, and its Markov blanket contains all the parent nodes, child nodes, and spouse nodes that directly affect the target variable;

[0026] Step 2.1.3: Find the node set X = {x1, x2,..., x n}, where X is the set of candidate adjacent nodes of the target node T, and x i represents a single node in the set; score the relevance between each node in the node set X and the target node T using the scoring function scoreI(X, T, G); where scoreI(X, T, G) is a function used to evaluate the relevance between the node set X and the target node T in the graph structure G, and G represents the current graph structure; if scoreI(X, T, G)>0, then add an edge between the target node T and the node x i to represent the directed relationship that the node x i has a direct impact on the target node T; similarly, find the node set Y = {y1, y2,..., y m} related to the node set X, where Y is the set of nodes that have a potential association relationship with the node set X, and y jRepresents a single node in the set; score the relevance between each node in node set Y and the nodes in X using the scoring function scoreI(X, Y, G); where scoreI(X, Y, G) is a function for evaluating the relevance of node sets X and Y in the graph structure G; if the scoring result scoreI(X, Y, G) > 0, then add an edge between node x i and node y j to represent the influence relationship of node y j on node x i and improve the node connectivity in the graph structure;

[0027] Step 2.1.4: Based on the graph structure generated in Step 2.1.3, determine the preliminary Markov blanket set of the target node T, and return the graph structure and the preliminary Markov blanket set containing the target variable.

[0028] Furthermore, the Step S3 specifically includes:

[0029] Step 3.1: Based on the Markov blanket of the target node extracted in Step S2, use its included parent nodes, child nodes, and spouse nodes as input variables to construct a knowledge tracing model;

[0030] Step 3.2: Select and apply three machine learning algorithms, namely logistic regression, decision tree, and random forest, to construct, train, and optimize the model;

[0031] Step 3.3: According to the causal effect of the nodes in the Markov blanket on the target node, use machine learning algorithms to adjust and optimize the model parameters to improve the prediction performance and interpretability of the model, and establish a logistic regression knowledge tracing model based on the Markov blanket, a decision tree knowledge tracing model based on the Markov blanket, and a random forest knowledge tracing model based on the Markov blanket;

[0032] Step 3.4: During the model training process, further adjust the model by calculating the regression coefficients or feature importances to better reflect the dynamic changes in students' learning behaviors and knowledge states, and ensure the improvement of the accuracy of knowledge tracing in the time series;

[0033] Step 3.5: Validate and evaluate the trained model, verify the performance of the model through experiments, ensure that the model has good interpretability and prediction ability, and apply the final model to the knowledge tracing task to analyze and predict students' learning results.

[0034] Furthermore, the Step 3.3 specifically includes:

[0035] Step 3.3.1: Calculate the mutual information between data nodes, and determine the priority order of target nodes; Set a threshold, and determine the maximum number of parent nodes for each node according to the number of correlation features exceeding this threshold; Calculate the Markov blanket by using the fast greedy equivalence search algorithm in Step S2, and obtain the feature set that affects the target node;

[0036] Step 3.3.2: When the loss value is greater than or equal to 0.015, select and utilize three machine learning models: logistic regression, decision tree, and random forest; The loss value is a metric for evaluating the prediction accuracy of the model, and the lower the value, the better the model prediction effect; The parameters of logistic regression include the bias term and feature weights, which are used to linearly combine the probability of predicting the knowledge state of the target node, denoted as P(True|β), where P(True|β) represents the probability that the target node is in the correct state under the given parameters; The parameters of the decision tree and random forest include feature importance, splitting rules, or relevant hyperparameters, which are used to optimize the tree structure and prediction results;

[0037] Step 3.3.3: When new learning response data of students is obtained, based on the model parameters and response data of the previous round, iteratively optimize the model parameters of the logistic regression, decision tree, and random forest models and recalculate the loss value; By adjusting the parameters to gradually reduce the loss value until the loss value is less than 0.015 or reaches the convergence criterion, output the result of the finally optimized knowledge tracing model for predicting the knowledge mastery state of students.

[0038] Furthermore, the improvement in Step 3.4 for the accuracy of knowledge tracing in the time series specifically includes:

[0039] Step 3.4.1: Combine the model parameters with the learning process, and redefine the variable meanings in the model formula using different causal features from the Markov blanket feature set to ensure that the model can better capture the dynamic changes of learners in the time series;

[0040] Step 3.4.2: Model the knowledge tracing problem as a binary classification task, only consider the correctness of learners in objective questions, and map the learner's answer R ij = 1 as the learning result, and map it to the target node in the Markov blanket; where, R ij = 1 represents the learner's i-th answer to the j-th question; When R ij = 1, it means that the answer is correct;

[0041] Step 3.4.3: Define the independent variable Z as the set of all features that have a causal effect on the target node, and calculate the regression coefficient β of each feature through the training process n, to fit the potential relationships in the learner behavior data and use the regression coefficients for model optimization; the fitting process is used to adjust the model to accurately capture the potential causal relationships in the data, so as to improve the prediction accuracy and interpretability of the model;

[0042] Step 3.4.4: As the learning process progresses, dynamically adjust the model parameters of the model so that the model can automatically adapt to the changes in students' learning behaviors, gradually improve the knowledge tracking performance of the model in the time series, and ensure its accuracy.

[0043] Furthermore, the verification and evaluation of the trained model described in Step 3.5 specifically include:

[0044] Step 3.5.1: Preprocess the test data set, calculate the Pearson correlation coefficient between nodes to determine the priority order of target nodes, and the calculation formula is: where xi and yi are the values of two nodes in the test data set, and are the corresponding means, and r represents the correlation coefficient between nodes; set a threshold based on the correlation coefficient, and based on this, determine the maximum number of parent nodes for each node, and extract the feature set that affects the target node;

[0045] Step 3.5.2: Use the test score of the student learning benefit index or the knowledge mastery rate as the target node, use the feature set extracted in Step 3.5.1 and the student learning data in the test data set as prior knowledge, and construct a feature network graph containing the target node through the Fast Greedy Equivalence Search - Markov Blanket (FGES - MB) algorithm to obtain the Markov blanket; select machine learning models logistic regression and random forest, and iteratively train the model until the loss value is less than the preset threshold of 0.015;

[0046] Step 3.5.3: Evaluate the model performance, compare the model performances including the Markov blanket features, mutual information features, and all features through ablation experiments, and the measurement indicators include accuracy, F1 score, and area under the curve (AUC); among them, accuracy measures the overall correctness of the prediction, the F1 score is the harmonic mean of precision and recall, and AUC evaluates the model classification effect; compare different feature sets according to the AUC value to verify the role of the Markov blanket in improving the interpretability and prediction ability of the model.

[0047] Furthermore, the construction of the personalized learning recommendation algorithm described in Step S4 specifically includes:

[0048] First, analyze the learning performance of each student using the prediction results of the knowledge tracing model to determine the student's mastery level of specific knowledge points. This analysis is not limited to predicting the student's performance in the next test, but also includes deeply exploring the student's learning needs and weak links. According to these analysis results, personalized learning resource recommendations will be automatically generated, including learning materials, practice question banks, and video explanations that match the current learning progress.

[0049] Subsequently, dynamically adjust the recommended content to adapt to the student's learning progress. When a student shows strong mastery ability in a certain knowledge point, increase the difficulty of that knowledge point and recommend higher-level learning materials. If a student repeatedly makes mistakes in certain content, prioritize recommending the corresponding basic knowledge and supplementary materials to consolidate the learning foundation. In addition, it can also adjust the push frequency and form of learning resources according to the student's learning rhythm and preferences to ensure that the learning process is both efficient and in line with the student's individual needs.

[0050] A computer-readable storage medium stores multiple instructions, characterized in that the instructions are adapted to be loaded and executed by a processor of a terminal device for the described method of interpretable knowledge tracing optimized by Markov blanket.

[0051] The present invention has the following beneficial effects:

[0052] The method of interpretable knowledge tracing optimized by Markov blanket proposed by the present invention not only improves the prediction accuracy but also enhances the interpretability and generalization ability of the model. By introducing Markov blanket, the model can effectively capture the features closely related to the learning results and incorporate these features into the core parameters of the knowledge tracing model. This method not only improves the prediction accuracy of the model but also makes the model have a clearer explanatory power for the key factors in the student's learning process. Compared with the traditional knowledge tracing model, the present invention can more accurately identify the causal relationships affecting the learning results, thereby providing more personalized and accurate learning suggestions. In addition, the present invention verifies the performance of the proposed model through ablation experiments. The experimental results on multiple data sets show that the Markov blanket knowledge tracing model is significantly better than the existing models, especially in performance indicators such as the area under the curve AUC. At the same time, the essential interpretability of the model brings great advantages to the applications in the education field, enabling educators to more intuitively understand and apply these intelligent analysis tools, reducing the ethical risks of artificial intelligence in education. Generally speaking, the present invention not only improves the effectiveness and reliability of the personalized learning system but also provides an innovative solution with broad application prospects for future intelligent education. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is the flowchart of the present invention;

[0054] Figure 2 It is the Markov blanket learning result graph of the present invention based on the Junyi dataset;

[0055] Figure 3 It is the Markov blanket learning result graph of the present invention based on the ASSISTments2009 - 2010 dataset. Specific implementation manners

[0056] Next, the technical solutions in one or more embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in one or more embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on one or more embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present invention.

[0057] Referring to the figure, a knowledge tracing method based on Markov blanket of the present invention specifically includes:

[0058] Step S1: Step S1 of the present invention aims to optimize the score - based Markov blanket learning algorithm by integrating prior knowledge, thereby enhancing the performance and interpretability of the knowledge tracing model. In this step, the calculation of the mutual information value between all nodes is first performed. This calculation process is crucial for evaluating and determining the mutual dependence between nodes in the dataset. The level of the mutual information value reflects the closeness of information sharing between nodes and is the basis for subsequent node ranking and scoring.

[0059] After the mutual information calculation is completed, according to the obtained mutual information value, a node ranking algorithm based on mutual information is used to optimize the priority order of each node. This ranking process not only considers the mutual information statistical results but also incorporates experts' knowledge, relevant physical laws, and other objective facts, thereby ensuring the scientificity and practicality of the node priority. Through such optimization, the maximum number of parent nodes of each node is reasonably set, effectively restricting the scoring function of the learning algorithm, reducing the search space and calculation cost.

[0060] In addition, the optimized node priority order and the setting of the maximum number of parent nodes directly affect the execution efficiency and the accuracy of the results of the Markov blanket learning algorithm. By precisely controlling these parameters, the algorithm can more efficiently explore the potential connections between nodes, thereby providing a solid foundation for constructing a high - performance knowledge tracing model. This method not only improves the prediction accuracy of the model but also enhances the application value of the model in the actual educational environment, enabling it to provide more accurate and personalized learning support for educators and learners.

[0061] Step S2: Based on the node priority order and the maximum number of parent nodes determined in Step S1, a constraint condition is imposed on the search process. This constraint condition helps to limit the search scope and reduce the computational amount, enabling a more efficient determination of the preliminary Markov blanket set containing the target variable during the learning process of the Markov blanket. Through this step, it is determined which nodes are most likely to have direct or indirect relevance to the target node, laying a foundation for subsequent network structure optimization.

[0062] Next, the search strategy and scoring function of the global Bayesian network structure learning algorithm are used to evaluate the relationships between nodes one by one. In this process, by scoring the relevance between each candidate node and the target node, the initial graph structure is updated. The application of the scoring function enables the model to dynamically evaluate the connection relationship between each pair of nodes, retain those connections with higher scores and stronger relevance, and eliminate low-relevance or irrelevant nodes. This process ensures that the connections in the graph structure are more reasonable, further optimizing the network structure of the Markov blanket.

[0063] After the preliminary graph structure is determined, forward search begins. The purpose of forward search is to gradually add directed edges that meet the scoring requirements on the basis of the identified direct association structure, and connect the eligible candidate nodes to the target node through these edges. In this way, the Markov blanket of the target node is expanded to include more nodes directly related to the target node, thereby enhancing the prediction ability of the model.

[0064] After the forward search is completed, backward search follows. The task of backward search is to check each directed edge added during the forward search one by one and evaluate the effectiveness of each edge. If the score of some edges is lower than the set threshold, they will be deleted to clean up redundant or unnecessary connections. This step helps to maintain the simplicity of the graph structure and avoid overfitting of the model due to redundant connections.

[0065] Finally, after the forward and backward searches are completed, the final graph structure is pruned. The pruning process will further remove the possible redundant edges in the graph to ensure that the Markov blanket of the target node only contains the most relevant nodes. Through this series of operations, an optimal feature subset is finally formed, providing reliable input features for the subsequent knowledge tracing model and ensuring the efficiency and interpretability of the model in predicting students' learning behaviors.

[0066] Step S3: In step S3, based on the Markov blanket feature set obtained in step S2, a suitable machine learning algorithm is selected to construct and train the knowledge tracing model. The implementation of this step is not limited to theoretical derivation, but also combines the application of actual datasets, and the knowledge tracing processes of the Junyi dataset and the ASSISTments2009 - 2010 dataset are analyzed through examples. The following are the specific operation steps:

[0067] Sub - step 1: First, based on the Markov blanket feature set determined in step S2, start constructing the knowledge tracing model. For different datasets, respective target variables are selected. For example, in the Junyi dataset, earned_proficiency is selected as the target variable, while in the ASSISTments2009 - 2010 dataset, correct is selected as the target variable. These target variables are determined according to the feature correlation and causal relationship of the datasets. In the feature selection process, the Pearson correlation coefficient method is used to calculate the mutual information value between nodes, and a threshold of 0.01 is used to establish the priority order of feature correlation coefficients.

[0068] In the Junyi dataset, by applying the Markov blanket algorithm, a network graph structure as shown in Figure 2 is finally formed, where earned_proficiency is the target node, and the blue nodes in the graph represent the features directly associated with the target node. In the ASSISTments2009 - 2010 dataset, the same algorithm is used, and finally the network graph structure in Figure 3 is generated, with correct as the target node, connecting various related features.

[0069] Sub - step 2: After feature selection and the determination of the preliminary model structure, select an appropriate machine learning algorithm to construct and train the model. The available algorithms include logistic regression, decision tree, and random forest, etc. For the Junyi dataset and the ASSISTments dataset, multiple experiments are conducted respectively to verify the performance of different algorithms on different datasets.

[0070] In the specific training process, the nodes in the Markov blanket feature subset, including parent nodes, child nodes, and spouse nodes, are used as the input features of the model. By training these features, an efficient knowledge tracing model can be constructed, which can accurately predict the learning status of students. For example, in Figure 3 , the correct node is used as the target node, and the information of its parent nodes, child nodes, and spouse nodes is used to train the model, and the respective regression coefficients or decision tree structures are generated.

[0071] Sub-step 3: During the model training process, according to the causal effects of the nodes in the Markov blanket feature subset on the target node, further adjust and optimize the model parameters. By calculating the regression coefficients or feature importance, the model can more accurately reflect the dynamic changes in students' learning behaviors and knowledge states. In specific operations, for earned_proficiency in the Junyi dataset, features without direct causal relationships, such as hint_used and correct, were excluded, while for the ASSISTments dataset, a feature with low correlation but significant impact, hint_total, was discovered.

[0072] By continuously optimizing the parameters, these models can be dynamically adjusted to adapt to changes in students' behaviors, thus ensuring the accuracy and interpretability of predictions. Finally, after multiple iterations, the model demonstrated good knowledge tracing performance in the time series and was successfully applied to an actual personalized learning recommendation system.

[0073] Step S4: Based on the knowledge tracing model constructed in Step S3, the present invention deeply analyzes students' learning behavior data, predicts their knowledge mastery status, and generates personalized learning paths and resource recommendations. The algorithm aims to closely integrate students' learning behaviors with personalized educational suggestions, thereby optimizing learning effects and ensuring the personalization and adaptability of the learning process.

[0074] First, through the prediction results of the knowledge tracing model, the student learning situation analysis algorithm comprehensively analyzes the learning behavior data of each student to accurately predict the mastery level of students on different knowledge points. This analysis includes predicting students' performance in the next test and mining students' learning needs and weak links; by analyzing data such as the correctness of students' answers, the frequency of using hints, and the answering time, identify which knowledge points students are weak in and need further consolidation.

[0075] Based on the above analysis results, the recommendation algorithm generates personalized learning resource recommendations, including learning materials, practice question banks, video explanations, etc. that match the current learning progress and knowledge mastery status. When a student has a good grasp of a certain knowledge point, the algorithm increases the difficulty of that knowledge point and recommends higher-level learning materials; if a student makes mistakes repeatedly on certain knowledge points, basic knowledge supplementary materials or detailed explanations are preferentially recommended to consolidate knowledge.

[0076] The recommendation algorithm also dynamically adjusts the recommended content according to the student's real-time learning progress; if the student shows significant progress during the learning process, the algorithm increases the difficulty of the knowledge points and recommends more complex learning content; if the student makes repeated mistakes in certain content, the difficulty is reduced and basic learning materials are provided to help the student master step by step. The algorithm further adjusts the push frequency and form of learning resources according to the student's learning rhythm and preferences to ensure that the learning process is efficient and meets individual needs.

[0077] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system or a computer program product. Therefore, the present invention can adopt a completely software embodiment or a form combining software and hardware embodiments. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media including, but not limited to, disk storage, CD-ROM, optical storage, etc., containing computer-usable program code.

Claims

1. A Markov blanket optimized interpretable knowledge tracking method, characterized in that: The method comprises: Step S1: Use prior knowledge to constrain the scoring-based Markov blanket learning algorithm. First, calculate the mutual information value between all nodes, determine the priority order of each node based on the mutual information, physical laws, and expert knowledge elements, and set the maximum number of parent nodes for each node according to the priority of the node; Step S2: Using a fast greedy equivalent search algorithm, combined with a scoring function and the Bayesian information criterion, forward search and backward search are performed on the target node, the Markov blanket of the target node is extracted, the parent node, child node and spouse node corresponding to the target node are identified, and the Markov blanket of the target variable is learned using an independence test and a scoring function; Step S3: According to the Markov blanket obtained in step S2, three machine learning algorithms, namely logistic regression, decision tree and random forest, are selected and applied for model construction and training to create a knowledge tracking model; Step S4: Based on the knowledge tracking model constructed in step S3, analyze the students' learning behavior data, predict their knowledge mastery status, build a personalized learning recommendation algorithm, and generate learning path adjustment suggestions to optimize subsequent learning effects and tracking.

2. The Markov blanket optimized interpretable knowledge tracking method according to claim 1, characterized in that: The step S1 specifically includes: Step 1.1: Calculate the mutual information value between all nodes, determine the priority order of each node through the node sorting algorithm based on mutual information, and further optimize the node priority arrangement by combining expert knowledge, physical laws, and knowledge dependencies recognized in the education field; Step 1.2: According to the node priority order, set the maximum number of parent nodes for each node to limit the scoring function of the Markov blanket learning algorithm, reduce the search space and computational cost, and provide an effective prior knowledge basis for the subsequent Markov blanket learning process; The maximum number of parent nodes is determined based on the mutual information value between nodes, and the priority order will affect the subsequent search process.

3. The Markov blanket optimized interpretable knowledge tracking method according to claim 1, characterized in that: The step S2 specifically includes: Step 2.1: Apply the fast greedy equivalent search algorithm, combine the node priority order and the maximum number of parent nodes determined in step S1, add a constraint to limit the search range, and determine the Markov blanket set that initially contains the target variable; Step 2.2: Use the search strategy and scoring function in the global Bayesian network structure learning algorithm to evaluate the relationship between nodes one by one, update the initial graph structure, and score the relevance of candidate nodes to optimize the network structure of the Markov blanket set; Step 2.3: Evaluate the direct association structure between the target node and the candidate nodes, identify and retain the parent nodes, child nodes, and spouse nodes with high scores, remove irrelevant or low-association nodes, and update the graph structure; Step 2.4: Based on the evaluated direct association structure, forward search is performed to gradually increase the directed edges that meet the scoring requirements, connect the qualified candidate nodes with the target node through edges, and expand the Markov blanket of the target node; Step 2.5: Perform a backward search to check the validity of each directed edge added in the forward search, and delete the edges with scores below the threshold to clean up redundant or unnecessary connections and keep the graph structure simple. Step 2.6: Based on the backward search, the final graph structure is pruned to remove redundant edges, ensuring that the Markov blanket of the target node only contains nodes with the highest correlation coefficient and forming the final feature subset.

4. The Markov blanket optimized interpretable knowledge tracking method according to claim 3, characterized in that: The step 2.1 specifically includes: Step 2.1.1: Calculate the mutual information value between all nodes, determine the priority order of each node, and set the maximum number of parent nodes for each node based on the mutual information value and priority order; Step 2.1.2: Use the set scoring function to evaluate the association between nodes and find all possible adjacent node pairs of the target node T, where T is the representative node of the target variable and its Markov blanket contains all parent nodes, child nodes, and spouse nodes that directly affect the target variable; Step 2.1.3: Find the node set X = {x1, x2, ..., x n }, where X is the set of candidate adjacent nodes of the target node T, x i represents a single node in the set; scores the association between each node in the node set X and the target node T using the scoring function scoreI(X,T,G); where scoreI(X,T,G) is a function used to evaluate the association between the node set X and the target node T in the graph structure G, where G represents the current graph structure; if scoreI(X,T,G)>0, then the target node T and node x i Add an edge between to represent node x i There is a directed relationship that directly affects the target node T; similarly, find the node set Y related to the node set X = {y1, y2, ..., y m }, where Y is the set of nodes that have a potential association relationship with the node set X, y j represents a single node in the set; scores the association between each node in the node set Y and the nodes in X using the scoring function scoreI(X,Y,G); where scoreI(X,Y,G) is a function used to evaluate the association between the node sets X and Y in the graph structure G; if the scoring result scoreI(X,Y,G)>0, then at node x i and node y j Add an edge between them to represent node y j For node x i influence relationship and improve the node connectivity in the graph structure; Step 2.1.4: Determine the preliminary Markov blanket set of the target node T based on the graph structure generated in step 2.1.3, and return the graph structure and the preliminary Markov blanket set containing the target variable.

5. The Markov blanket optimized interpretable knowledge tracking method according to claim 1, characterized in that: The step S3 specifically includes: Step 3.1: Based on the Markov blanket of the target node extracted in step S2, the parent nodes, child nodes and spouse nodes contained in it are used as input variables to construct a knowledge tracking model; Step 3.2: Select and apply three machine learning algorithms: logistic regression, decision tree, and random forest to build, train, and optimize the model; Step 3.3: According to the causal effects of the nodes in the Markov blanket on the target nodes, the model parameters are adjusted and optimized using machine learning algorithms to improve the prediction performance and interpretability of the model. A logistic regression knowledge tracking model based on the Markov blanket, a decision tree knowledge tracking model based on the Markov blanket, and a random forest knowledge tracking model based on the Markov blanket are established. Step 3.4: During the model training process, further adjust the model by calculating the regression coefficient or feature importance to better reflect the dynamic changes of students' learning behavior and knowledge status, and ensure the accuracy of knowledge tracking in time series; Step 3.5: Verify and evaluate the trained model, verify the performance of the model through experiments, ensure that the model has good interpretability and predictive ability, and apply the final model to the knowledge tracing task to analyze and predict students' learning outcomes.

6. The Markov blanket optimized interpretable knowledge tracking method according to claim 5, characterized in that: The step 3.3 specifically includes: Step 3.3.1: Calculate the mutual information between data nodes and determine the priority order of the target node; set a threshold and determine the maximum number of parent nodes for each node based on the number of correlation features exceeding the threshold; calculate the Markov blanket by using the fast greedy equivalent search algorithm of step S2 and obtain the feature set that affects the target node; Step 3.3.2: When the loss value is greater than or equal to 0.015, select and use three machine learning models: logistic regression, decision tree, and random forest; the loss value is a metric used to evaluate the accuracy of model predictions, and the lower the value, the better the model prediction effect; the parameters of logistic regression include bias terms and feature weights, which are used to linearly combine and predict the probability of the knowledge state of the target node, expressed as P(True|β), where P(True|β) represents the probability that the target node is in the correct state under given parameters; the parameters of decision trees and random forests include feature importance, splitting rules or related hyperparameters, which are used to optimize the tree structure and prediction results; Step 3.3.3: When obtaining new learning response data of students, based on the model parameters and response data of the previous round, iteratively optimize the model parameters of the logistic regression, decision tree and random forest models to recalculate the loss value; adjust the parameters to gradually reduce the loss value until the loss value is less than 0.015 or reaches the convergence standard, and output the final optimized knowledge tracking model results to predict the students' knowledge mastery status.

7. The Markov blanket optimized interpretable knowledge tracking method according to claim 5, characterized in that: Step 3.4 improves the accuracy of knowledge tracking in time series, including: Step 3.4.1: Combine model parameters with the learning process and redefine the meaning of variables in the model formula using different causal features from the Markov blanket feature set to ensure that the model can better capture the dynamic changes of the learner in time series; Step 3.4.2: Model the knowledge tracking problem as a binary classification task, considering only the correctness of the learner in the objective questions and classifying the learner's answer R ij = 1 is regarded as the learning result and mapped to the target node in the Markov blanket; where R ij =1 indicates the learner's i-th answer to the j-th question; when R ij =1, indicating that the answer is correct; Step 3.4.3: Define the independent variable Z as the set of all features that have a causal effect on the target node, and calculate the regression coefficient β of each feature through the training process n , to fit the potential relationships in the learner behavior data and use the regression coefficients for model optimization; the fitting process is used to adjust the model to accurately capture the potential causal relationships in the data to improve the model's predictive accuracy and interpretability; Step 3.4.4: As the learning process progresses, dynamically adjust the model parameters of the model so that the model can automatically adapt to changes in students' learning behavior, gradually improve the model's knowledge tracking performance in time series, and ensure its accuracy.

8. The Markov blanket optimized interpretable knowledge tracking method according to claim 5, characterized in that: Step 3.5 verifies and evaluates the trained model, including: Step 3.5.1: Preprocess the test data set and calculate the Pearson correlation coefficient between nodes to determine the priority order of the target nodes. The calculation formula is: Among them, xi and yi are the values ​​of two nodes in the test data set, and is the corresponding mean, r represents the correlation coefficient between nodes; a threshold is set according to the correlation coefficient, and based on this, the maximum number of parent nodes of each node is determined to extract the feature set that affects the target node; Step 3.5.2: Take the student learning benefit index test score or knowledge point mastery rate as the target node, use the feature set extracted in step 3.5.1 and the student learning data in the test data set as prior knowledge, and construct a feature network graph containing the target node through the fast greedy equivalent search-Markov blanket (FGES-MB) algorithm to obtain the Markov blanket; select the machine learning model logistic regression and random forest, and iteratively train the model until the loss value is less than the preset threshold of 0.015; Step 3.5.3: Evaluate model performance. Compare the performance of models including Markov blanket features, mutual information features, and all features through ablation experiments. The measurement indicators include accuracy, F1 score, and area under the curve (AUC). Accuracy measures the overall correctness of the prediction, F1 score is the harmonic mean of precision and recall, and AUC evaluates the classification effect of the model. Compare different feature sets based on the AUC value to verify the role of Markov blanket in improving the interpretability and predictive ability of the model.

9. The Markov blanket optimized interpretable knowledge tracking method of claim 1, characterized in that: Step S4 of constructing a personalized learning recommendation algorithm specifically includes: First, the prediction results of the knowledge tracking model are used to analyze each student's learning performance and determine the student's mastery of specific knowledge points. This analysis is not limited to predicting the student's performance in the next test, but also includes in-depth exploration of the student's learning needs and weak links. Based on these analysis results, personalized learning resource recommendations will be automatically generated, including learning materials, exercise question banks, and video explanations that match the current learning progress. Subsequently, the recommended content will be dynamically adjusted to suit the students' learning progress; when students show strong mastery of a certain knowledge point, the difficulty of that knowledge point will be increased and higher-level learning materials will be recommended; if students repeatedly make mistakes on certain content, the corresponding basic knowledge and supplementary materials will be recommended first to consolidate the learning foundation; in addition, the frequency and form of learning resource push can be adjusted according to the students' learning pace and preferences to ensure that the learning process is both efficient and meets the students' individual needs.

10. A computer-readable storage medium storing a plurality of instructions, characterized in that: The instructions are suitable for being loaded by a processor of a terminal device and executing a Markov blanket optimized explainable knowledge tracing method as described in any one of claims 1-9.

Citation Information

Cited By

  • Automatic searching customer obtaining method and system

    CN122045526A