Knowledge tracking prediction method and system oriented to group feature fusion and application
By constructing a multi-layer GRU network prediction model based on group feature fusion, the problem of traditional knowledge tracing methods failing to utilize student group characteristics is solved, achieving efficient knowledge tracing and prediction both online and offline, and improving prediction accuracy and interpretability.
Patent Information
- Application Number
- CN202510843957.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-31
AI Technical Summary
Traditional knowledge tracking and prediction methods fail to effectively utilize student group characteristics, resulting in an inability to accurately capture students' knowledge mastery in offline teaching scenarios, thus limiting the accuracy and interpretability of prediction results.
We design a multi-layer GRU network prediction model based on group feature fusion. By constructing a group feature embedding structure and combining an IRT model with a relative position encoding attention mechanism, we enhance the model's ability to capture input features. We use binary cross-entropy loss and L2 regularization to optimize the model parameters.
This improves the accuracy and interpretability of the knowledge tracing model's predictions both online and offline, enabling it to better capture group information among students and enhance the accuracy and reliability of prediction results.
Smart Images

Figure CN120876175A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of knowledge tracking technology, and relates to a knowledge tracking and prediction method, system and application oriented to group feature fusion. Background Technology
[0002] In real-world scenarios such as online and offline teaching, there is often a wealth of student information that remains to be utilized. Traditional knowledge tracking and prediction methods often focus only on basic student performance data, neglecting to utilize additional valuable data such as information reflecting group characteristics. Consequently, they fail to accurately capture students' knowledge acquisition status, limiting the accuracy and interpretability of prediction results.
[0003] Knowledge tracking is a task that dynamically predicts students' knowledge mastery by combining their historical test-taking data and using probabilistic inference or machine learning models. Early knowledge tracking models based on probabilistic inference had relatively simple structures, focusing on the knowledge points themselves while ignoring individual differences among students and the correlations between knowledge points.
[0004] In knowledge tracing and prediction tasks, the DIMKT model performs exceptionally well. This model is a deep learning-based knowledge tracing and prediction method based on difficulty feature fusion. It designs and constructs separate problem difficulty and knowledge point difficulty models, using a GRU network structure to simulate students' knowledge mastery and forgetting processes, helping the model better understand the differences in students' performance on questions of varying difficulty. However, the original DIMKT model only incorporates supplementary features such as students' problem-solving behaviors. The limitation of behavioral information fusion is that it is only suitable for online education platforms where behavioral data can be collected. In offline teaching scenarios, it cannot accurately capture and collect students' problem-solving behaviors, thus failing to fully realize its value. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention aims to provide a knowledge tracking prediction method, system, and application oriented towards group feature fusion. The prediction method designs a multi-layer GRU network prediction model based on the fusion of group features and difficulty, effectively capturing group information among students and combining it with difficulty features to construct a prediction task. The core idea of the implementation process of this invention's prediction method is to fully utilize the easily obtainable group feature information in both online and offline scenarios to improve the prediction of the knowledge tracking model. Based on the question, knowledge point, and answer information data required for basic knowledge tracking problems, group information data is extracted from the original dataset as a supplement to construct a group feature embedding structure. A feature embedding of the average ability of the student group is dynamically constructed based on the IRT model. The feature embeddings are concatenated to form the final input, and then an attention mechanism based on relative position encoding is combined to enhance the model's capture of key information of the input features. Finally, a multi-layer GRU network is used to generate the final correct answer probability output.
[0006] The specific execution process for achieving the purpose of this invention includes the following steps:
[0007] Step 1: Data Collection and Preprocessing of Student Problem-Solving Sequences
[0008] The collected raw dataset includes student problem-solving sequence data. Each sequence contains student personal information and problem-solving behavior information; specifically, it includes features such as student ID, student school ID, student class ID, student teacher ID, question, knowledge point, and answer. The preprocessing includes: first, sorting the students sequentially by group according to time attributes to ensure the correctness of the problem-solving sequence; removing null values from each feature in the problem-solving sequence data, and encoding the student, question, knowledge point, student school, student class, and student teacher respectively; deleting records with fewer than 30 answers to questions and records with fewer than 30 answers to knowledge points, finally constructing a complete model training sequence dataset.
[0009] Step 2: Construct the GDIMKT model and perform training and optimization. Existing DIMKT models only consider difficulty feature inputs, while the GDIMKT model in this invention simulates the impact of problem-solving behaviors with different difficulty features on students' knowledge mastery status by using one or more GRU network structures.
[0010] By combining student problem-solving sequence data with group information, this invention proposes a GDIMKT model based on group feature fusion to optimize the training of a knowledge tracking model. The specific modeling process is as follows: the preprocessed model training sequence dataset is used as the input of the GDIMKT model, feature embeddings are constructed separately and concatenated to obtain the input feature embedding; the input feature embedding is used to capture key features through an attention mechanism based on relative position encoding, and then the key features are trained using a one- or multi-layer GRU network structure.
[0011] For the GDIMKT model established in step two, a binary cross-entropy loss function is used in conjunction with L2 regularization to enhance the model's ability to fit complex data. Binary cross-entropy is suitable for situations where the output is a predicted probability, while L2 regularization helps prevent overfitting and ensures the model's generalization ability. The adaptive moment estimation optimization algorithm (Adam) is used to optimize the model parameters. The optimal parameters are learned by maximizing the log-likelihood. After each parameter update, the loss function value of the student's answer sequence on the training set is calculated, and training is repeated until the loss function value converges to the optimum. After training, the model outputs optimized model parameters, which will be used for subsequent predictions of student answers to improve the model's prediction accuracy.
[0012] Step 3: Using the optimized model trained in Step 2, predict how students will answer questions.
[0013] In response to the above steps, the specific measures taken also include:
[0014] In step one above, the present invention constructs model training sequence data based on the data extracted and organized from the student's personal information and problem-solving behavior information in the original dataset. The basic form is as follows:
[0015] (GROUP,kc,p,a)1
[0016] (GROUP,kc,p,a)2
[0017] …
[0018] (GROUP,kc,p,a) t
[0019] The GROUP contains student ID, student school ID, student class ID, student teacher ID, kc represents knowledge point, p represents question, a represents answer, and t represents data number, corresponding to each student.
[0020] In step two above, the GDIMKT model is a knowledge tracking model based on group feature fusion, suitable for judging students' mastery of knowledge points in large-scale online or offline learning. This represents the hidden state of knowledge points for student i at time step t-1. Let represent the question that student i performs at time step t. Indicates that student i is at time step t for the problem. The probability of answering a question correctly at a future time step is calculated. Using historical question-answering sequence data from time step 1 to t-1, the probability of answering a question at a future time step is predicted. The constructed model is as follows:
[0021]
[0022] See the GDIMKT model framework diagram. Figure 1 The model comprises an input layer, an attention mechanism layer, and a knowledge state update layer. The knowledge state update layer further includes one or more GRU network layers and an output layer. Data enters the attention mechanism layer through the input layer. The attention mechanism layer captures key feature information by calculating attention weights, thus obtaining the model input. Then, one or more GRU network layers determine how much of the current input information and the hidden state information from the previous time step are retained to update the new hidden state, controlling the degree of hidden state update, and finally outputting the prediction result, forming a complete and efficient prediction model framework.
[0023] The construction and training of the GDIMKT model specifically includes the following three steps:
[0024] Step 2.1 Constructing Model Feature Inputs
[0025] The input to the GDIMKT model includes group information features, primarily combining group information from the original dataset, such as student schools, classes, and teachers. This group information is transformed into categorical features through one-hot encoding and used to construct feature embeddings.
[0026] The difficulty features of knowledge points and the difficulty features of questions are constructed, and the calculation of the difficulty of knowledge points and the difficulty of questions are as follows:
[0027]
[0028] Among them, S i Let a represent the set of students who answered knowledge point i. ij This refers to student j's answer to knowledge point i, a im This refers to student m's answer to question i.
[0029] For the group average ability embedding, the student's ability is derived by inversely using the calculation formula of the IRT model. The ability is dynamically calculated based on the student's knowledge mastery status in each training iteration, while simultaneously incorporating group characteristic information to calculate the group average ability. The specific method is as follows:
[0030]
[0031] Wherein, P(X) ij =1) is the probability that student i answers correctly on knowledge point j, B i D is the ability parameter of student i. j G is the difficulty parameter of knowledge point j. i This indicates the total number of group categories to which it belongs.
[0032] The training sequence data feature input and group feature input of the knowledge tracking model are transformed into embeddings and concatenated according to the feature order to obtain a high-dimensional feature input vector, which can be used for learning and processing by subsequent neural network layers.
[0033] Step 2.2 Attention Mechanism Layer Based on Relative Position Encoding
[0034] After the model feature input is constructed, the model input is passed through a multi-head attention mechanism layer with relative position encoding. The attention mechanism layer first obtains the query, key, and value through a linear transformation and calculates the attention score. Simultaneously, a one-hot encoding tensor is created to extract the relative position, and the relative position encoding is calculated using a dot product and added to the attention score. Then, the attention weights are calculated, the residuals are concatenated with the layer normalization, and the key-value pair of the current segment is retained for use in the next segment. The specific implementation of the attention mechanism layer is as follows:
[0035] First, given the input embedding vector X, a linear transformation is used to obtain the query matrix Q, the key matrix K, and the value matrix V:
[0036] Q = XW Q
[0037] K = XW K
[0038] V = XW V
[0039] in, is a learnable weight matrix, and d is the hidden layer dimension.
[0040] Divide the above matrix into heads according to the number of heads h, and obtain the query, key, and value vector for each head:
[0041] Q i =split(Q)
[0042] K i =split(K)
[0043] V i =split(V)
[0044] Create a learnable relative position encoding matrix Calculate the relative position code:
[0045] rel_pos=R·I
[0046] Where m is the sequence length and I is the one-hot encoded tensor used to extract relative position information.
[0047] The attention score is calculated by combining the dot product of the query and the key, and the relative position encoding. If there is a key value K in the previous segment... prev Then concatenate it with the current key value and calculate the attention score:
[0048]
[0049] Calculate the attention weights and apply them to the value vector, then concatenate the outputs of all heads and perform a linear transformation to obtain the final output:
[0050] A i = softmax(Attention) i )
[0051] O i =A i V i
[0052] O=LayerNorm(dropout(concat(O 1O 2 ,…,O h ),rate)+X)
[0053] W o =LayerNorm(O+dropout(dense(O),rate))
[0054] in, It is the output weight matrix. Dropout means randomly setting some elements to 0 with a probability of rate. LayerNorm means layer normalization. Its specific calculation formula is as follows: μ and σ are the mean and standard deviation of the input X, respectively. γ and β are the learnable scale and offset parameters.
[0055]
[0056] In this invention, the hidden state vector for each time step is calculated by training the model step by step.
[0057] Step 2.3 Build one or more GRU networks and output latent vectors
[0058] The hidden state at the current time step is concatenated with the historical time steps, denoted as H. A sigmoid activation function is applied to the hidden state, transforming the output into a probability distribution. The calculation method is as follows:
[0059] H = concat(h1, h2, ..., h n )
[0060]
[0061] Where h represents the hidden knowledge state at each time step, t represents the target data, and σ is the sigmoid function, specifically σ(x) = 1 / (1 + e^(-x)). This is the predicted probability value.
[0062] In the process of optimizing the model, binary cross-entropy is used as the main loss function, along with an L2 regularization term. The loss function is expressed as follows:
[0063]
[0064] Where N is the number of samples, λ is the regularization coefficient, and y i For the true value, represents the predicted probability value, and v represents the model parameters.
[0065] The constructed loss function consists of two parts: the first part is the binary cross-entropy loss term, which measures the difference between the probability distribution predicted by the model and the true label distribution; the binary cross-entropy loss calculates the error of each term by comparing the predicted probability output by the model with the actual binary labels (e.g., 0 and 1), and accumulates and averages these errors to reflect the overall accuracy of the model prediction.
[0066] The second part is the L2 regularization term, which calculates the L2 norm of the trainable parameters in the model and adds it to the overall loss in a certain proportion, thereby encouraging the model parameters to converge to smaller values, making the weight distribution of the model smoother and reducing the complexity of the model; its purpose is to prevent the model from overfitting and improve the stability of model training through regularization constraints.
[0067] Finally, the model uses the Adaptive Moments Estimation Optimization (Adam) algorithm to minimize the loss function, defines the learning rate and adjusts it dynamically to further improve the accuracy of predictions;
[0068] The output model parameters that minimize the loss function are used as the model parameters for the optimized GDIMKT model.
[0069] The prediction output of the model in this invention is a continuous probability value between 0 and 1, with 0.5 as the boundary to divide correct and incorrect results as the final output;
[0070] The training range of the model in this invention is [1, T-1], the prediction time step is T, and T represents the time step; the model input is the question p completed by the student. i The question corresponds to the knowledge point kc i Difficulty level of the questions (pd) i Knowledge point difficulty kcd i School i Class C i Teacher t i Group average ability abliity i The splicing and embedding constitutes the structure;
[0071] And / or,
[0072] At each time step, use the student's historical problem-solving performance from the previous time step (z). t-1 The hidden state h of the multi-layer GRU network at the previous time step t-1 As input, we obtain the network hidden state h updated at the current time step. t ;
[0073] And / or,
[0074] H1, h2, ... h1 of all time steps t-1 splicing, denoted as H t-1The spliced H t-1 The predicted probability value is obtained by applying the sigmoid activation function.
[0075] Step 3: Using the optimized model trained in Step 2, predict the students' accuracy rate in answering questions.
[0076] The present invention also provides a prediction system for implementing the above-described sequence prediction method, the prediction system comprising: a data acquisition module, a model building module, an optimization module, and a prediction module;
[0077] The data acquisition module is used to collect students' problem-solving sequence data and perform data preprocessing to generate a problem-solving sequence dataset containing student group information, problem-solving information, and difficulty information.
[0078] The model building module is used to design and build the GDIMKT model, and combined with the group feature fusion module, it performs knowledge tracking and prediction on the students' knowledge mastery status in the next time step.
[0079] The optimization module is used to calculate a loss function that combines binary cross-entropy loss and L2 regularization to optimize model parameters.
[0080] The prediction module is used to predict the student's answer to the question in the next time step based on the optimized feature parameters and the student's knowledge mastery status at the current time step, and output the probability of whether the student's answer is correct.
[0081] This invention also provides the application of the above-mentioned prediction method or prediction system in online and / or offline teaching scenarios, such as question recommendation and personalized teaching.
[0082] The beneficial effects of this invention include: by constructing a GDIMKT model based on group feature fusion, it effectively combines underutilized group information in the original dataset, fully capturing the relationships between students at the group level, and effectively avoiding the limitation of not being able to extract student behavioral feature information in offline teaching scenarios. Compared with existing knowledge tracing models that are mainly designed for online teaching scenarios, this invention can process student problem-solving data in both online and offline teaching scenarios, making it a universal knowledge tracing model. Furthermore, based on the original model, this invention integrates student group features to achieve knowledge tracing tasks, and its performance is superior to the original model in metrics such as Acc and Auc.
[0083] This invention improves the accuracy and interpretability of model predictions in knowledge tracing tasks by more effectively predicting test results based on test sequence data by integrating group characteristics. This method provides an innovative and efficient solution for the field of knowledge tracing and provides important support for teachers to understand students' learning and provide personalized teaching in actual teaching scenarios, and has broad application value. Attached Figure Description
[0084] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0085] Figure 1 This is a diagram of the GDIMKT model framework of the present invention.
[0086] Figure 2 This is a flowchart of the prediction method of the present invention. Detailed Implementation
[0087] The present invention will be further described in detail below with reference to the specific embodiments and accompanying drawings. Except for the contents specifically mentioned below, the processes, conditions, and experimental methods for implementing the present invention are all common knowledge and general knowledge in the art, and the present invention does not have any particular limitations.
[0088] This invention provides a knowledge tracking method, system, and application oriented towards group feature fusion. By combining the group information characteristics of students, this invention can reveal the inherent relationships between students, enabling the model to make more accurate knowledge tracking predictions. In implementation, based on the understanding of group information in the question-solving sequence data, a group feature extraction method is designed to incorporate student group information into the input of the knowledge tracking model training, and an attention mechanism is used to filter key information, providing more accurate model predictions. The introduction of group feature integration in this invention helps the model extract potential group relationships between students, while also understanding the knowledge point learning status between groups. Compared with traditional knowledge tracking models that focus on individuals, the knowledge tracking model incorporating group features is more in line with actual offline teaching scenarios and has greater interpretability in prediction.
[0089] This invention addresses the limitations of knowledge tracing tasks in fully integrating offline teaching scenarios by proposing a universally applicable knowledge tracing model based on group feature fusion (GDIMKT) that can be used both online and offline. Furthermore, this model improves the accuracy and interpretability of knowledge tracing predictions. Figure 2As shown, this method improves the model's prediction accuracy by constructing group feature inputs and supplementing potential group relationships among students through feature fusion. The method includes: Step 1: Extracting corresponding features from the original student problem-solving dataset and constructing a student problem-solving dataset after preprocessing; Step 2: Designing a GDIMKT model with a multi-network structure, using multiple GRU networks to control feature information output to simulate forgetting and remembering behaviors during student learning; designing a loss function module, combining binary cross-entropy loss with L2 regularization to enhance the model's fitting effect on complex data while preventing overfitting; Step 3: Predicting student problem-solving results using the knowledge tracking model trained and optimized in Step 2. The group feature fusion knowledge tracking model constructed in this invention can effectively predict student learning status in online and offline teaching scenarios, providing assistance to teachers in understanding student learning and providing personalized instruction.
[0090] The present invention will be further described in detail through the following specific embodiments.
[0091] Example 1
[0092] The Assistment09 dataset was obtained from an online education data platform. This dataset consists of student test-taking data collected from the Assistments online tutoring system between 2009 and 2010. Specifically, the features include test-taking data from students of different schools, classes, and teachers. It also includes features such as the knowledge points corresponding to the questions and whether the students answered correctly. The preprocessing included removing missing data, sorting by the order in which the questions were answered, and calculating the difficulty features of the students' abilities and the knowledge points of the questions. A total of 282,605 student test-taking sequences were obtained. The data format is (GROUP, kc, p, a, pd, kd), where (kc, p, a) t This dataset contains the problem-solving data for a specific student at time t. The Group element includes the student's school, class, teacher, and average group ability. The sequence data is statistically analyzed based on the frequency of each problem, spanning from 2009 to 2010, and includes 3852 students, 123 knowledge points, and 173109 problems. The training, validation, and test sets are divided according to the students in a 7:2:1 ratio.
[0093] Next, the GDIMKT model is constructed. The batch size is set to 128, the learning rate to 0.001, the number of hidden layer nodes to 128, and the total number of training epochs to 100. Early stopping is used to control iterative training. The pre-processed training set data is input into the GDIMKT model in sequence. Key feature information is captured through a relative position encoding attention mechanism, and training is performed using a multi-layer GRU network. The hidden state vectors obtained from model training are concatenated. A binary cross-entropy loss function combined with L2 regularization is used to enhance the fitting ability to complex data; the adaptive moment estimation optimization algorithm (Adam) is used to optimize the model parameters. After each parameter update, the loss function value for each type of product on the training set is calculated, and training is repeated until the loss function value converges to the optimum. After training, the optimized model parameters are output and saved.
[0094] Finally, a prediction scheme is constructed. Using a student performance prediction model trained with GDIMKT, the probability of a student answering the next question correctly is predicted, converting probabilities higher than 0.5 to 1. Probabilities higher than 0.5 are converted to 1, and probabilities lower than 0.5 are converted to 0, representing whether the student answered correctly.
[0095] The GDIMKT model implemented in this invention is compared with the original DIMKT model and several other knowledge tracing models, using Auc and Acc as evaluation metrics. The report presents the evaluation of the experimental prediction results on the test set. Auc refers to the area under the ROC curve, used to measure the model's discriminative ability; Acc refers to the proportion of correctly predicted samples out of the total sample size, mainly used to measure the overall predictive ability of the model. The results are shown in Table 1.
[0096] Table 1 Comparison of prediction results under different models
[0097]
[0098] Higher values for each evaluation metric indicate better predictive performance. The GDIMKT model implemented in this invention outperforms other baseline models across all metrics. GDIMKT's AUC is 2.41% higher than DIMKT, significantly outperforming other models in these metrics. This improvement not only reflects GDIMKT's enhanced ability to distinguish between positive and negative samples but also signifies its ability to accurately predict students' mastery of knowledge points in a wider range of scenarios. Furthermore, GDIMKT also improves its ACC by 1.60%, demonstrating its ability to more accurately predict students' learning progress. Whether the predictions are correct or incorrect, they are more realistic, providing a more reliable basis for educational decision-making.
[0099] Example 2
[0100] The Assistment12 dataset was obtained from an online education data platform. This data, sourced from the Assistments online tutoring system, covers student interaction data from the 2012-2013 academic year. Specifically, feature extraction involved numbering student problem-solving sequences according to categories such as student, question, and knowledge point, forming features. Each feature number uniquely represents a student, question, and knowledge point. Preprocessing included sequence sorting, null value removal, calculation of question and knowledge point difficulty, and group average ability features, involving a total of 2,711,813 student problem-solving sequences. The data format was (CROUP, kc, p, a, pd, kd), (kc, p, a). t The dataset contains sequence data of students' problem-solving activities at time step t. The problem-solving sequences are organized by student and sorted by the time they were solved. The overall data spans from 2012 to 2013, and includes 29,018 students, 53,091 problems, and 265 knowledge points. The training set, validation set, and test set are divided according to the number of students in a 7:2:1 ratio.
[0101] Next, the GDIMKT model is constructed. The batch size is set to 128, the learning rate to 0.001, the number of hidden layer nodes to 128, and the total number of training epochs to 100. Early stopping is used to control iterative training, preventing overfitting and saving training time and resources. The pre-processed training set data is input into the GDIMKT model in sequence. Key feature information is captured through a relative position encoding attention mechanism, and training is performed using a multi-layer GRU network. The hidden state vectors obtained from model training are concatenated. A binary cross-entropy loss function combined with L2 regularization is used to enhance the fitting ability to complex data; the adaptive moment estimation optimization algorithm (Adam) is used to optimize the model parameters. After each parameter update, the loss function value for each type of product on the training set is calculated, and training is repeated until the loss function value converges to the optimum. After training, the optimized model parameters are output and saved.
[0102] Finally, a prediction scheme is constructed. A student knowledge tracking prediction model trained using the GDIMKT model is used to predict the probability of a student answering a question correctly or incorrectly at the next time step. Probabilities higher than 0.5 are converted to 1, and probabilities lower than 0.5 are converted to 0, representing whether the student answered correctly or not, respectively.
[0103] The GDIMKT model implemented in this invention is compared with the original DIMKT model and several other knowledge tracing models, with Auc and Acc as evaluation metrics. The results are shown in Table 2.
[0104] Table 2 Comparison of prediction results under different models
[0105]
[0106] Higher values for each evaluation metric indicate better predictive performance. The GDIMKT model implemented in this invention outperforms other baseline models across all metrics. GDIMKT's AUC is 0.89% higher than DIMKT, showing better performance than other models. This means that the GDIMKT model can more accurately distinguish between positive and negative samples when dealing with complex data, and is more adept at capturing key information within the data. Furthermore, GDIMKT also improves its ACC by 0.55%, indicating higher accuracy in its predictions. This reflects the close agreement between the model's predictions and actual results. This improvement means that the GDIMKT model can more accurately predict students' mastery of knowledge points, providing a more reliable basis for personalized learning and optimizing teaching strategies.
[0107] References
[0108] [1]Salinas D,Bohlke-Schneider M,Callot L,et al.High-DimensionalMultivariate Forecasting with Low-Rank Gaussian Copula Processes[C] / / Proceedings of the 33rd International Conference on Neural InformationProcessing Systems.2019:6827–6837.
[0109] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented using the object-oriented programming language Python.
[0110] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0111] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0112] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0113] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0114] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
[0115] The scope of protection of this invention is not limited to the above embodiments. Any variations and advantages that can be conceived by those skilled in the art without departing from the spirit and scope of this invention are included in this invention and are protected by the appended claims.
Claims
1. A knowledge tracking and prediction method oriented towards group feature fusion, characterized in that, The knowledge tracking and prediction method is based on a multi-layer GRU network that integrates group characteristics and difficulty, capturing group information among students and combining it with difficulty features to predict the accuracy of answering questions; including: Step 1: Collect the original dataset of students' problem-solving, extract the problem-solving sequence data and preprocess it, and build the model training sequence data; Step 2: Construct the GDIMKT model and use the training sequence data of the model constructed in Step 1 to train and optimize the GDIMKT model; Step 3: Using the optimized model trained in Step 2, predict the students' accuracy rate in answering questions.
2. The method as described in claim 1, characterized in that, In step one, the question-solving sequence data extracted from the original dataset includes the following features: student ID, student school ID, student class ID, student teacher ID, question, knowledge point, and answer; The preprocessing includes: sorting students by group according to time attributes to ensure the correct sequence of students answering questions; removing null values from each feature in the question-answering sequence data, and encoding students, questions, knowledge points, student schools, student classes, and student teachers respectively; deleting records with fewer than 30 answers to questions and records with fewer than 30 answers to knowledge points, and constructing a model training sequence dataset. The obtained model training sequence data is in the following format: (GROUP,kc,p,a) t , The GROUP contains student ID, student school ID, student class ID, student teacher ID, kc represents knowledge point, p represents question, a represents answer, and t represents data number, corresponding to each student.
3. The method as described in claim 1, characterized in that, In step two, the mathematical form of the constructed GDIMKT model is as follows: in, This represents the hidden state of knowledge points for student i at time step t-1. Let represent the question that student i performs at time step t. Indicates that student i is at time step t for the problem. The probability of doing something correctly; And / or, The GDIMKT model includes: an input layer, an attention mechanism layer, and a knowledge state update layer; The knowledge state update layer further includes: one or more GRU network layers and an output layer.
4. The method as described in claim 1, characterized in that, In step two, the training process of the GDIMKT model includes the following: Step a: Construct and concatenate feature embeddings based on the training sequence data of the model to obtain the input feature embeddings; Step b: Embed the input features to capture key features through an attention mechanism based on relative position encoding; Step c: Train the key features using a one- or multi-layer GRU network structure.
5. The method as described in claim 4, characterized in that, In step a, the group information in the original dataset is transformed into category features through one-hot encoding, and feature embeddings are constructed; And / or, The difficulty features of knowledge points and the difficulty features of questions are constructed, and the calculation of the difficulty of knowledge points and the difficulty of questions are as follows: Among them, S i Let a represent the set of students who answered knowledge point i. ij This refers to student j's answer to knowledge point i, a im This refers to student m's answer to question i; And / or, Based on the students' knowledge mastery status in each training iteration and combined with the students' group characteristics, the average group ability embedding is calculated as follows: Wherein, P(X) ij =1) is the probability that student i answers correctly on knowledge point j, B i D is the ability parameter of student i. j G is the difficulty parameter of knowledge point j. i This indicates the total number of group categories to which it belongs.
6. The method as described in claim 4, characterized in that, The attention mechanism is implemented by obtaining the query, key, and value through linear transformation, and calculating the attention score; at the same time, a one-hot encoding tensor for extracting relative positions is created, and the relative position encoding is calculated using dot product and added to the attention score; Then, the attention weights are calculated, the residuals and layer normalization are connected, and the key values of the current segment are retained for use in the next segment. include: Given an input embedding vector X, a linear transformation yields a query matrix Q, a key matrix K, and a value matrix V: Q=XW Q , K=XW K , V=XW V , in, It is a learnable weight matrix, and d is the hidden layer dimension; Divide the query matrix Q, key matrix K, and value matrix V into heads according to the number of heads h, obtaining the query, key, and value vectors for each head: Q i =split(Q), K i =split(K), V i =split(V), Create learnable relative position codes Calculate the relative position code: rel_pos = R·I, Where m is the sequence length and I is the one-hot encoded tensor used to extract relative position information; The attention score is calculated by combining the dot product of the query and the key, and the relative position encoding. If there is a key value K in the previous segment... prev Then concatenate it with the current key value and calculate the attention score: Calculate the attention weights and apply them to the value vector, then concatenate the outputs of all heads and perform a linear transformation to obtain the final output: A i =softmax(Attention i ), About i =A i In i , O=LayerNorm(dropout(concat(O 1 ,ZERO 2 ,…,ZERO h ),rate)+X), W o =LayerNorm(O+dropout(dense(O),rate)), in, It is the output weight matrix. Dropout means randomly setting some elements to 0 with a probability of rate. LayerNorm means layer normalization.
7. The method as described in claim 4, characterized in that, In step c, the hidden state of the current time step is concatenated with the historical time step, denoted as H; a sigmoid activation function is applied to the hidden state, transforming the output into a probability distribution; the calculation method is shown below: H=concat(h1,h2,…,h n ), Where h represents the hidden knowledge state at each time step, t represents the target data, and σ is the sigmoid function, with the formula σ(x)=1 / (1+e^(-x)). This is the predicted probability value.
8. The method as described in claim 1, characterized in that, In step two, the GDIMKT model is optimized by minimizing the loss function that combines the binary cross-entropy and L2 regularization terms; The loss function is expressed as follows: Where N is the number of samples, λ is the regularization coefficient, and y i For the true value, v represents the predicted probability value, and v represents the model parameters; And / or, The learning rate is defined and dynamically adjusted by maximizing the optimal parameters of the log-likelihood learning model through an adaptive moment estimation optimization algorithm to minimize the loss function. The output model parameters that minimize the loss function are used as the model parameters for the optimized GDIMKT model.
9. A prediction system that implements the prediction method according to any one of claims 1-8, characterized in that, The prediction system includes: a data acquisition module, a model building module, an optimization module, and a prediction module; The data acquisition module is used to collect students' problem-solving sequence data and perform data preprocessing to generate a problem-solving sequence dataset containing student group information, problem-solving information, and difficulty information. The model building module is used to design and build the GDIMKT model, and combined with the group feature fusion module, it performs knowledge tracking and prediction on the students' knowledge mastery status in the next time step. The optimization module is used to calculate a loss function that combines binary cross-entropy loss and L2 regularization to optimize model parameters. The prediction module is used to predict the student's answer to the question in the next time step based on the optimized feature parameters and the student's knowledge mastery status at the current time step, and output the probability of whether the student's answer is correct.
10. The prediction method as described in any one of claims 1-8, or the prediction system as described in claim 9, applied in online and / or offline teaching for question recommendation and personalized instruction.