Dual-channel knowledge tracing method and model based on adaptive decay attention

CN121723374BActive Publication Date: 2026-09-25HUAZHONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511840830.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-09-25
Estimated Expiration
2045-12-08

AI Technical Summary

Technical Problem

[0004]遗忘因子泛化性不足:多数模型采用固定或线性衰减函数模拟遗忘过程,无法适应不同学生、不同知识点的动态遗忘规律;

Benefits of technology

[0045]本申请提供一种基于自适应衰减注意力的双渠道知识追踪方法,包括以下步骤:获取学生的历史学习的交互序列,所述交互序列包括按时间顺序排列的多个交互数据,每个交互数据包括问题标识、概念标识和响应结果;在练习表示阶段,基于所述交互序列为每个交互数据分别生成概念交互表征和问题交互表征;在知识状态建模阶段,将所述概念交互表征序列输入至概念渠道时间衰减注意力模块,得到概念渠道知识状态序列;将所述问题交互表征序列输入至问题渠道时间衰减注意力模块,得到问题渠道知识状态序列;所述概念渠道时间衰减注意力模块和问题渠道时间衰减注意力模块结构相同,均采用可学习的衰减参数矩阵实现非均匀指数衰减;通过动态融合机制,融合所述概念渠道知识状态序列和问题渠道知识状态序列,得到融合知识状态序列;利用待预测练习的表示作为查询,对所述融合知识状态序列进行时间衰减注意力运算,得到与所述待预测练习相关联的最终知识状态;在预测阶段,基于所述最终知识状态和所述待预测练习的表示,计算学生正确回答所述待预测练习的概率。本申请公开了一种基于自适应衰减注意力的双渠道知识追踪方法及模型,旨在解决现有知识追踪模型在遗忘因子泛化性不足、概念与问题特征耦合的缺陷,通过构建时间敏感的注意力机制、独立的概念与问题表征通道,以及动态融合预测架构模型,实现对学生知识状态的精细化建模,显著提升预测精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121723374B_ABST
    Figure CN121723374B_ABST
Patent Text Reader

Abstract

The application discloses a double-channel knowledge tracking method and model based on adaptive decay attention, aiming to solve the defects of the existing knowledge tracking model in the insufficient generalization of the forgetting factor and the coupling of the concept and the problem characteristics, and realize fine modeling of the student knowledge state by constructing a time-sensitive adaptive decay attention mechanism, an independent concept and problem representation channel, and a dynamically fused prediction architecture model, thereby significantly improving the prediction accuracy of the knowledge tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence education technology, specifically to a dual-channel knowledge tracking method and model based on adaptive decaying attention. Background Technology

[0002] With the accelerated pace of urbanization and the digital transformation of education, personalized learning has become a core development direction for intelligent education systems. Knowledge tracking, as a key supporting technology for Intelligent Tutoring Systems (ITS), has the core task of inferring a student's mastery of various knowledge points in real time based on their historical answer sequences, and providing decision support such as exercise recommendations and learning path planning accordingly. In recent years, knowledge tracking technology has been widely applied in online education platforms and smart school teaching systems.

[0003] Early knowledge tracing models, represented by Bayesian Knowledge Tracing (BKT), modeled students' knowledge state transitions using Hidden Markov Models. However, BKT assumes single skills and ignores differences in question difficulty, making it difficult to handle complex multi-skill scenarios. Subsequently, deep learning methods were introduced into the field of knowledge tracing, such as Dynamic Key-Value Memory Networks (DKVMN) and Deep Knowledge Tracing (DKT), significantly improving modeling capabilities. In recent years, attention-based models have gradually become mainstream due to their powerful sequence modeling capabilities and parallel computing advantages. Nevertheless, existing models still face three core bottlenecks in practical applications:

[0004] The forgetting factor has insufficient generalization: most models use fixed or linear decay functions to simulate the forgetting process, which cannot adapt to the dynamic forgetting patterns of different students and different knowledge points.

[0005] The coupling between concepts and problem characteristics is severe: the question embedding usually includes both knowledge points and question characteristics (such as the design of distractors and the style of expression), making it difficult for the model to distinguish whether the concept has been mastered and whether the question has been understood.

[0006] Furthermore, existing model evaluations often focus on a single metric—prediction accuracy—lacking specific assessments of the effectiveness of forgetting modeling and feature decoupling. In long-sequence, large-scale learning scenarios, the adaptability, stability, and interpretability of models still need improvement, further hindering the practical application of knowledge tracing technology in high-concurrency, high-requirement intelligent tutoring systems. Summary of the Invention

[0007] The main objective of this application is to provide a dual-channel knowledge tracing model based on adaptive decaying attention, comprising the following steps:

[0008] Step S10: Obtain the student's historical learning interaction sequence, which includes multiple interaction data arranged in chronological order, and each interaction data includes a question identifier, a concept identifier, and a response result;

[0009] Step S20: In the practice representation phase, a concept interaction representation and a question interaction representation are generated for each interaction data based on the interaction sequence;

[0010] Step S30: In the knowledge state modeling stage, the concept interaction representation sequence is input into the concept channel time decay attention module to obtain the concept channel knowledge state sequence; the question interaction representation sequence is input into the question channel time decay attention module to obtain the question channel knowledge state sequence; the concept channel time decay attention module and the question channel time decay attention module have the same structure and both use a learnable decay parameter matrix to achieve non-uniform exponential decay;

[0011] Step S40: Through a dynamic fusion mechanism, the concept channel knowledge state sequence and the problem channel knowledge state sequence are fused to obtain a fused knowledge state sequence;

[0012] Step S50: Using the representation of the exercise to be predicted as a query, perform time decay attention operation on the fused knowledge state sequence to obtain the final knowledge state associated with the exercise to be predicted;

[0013] Step S60: In the prediction phase, based on the final knowledge state and the representation of the exercise to be predicted, calculate the probability that the student answers the exercise to be predicted correctly.

[0014] In one embodiment, the step of obtaining a student's historical learning interaction sequence, the interaction sequence including multiple interaction data arranged in chronological order, each interaction data including a question identifier, a concept identifier, and a response result, includes:

[0015] The system retrieves raw interaction logs from its backend database and performs data cleaning and preprocessing on these logs. The preprocessing includes:

[0016] Filter out interaction records that are missing key features or contain outliers;

[0017] Remove data whose interaction sequence length is less than a preset threshold;

[0018] Interaction sequences longer than the preset maximum length are divided into multiple subsequences, and question identifiers and concept identifiers are mapped to consecutive integer indices;

[0019] The student interaction records are sorted according to timestamps to construct the historical learning interaction sequence.

[0020] In one embodiment, the step of generating a conceptual interaction representation and a question interaction representation for each interaction data based on the interaction sequence includes:

[0021] The problem identifier, concept identifier, and response result are respectively mapped into a problem embedding vector, a concept embedding vector, and a response embedding vector through an embedding layer;

[0022] The concept interaction representation is obtained by adding the concept embedding vector to the response embedding vector;

[0023] The problem interaction representation is obtained by adding the problem embedding vector to the response embedding vector.

[0024] In one embodiment, the implementation process of the concept channel time decay attention module and the problem channel time decay attention module includes:

[0025] Construct a learnable decay parameter matrix;

[0026] Based on the attenuation parameter matrix, calculate the cumulative attenuation factor between any two interaction positions and the original attention score between the query vector and the key vector;

[0027] Multiplying the original attention score by the corresponding cumulative decay factor yields an attention score subject to a non-uniform exponential decay constraint.

[0028] Based on the attention score with the applied non-uniform exponential decay constraint, the value vector is weighted and summed to obtain the output.

[0029] In one embodiment, constructing a learnable attenuation parameter matrix specifically includes:

[0030] A low-dimensional parameter matrix is ​​used to simulate a learnable decay parameter matrix, the low-dimensional parameter matrix having a size of num_intervals×n, where num_intervals is the preset number of time intervals and n is the maximum length of the sequence;

[0031] For positions i and j in the sequence, calculate the time steps Δ = [n / num_intervals] contained in each interval, then determine the interval index of position i based on Δ, and obtain the corresponding decay parameter from the low-dimensional parameter matrix.

[0032] In one embodiment, the step of fusing the concept channel knowledge state sequence and the problem channel knowledge state sequence through a dynamic fusion mechanism to obtain a fused knowledge state sequence includes:

[0033] For each time step t, the fused knowledge state sequence h_fused is calculated according to the following formula, expressed as:

[0034] h_fused = h_concept + (w + m)⊙h_question

[0035] Where h_fused represents the fused knowledge state, h_concept represents the concept channel knowledge state, h_question represents the question channel knowledge state, w is the learnable weight vector, m is the content modulation vector dynamically calculated based on the current input content, and ⊙ represents element-wise multiplication.

[0036] In one embodiment, the step of using the representation of the exercise to be predicted as a query to perform a time-decayed attention operation on the fused knowledge state sequence to obtain the final knowledge state associated with the exercise to be predicted includes:

[0037] The query vector is used as a representation of the exercise to be predicted, and the key vector and value vector are used as a sequence of fused knowledge states.

[0038] The query vector, key vector, and value vector are input into the time decay attention module for calculation;

[0039] The output of the time decay attention module is the final knowledge state h_t associated with the exercise to be predicted.

[0040] In one embodiment, the step of calculating the probability that a student correctly answers the exercise to be predicted based on the final knowledge state and the representation of the exercise to be predicted during the prediction phase includes:

[0041] The final knowledge state h_t is concatenated with the representation e_{t+1} of the exercise to be predicted to obtain a combined feature vector;

[0042] The combined feature vector is input into the prediction module, which includes at least one fully connected layer.

[0043] The prediction module processes the data and finally outputs a probability value p_{t+1} between 0 and 1 through an activation function. This probability value represents the probability that the model predicts the student will correctly answer the exercise to be predicted.

[0044] Therefore, this application has the following beneficial effects:

[0045] This application provides a dual-channel knowledge tracking method based on adaptive decaying attention, comprising the following steps: acquiring a student's historical learning interaction sequence, the interaction sequence including multiple interaction data arranged in chronological order, each interaction data including a question identifier, a concept identifier, and a response result; in the practice representation stage, generating a concept interaction representation and a question interaction representation for each interaction data based on the interaction sequence; in the knowledge state modeling stage, inputting the concept interaction representation sequence into a concept channel time decaying attention module to obtain a concept channel knowledge state sequence; inputting the question interaction representation sequence into a question channel time decaying attention module to obtain a question channel knowledge state sequence; the concept channel time decaying attention module and the question channel time decaying attention module have the same structure, both using a learnable decaying parameter matrix to achieve non-uniform exponential decay; fusing the concept channel knowledge state sequence and the question channel knowledge state sequence through a dynamic fusion mechanism to obtain a fused knowledge state sequence; using the representation of the exercise to be predicted as a query, performing time decaying attention operation on the fused knowledge state sequence to obtain the final knowledge state associated with the exercise to be predicted; in the prediction stage, calculating the probability that the student correctly answers the exercise to be predicted based on the final knowledge state and the representation of the exercise to be predicted. This application discloses a dual-channel knowledge tracking method and model based on adaptive decaying attention, which aims to solve the shortcomings of existing knowledge tracking models in terms of insufficient generalization of forgetting factors and coupling of concept and problem features. By constructing a time-sensitive attention mechanism, independent concept and problem representation channels, and a dynamic fusion prediction architecture model, it achieves refined modeling of students' knowledge status and significantly improves prediction accuracy. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a system flowchart for a dual-channel knowledge tracing method based on adaptive decaying attention;

[0048] Figure 2 This is an architecture diagram of a dual-channel knowledge tracing method based on adaptive decaying attention;

[0049] Figure 3 This is a visualization of the decay parameters of the concept channel (left) and question channel (right) on the ASSISTments2017 dataset for a dual-channel knowledge tracing method based on adaptive decaying attention.

[0050] Figure 4 This is a visualization of the decay parameters of the concept channel (left) and question channel (right) on the AAAI2023 dataset for a dual-channel knowledge tracing method based on adaptive decaying attention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0052] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application.

[0053] To address the shortcomings of existing technologies, this application provides a dual-channel knowledge tracking method based on adaptive decaying attention. It achieves individualized forgetting modeling by constructing a time-sensitive attention mechanism, decoupling student learning interaction sequences into two independent channels: a concept dimension and a question dimension. A dynamic fusion strategy integrates the knowledge states from both channels, accurately capturing the differentiated forgetting patterns of students across different knowledge points and effectively distinguishing between knowledge mastery and question-answering experience. This method uses a learnable decay parameter matrix instead of a fixed decay pattern, enabling the model to adapt to the memory characteristics of different students. The independent concept and question channel modeling architecture avoids the coupling interference between concept and question features in traditional methods. Combined with an adaptive prediction module and cross-entropy objective function optimization, it achieves a closed-loop process of "accurate forgetting modeling - independent feature representation - dynamic prediction fusion," ultimately providing high-precision and highly generalized learning assessment capabilities in multi-disciplinary scenarios.

[0054] This application provides a dual-channel knowledge tracing model based on adaptive decaying attention, including steps S10-S50, as described above. Figure 1 , Figure 1 This is a system flowchart for a dual-channel knowledge tracing method based on adaptive decaying attention.

[0055] Step S10: Obtain the student's historical learning interaction sequence, which includes multiple interaction data arranged in chronological order, and each interaction data includes a question identifier, a concept identifier, and a response result;

[0056] Step S20: In the practice representation phase, a concept interaction representation and a question interaction representation are generated for each interaction data based on the interaction sequence;

[0057] Step S30: In the knowledge state modeling stage, the concept interaction representation sequence is input into the concept channel time decay attention module to obtain the concept channel knowledge state sequence; the question interaction representation sequence is input into the question channel time decay attention module to obtain the question channel knowledge state sequence; the concept channel time decay attention module and the question channel time decay attention module have the same structure and both use a learnable decay parameter matrix to achieve non-uniform exponential decay;

[0058] Step S40: Through a dynamic fusion mechanism, the concept channel knowledge state sequence and the problem channel knowledge state sequence are fused to obtain a fused knowledge state sequence;

[0059] Step S50: Using the representation of the exercise to be predicted as a query, perform time decay attention operation on the fused knowledge state sequence to obtain the final knowledge state associated with the exercise to be predicted;

[0060] Step S60: In the prediction phase, based on the final knowledge state and the representation of the exercise to be predicted, calculate the probability that the student answers the exercise to be predicted correctly.

[0061] Specifically, this embodiment proposes a dual-channel knowledge tracking method based on adaptive decaying attention. This method achieves accurate modeling and prediction of students' knowledge states through six core steps, as follows:

[0062] Step S10: Obtain historical learning interaction sequences

[0063] Historical learning data of students is obtained from the backend database of the intelligent education system. This data includes chronologically arranged interaction sequences, with each interaction containing three core elements: question identifier, concept identifier, and response result. The raw data is first preprocessed, including data cleaning, outlier handling, and removal of interaction records with missing or outlier key features (question identifier, concept identifier, and response result). Student records with too few interactions are filtered out to ensure the sequence length meets modeling requirements. Question identifiers and concept identifiers are mapped to continuous numerical indices for subsequent vectorization. Finally, a well-organized historical learning interaction sequence is formed, providing a high-quality data foundation for subsequent modeling.

[0064] Step S20: Generate dual-channel interaction representation

[0065] In the practice representation phase, conceptual interaction representations and question interaction representations are generated for each interaction data. First, a dense vector is mapped from the question identifier, concept identifier, and response result through an embedding layer. The conceptual interaction representation is obtained by adding the concept embedding vector and the response embedding vector, primarily reflecting the student's mastery of the knowledge points. The question interaction representation is obtained by adding the question embedding vector and the response embedding vector, mainly reflecting the student's ability to solve specific questions. This dual-channel representation design effectively decouples conceptual knowledge from question-specific knowledge, laying the foundation for subsequent independent modeling. For the practice to be predicted, its representation is generated by element-wise multiplication of the question embedding and the concept embedding, and then superimposed with the concept embedding vector (i.e., yt+1=e). q t +1 ⊙e c t +1 +e c t +1 ).

[0066] Step S30: Dual-channel knowledge state modeling

[0067] In the knowledge state modeling phase, a dual-channel architecture is adopted to process the concept sequence and the question sequence separately. The concept interaction representation sequence is input into the concept channel time decay attention module, and the question interaction representation sequence is input into the question channel time decay attention module. The two modules have the same structure but independent parameters, both using a learnable decay parameter matrix to achieve non-uniform exponential decay. The time decay attention mechanism avoids overfitting through a parameter grouping strategy, enhances flexibility through a content modulation mechanism, and efficiently calculates the cumulative decay factor through the logarithmic prefix sum difference method. This design enables the model to adaptively learn the personalized forgetting patterns of different students and different knowledge points, achieving accurate modeling of the evolution of knowledge state. The specific logic of the parameter grouping strategy is to calculate the time step Δ=[n / num_intervals] for each interval, determine the interval to which the position belongs based on Δ, and simulate the high-dimensional decay matrix with a low-dimensional parameter matrix.

[0068] Step S40: Dynamically integrate knowledge states

[0069] A dynamic fusion mechanism effectively integrates the knowledge states from the concept channel and the question channel. This mechanism not only includes learnable baseline weight parameters but also introduces a modulation vector based on dynamic content computation. Specifically, the fusion formula adds the knowledge state from the concept channel to the weighted knowledge state from the question channel. The weight vector dynamically adjusts with the input content, enabling the model to adaptively balance the contributions of the two channels based on different dataset characteristics and learning contexts. This dynamic fusion strategy ensures both the model's generalization ability and its adaptability to different educational scenarios.

[0070] Step S50: Optimize Knowledge Status Query

[0071] The model uses the representation of the exercise to be predicted as the query, where both the query vector and key vector are representations of the exercise, and the value vector serves as the fused knowledge state sequence. A time-decay attention operation is then performed on this fused knowledge state sequence. This step refines the knowledge from the global knowledge state to knowledge relevant to a specific exercise. Through the attention mechanism, the model can selectively extract the most relevant information from historical knowledge states based on the content features of the exercise to be predicted. This query mechanism effectively improves the relevance and usability of the knowledge state, providing a more accurate input for the final prediction.

[0072] Step S60: Generate prediction results

[0073] In the prediction phase, the final knowledge state is combined with the representation of the practice question to be predicted, and input into the prediction network to calculate the probability of the student answering correctly. The prediction network typically employs a multilayer perceptron structure, using fully connected layers and activation functions to achieve nonlinear transformations. The prediction network has two fixed fully connected layers, with the activation function being ReLU (first layer) + Sigmoid (second layer), and the loss function being the predicted value p. t+1 Compared with the actual response r t+1 The cross-entropy log loss.

[0074] Finally, the activation function outputs a probability value between 0 and 1. Model training uses a cross-entropy loss function, and all parameters are optimized through backpropagation. This entire process forms an end-to-end knowledge tracking solution, providing data support for personalized teaching. Figure 2 This is an architecture diagram of a dual-channel knowledge tracking method based on adaptive decaying attention. First, student learning interactions (such as answer records) are transformed into vectors. Then, the model uses a dual-channel time-decaying attention mechanism to split information into two independent channels: "concepts" and "questions," which are processed in parallel to simulate the persistent memory of concepts and the easily forgotten memory of questions in the human brain, respectively. The knowledge states from the two channels are integrated through a dynamic fusion mechanism that adaptively adjusts the weights of information from different channels. The integrated knowledge state then undergoes another attention query, focusing on the information most relevant to the question to be predicted. Finally, the prediction module outputs a probability value, accurately predicting the student's accuracy in answering the next question.

[0075] Furthermore, in this embodiment, the step of obtaining the student's historical learning interaction sequence, wherein the interaction sequence includes multiple interaction data arranged in chronological order, and each interaction data includes a question identifier, a concept identifier, and a response result, includes:

[0076] The system retrieves raw interaction logs from its backend database and performs data cleaning and preprocessing on these logs. The preprocessing includes:

[0077] Filter out interaction records that are missing key features or contain outliers;

[0078] Remove data whose interaction sequence length is less than a preset threshold;

[0079] Interaction sequences longer than the preset maximum length are divided into multiple subsequences, and question identifiers and concept identifiers are mapped to consecutive integer indices;

[0080] The student interaction records are sorted according to timestamps to construct the historical learning interaction sequence.

[0081] Specifically, in this embodiment, the data cleaning and preprocessing process is performed according to the following steps:

[0082] First, the system verifies the integrity of the raw interaction logs retrieved from the intelligent tutoring system's backend database. It scans all data records, identifying and removing those with missing values ​​in key fields such as question identifiers, concept identifiers, or response results. Simultaneously, it detects and filters records containing outliers, including but not limited to answer times exceeding reasonable limits, invalid question or concept identifiers, and response data that does not conform to a preset format.

[0083] Secondly, based on the characteristics of the knowledge tracing task, the interaction sequences are quality-screened. A lower bound threshold for sequence length is set, typically 3 or 4, and all data containing students with fewer than this threshold of interactions are removed. This measure ensures that each student's interaction sequence contains sufficient information to support effective knowledge state modeling. For long sequences in the dataset, a maximum sequence length of 200 is set, and interaction sequences exceeding this length are divided into multiple consecutive subsequences with a length not exceeding this maximum value, while each subsequence maintains its original temporal order.

[0084] Next, feature engineering preprocessing is performed. Categorical variables such as question identifiers and concept identifiers in the original data are converted into continuous integer indices starting from 0 by building a mapping table. This process not only reduces data sparsity but also improves the learning efficiency of subsequent embedding layers. The construction of the mapping table ensures that each unique question identifier and concept identifier corresponds to a unique integer value, and these integer values ​​are continuously distributed in the corresponding index space.

[0085] Finally, the interaction records for each student are rigorously sorted according to the time dimension. Based on the timestamp field in the records, all interactions of the same student are arranged in chronological order, constructing a historical learning interaction sequence with a clear temporal relationship. Maintaining this temporal structure is crucial for subsequent modeling of the dynamic evolution of knowledge states, as it ensures that the model can accurately capture the patterns of knowledge mastery changes over time during the learning process. Through the above data preprocessing, the original teaching system logs are transformed into a well-organized, clean dataset with a clear spatiotemporal structure, providing a high-quality input foundation for subsequent knowledge tracing modeling.

[0086] Further, in this embodiment, the step of generating a concept interaction representation and a question interaction representation for each interaction data based on the interaction sequence includes:

[0087] The problem identifier, concept identifier, and response result are respectively mapped into a problem embedding vector, a concept embedding vector, and a response embedding vector through an embedding layer;

[0088] The concept interaction representation is obtained by adding the concept embedding vector to the response embedding vector;

[0089] The problem interaction representation is obtained by adding the problem embedding vector to the response embedding vector.

[0090] Specifically, in this embodiment, the step of generating a concept interaction representation and a question interaction representation for each interaction data based on the interaction sequence is performed according to the following detailed steps:

[0091] First, three independent embedding layers are initialized: a question embedding layer, a concept embedding layer, and a response embedding layer. The question embedding layer has a dimension of (total_problems, d), where total_problems represents the total number of unique question identifiers in the dataset, and d is the dimension of the embedding vector; in this embodiment, it is preferably set to 128 dimensions. The concept embedding layer has a dimension of (total_skills, d), where total_skills represents the total number of unique concept identifiers in the dataset. The response embedding layer has a dimension of (2, d), corresponding to the correct and incorrect response states, respectively. The weight parameters in these embedding layers are randomly initialized at the start of model training and optimized and updated using the backpropagation algorithm during subsequent training. For each interaction data (q_t, c_t, r_t) in the sequence, the following vectorization operation is performed:

[0092] The question embedding layer is used to find the d-dimensional embedding vector corresponding to the question identifier q_t, denoted as e_q^t.

[0093] The concept embedding layer is used to find the d-dimensional embedding vector corresponding to the concept identifier c_t, denoted as e_c^t.

[0094] The d-dimensional embedding vector corresponding to the response result r_t is found through the response embedding layer, denoted as e_r^t.

[0095] After obtaining the basic embedding vectors, two different types of interaction representations are constructed: the concept interaction representation x_c^t is obtained by performing a vector addition operation between the concept embedding vector e_c^t and the response embedding vector e_r^t, i.e., x_c^t = e_c^t + e_r^t. This operation integrates knowledge point features with answer result features to form a state representation that reflects the student's mastery of a specific concept.

[0096] The question-interaction representation x_q^t is obtained by adding the question embedding vector e_q^t and the response embedding vector e_r^t, i.e., x_q^t = e_q^t + e_r^t. This operation integrates specific question features with answer result features to form a state representation reflecting the student's proficiency in answering specific questions. The key to this dual-channel representation design is that it avoids the traditional approach of prematurely merging question features and concept features into a single practice representation. Instead, it constructs independent interaction representations, providing a foundation for subsequent dual-channel knowledge state modeling. This design allows the model to perform sequence modeling in both the concept semantic space and the question-specific space, which is more in line with the theoretical basis in cognitive science that semantic memory and episodic memory have different characteristics.

[0097] For the next exercise to be predicted (q_{t+1}, c_{t+1}), a different strategy is used for representation generation: first, e_q^{t+1} and e_c^{t+1} are obtained through an embedding layer; then, feature interaction is performed through element-wise multiplication; finally, a learnable bias vector μ is added, i.e., e_{t+1} = e_q^{t+1} ⊙ e_c^{t+1} + μ. This representation method preserves the association information between the question and the concept, and provides rich feature representations for subsequent prediction tasks. Through the above technical solution, the original low-dimensional discrete labels are transformed into high-dimensional dense vector representations rich in semantic information, providing high-quality input features for subsequent deep knowledge state modeling.

[0098] Furthermore, in this embodiment, the implementation process of the concept channel time decay attention module and the problem channel time decay attention module includes:

[0099] Construct a learnable decay parameter matrix;

[0100] Based on the attenuation parameter matrix, calculate the cumulative attenuation factor between any two interaction positions and the original attention score between the query vector and the key vector;

[0101] Multiplying the original attention score by the corresponding cumulative decay factor yields an attention score subject to a non-uniform exponential decay constraint.

[0102] Based on the attention score with the applied non-uniform exponential decay constraint, the value vector is weighted and summed to obtain the output.

[0103] Specifically, in this embodiment, the implementation process of the concept channel time decay attention module and the problem channel time decay attention module is performed according to the following steps:

[0104] First, a learnable decay parameter matrix is ​​constructed. To avoid overfitting caused by directly learning the complete n×n decay matrix, a parameter grouping mechanism is used to construct a low-dimensional parameter matrix P with dimensions [num_intervals, n], where num_intervals is the preset number of time intervals (preferred to be 16 in this embodiment), and n is the maximum sequence length. For any position pair (i, j), its interval index is determined by calculating i / (n / num_intervals), and the corresponding basic decay parameter p_{i,j} is obtained from matrix P.

[0105] Next, the original attention score and cumulative decay factor are calculated. For the input sequence X, a query matrix Q, a key matrix K, and a value matrix V are generated through a linear transformation:

[0106] Q = XW^Q, K = XW^K, V = XW^V,

[0107] Where W^Q, W^K, and W^V are learnable projection matrices. Q, K, and V are divided into h heads (h=8 in this embodiment), with the dimension d_k = d / h for each head. The original attention score A_raw is calculated in each attention head. A content modulation mechanism is introduced to enhance flexibility when calculating the cumulative decay factor. For a query Q_i, it is mapped to a scalar modulation amount m_i through a small feedforward network. m_i is added to the basic decay parameter p_{i,j} and constrained to the (0,1) interval by the Sigmoid function to obtain the final decay element.

[0108] δ_{i,j} = σ(p_{i,j} + m_i)

[0109] Simultaneously, a causal lower triangle mask is applied to ensure temporal causality. An efficient logarithmic prefix sum difference method is used to calculate the cumulative decay factor D_{i,j}. Finally, the cumulative decay factor D_{i,j} = exp(log(D_{i,j})) is obtained through exponential operation.

[0110] Next, the original attention scores are multiplied by the cumulative decay factor, and a non-uniform exponential decay constraint A = A_raw⊙D_{i,j} is applied. A causal mask and the Softmax function are then applied to the resulting matrix A to obtain the normalized attention weights.

[0111] Finally, the value vector V is weighted and summed using normalized attention weights to obtain the output of a single attention head. The outputs of all attention heads are concatenated and processed through a linear transformation to form the final multi-head attention output. The core of this time-decaying attention mechanism lies in achieving non-uniform exponential decay through a learnable decay parameter matrix, breaking the limitations of traditional fixed decay coefficients; employing a parameter grouping mechanism to balance model capacity and generalization ability; utilizing a content modulation mechanism to enhance the flexibility of the decay mode; and achieving efficient computation through a logarithmic prefix sum difference method.

[0112] Furthermore, in this embodiment, constructing a learnable attenuation parameter matrix specifically includes:

[0113] A low-dimensional parameter matrix is ​​used to simulate a learnable decay parameter matrix, the low-dimensional parameter matrix having a size of num_intervals×n, where num_intervals is the preset number of time intervals and n is the maximum length of the sequence;

[0114] For positions i and j in the sequence, calculate the time steps Δ = [n / num_intervals] contained in each interval, then determine the interval index of position i based on Δ, and obtain the corresponding decay parameter from the low-dimensional parameter matrix.

[0115] Specifically, in this embodiment, the technical solution for constructing a learnable attenuation parameter matrix is ​​implemented according to the following detailed steps:

[0116] First, the key hyperparameters of the model are determined. The maximum sequence length n is set to 200, a reasonable value determined based on statistical analysis of the length of student interaction sequences in typical educational datasets. The number of time intervals num_intervals, as an important model hyperparameter, is preferably set to 16 in this embodiment after grid search validation. This value is determined based on sufficient experimental verification: when num_intervals is too small (e.g., 4 or 8), the model's expressive power is limited; when it is too large (e.g., 32 or 64), it is prone to overfitting.

[0117] Secondly, a low-dimensional parameter matrix P is constructed. This matrix has dimensions [num_intervals, n], i.e., 16×200. Each element P[k,m] in the matrix is ​​a learnable parameter, initialized during model initialization using a uniform distribution U(0.6,0.9). This initialization strategy is based on considerations of the physical meaning of the decay parameter: the initial value is within the range of 0.6-0.9, which avoids the gradient vanishing problem and aligns with the prior knowledge that historical information should have a certain degree of continuity in knowledge tracing tasks.

[0118] In the specific calculation process, for any position pair (i,j) in the sequence, where i represents the current query position and j represents the historical key position, the corresponding decay parameter is obtained as follows:

[0119] Calculate the interval index: index = i / (n / num_intervals), time step Δ = [n / num_intervals]. Since i is an integer, integer division is used in the actual calculation: index = i / 13. Obtain the basic decay parameters from the parameter matrix P: p_base = P[index, j]. The core advantage of this parameter grouping mechanism is that it effectively simulates the 200×200 high-dimensional decay parameter space with a low-dimensional matrix of 16×200, reducing the number of parameters to be learned from 40,000 to 3,200, reducing parameter complexity by 87.5%. This design significantly improves the generalization performance of the model while maintaining its expressive power, effectively preventing overfitting.

[0120] To verify the effectiveness of the parameter grouping mechanism, this embodiment conducted detailed ablation experiments. Experimental results show that when num_intervals=1, the model degenerates into a globally shared decay parameter, with the AUC on the ASSISTments2017 dataset dropping to 0.765. When num_intervals=200, although the model achieves an AUC of 0.788, significant overfitting occurs on the validation set. When num_intervals=16, the model achieves the best AUC performance of 0.7938 on the test set while maintaining good training stability. Furthermore, this embodiment also designs a dynamic adjustment mechanism, allowing the model to adjust the value of num_intervals according to different application scenarios. For tasks with small datasets or short sequences, the value of num_intervals can be appropriately decreased; while for tasks with abundant data and complex sequence patterns, the hyperparameter value can be appropriately increased.

[0121] Furthermore, in this embodiment, the step of fusing the concept channel knowledge state sequence and the problem channel knowledge state sequence through a dynamic fusion mechanism to obtain a fused knowledge state sequence includes:

[0122] For each time step t, the fused knowledge state sequence h_fused is calculated according to the following formula, expressed as:

[0123] h_fused = h_concept + (w + m)⊙h_question

[0124] Where h_fused represents the fused knowledge state, h_concept represents the concept channel knowledge state, h_question represents the question channel knowledge state, w is the learnable weight vector, m is the content modulation vector dynamically calculated based on the current input content, and ⊙ represents element-wise multiplication.

[0125] Specifically, in this embodiment, a learnable weight vector w is initialized. The dimension of this vector is consistent with the knowledge state dimension d, which is 128 in this embodiment. The weight vector w adopts a zero-initialization strategy, that is, the initial value is set to an all-zero vector. This initialization method ensures that the model will not be biased towards the problem channel in the early stage of training, enabling the model to autonomously learn a suitable basic weight distribution based on data features.

[0126] Secondly, a computational network for the content modulation vector m is constructed. This network employs a two-layer fully connected structure: the first layer maps the input feature dimension from 2d (composed of h_concept and h_question) to an intermediate dimension d_m=64, using the ReLU activation function to introduce non-linearity; the second layer maps the intermediate dimension back to the output dimension d=128, without using an activation function to maintain the linearity of the output. The mathematical expression of this modulation network is:

[0127] m = W_2 · ReLU(W_1 · [h_concept || h_question] + b_1) + b_2

[0128] Where W_1 ∈ R^(64×256) and W_2 ∈ R^(128×64) are weight matrices, and b_1 ∈ R^64 and b_2 ∈ R^128 are bias vectors.

[0129] Calculate the base fusion weights by adding the learnable weight vector w to the content modulation vector m, resulting in a comprehensive weight vector w_total = w + m. Then, perform a Hadamard product operation on the comprehensive weight vector w_total and the question channel knowledge state h_question to obtain the weighted question channel contribution.

[0130] h_question_weighted = w_total ⊙ h_question

[0131] Add the concept channel knowledge state h_concept to the weighted problem channel contribution to obtain the fused knowledge state:

[0132] h_fused = h_concept + h_question_weighted

[0133] The core of this dynamic fusion mechanism lies in providing a global bias through fixed, learnable weights w, and achieving fine-grained adjustment based on specific contexts through a dynamic modulation vector m, thus balancing global consistency and local adaptability. The calculation of the modulation vector m depends on the dual-channel knowledge state at the current moment, enabling the fusion weights to adaptively adjust according to the characteristics of the student's current knowledge state. Element-wise multiplication allows the model to employ different fusion strategies on different feature dimensions, achieving fine-grained information fusion.

[0134] Experimental results show that the mean baseline weight w of the problem channels learned by this dynamic fusion mechanism is 1.51 on the ASSISTments2017 dataset, while it is 0.38 on the AAAI2023 dataset. This difference accurately reflects the actual distribution of the importance of problem features in different datasets, demonstrating the mechanism's good adaptability. Furthermore, ablation experiments show that removing the dynamic modulation vector m reduces the model performance by 0.8 percentage points, fully validating the effectiveness of the dynamic adjustment mechanism.

[0135] To verify whether the dual-channel modeling mechanism can effectively achieve the difference in sequence weight decay scale between the problem channel and the concept channel, this study designed and analyzed the following experimental scheme. In existing time-decay attention mechanisms, the degree of sequence weight decay is mainly determined by decay parameters, which are accumulated to form a decay factor. When the decay parameter approaches 1, the decay factor decreases slowly, leading to a longer retention time of historical information; conversely, when the decay parameter approaches 0, the decay factor decreases sharply, causing historical information to fade rapidly. Therefore, if a significant difference can be observed between the decay parameters of the problem channel and the concept channel, it can be determined that the dual-channel modeling mechanism has successfully achieved the differentiation of the sequence weight decay scale between the two channels. To eliminate the interference of other factors, this application makes specific simplifications to the time-decay attention mechanism, removing the dot product attention calculation and content modulation mechanism, so that the decay factor of the two channels is completely determined by the decay parameter. After these modifications, the method of this application was retrained on the ASSISTments2017 and AAAI2023 datasets, with the parameter num_intervals set to the sequence length to facilitate clear visualization and analysis of the experimental results. After training, decay parameters of temporal decay attention were extracted from both channels, and visual analysis was performed to investigate whether there were significant differences in these decay parameters in the dual-channel modeling mechanism. Figure 3 As shown, Figure 3 This is a visualization of the decay parameters of the concept channel (left) and question channel (right) on the ASSISTments2017 dataset for a dual-channel knowledge tracing method based on adaptive decaying attention. Figure 4 This is a visualization of the decay parameters of the concept channel (left) and question channel (right) on the AAAI2023 dataset for a dual-channel knowledge tracing method based on adaptive decaying attention. Figure 3 and Figure 4 The decay parameter distributions (parameter indices ranging from 0 to 18) of the proposed method on the ASSISTments2017 and AAAI2023 datasets are presented respectively. Visual analysis clearly shows a significant difference in decay parameters between the two channels of the dual-channel modeling mechanism. Specifically, in both datasets, the concept channel consistently exhibits a higher decay parameter value, while the problem channel shows a relatively lower decay parameter value. A higher decay parameter value implies a slower decay rate, allowing historical information to be retained for a longer period, thus indicating that the concept channel is superior to the problem channel in preserving long-term historical information. These visualization results clearly validate that the dual-channel modeling mechanism can effectively achieve differentiated performance between the problem and concept channels in terms of sequence weight decay scale.

[0136] Further, in this embodiment, the step of using the representation of the exercise to be predicted as a query to perform time-decay attention operation on the fused knowledge state sequence to obtain the final knowledge state associated with the exercise to be predicted includes:

[0137] The query vector is used as a representation of the exercise to be predicted, and the key vector and value vector are used as a sequence of fused knowledge states.

[0138] The query vector, key vector, and value vector are input into the time decay attention module for calculation;

[0139] The output of the time decay attention module is the final knowledge state h_t associated with the exercise to be predicted.

[0140] Specifically, in this embodiment, the knowledge state refinement step aims to extract information highly relevant to the upcoming specific exercise from the global, historical knowledge state. The implementation process is as follows:

[0141] First, the representation vector e_{t+1} of the exercise to be predicted is used as the query vector in the attention mechanism. This vector encapsulates the core features of the next exercise. Simultaneously, the fused knowledge state sequence H_fused, previously obtained through a dynamic fusion mechanism and covering the entire learning history, is used. The key vector is used to calculate the matching degree with the query, while the value vector carries the substantive knowledge information that needs to be extracted and aggregated.

[0142] Next, the query, key, and value are input into a dedicated time-decay attention module for computation. Within this module, the raw relevance score between the query vector e_{t+1} and each historical key vector in the sequence (i.e., each time step of H_fused) is first calculated. Then, the module applies its core non-uniform exponential decay constraint: it uses an internally learnable decay parameter matrix to calculate the cumulative decay factor from the current query position (virtually pointing to the end of the sequence) back to each historical key vector position. This factor simulates the effects of knowledge fading over time and memory decay, assigning lower weights to distant and irrelevant historical information.

[0143] Then, the original relevance scores are multiplied by the corresponding cumulative decay factors to obtain the final attention weights after time decay modulation. These weights determine the amount of information contributed by each historical knowledge state when predicting the current specific exercise. Finally, these normalized weights are used to perform a weighted summation of the corresponding value vector (i.e., H_fused itself). This summation result is the output of the time decay attention module, which is defined as the final knowledge state h_t closely related to the exercise to be predicted, e_{t+1}. This state h_t is no longer a simple mixture of all historical information, but rather the essence after targeted filtering and focusing, providing the most direct and relevant state input for subsequent high-precision prediction.

[0144] Further, in this embodiment, the step of calculating the probability that a student correctly answers the exercise to be predicted based on the final knowledge state and the representation of the exercise to be predicted during the prediction phase includes:

[0145] The final knowledge state h_t is concatenated with the representation e_{t+1} of the exercise to be predicted to obtain a combined feature vector;

[0146] The combined feature vector is input into the prediction module, which includes at least one fully connected layer.

[0147] The prediction module processes the data and finally outputs a probability value p_{t+1} between 0 and 1 through an activation function. This probability value represents the probability that the model predicts the student will correctly answer the exercise to be predicted.

[0148] Specifically, in this embodiment, the prediction stage, as the final output of the model, has the core task of deeply fusing the knowledge state refined through the aforementioned steps with the features of the exercise to be predicted, and transforming it into an intuitive prediction value with clear probabilistic meaning. This process begins with feature combination, which involves concatenating the final knowledge state vector h_t with the representation vector e_{t+1} of the exercise to be predicted along the feature dimension. This concatenation operation is a crucial information fusion step, placing h_t, representing the student's dynamic internal knowledge level, and e_{t+1}, representing the external exercise characteristics and examination content, in the same feature space, forming a more information-rich combined feature vector. The dimension of this vector is the sum of the dimensions of the two vectors, ensuring that all relevant information is completely preserved, providing a comprehensive input foundation for subsequent complex nonlinear mappings.

[0149] Next, this combined feature vector is fed into a specially designed prediction module. This module is essentially a small feedforward neural network, whose core structure consists of at least one fully connected layer. In a preferred embodiment of this example, a two-layer fully connected layer of a multilayer perceptron is used. The first fully connected layer is responsible for nonlinear transformation and dimensionality reduction of high-dimensional features. It linearly projects the input through a weight matrix and a bias vector, and then applies the ReLU activation function to introduce a nonlinear decision boundary, enabling it to capture complex, nonlinear patterns in student answering behavior. To avoid overfitting the model to the training data, a dropout technique is usually introduced after this layer, randomly masking the output of some neurons to force the model to learn more robust features.

[0150] The second fully connected layer acts as the final predictor, mapping the intermediate features from the first layer's output to a single scalar value. This scalar value is then fed into the Sigmoid activation function. The Sigmoid function compresses and maps real numbers in any range to the interval (0,1), and its output value p_{t+1} is strictly bounded within this range, thus possessing a clear probabilistic interpretation. This probability value p_{t+1} is the model's final prediction output, quantitatively representing the expected probability that the student can correctly answer the predicted exercise e_{t+1} based on all historical interaction sequences and the current knowledge state. The closer the value is to 1, the higher the confidence of the model in predicting the student's correct answer; the closer the value is to 0, the greater the probability of an incorrect answer. The parameters of the entire model, including the weights and biases in this prediction module, are jointly optimized end-to-end on the training set by minimizing the binary cross-entropy loss between the predicted probability and the actual answer, thereby ensuring the accuracy of the prediction.

[0151] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0152] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0153] It should be particularly noted that, through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, or of course, by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A dual-channel knowledge tracing method based on adaptive decaying attention, characterized in that, Includes the following steps: Step S10: Obtain the student's historical learning interaction sequence, which includes multiple interaction data arranged in chronological order, and each interaction data includes a question identifier, a concept identifier, and a response result; Step S20: In the practice representation phase, a concept interaction representation and a question interaction representation are generated for each interaction data based on the interaction sequence; Step S30: In the knowledge state modeling stage, the concept interaction representation sequence is input into the concept channel time decay attention module to obtain the concept channel knowledge state sequence; the question interaction representation sequence is input into the question channel time decay attention module to obtain the question channel knowledge state sequence; the concept channel time decay attention module and the question channel time decay attention module have the same structure and both use a learnable decay parameter matrix to achieve non-uniform exponential decay; Step S40: Through a dynamic fusion mechanism, the concept channel knowledge state sequence and the problem channel knowledge state sequence are fused to obtain a fused knowledge state sequence; Step S50: Using the representation of the exercise to be predicted as a query, perform time decay attention operation on the fused knowledge state sequence to obtain the final knowledge state associated with the exercise to be predicted; Step S60: In the prediction phase, based on the final knowledge state and the representation of the exercise to be predicted, calculate the probability that the student correctly answers the exercise to be predicted; The step of generating conceptual interaction representations and question interaction representations for each interaction data based on the interaction sequence includes: The problem identifier, concept identifier, and response result are respectively mapped into a problem embedding vector, a concept embedding vector, and a response embedding vector through an embedding layer; The concept interaction representation is obtained by adding the concept embedding vector to the response embedding vector; The problem interaction representation is obtained by adding the problem embedding vector to the response embedding vector; The implementation process of the conceptual channel time decay attention module and the problem channel time decay attention module includes: Construct a learnable decay parameter matrix; Based on the attenuation parameter matrix, calculate the cumulative attenuation factor between any two interaction positions and the original attention score between the query vector and the key vector; Multiplying the original attention score by the corresponding cumulative decay factor yields an attention score subject to a non-uniform exponential decay constraint. Based on the attention score with the applied non-uniform exponential decay constraint, the value vector is weighted and summed to obtain the output.

2. The method according to claim 1, characterized in that, The step of obtaining the student's historical learning interaction sequence, wherein the interaction sequence includes multiple interaction data arranged in chronological order, and each interaction data includes a question identifier, a concept identifier, and a response result, includes: The system retrieves raw interaction logs from its backend database and performs data cleaning and preprocessing on these logs. The preprocessing includes: Filter out interaction records that are missing key features or contain outliers; Remove data whose interaction sequence length is less than a preset threshold; Interaction sequences longer than the preset maximum length are divided into multiple subsequences, and question identifiers and concept identifiers are mapped to consecutive integer indices; The student interaction records are sorted according to timestamps to construct the historical learning interaction sequence.

3. The method according to claim 1, characterized in that, The construction of a learnable attenuation parameter matrix specifically includes: A low-dimensional parameter matrix is ​​used to simulate a learnable decay parameter matrix, the low-dimensional parameter matrix having a size of num_intervals×n, where num_intervals is the preset number of time intervals and n is the maximum length of the sequence; For positions i and j in the sequence, calculate the time steps Δ = [n / num_intervals] contained in each interval, then determine the interval index of position i based on Δ, and obtain the corresponding decay parameter from the low-dimensional parameter matrix.

4. The method according to claim 1, characterized in that, The step of fusing the concept channel knowledge state sequence and the problem channel knowledge state sequence through a dynamic fusion mechanism to obtain a fused knowledge state sequence includes: For each time step t, the fused knowledge state sequence h_fused is calculated according to the following formula, expressed as: h_fused = h_concept + (w + m)⊙h_question Where h_fused represents the fused knowledge state, h_concept represents the concept channel knowledge state, h_question represents the question channel knowledge state, w is the learnable weight vector, m is the content modulation vector dynamically calculated based on the current input content, and ⊙ represents element-wise multiplication.

5. The method according to claim 1, characterized in that, The step of using the representation of the exercise to be predicted as a query to perform time-decayed attention operation on the fused knowledge state sequence to obtain the final knowledge state associated with the exercise to be predicted includes: The query vector is used as a representation of the exercise to be predicted, and the key vector and value vector are used as a sequence of fused knowledge states. The query vector, key vector, and value vector are input into the time decay attention module for calculation; The output of the time decay attention module is the final knowledge state h_t associated with the exercise to be predicted.

6. The method according to claim 1, characterized in that, The step of calculating the probability that a student correctly answers the exercise to be predicted, based on the final knowledge state and the representation of the exercise to be predicted, in the prediction phase includes: The final knowledge state h_t is concatenated with the representation e_{t+1} of the exercise to be predicted to obtain a combined feature vector; The combined feature vector is input to the prediction module, which includes at least one fully connected layer. The prediction module processes the vector and finally outputs a probability value p_{t+1} between 0 and 1 through an activation function. This probability value represents the probability that the model predicts that the student will correctly answer the exercise to be predicted.

7. A dual-channel knowledge tracking device based on adaptive decaying attention, characterized in that, The knowledge tracking device is constructed and generated using the knowledge tracking method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Knowledge tracking method based on knowledge space

    CN117371528A

  • Methods and systems for creating personalized math answer keys for top students

    KR102878631B1