A knowledge tracking method based on contrastive learning and memory mechanism

By comparing learning technology and information processing models, the problem of difficulty in extracting local and global information in students' learning history in the existing technology is solved, accurate modeling of students' knowledge status and prediction of future learning performance is achieved, and large-scale personalized teaching is supported.

CN115906997BActive Publication Date: 2025-06-06HUAZHONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211312281.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2025-06-06
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

The existing knowledge tracking method based on human memory mechanism is difficult to effectively extract local and global information in student learning history, and is affected by the sparse characteristics of educational data, resulting in difficulty in accurately inferring knowledge status.

Method used

Using a knowledge tracking method based on contrast learning and memory mechanism, we learn transferable sequence enhancement representation from sparse learning history through contrast learning technology, and combine information processing models, model knowledge update process to extract local and global information, and predict students' future learning performance.

Benefits of technology

It realizes scientific and comprehensive modeling of students' knowledge status during learning, can accurately predict students' future learning performance, and supports large-scale personalized teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115906997B_ABST
    Figure CN115906997B_ABST
Patent Text Reader

Abstract

The present invention relates to the fields of educational big data mining, contrastive learning and student behavior modeling, and provides a knowledge tracking method based on contrastive learning and memory mechanism, including: (1) realizing sequence enhancement representation; (2) modeling knowledge update process; (3) predicting students' future learning performance. The present invention uses contrastive learning, natural language processing, convolutional neural network, time series modeling and other technical methods, based on information processing model, to systematically and deeply mine students' behavior patterns, and can scientifically and comprehensively model the change process of students' knowledge status and predict students' learning situation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of educational big data mining, contrastive learning and student behavior modeling, and in particular to a knowledge tracking method based on contrastive learning and memory mechanism. Background Art

[0002] The continuous development of intelligent tutoring systems and educational big data technology has provided strong technical support for large-scale personalized teaching. By analyzing each student's learning history data, personalized feedback and learning resource recommendations are provided to them. A key issue in achieving personalized analysis of students is to track the changes in students' knowledge levels over time based on their historical learning trajectories, with students as the center, so as to accurately predict their performance in future learning. This is also called the knowledge tracking problem. The main task of knowledge tracking is to model the changes in students' knowledge mastery during the learning process and predict their future learning performance.

[0003] The goal of knowledge tracking is to simulate the representation of students' knowledge states by analyzing their learning history trajectories. Knowledge states represent the degree of mastery of skills during the learning process; however, the learning process is affected by many cognitive factors, especially human memory; the existing knowledge tracking method HMN based on human memory mechanism, although it simulates the working memory model, assumes that the sum of the capacity of working memory and long-term memory is fixed, and the capacity of the two is in a state of increase and decrease, which is inconsistent with the working memory and long-term memory in actual research; and HMN cannot effectively extract local and global information from students' learning history. Whether this information can be extracted plays an important role in modeling human memory mechanisms. In addition, HMN is still affected by the sparse characteristics of educational data. The representation learned from sparse data sets is prone to bias or overfitting, which hinders the accurate inference of the underlying knowledge state. In order to remedy this problem, the present invention intends to use contrastive learning technology to achieve enhanced representation of sequence data. This method can learn generalizable representations from sparse learning history.

[0004] At present, there are two studies on contrastive learning knowledge tracking: Bi-CLKT and CL4KT. Among them, Bi-CLKT first introduced contrastive learning into the field of knowledge tracking. It designed an end-to-end contrastive learning framework to perform "question to question" (E2E) and "knowledge point to knowledge point" (C2C) association information discrimination at the global and local levels. CL4KT adopts an end-to-end architecture, combines contrastive learning with Transformer, and proposes four data enhancement methods for the field of knowledge tracking. Both studies adopt an end-to-end contrastive learning framework. The application effect of this architecture is affected by the number of negative samples, which is limited by the batch size. If the model is to achieve the expected effect, a larger batch size needs to be set, which will increase the amount of calculation. Therefore, this architecture has high requirements for training equipment, which is not conducive to the purpose of large-scale personalized teaching. In order to solve this problem, this study intends to adopt a contrastive learning framework based on negative sample queues and generative architectures. These two architectures effectively solve the problem that the number of negative samples is limited by the batch size, providing support for the realization of large-scale personalized teaching. Summary of the invention

[0005] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and to provide a knowledge tracking method based on contrastive learning and memory mechanism. It comprehensively utilizes contrastive learning, natural language processing, convolutional neural network, time series modeling and other technical methods, and based on the information processing model, systematically conducts in-depth mining of student behavior patterns, which can scientifically and comprehensively model the changing process of students' knowledge status and predict students' learning situation.

[0006] The purpose of the present invention is achieved through the following technical measures.

[0007] A knowledge tracking method based on contrastive learning and memory mechanism comprises the following steps:

[0008] (1) Realizing sequence enhancement representation: Based on contrastive learning technology, a transferable sequence enhancement representation is learned from sparse learning history, including data enhancement of student learning history, vectorized representation of student learning history, updating of negative sample queues and parameters, and calculation of contrast loss;

[0009] (2) Modeling knowledge updating process: Based on the information processing model, the updating process of modeling knowledge effectively extracts local and global information from students' learning history, including inputting information into the perceptual memory module, processing and storing information in the working memory module, and storing and retrieving information in the long-term memory module;

[0010] (3) Predicting students’ future learning performance: Based on modeling the changes in students’ knowledge status, students’ future learning performance can be predicted.

[0011] In the above technical solution, the step (1) of achieving sequence enhancement characterization is specifically as follows:

[0012] (1-1) Data enhancement of students’ learning history: Four data enhancement methods, namely, question masking, question replacement, interaction sequence interception, and interaction sequence shuffling, are used in combination. A random data enhancement strategy is adopted to perform data enhancement on students’ question sequences and learning interaction sequences (calculated by combining question sequences and answer sequences). Two sequences amplified from the same sequence are considered as positive sample pairs, and two sequences amplified from different sequences are considered as negative sample pairs.

[0013] (1-2) Vectorized representation of student learning history: For the question dimension, the question sequence after data enhancement is input into the encoding layer to obtain the vectorized representation corresponding to the question in the student’s learning history; for the student dimension, the learning interaction sequence after data enhancement is input into the encoding layer and the projection layer in turn to obtain the vectorized representation corresponding to the learning interaction in the student’s learning history;

[0014] (1-3) Update of negative sample queue and parameters and calculation of contrast loss: The negative sample queue is updated through the queue entry and exit functions, and a loss function called InfoNCE is used to calculate the loss value of contrastive learning. It is used to measure the similarity of sample pairs in the representation space, and then the momentum update of parameters is achieved through gradient calculation and back propagation.

[0015] In the above technical solution, the modeling knowledge updating process in step (2) is specifically as follows:

[0016] (2-1) Input information into the perceptual memory module: This module simulates perceptual memory through a convolutional neural network and extracts local information of the student’s learning history based on a sliding window, where the sliding window size will be set as a parameter to be optimized during the training process; first, the student’s test question sequence and learning interaction sequence are encoded using the Embedding method to obtain the test question vector and interaction vector, and then the obtained test question vector and learning interaction vector are input into the convolutional neural network to complete the transformation operation, and finally the test question vector and learning interaction vector that aggregate the local information are obtained;

[0017] (2-2) Working memory module processes and stores information: This module simulates working memory through the Transformer neural network to realize the functions of information processing and storage. The question encoder and knowledge encoder based on the attention mechanism will respectively realize the global information extraction of the question vector and the learning interaction vector. First, the question vector and the learning interaction vector that aggregate local information are input into the question encoder and the knowledge encoder respectively. Secondly, different weights are assigned to the elements in each vector and they are fused. Finally, the question vector and the learning interaction vector that aggregate global information are obtained.

[0018] (2-3) The long-term memory module realizes the storage and retrieval of information: This module simulates long-term memory through a matrix structure to realize the storage function of long-term memory; first, the learning interaction vector is written into the long-term memory matrix through a multi-layer perceptron to update the long-term memory matrix, and then the content in the long-term memory matrix is ​​retrieved through another multi-layer perceptron. Finally, the retrieved memory vector is fused with the learning interaction vector in the Transformer to obtain the student’s current knowledge state vector.

[0019] In the above technical solution, the prediction of students' future learning performance described in step (3) is specifically as follows: after extracting local and global information from all students' learning histories, firstly, the question vector and knowledge state vector finally obtained in step (2) are input into the knowledge retriever (composed of Transformer blocks, which implement the retrieval of question vectors and knowledge state vectors based on the attention mechanism), then the retrieved new knowledge state vector and the embedding vector of the current question are concatenated to obtain a new embedding vector of the current question and the knowledge state, and finally, they are input into a fully connected network and a sigmoid function in sequence to generate a predicted probability of the student correctly answering the current question and predict the student's answer to the current question.

[0020] The present invention is based on a knowledge tracking method of contrastive learning and memory mechanism, adopts contrastive learning technology in the field of deep learning to learn transferable representations from sparse learning history, and at the same time, combines Gagne's information processing theory model to build a knowledge tracking model, simulate the information processing process of learning and memory, realize the modeling of the knowledge state of students in the learning process, and predict students' future learning performance. The present invention can scientifically and comprehensively predict students' learning situation, and provide support for intelligent tutoring systems to carry out large-scale personalized teaching. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a framework diagram of the knowledge tracking model in an embodiment of the present invention.

[0022] Figure 2 This is an example diagram of the structure of the negative sample queue.

[0023] Figure 3 This is the information processing model diagram. DETAILED DESCRIPTION

[0024] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0025] like Figure 1 As shown, an embodiment of the present invention provides a knowledge tracking method based on contrastive learning and memory mechanism, comprising the following steps:

[0026] (1) Realizing sequence enhancement representation

[0027] First, define the important network structures of the contrastive learning part: negative sample queue, encoding layer, projection layer and prediction layer. The negative sample queue is a queue structure used to store negative sample features, so it has a first-in-first-out feature and can be updated during training; the encoding layer consists of a basic encoder and a momentum encoder. Each encoder is composed of a Transformer block, which shares parameters with the encoder in the modeling knowledge update module; the projection layer consists of a basic projection block and a momentum projection block. Each projection block is composed of a linear function, an activation function, and a regularization function; the prediction layer is composed of a linear function and an activation function. Among them, the parameters of the momentum encoder and the momentum projection block are initialized by the basic encoder and the basic projection block; during the training process, the parameters are updated by momentum update. The formula for the momentum update parameters is as follows:

[0028] param T =base_param T *(1-m)+param T-1 *m

[0029] Among them, param T is the momentum encoder or momentum projection block parameter at the current time T, param T-1 is the momentum encoder or momentum projection block parameter at the previous time T-1, base_param T is the basic encoder or basic projection block parameter at the current time T, and m is the hyperparameter of momentum update.

[0030] For example: The structure of the negative sample queue is as follows Figure 2 For example, Figure 2 In the figure, the dimension of the negative sample queue is (k, d), where k represents the length of the queue and d is the dimension of the negative sample features. During the training process, the earliest batch of negative sample features are dequeued each time, and the latest batch of negative sample features are queued.

[0031] The problem to be solved in realizing sequence enhanced representation is to train the encoder in the modeling knowledge update module so that it can better learn transferable representations from sparse learning history. The problem is specifically stated as follows: given a student's learning history sequence, a contrastive learning model is constructed. The model is trained to obtain an encoder that can well learn transferable sequence representations, and this encoder will be used in the modeling knowledge update module.

[0032] The steps to build a contrastive learning framework include data augmentation of student learning history, vectorized representation of learning history, updating of negative sample queues and parameters, and calculation of contrastive loss.

[0033] (1-1) Data enhancement of student learning history

[0034] First, the student's test sequence q_seq and answer sequence r_seq are used to obtain the student's learning interaction sequence through the following formula:

[0035] x_seq=q_seq+q_num*r_seq

[0036] Among them, x_seq is the student's learning interaction sequence, q_seq is the student's test question sequence, r_seq is the student's answer sequence, and q_num is the total number of knowledge points involved in the dataset.

[0037] Secondly, four data augmentation methods, namely, question masking, question replacement, interactive sequence interception, and interactive sequence shuffling, are used to enhance the question sequence q_seq to obtain a new question sequence q_seq1. Then, the above data augmentation methods are repeatedly used on the question sequence to obtain the question sequence q_seq2, and q_seq1 and q_seq2 are positive sample pairs.

[0038] Finally, repeat the operation of the test sequence q_seq on the student's learning interaction sequence x_seq to obtain new learning interaction sequences x_seq1 and x_seq2, and x_seq1 and x_seq2 are positive sample pairs.

[0039] (1-2) Vectorized representation of learning history

[0040] First, the test question sequence q_seq1 is input into the basic encoder to obtain the vectorized representation q_query of the test question, and q_seq2 is input into the momentum encoder to obtain q_key. Since q_seq1 and q_seq2 are positive sample pairs, the vectorized representations q_query and q_key of the test question are also positive sample pairs, and the vectorized features stored in q_query and the negative sample queue are negative sample pairs.

[0041] Secondly, the learning interaction sequence x_seq1 is input into the basic encoder to obtain the vectorized representation x_query of the learning interaction, and x_seq2 is input into the momentum encoder to obtain x_key.

[0042] Finally, the vectorized representation x_query of the learning interaction is input into the basic projection block and mapped into the student’s knowledge state vector ks1, and x_key is input into the basic projection block and mapped into ks2.

[0043] (1-3) Negative sample queue and parameter update and comparison loss calculation

[0044] First, the cosine similarity of the vectorized representations q_query and q_key of the test question is calculated to obtain sim1_1. The cosine similarity of q_query and all features in the negative sample queue is calculated to obtain sim1_2. Sim1_1 and sim1_2 are concatenated to obtain sim1, which is input into the loss function InfoNCE to obtain the loss value CL_loss1 of the test question.

[0045] Secondly, the knowledge state vector ks1 is input into the prediction layer to obtain the predicted knowledge state vector ks3, the knowledge state vector ks2 is predicted, the cosine similarity of the knowledge state vectors ks2 and ks3 is calculated to obtain sim2, which is input into the loss function InfoNCE to obtain the loss value CL_loss2 of the test question.

[0046] Finally, the batch test feature x_key is used to update the negative sample queue through the queue entry and exit functions; and the momentum of the model parameters is updated through back propagation and gradient calculation based on the momentum parameter update formula.

[0047] (2) Modeling knowledge updating process

[0048] In order to better model the changes in students' knowledge status during learning and memorization, a neural network model is constructed based on the information processing model to track the changes in students' knowledge status through their learning history.

[0049] (2-1) Input information into the perception memory module

[0050] First, the convolutional neural network structure CNN is defined, and the sliding window size is set as a parameter and optimized during training. The convolutional neural network is used to simulate the perceptual memory module and extract local information of the student's learning history based on the sliding window.

[0051] Secondly, read the student's test question sequence q_seq and answer sequence r_seq, generate the learning interaction sequence x_seq by combining them, and use the Embedding encoding method to obtain the test question embedding vector q and the learning interaction embedding vector x for the test question sequence q_seq and the learning interaction sequence x_seq.

[0052] Finally, the test question embedding vector q and the learning interaction embedding vector x are respectively input into the convolutional neural network to obtain the test question vector Q and the learning interaction vector X that aggregate local information.

[0053] (2-2) Working memory modules process and store information

[0054] First, define the Transformer block and use it to build the question encoder, knowledge encoder, and knowledge retriever. Use the Transformer neural network to simulate the working memory module to realize the information processing and storage functions, as well as the global information extraction of the question and learning interaction vector.

[0055] Secondly, the question vector Q that aggregates local information is input into the question encoder, where the vector Q is used as the query, key, and value parameters in the attention mechanism to calculate the attention weight, thereby assigning different weights to each element in the vector Q to obtain the question vector Q′ that aggregates global information.

[0056] Finally, the interaction vector X that aggregates local information is input into the knowledge encoder, where vector X is used as the query, key, and value parameters in the attention mechanism to calculate the attention weight, thereby assigning different weights to each element in vector X to obtain the interaction vector X′ that aggregates global information.

[0057] (2-3) Long-term memory module realizes information storage and retrieval

[0058] First, construct a long-term memory matrix with a matrix dimension of (w, q_num). Among them, w is the long-term memory capacity, q_num is the total number of knowledge points involved in the data set, and the long-term memory matrix is ​​used to store the students' knowledge mastery status at different times. In addition, construct write_heads and read_heads function blocks, which are used to write or read the learning interaction vector into the long-term memory matrix respectively. Both function blocks are composed of linear functions and activation functions.

[0059] Secondly, the learning interaction vector X′ that aggregates the global information is written into the long-term memory matrix through the write_heads function to update the long-term memory matrix.

[0060] Finally, for the updated long-term memory matrix, use read_heads to read the memory vector M of the matrix at the current time, concatenate the learning interaction vector X′ with the memory vector M, and remap its dimension to the dimension of X′ through the multi-layer perceptron, and finally obtain the current student’s knowledge state vector X″.

[0061] (3) Predicting students’ future learning performance

[0062] First, after extracting local and global information from all students’ learning histories, the final test question vector Q′ and knowledge state vector X″ are input into the knowledge retriever. The test question vector Q′ is used as the query and key parameters in the attention mechanism, and the knowledge state vector X″ is used as the value parameter to obtain the retrieved knowledge state vector H.

[0063] Secondly, the currently retrieved knowledge state vector H and the embedding vector of the current question are q Perform concatenation operation to obtain a new vector H′.

[0064] Finally, the new vector H′ is input into a fully connected network and finally passes through the sigmoid activation function to generate the predicted probability that the student will correctly answer the current question. It represents the probability that the student will successfully answer the current question at time T.

[0065] The contents not described in detail in this specification belong to the prior art known to those skilled in the art.

[0066] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A knowledge tracking method based on contrastive learning and memory mechanism, Features The method comprises the following steps: (1) Realize sequence enhancement representation: Based on contrastive learning technology, learn transferable sequence enhancement representation from learning history, including data enhancement of student learning history, vectorized representation of student learning history, update of negative sample queue and parameters, and calculation of contrast loss; specifically: Define the network structure of the contrastive learning part: negative sample queue, encoding layer, projection layer and prediction layer; the negative sample queue is a queue structure used to store negative sample features. It has a first-in-first-out feature and can update the queue during training; the encoding layer consists of a basic encoder and a momentum encoder. Each encoder is composed of a Transformer block, which shares parameters with the encoder in the modeling knowledge update module; the projection layer consists of a basic projection block and a momentum projection block. Each projection block is composed of a linear function, an activation function and a regularization function; the prediction layer is composed of a linear function and an activation function; the parameters of the momentum encoder and the momentum projection block are initialized by the basic encoder and the basic projection block; during training, the parameters are updated by momentum updating; (1-1) Data enhancement of students’ learning history: Four data enhancement methods, namely, question masking, question replacement, interaction sequence interception, and interaction sequence shuffling, are used in combination. A random data enhancement strategy is adopted to enhance the students’ question sequence and learning interaction sequence. The learning interaction sequence is calculated by combining the question sequence and the answer sequence. Two sequences amplified from the same sequence are considered as positive sample pairs, and two sequences amplified from different sequences are considered as negative sample pairs. (1-2) Vectorized representation of student learning history: For the question dimension, the question sequence after data enhancement is input into the encoding layer to obtain the vectorized representation corresponding to the question in the student’s learning history; for the student dimension, the learning interaction sequence after data enhancement is input into the encoding layer and the projection layer in turn to obtain the vectorized representation corresponding to the learning interaction in the student’s learning history; (1-3) Update of negative sample queue and parameters and calculation of contrast loss: The negative sample queue is updated through the queue entry and exit functions, and the loss function InfoNCE is used to calculate the loss value of contrastive learning, which is used to measure the similarity of sample pairs in the representation space. Then, the momentum update of parameters is achieved through gradient calculation and back propagation. (2) Modeling knowledge updating process: Based on the information processing model, the updating process of modeling knowledge extracts local and global information from students’ learning history, including inputting information into the perceptual memory module, processing and storing information in the working memory module, and storing and retrieving information in the long-term memory module; (3) Predicting students’ future learning performance: Based on modeling the changes in students’ knowledge status, students’ future learning performance can be predicted.

2. The knowledge tracking method based on contrastive learning and memory mechanism according to claim 1, Features The modeling knowledge updating process described in step (2) is specifically as follows: (2-1) Input information into the perceptual memory module: This module simulates perceptual memory through a convolutional neural network and extracts local information of the student’s learning history based on a sliding window, where the sliding window size will be set as a parameter to be optimized during the training process; first, the student’s test question sequence and learning interaction sequence are encoded using the Embedding method to obtain the test question vector and interaction vector, and then the obtained test question vector and learning interaction vector are input into the convolutional neural network to complete the transformation operation, and finally the test question vector and learning interaction vector that aggregate the local information are obtained; (2-2) Working memory module processes and stores information: This module simulates working memory through the Transformer neural network to realize the functions of information processing and storage. The question encoder and knowledge encoder based on the attention mechanism will respectively realize the global information extraction of the question vector and the learning interaction vector. First, the question vector and the learning interaction vector that aggregate local information are input into the question encoder and the knowledge encoder respectively. Secondly, different weights are assigned to the elements in each vector and they are fused. Finally, the question vector and the learning interaction vector that aggregate global information are obtained. (2-3) The long-term memory module realizes the storage and retrieval of information: This module simulates long-term memory through a matrix structure to realize the storage function of long-term memory; first, the learning interaction vector is written into the long-term memory matrix through a multi-layer perceptron to update the long-term memory matrix, and then the content in the long-term memory matrix is ​​retrieved through another multi-layer perceptron. Finally, the retrieved memory vector is fused with the learning interaction vector in the Transformer to obtain the student’s current knowledge state vector.

3. The knowledge tracking method based on contrastive learning and memory mechanism according to claim 1, Features The prediction of students' future learning performance described in step (3) is specifically as follows: after extracting local and global information from all students' learning histories, firstly, the question vector and knowledge state vector finally obtained in step (2) are input into the knowledge retriever, and the knowledge retriever is composed of Transformer blocks, which realizes the retrieval of question vectors and knowledge state vectors based on the attention mechanism. Secondly, the retrieved new knowledge state vector and the embedding vector of the current question are concatenated to obtain a new embedding vector of the current question and the knowledge state. Finally, they are input into a fully connected network and a sigmoid function in sequence to generate a predicted probability of the student correctly answering the current question and predict the student's answer to the current question.

Citation Information

Patent Citations

  • Knowledge tracking method based on test question heterogeneous graph representation and learner embedding

    CN113344053A

  • Artificial cognitive declarative-based memory model to dynamically store, retrieve, and recall data derived from aggregate datasets

    US20180240015A1