Large model knowledge tracking method based on plug-and-play instruction

By using a large-modal method of plug-and-play instructions in knowledge tracking, integrating multimodal information, the problem of insufficient accuracy and efficiency of knowledge tracking in the existing technology is solved, and more accurate prediction of students' knowledge status is achieved.

CN119938869APending Publication Date: 2025-05-06EAST CHINA NORMAL UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510166090.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing knowledge tracking methods are difficult to make full use of rich world knowledge for reasoning, and are difficult to capture student behavior patterns, resulting in insufficient accuracy and efficiency of knowledge tracking.

Method used

The large-model knowledge tracking method based on plug-and-play instructions is adopted. Through the plug-in context and plug-in sequence module, the large language model and the traditional sequence model are effectively integrated, and long context knowledge and interactive sequences are captured to achieve flexible integration of multimodal information.

Benefits of technology

It realizes more accurate prediction of students' knowledge status, highlights the ability to understand students' learning behavior and knowledge status, and significantly improves the efficiency and accuracy of knowledge tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938869A_ABST
    Figure CN119938869A_ABST
Patent Text Reader

Abstract

The invention discloses a large model knowledge tracking method based on a plug-and-play instruction, which is characterized in that the plug-and-play instruction is constructed by adopting an instruction template containing a specific problem and a specific concept slot, the integration of multi-modal information in a large language model is realized, and LLMs (Logical Language Models) utilizes rich knowledge and strong reasoning ability of the LLMs to realize the knowledge tracking of the large model. The student knowledge state prediction method specifically comprises the steps of plug-and-play instruction construction, context insertion module processing, sequence insertion module processing, model training and reasoning and the like. Compared with the prior art, the method has the advantages that multi-modal information is integrated, the knowledge tracking effect is improved by utilizing rich knowledge and strong reasoning ability of a large language model, the application potential in an intelligent education system is huge, teachers can be assisted to make personalized teaching plans, a powerful tool is provided for education research, personalized education development is promoted, and the method is worthy of popularization and application. The method improves the efficiency of computer-aided education, and has a remarkable application value and a good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of personalized education and artificial intelligence technology, and more specifically, to a large model knowledge tracking method based on plug-and-play instructions. Background Art

[0002] In today's education field, knowledge tracking, as a key technology for personalized education, plays a vital role in improving the quality and efficiency of education. Its core goal is to accurately predict the accuracy of students' answers to subsequent questions based on their past question and answer records. This technology can help teachers and education systems gain in-depth insights into students' knowledge mastery, skill levels, and forgetting patterns, thereby tailoring teaching plans for students, providing accurately adapted learning resources, significantly enhancing the effectiveness of computer-assisted education, and achieving personalized and efficient education.

[0003] Traditional knowledge tracking methods have many limitations. Models based on deep learning mainly rely on question IDs to learn interaction behavior sequences. Although models such as LSTM and Transformer have achieved certain results in capturing the representation of problem-solving records, and some studies have also introduced time factors and used graph neural networks to strengthen the interactive representation of ID sequences, these models generally do not adequately mine the rich semantic knowledge contained in the question text. Pre-trained language models (PLMs), such as BERT, have been introduced into knowledge tracking to process question text information, but due to their limited reasoning ability and world knowledge acquisition ability, it is difficult to effectively mine the logic and reasons behind the question sequence. Although large language models (LLMs) have shown strong capabilities in natural language processing tasks and have the advantages of generation, reasoning, and instruction following, they face severe challenges in knowledge tracking applications. On the one hand, they find it difficult to capture sequential interaction behaviors from ID sequences that reflect students' knowledge status, and often split IDs into multiple tags, resulting in a large loss of semantic information; on the other hand, when processing the long text context of comprehensive problem-solving records, even powerful models such as LLaMA and GPT-4o cannot effectively learn students' knowledge status from them.

[0004] In summary, existing knowledge tracking methods have obvious deficiencies in making full use of knowledge for reasoning and accurately capturing student behavior patterns. This urgently requires an innovative solution that can integrate the advantages of multiple technologies to overcome these difficulties, thereby achieving more accurate and efficient knowledge tracking and meeting the evolving needs of personalized education. Summary of the invention

[0005] The purpose of the present invention is to design a large model knowledge tracking method based on plug-and-play instructions in view of the deficiencies of the prior art. The plug-in context and plug-in sequence modules are used to construct a large model of plug-and-play instructions, and the large language model and the traditional sequence model are effectively integrated into multimodal information to capture long context knowledge and interaction sequences, so that LLMs can make full use of their rich knowledge and powerful reasoning ability to achieve accurate prediction of students' knowledge status. The method is simple and has outstanding ability to understand students' learning behavior and knowledge status. It plays an important role in knowledge tracking tasks and provides strong support for accurately predicting students' learning performance. It can be applied to intelligent education systems to assist teachers in accurately grasping students' knowledge status, formulating personalized teaching plans, and improving the efficiency of computer-assisted education. It effectively solves the problems of not being able to fully utilize rich world knowledge for reasoning and difficulty in capturing students' behavior patterns, and achieves more accurate prediction of students' knowledge status. It has significant application value and good application prospects.

[0006] The specific technical solution for achieving the purpose of the present invention is: a large model knowledge tracing method (LLM-KT) based on plug-and-play instructions, which is characterized in that the specific operation of the method is implemented according to the following steps:

[0007] Step 1: Plug and Play Instructions Build

[0008] Plug-and-play instructions are an important component of the LLM-KT framework, which is designed to closely align large language models (LLMs) with knowledge tracing tasks at the task level, achieve flexible integration of multimodal information in LLMs, and then guide LLMs to accurately capture changes in students' learning behaviors and knowledge states.

[0009] The design of building plug-and-play instructions includes instruction templates for specific questions and specific concept slots. The input part of the instruction template integrates the student's historical question and answer records and the target question information that currently needs to be predicted, and the output part is the prediction result of the LLM-KT instance model on the correctness of the student's answer to the target question.

[0010] Specifically, the input content of the build Contains historical question and answer records of students in chronological order and target issues Related information. Taking a record in the historical question-answering record as an example, its format is "questionQID=38 [QuesEmbed38] involving concept CID=219 [ConcEmbed219] correctly", where QID stands for question ID, CID stands for concept ID, [QuesEmbed38] and [ConcEmbed219] are specific tags for question 38 and concept 219, and the embeddings of these tags are learned through plugin context and plugin sequence. The format of the target question information is similar, but does not contain information about whether the answer is correct, such as "question QID=55 [QuesEmbed55] involving concept CID=245 [ConcEmbed245]". Organizing this information in a specific order to form a complete input string allows LLMs to understand the context of each question and the student's performance related to it.

[0011] Step 2: Context Insertion Module Processing

[0012] In the entire technical solution, the context insertion module is responsible for deep processing of the text of questions and knowledge concepts to help the model better capture the semantic information and provide support for accurate knowledge tracking. The processing of this module mainly includes two steps: context encoder encoding and context adapter mapping:

[0013] (1) Context encoder encoding: Select a suitable pre-trained language model as the context encoder, such as LLaMA2, BERT or all-mpnet-base-v2. These pre-trained language models are trained on large-scale text data and have powerful text feature extraction capabilities. Take the question text as an example, input it into the context encoder, and the encoder will parse and encode the vocabulary, grammar, semantics and other information in the text. In this process, the model will convert the text into a vector representation with a specific dimension based on the language knowledge and patterns it has learned, that is, the text question representation For concept text, the concept representation is also obtained through the context encoder The specific formula is as follows:

[0014] ; .

[0015] Different pre-trained language models may have different encoding methods and vector dimensions. For example, the BERT model encodes the input text through a multi-layer Transformer structure, and its hidden layer dimension is 768; the LLaMA2 model encodes based on its own architecture, and the dimension setting may vary depending on the specific version; the all-mpnet-base-v2 model is pre-trained with a specially designed task, and the vector dimension obtained is 4096. These vector representations contain rich semantic information, which can reflect the potential relationship between questions and concepts, and provide an important data foundation for subsequent knowledge tracking tasks.

[0016] (2) Context adapter mapping: Since the text representation space obtained by the context encoder may be different from the semantic space of LLMs, in order to ensure that these representations can be effectively integrated into LLMs, a context adapter is needed for mapping. Here, a simple multi-layer perceptron (MLP) layer is used as a context adapter. The MLP layer consists of multiple neurons, and by setting different weights and biases, it can perform nonlinear transformations on the input vector. The text question representation obtained by the context encoder is mapped to and conceptual representation They are input into the context adapter respectively, processed by the MLP layer, and output a representation compatible with the LLMs semantic space and , the formula is as follows:

[0017]

[0018]

[0019] In this process, the context adapter learns how to transform the original text representation into a form more suitable for LLMs processing, so that LLMs can better understand and utilize this information. For example, if the embedding layer dimension of LLMs is (assuming 4096), and the vector dimension of the context encoder output is (such as BERT's 768), the context adapter maps the 768-dimensional vector to a 4096-dimensional space through a series of linear transformations and nonlinear activation functions to ensure effective transmission and fusion of information. Through the mapping of the context adapter, the text information of questions and concepts can be better combined with the knowledge system of LLMs, providing strong support for the model to accurately capture the knowledge status of students in knowledge tracking tasks.

[0020] Step 3: Sequence Insertion Module Processing

[0021] The sequence insertion module is another important part of the LLM-KT technical solution. It is committed to effectively integrating new modalities such as ID sequences into large language models (LLMs), so that the model can fully learn semantic and interactive behavior information in knowledge tracking tasks. The processing of this module mainly covers two important links: sequence encoder learning and sequence adapter conversion:

[0022] (1) Sequence encoder learning: The sequence composed of question ID and concept ID is regarded as a new modality, and a traditional sequence learning model such as deep knowledge tracking (DKT) or attention knowledge tracking (AKT) is used as the sequence encoder. These traditional models have unique advantages in processing ID-based sequence interactions and can mine the potential relationship and interaction pattern between question ID and concept ID by learning from historical data. Taking the question ID sequence as an example, the sequence encoder is learned to obtain the question ID embedding ; Similarly, for the concept ID sequence, we get the concept ID embedding , the specific calculation formula is as follows:

[0023] ;

[0024] .

[0025] in, Represents the vector dimension of the sequence encoder output, which reflects the model's extraction dimension of the ID sequence features. Different sequence encoders have different network structures and learning capabilities, and the output vector dimensions may also vary. For example, the DKT model is based on the recurrent neural network (RNN) structure, and learns the feature representation of the question ID at different time steps by gradually processing the time series data; the AKT model introduces an attention mechanism, which can pay more attention to the key information in the sequence, thereby obtaining a more representative ID embedding. These embedding vectors contain rich sequence interaction information, which provides an important basis for subsequent models to understand students' learning behavior.

[0026] (2) Sequence adapter conversion: Since the semantic space of traditional sequence models is significantly different from that of large language models, directly inputting the ID embedding obtained by the sequence encoder into LLMs may cause information mismatch and affect model performance. Therefore, it is necessary to introduce a sequence adapter to convert the ID representation learned by the traditional model into the semantic space of LLMs. Specifically, the question ID is embedded in and concept ID embedding They are input into the sequence adapter respectively, and after a series of transformation operations, a representation aligned with the LLMs semantic space is obtained. and , the formula is as follows:

[0027] ;

[0028] .

[0029] The sequence adapter can usually be designed as a structure containing multiple layers of neural networks, which learn how to effectively convert between different semantic spaces through training. During the conversion process, it will adjust and optimize the input ID embedding according to the characteristics and requirements of LLMs, so that the converted representation can be better integrated with LLMs, thereby enhancing the performance of the model in knowledge tracking tasks. Through the conversion of the sequence adapter, LLMs can better utilize the sequence interaction information learned by the traditional sequence model and improve the understanding and prediction of students' knowledge status.

[0030] Step 4: Model training and inference

[0031] During the training phase, the constructed content Input to Training is carried out in is a trainable parameter of the model. The training objective uses the cross entropy loss of the language model. By minimizing this loss, the model parameters are adjusted so that the model can more accurately predict the correctness of the student's answer to the target question. In order to improve the training efficiency and reduce the number of parameter updates, the low-rank adaptation (LoRA) method is adopted. The core idea of ​​LoRA is to adjust the model parameters through low-rank decomposition. It adds two low-rank matrices to the model's weight matrix and , during fine-tuning, only these two low-rank matrices are updated , while keeping the original weight matrix unchanged. This method significantly reduces the computational and storage costs, making it possible to efficiently train the model even with limited resources.

[0032] In the inference phase, input the constructed content , the instance model will output a prediction of the correctness of the student’s answer to the target question. Specifically, by calculating the probability of predicting “Yes” To determine the probability that a student answers correctly, the calculation formula is:

[0033] .

[0034] in, and are the probabilities of the two tags “Yes” and “No” output by the instance model respectively.

[0035] when When it is greater than 0.5, the model is considered to predict that the student answers the target question correctly; otherwise, it predicts that the student answers incorrectly. Through this plug-and-play instruction construction method, LLMs can make full use of their rich knowledge and powerful reasoning ability, play an important role in knowledge tracking tasks, and provide strong support for accurately predicting students' learning performance.

[0036] Compared with the prior art, the present invention has the ability to effectively integrate multimodal information with the help of plug-in context and plug-in sequence modules, capture long context knowledge and interaction sequences, and understand students' learning behaviors and knowledge states. The ablation experiment verifies the key role of each component in improving performance. It has good adaptability under different sequence lengths, and the performance improves significantly with the increase of sequence length. It is better than traditional models and has excellent results in the field of knowledge tracking. Experiments show that its performance on four benchmark data sets exceeds that of about 20 strong baseline models, with high prediction accuracy. It has great application potential in intelligent education systems, can help teachers develop personalized teaching plans, provide powerful tools for educational research, promote the development of personalized education, and improve the efficiency of computer-assisted education. It has significant application value and good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a process framework diagram of the present invention. DETAILED DESCRIPTION

[0038] The present invention is further described in detail in conjunction with the following specific examples and drawings. This example is based on the knowledge tracking dataset composed of ASSISTments2009, ASSISTments2015, Junyi Academy and NeurIPS 2020 EducationChallenge, and aims to show in detail the specific operation process and application effect of the LLM-KT method. The process, conditions, experimental methods, etc. for implementing the present invention, except for the contents specifically mentioned below, are common knowledge and common common sense in the field, and the present invention does not specifically limit the contents.

[0039] Example 1

[0040] See also Figure 1 The present invention integrates a large language model with a traditional sequence model and uses plug-and-play instructions to achieve accurate prediction of student knowledge status. The specific steps are as follows:

[0041] Step 1: Plug and Play Instructions Build

[0042] Take a student's answer record on the ASSISTments2009 dataset as an example to build plug-and-play instructions. Assume that the student has answered question QID=38 (involving concept CID=219) correctly, question QID=40 (involving concept CID=230) correctly, and question QID=57 (involving concept CID=204) incorrectly, and the current question to be predicted is QID=55 (involving concept CID=245). According to the construction rules of plug-and-play instructions, the generated input instruction is: "The student has previously, in chronological order, answered question QID=38 [QuesEmbed38] involving concept CID=219 [ConcEmbed219] correctly, question QID=40[QuesEmbed40] involving concept CID=230 [ConcEmbed230] correctly, questionQID=57 [QuesEmbed57] involving concept CID=204 [ConcEmbed204] incorrectly. Please predict whether the student will answer the next question QID=55[QuesEmbed55] involving concept CID=245 [ConcEmbed245] correctly. Responsewith 'Yes' or 'No'." Among them, the embeddings of tags such as [QuesEmbed38] and [ConcEmbed219] will be learned later through the plug-in context and plug-in sequence modules.

[0043] Step 2: Context Insertion Module Processing

[0044] (1) Context encoder encoding: BERT is selected as the context encoder. The text of question QID=38 and the text of concept CID=219 are input into LLaMA2 respectively. According to the formula and , get the text problem representation and conceptual representation Due to the specific architecture and training method of LLaMA2, it performs deep semantic analysis on the input text and converts the text into a vector representation containing rich semantic information.

[0045] (2) Context Adapter Mapping: Multilayer Perceptron (MLP) is used as the context adapter. Assume that the embedding layer dimension of the large language model used is is 4096, and the vector dimension of BERT output (here assumed) is 768. and Enter them into the context adapter respectively, according to the formula and , after a series of linear transformations and nonlinear activation function processing in the MLP layer, the 768-dimensional vector is mapped to a 4096-dimensional space, obtaining a representation that is compatible with the semantic space of the large language model and , ensuring that the textual information of questions and concepts can be effectively integrated into the large language model.

[0046] Step 3: Sequence Insertion Module Processing

[0047] (1) Sequence encoder learning: AKT is selected as the sequence encoder. The QID sequence [38, 40, 57] and CID sequence [219, 230, 204] of the student's answer are input into the AKT model respectively. According to the formula and , the AKT model uses its attention mechanism to learn the question ID embedding and concept ID embedding The AKT model can focus on the key information in the sequence, so that the obtained embedding vector can better reflect the potential relationship and interaction pattern between the question ID and the concept ID.

[0048] (2) Sequence adapter conversion: Design a sequence adapter containing a multi-layer neural network. and Enter them into the sequence adapter respectively, according to the formula and The sequence adapter learns how to convert between different semantic spaces through training, converting the ID representation learned by the traditional AKT model into a representation that is aligned with the semantic space of the large language model. and , so that large language models can effectively utilize these sequence interaction information.

[0049] Step 4: Model training and inference

[0050] Model training: The plug-and-play instructions obtained through the above processing are input into the instance model based on the large language model (such as LLaMA2) for training. The low-rank adaptation (LoRA) method is used for fine-tuning, and the rank of LoRA is set to 32, alpha is set to 32, and dropout is set to 0.1. The Adam optimizer is used, and the learning rate is set to 3× , the weight decay is 1× , and the cosine learning rate scheduler is used to adjust the learning rate. The training is performed for up to 10 epochs, the batch size is set to 32, and the gradient accumulation strategy is used. During the training process, the cross entropy loss of the language model is used as the training target, and the model parameters are continuously adjusted. ( , A and B are two low-rank matrices added by LoRA), which enables the model to more accurately predict the correctness of students’ answers to target questions.

[0051] Model reasoning: In the reasoning stage, the built plug-and-play instructions are input, and the instance model will output a prediction of the correctness of the student's answer to the target question. The probability of predicting "Yes" is calculated according to the following formula :

[0052] .

[0053] Assuming the calculated The value is 0.7, which is greater than 0.5, and the instance model predicts that the student answers the target question (QID=55, ​​involving concept CID=245) correctly.

[0054] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.

Claims

1. A large model knowledge tracking method based on plug-and-play instructions, characterized in that: The plug-and-play instructions are constructed by using instruction templates containing specific questions and specific concept slots to realize the integration of multimodal information in the large language model, so that the large language model can use its rich knowledge and powerful reasoning ability to predict the student's knowledge status. The large model knowledge tracking method includes the following specific steps: Step 1: Plug and Play Instructions Build A plug-and-play instruction is constructed using an instruction template containing specific questions and specific concept slots to achieve the integration of multimodal information in a large language model. The input of the instruction template is the student's historical question and answer records and the target question that needs to be predicted at present, and the output is the prediction result of the large language instance model on the correctness of the student's answer to the target question. Step 2: Context Insertion Module Processing A context insertion module consisting of a context encoder and a context adapter is used, and its processing specifically includes: 2-1: Context Encoder Encoding Select LLaMA2, BERT or all-mpnet-base-v2 pre-trained language model as the context encoder, input the question text and concept text into the context encoder, parse and encode the vocabulary, grammar and semantics in the text, and obtain the text question representation expressed as follows: and conceptual representation : ; ; 2-2: Context Adapter Mapping Using context adapters to represent text questions and conceptual representation Mapping is performed to obtain the following text question representation that is compatible with the semantic space of the large language model: and conceptual representation : ; ; The context adapter is a multi-layer perceptron composed of multiple neurons, which performs nonlinear transformation on the input vector by setting different weights and biases; Step 3: Sequence Insertion Module Processing A sequence insertion module consisting of a sequence encoder and a sequence adapter is used, and its processing specifically includes: 3-1: Sequence Encoder Learning The sequence learning model using deep knowledge tracking or attention knowledge tracking is the sequence encoder. The question ID sequence and concept ID sequence are input into the sequence encoder for learning, and the question ID embedding expressed as follows is obtained: and concept ID embedding : ; ; 3-2: Sequence adapter conversion Use the Sequence Adapter to embed the question ID And the concept ID embedding is transformed to obtain the following question ID representation aligned with the semantic space of the large language model and concept ID : ; ; The sequence adapter is a structure including a multi-layer neural network, which is trained to learn the conversion between different semantic spaces; Step 4: Model training and inference 4-1: Training phase What will be built enter The model is trained using a low-rank adaptation method. Its training objective is to use the cross-entropy loss of the language model and adjust the model parameters by minimizing the cross-entropy loss. is a large language instance model using the LLM-KT framework, where is the training parameter; 4-2: Reasoning stage What will be built Input the large language instance model trained above to predict the correctness of the student's answer to the target question, and the probability of predicting "Yes" Calculated by the following formula: ; in, and The probabilities of "Yes" and "No" output by the large language instance model respectively; 4-3: Prediction stage When the predicted probability If it is greater than 0.5, the large language instance model predicts that the student’s answer to the target question is correct; otherwise, it predicts that the student’s answer is incorrect.

2. The large model knowledge tracking method based on plug-and-play instructions according to claim 1 is characterized in that In, The BERT model encodes the input text through a multi-layer Transformer structure, and its hidden layer dimension is 768; the LLaMA2 model encodes based on its own architecture; the all-mpnet-base-v2 model obtains a vector dimension of 4096 after pre-training tasks.

3. The large model knowledge tracking method based on plug-and-play instructions according to claim 1 is characterized in that: The low-rank adaptation method adjusts the model parameters by low-rank decomposition, which adds two low-rank matrices to the model's weight matrix and , to update these two low-rank matrices during fine-tuning , while the original weight matrix remains unchanged.

4. The large model knowledge tracking method based on plug-and-play instructions according to claim 1 is characterized in that: Contents of the build Contains: Historical question and answer records of students in chronological order and target issues Related information.

Citation Information

Cited By

  • Space-time fusion knowledge tracking method based on large model annotation enhancement

    CN120278253A

  • Knowledge tracking method for student and question joint modeling based on large language model

    CN121327154A

  • Knowledge tracking method based on large language model for joint modeling of students and questions

    CN121327154B