A Few-Shot Exercise Link Knowledge Concept Method Based on Prompt Tuning

Through the method based on prompt tuning, the correlation between large pre-trained language models and unified template calculation exercises and knowledge concepts is solved, and the problem of lack of training data in the intelligent education system is achieved, and efficient knowledge tracking is achieved in the case of few samples.

CN115311113BActive Publication Date: 2025-07-08YANGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210815834.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-12
Publication Date
2025-07-08
Estimated Expiration
2042-07-12

AI Technical Summary

Technical Problem

The existing intelligent education system lacks training data when linking exercises and knowledge concepts, which makes it difficult to effectively classify multi-label texts with few samples, affecting the accuracy of knowledge tracking.

Method used

Using a prompt tuning-based method, the correlation between exercises and knowledge concepts is calculated through large pre-trained language models and unified templates, and the threshold mechanism and learning rate fine-tuning are used to optimize the parameters of the pre-trained language model to improve the accuracy of label marking.

Benefits of technology

With few samples, the effect of labeling the exercises with knowledge concepts is significantly improved, providing a data basis for online learning, and improving the efficiency and accuracy of knowledge tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311113B_ABST
    Figure CN115311113B_ABST
Patent Text Reader

Abstract

The present invention discloses a few-shot exercise-link knowledge concept method based on prompt tuning. The method includes collecting and organizing exercise resources corresponding to a certain course and knowledge concepts to form a data set; using an instant tuning method with a unified template to calculate the relevance between exercises and knowledge concepts, and independently predicting the probability of each concept; through a threshold mechanism, determining the knowledge concepts related to the exercises, tagging the corresponding knowledge concepts for the exercises, and at the same time, by fine-tuning the task model, cyclically training the pre-trained language model to adapt it to the current task, so as to tag the knowledge concepts for the exercises in the case of a small sample size, and improving the effect of few-shot exercise-link knowledge concepts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of adaptive learning and few-shot multi-label text classification, and particularly relates to a few-shot exercise-link knowledge concept method based on prompt tuning. Background Art

[0002] In recent decades, personalized learning has become the mainstream solution in intelligent education systems to improve students' learning interest and learning experience. One of the foundations and key tasks of personalized learning is knowledge tracing, whose purpose is to evaluate the learning status of students' knowledge concepts. Exercises play an important role in the knowledge tracking task and are one of the indicators to evaluate whether students have mastered specific knowledge concepts. Students in intelligent education systems select correct exercises according to their own needs and acquire specific knowledge concepts during the process of doing exercises. That is to say, the changes in the acquisition of knowledge concepts by students during the exercise process can be traced. From this perspective, knowledge tracing should consist of the student-exercise-knowledge concept hierarchy. However, most of the existing knowledge tracing methods are modeled between partial hierarchies (i.e., student-exercise or student-concept). This is because in some intelligent systems, there is a lack of connection between exercises and knowledge concepts.

[0003] Essentially, linking exercises with knowledge concepts is a multi-label text classification (MLTC) problem. The relationship between exercises and knowledge concepts is one-to-one or one-to-many, and the purpose is to assign one or more concepts to each input exercise in the dataset. At the same time, there is a high semantic correlation between exercises and knowledge concepts.

[0004] However, in the multi-label text classification task, deep learning-based methods require a large amount of training data for model optimization, which is usually time-consuming and labor-intensive in real scenarios. Unfortunately, when linking exercises with knowledge concepts, there is usually a lack of training data because for some knowledge concepts corresponding to few exercises or new courses, the labeled data may be scarce. Summary of the Invention

[0005] Object of the Invention: The object of the present invention is to provide a few-shot exercise-link knowledge concept method based on prompt tuning, so as to improve the effect of exercise-link knowledge concepts in the case of few samples.

[0006] Technical Solution: A few-shot exercise-link knowledge concept method based on prompt tuning provided by the present invention includes the following steps:

[0007] S1. Experts label the knowledge concepts corresponding to the course Course, represent the course Course as a label space C = {c1, c2,... c N} with N knowledge concepts, sort out and collect the corresponding exercises, represent them as an exercise instance space E, and for some instances in the exercise instance space E Mark the corresponding set of tags Form a data set

[0008] S2. Calculate the relevance between exercises and knowledge concepts through an instant optimization method with a unified template, and the specific implementation is as follows:

[0009] (2.1) Set M as a large pre-trained language model;

[0010] (2.2) Use the sequence e of unlabeled exercise e text resources encapsulated with a unified template prompt As the input of the pre-trained language model M, it is represented by the following formula:

[0011] e prompt =[CLS]Exercise Text[SEQ]Prompt[MASK]

[0012] Among them, [CLS] represents that the feature information obtained by passing through the pre-trained language model M is used as the semantic representation of the entire text; Exercise Text represents the content of the exercise text; [SEQ] is used to separate Exercise Text and Prompt; Prompt represents the unified template, and the specific content of the unified template is "Knowledge concept belongs to"; [MASK] is used to cover the word representing the corresponding knowledge concept label of the sentence.

[0013] (2.3) M uses the final hidden state h of the first token [CLS] as the representation of the entire sequence. Based on the pre-trained language model M, calculate the probability P that each knowledge concept c in the label space C fills the token [MASK] M ([MASK]=c|h), and use the mapping function Sigmod() to independently predict the probability of each concept, as shown in the following formula:

[0014] P(c|h)=Sigmod(P M ([MASK]=c|h))=Sigmod(Wh)

[0015] Among them, W is the parameter matrix.

[0016] S3. Add a threshold mechanism to determine the knowledge concepts that the exercise should be associated with, as shown in the following formula:

[0017] p(e)={c|P(c|e prompt )>t, c∈C}

[0018] Among them, t represents a fixed threshold. If P(c|e prompt )>t, then label the exercise e with the knowledge concept c;

[0019] S4. Fine-tune the parameters and parameter matrix W of the pre-trained language model M by maximizing the log-likelihood probability of the correct labels.

[0020] Further, in step S4, different learning rates are used for fine-tuning the parameters. The parameter θ is divided into {θ 1 , …, θ L}, where θ l represents the parameters of the l-th layer of the pre-trained model M. The parameter update expression is as follows:

[0021]

[0022] where η l represents the learning rate of the l-th layer.

[0023] Beneficial effects: Compared with the prior art, the remarkable feature of the present invention is that it can improve the effect of labeling knowledge concepts for exercises with a small sample size, and at the same time, provide a data basis for online learning through the method of quickly linking knowledge with exercises. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is the technical framework diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0026] Please refer to Figure 1 As shown, a few-shot exercise linking knowledge concept method based on prompt tuning provided by the present invention includes the following steps:

[0027] S1. Experts label the knowledge concepts corresponding to the course Course, represent the course Course as a label space C = {c1, c2, … c N} with N knowledge concepts, organize and collect the corresponding exercises, represent them as an exercise instance space E, and label some instances in the exercise instance space E with the corresponding label set to form a data set

[0028] S2. Calculate the relevance between the exercise and the knowledge concept through the method of instant tuning with a unified template. The specific implementation is as follows:

[0029] (2.1) Set M as a large pre-trained language model;

[0030] (2.2) Use the sequence e prompt of the unlabeled exercise e text resources encapsulated with a unified template as the input of the pre-trained language model M, which is represented by the following formula:

[0031] e prompt =[CLS]Exercise Text[SEQ]Prompt[MASK]

[0032] Among them, [CLS] represents the feature information obtained by it through the pre-trained language model M as the semantic representation of the entire text; Exercise Text represents the content of the exercise text; [SEQ] is used to separate Exercise Text and Prompt; Pr ompt represents a unified template, and the specific content of the unified template is "knowledge concept belongs to"; [MASK] is used to cover the word representing the corresponding knowledge concept label of the sentence.

[0033] (2.3) M uses the final hidden state h of the first token [CLS] as the representation of the entire sequence. Based on the pre-trained language model M, calculate the probability that each knowledge concept c in the label space C fills the token [MASK] as P M ([MASK]=c|h), and use the mapping function Sigmod() to independently predict the probability of each concept, as shown in the following formula:

[0034] P(c|h)=Sigmod(P M ([MASK]=c|h))=Sigmod(Wh)

[0035] Among them, W is the parameter matrix.

[0036] S3. Add a threshold mechanism to determine the knowledge concepts that the exercise should be associated with, as shown in the following formula:

[0037] p(e)={c|P(c|e prompt )>t, c∈C}

[0038] Among them, t represents a fixed threshold. If P(c|e prompt )>t, then label the exercise e with the knowledge concept c;

[0039] S4. Fine-tune the parameters of the pre-trained language model M and the parameter matrix W by maximizing the log-likelihood probability of the correct label. Use different learning rates to fine-tune the parameters. Divide the parameter θ into {θ 1 ,…,θ L}, θ l represents the parameter of the l-th layer of the pre-trained model M. The parameter update expression is shown below:

[0040]

[0041] Among them, η l represents the learning rate of the l-th layer.

[0042] Example 1

[0043] Compare the calculation results in two modes: the TagBert model and the mode described in the present invention. The method described in the present invention is abbreviated as PTMLTC, and TagBert is a model based on a large pre-trained model and a multi-label classification layer.

[0044] Select exercise samples. Micro F1 calculates the F1 scores of all exercise samples; Macro F1 calculates the average value of the F1 scores of all exercise samples obtained for each label category.

[0045] In this embodiment, a situation with 5 exercise sample control numbers is adopted. Each category refers to a different label. When tagging exercises, there is multi-label classification. 5-shot means that the number of samples in each category is controlled to 5. Since the multi-label classification problem cannot guarantee that the number of samples is exactly 5, the following principles are followed: (1) All labels appear at least 5 times in the support set; (2) If a piece of data is deleted and at least one label is less than 5, the average value of the F1 scores of all exercise samples obtained for each label category is shown in Table 1 below:

[0046] Table 1

[0047]

[0048] It can be seen from the table that the method PTMLTC described in the present invention finally obtains the F1 scores of all exercise samples and the average value of the F1 scores of all exercise samples obtained for each label category, both of which are greater than the effects produced by the TagBert model, and can improve the effect of tagging exercises with knowledge concepts in the case of few samples.

Claims

1. A few-shot exercise link knowledge concept method based on prompt tuning, characterized in that, Including the following steps: S1. Experts label the knowledge concepts corresponding to the course Course, and represent the course Course as a label space C = {c1, c2, … c N}, collect and organize the corresponding exercises, represent them as an exercise instance space E, and label some instances in the exercise instance space E with the corresponding label sets to form a data set S2. Calculate the relevance between the exercise and the knowledge concept by means of instant optimization with a unified template, and the specific implementation is as follows: (2.1) Set M as a large pre-trained language model; (2.2) Use a sequence e of unlabeled exercise e text resources encapsulated with a unified template prompt As the input of the pre-trained language model M, it is expressed by the following formula: e prompt = [CLS] Exercise Text [SEQ] Prompt [MASK] Among them, [CLS] represents the feature representation obtained by the pre-trained language model M as the feature representation of the entire text; Exercise Text represents the content of the exercise text; [SEQ] is used to separate Exercise Text and Prompt; Prompt represents the unified template, and the specific content of the unified template is "The knowledge concept belongs to"; [MASK] is used to cover the word representing the corresponding knowledge concept label of the sentence (2.3) M uses the final hidden state h of the token [CLS] as the representation of the entire sequence. Based on the pre-trained language model M, the probability that each knowledge concept c in the label space C fills the token [MASK] is calculated as P M ([MASK]=c|h). The probability of each concept is independently predicted using the mapping function Sigmod(), as shown in the following formula: P(c|h) = Sigmod(P M ([MASK] = c|h)) = Sigmod(Wh) Among them, W is the parameter matrix; S3. Add a threshold mechanism to determine the knowledge concept associated with the exercise, as shown in the following formula: p(e) = {c | P(c|e prompt ) > t, c ∈ C} where t represents a fixed threshold. If P(c|e prompt ) > t, then the exercise e is labeled with the knowledge concept c; S4. Use the data set D as the training data, and fine-tune the parameters of the pre-trained language model M and the parameter matrix W by maximizing the log-likelihood probability of the correct label, and repeat steps S2 - S3 to train the pre-trained language model M.

2. The few-shot exercise link knowledge concept method based on prompt tuning according to claim 1, wherein In step S4, different learning rates are used for fine-tuning the parameters of different layers, and the stochastic gradient descent method SGD is used to update all the parameters θ of the model. The parameter θ is divided into {θ 1 , …, θ L}, where θ l represents the parameters of the l-th layer of the pre-trained model M. The parameter update expression is as follows: Among them, η l represents the learning rate of layer l, is the gradient of the objective function of the layer l model.

3. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the computer program, it implements the steps of any one of claims 1 to 2.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of any one of claims 1 to 2.

Citation Information

Patent Citations

  • Knowledge tracking method based on learning migration

    CN113010580A

  • Cognitive description fused attention knowledge tracking method

    CN114021722A