Knowledge tracing cold start optimization method and system based on large language model
Through the large language model, a heterogeneous graph is constructed and a comparison learning training question-answer prediction model is solved, and the data sparsity problem of the knowledge tracking model is solved when the new problems and knowledge points are coldly started, achieving higher prediction accuracy and model adaptability.
Patent Information
- Application Number
- CN202510038695.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-10
AI Technical Summary
In the online learning platform, the ID-based knowledge tracking model faces data sparsity problems when new problems and knowledge points are coldly started, resulting in a decrease in prediction capabilities and system generalization performance.
The knowledge tracking cold start optimization method based on large language models is adopted, and the problem-solving steps and knowledge concepts of the problem-making problem are generated through the large language models, and the heterogeneous graph encoder and student status encoder are used, and the model cold start is completed by combining the comparison learning training and answering prediction model.
It effectively alleviates the data sparsity problem of the knowledge tracking model in the cold start stage, improves the prediction accuracy of new problems and knowledge points, and enhances the consistency and adaptability of the model's semantic understanding.
Smart Images

Figure CN119441508B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of knowledge tracing, and in particular relates to a knowledge tracing cold start optimization method and system based on a large language model. Background Art
[0002] Knowledge tracing is a core task in personalized education systems. Its goal is to dynamically predict students' mastery of knowledge points by analyzing the interaction records between students and questions on online learning platforms, and further assist in learning path planning and personalized recommendations. In recent years, with the development of deep learning, neural network-based knowledge tracing models (such as DKT, DKVMN, etc.) have made significant progress. These models can effectively improve the accuracy of predictions by encoding students' historical interaction data and learning the timing of knowledge points.
[0003] However, these ID-based methods usually rely on a large number of student-question interaction records on online learning platforms to train models, thus facing severe challenges of the cold start problem. Cold start mainly manifests itself in two aspects: first, new questions lack sufficient student interaction data, which makes it difficult for the model to accurately estimate their difficulty or knowledge point relevance; second, the introduction of new knowledge points makes it impossible for the model to fully learn the semantic characteristics related to them. This data sparsity not only reduces the model's ability to predict new exercises or knowledge points, but may also affect the generalization performance of the overall system. Therefore, for online learning, how to effectively alleviate the cold start problem has become a core problem that needs to be solved in the field of knowledge tracking. Summary of the invention
[0004] The purpose of the present invention is to solve the cold start problem faced by the model in the process of knowledge tracking on an online learning platform, and to provide a knowledge tracking cold start optimization method and system based on a large language model.
[0005] The specific technical solutions adopted by the present invention are as follows:
[0006] In a first aspect, the present invention provides a knowledge tracking cold start optimization method based on a large language model, which comprises:
[0007] S1. For each question extracted from the student online answer records on the online learning platform, the large language model first outputs a set of problem-solving steps for the question text, then the large language model annotates the set of knowledge concepts involved in the set of problem-solving steps for the question text and the set of problem-solving steps, and finally the large language model outputs a set of binary pairing relationships between the problem-solving steps and the related knowledge concepts for the question text, the set of problem-solving steps and the set of knowledge concepts under constraints;
[0008] S2. All questions and knowledge concepts on the online learning platform are regarded as two types of nodes in the heterogeneous graph, and according to the binary pairing relationship set of all questions, edge connections between knowledge concepts and questions are established in the heterogeneous graph, and according to the correlation between knowledge concepts, edge connections between knowledge concepts are established, and then the semantic embedding vector of each node is generated;
[0009] S3, taking the heterogeneous graph as the input of the heterogeneous graph encoder, encoding and aggregating the neighbor information of the graph nodes through different graph attention network layers in each encoding layer of the heterogeneous graph encoder, and aggregating the neighbor information into the node embedding while retaining the original embedding information of the node in combination with the skip knowledge mechanism, and finally outputting the final node embedding of the knowledge concept and the question by the last encoding layer of the heterogeneous graph encoder;
[0010] S4. Train a question answer prediction model through comparative learning to predict the target student's answer to the target question and complete the model cold start; the question answer prediction model is based on the heterogeneous graph output after encoding by the heterogeneous graph encoder, and the student state encoder encodes the student learning representation and student answer representation corresponding to all questions in the target student's online answer record, and then inputs the gated recurrent unit to model the target student's learning history to obtain the target student's current knowledge state, and finally inputs the target student's current knowledge state, the target question's semantic embedding vector and the student learning representation of the target question into the prediction head to obtain the student's answer prediction to the next question.
[0011] As a preferred embodiment of the first aspect, the specific implementation steps of S1 are as follows:
[0012] S11, obtaining the student's online answer record on the online learning platform, and extracting a question set consisting of all questions and the question text and question-related information corresponding to each question;
[0013] S12, constructing a first prompt template for prompting the large language model to output a solution step for the question text; then for each question in the question set, combining the corresponding question text with the first prompt template and inputting them into the large language model, so that the large language model outputs a solution step set for each question through chain reasoning;
[0014] S13, constructing a second prompt template, used to prompt the large language model to mark the knowledge concepts involved in the problem-solving step set for the question text and the problem-solving step set; then for each question in the problem set, combining the corresponding question text, the problem-solving step set and the second prompt template and inputting them into the large language model, so that the large language model outputs all the knowledge concepts included in the problem-solving step set of each question through chain reasoning, and obtains the knowledge concept set corresponding to each question;
[0015] S14. Construct a third prompt template with constraints, which is used to prompt the large language model to output the binary pairing relationship between the problem-solving steps and the related knowledge concepts for the problem text, the problem-solving step set and the knowledge concept set, wherein the constraints are used to prompt the large language model to meet the pairing principles when generating the binary pairing relationship; then, for each problem in the problem set, the corresponding problem text, the problem-solving step set and the knowledge concept set are combined with the third prompt template and input into the large language model, so that the large language model outputs the binary pairing relationship set contained in all the problem-solving steps in each problem through chain reasoning.
[0016] As a preferred embodiment of the first aspect, the question-related information includes the question type and the answer given by the student to the question.
[0017] As a preferred embodiment of the first aspect, the first prompt template, the second prompt template and the third prompt template all need to set placeholders for personalized fields, and automatically generate complete prompts that are actually input into the large language model by replacing the placeholders.
[0018] As a preferred embodiment of the above-mentioned first aspect, the constraints in the third prompt template are set to satisfy the following conditions at the same time: each problem-solving step is associated with at least one knowledge concept; each knowledge concept is associated with at least one problem-solving step; a many-to-many mapping is allowed between problem-solving steps and knowledge concepts; the large language model needs to combine the historical records that record the generated binary pairing relationships to sequentially generate new binary pairing relationships, and if the newly generated binary pairing relationship does not exist in the historical records, it will be updated to the historical records.
[0019] As a preferred embodiment of the above-mentioned first aspect, in each encoding layer of the heterogeneous graph encoder, encoding and aggregation are performed on the knowledge concept-problem edges through the first graph attention network layer, and the neighbor problem nodes connected with the central knowledge concept by edges are embedded and weightedly aggregated into the neighbor information of the central knowledge concept; encoding and aggregation are performed on the knowledge concept-knowledge concept edges through the second graph attention network layer, and the neighbor knowledge concept nodes connected with the central knowledge concept by edges are embedded and weightedly aggregated into the neighbor information of the central knowledge concept; encoding and aggregation are performed on the problem-knowledge concept edges through the third graph attention network layer, and the neighbor knowledge concept nodes connected with the central problem by edges are embedded and weightedly aggregated into the neighbor information of the central problem; and then for each problem node and knowledge concept node in the heterogeneous graph, the neighbor information is aggregated into its own node embedding through the skip knowledge mechanism.
[0020] As a preferred embodiment of the above-mentioned first aspect, in the student state encoder, the process of obtaining the current knowledge state of the target student is: traverse the online answer records in sequence to find the questions that have generated question-answer interactions with the target student, and at the same time determine the set of knowledge concepts contained in each interacted question, and average the final node embeddings of all knowledge concepts in the knowledge concept set as the student learning representation of the target student for this interacted question, and then combine the student learning representation and the correct or incorrect mark of the target student's answer to the interacted question to construct the student answer representation of the target student for this interacted question, and input the student learning representation and student answer representation corresponding to all the interacted questions of the target student into the gated recurrent unit in sequence for timing modeling, thereby outputting the current knowledge state of the target student.
[0021] As a preferred embodiment of the above-mentioned first aspect, when the answer prediction model is trained by contrastive learning, it is necessary to construct a first positive-negative sample pair for calculating the contrast loss of the question and a second positive-negative sample pair for calculating the contrast loss of the problem-solving steps; in the first positive-negative sample pair, the positive sample pair is a single question and knowledge concepts related to the problem, and the negative sample pair is a single question and knowledge concepts unrelated to the problem; in the second positive-negative sample pair, the positive sample pair is a single problem-solving step and a set of knowledge concepts related to the problem-solving step, and the negative sample pair is a single problem-solving step and a set of knowledge concepts unrelated to the problem-solving step.
[0022] As a preferred embodiment of the above-mentioned first aspect, when the answer prediction model is trained by contrastive learning, the total loss function is obtained by weighted calculation of three parts: the question contrast loss, the problem-solving step contrast loss, and the cross-entropy loss output by the prediction head in the answer prediction model.
[0023] In a second aspect, the present invention provides a knowledge tracking system based on a large language model, comprising:
[0024] An interactive module is used for users to specify target students and target questions for knowledge tracking;
[0025] A model cold start optimization module, used to obtain the heterogeneous graph output after encoding by the heterogeneous graph encoder and the trained answer prediction model according to the knowledge tracing cold start optimization method based on the large language model as described in any one of the first aspects above, and store them for calling;
[0026] The prediction module is used to call the heterogeneous graph and the answer prediction model stored in the model cold start optimization module according to the specified information in the interaction module, output the target student's answer to the target question, and push or visualize it according to the preset logic.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] 1) Traditional knowledge tracking models are highly dependent on manually annotated knowledge tags, which is inefficient and difficult to maintain consistency. In the present invention, a large language model is used to perform automatic semantic annotation and embedding generation modules to automatically generate semantic embeddings for questions and knowledge points, avoiding the limitations of manual annotation while ensuring the consistency of knowledge concepts. When generating embeddings for knowledge concepts, the present invention introduces contrastive learning to align the semantic features of questions and knowledge points, so that the model can deeply understand the knowledge content at a semantic level. Even in the absence of historical data in the cold start phase, the model can infer the meaning of new knowledge points or questions through automatically generated semantic embeddings. Moreover, contrastive learning further enhances the semantic alignment between knowledge points and questions, so that the model can predict the student's knowledge status based on the similarity of semantic embeddings. The automated semantic annotation and embedding generation method used in the present invention not only improves the prediction accuracy under cold start conditions, but also makes the model more consistent and adaptable in terms of knowledge semantic understanding, thereby still providing reliable knowledge tracking results in the absence of data.
[0029] 2) During the learning process of students on the online learning platform, there are often complex associations and dependencies between knowledge points. For example, mastering a specific knowledge point may require students to first understand the relevant basic knowledge. The present invention unifies the association between knowledge points and the knowledge content involved in the questions into the same structure by constructing a heterogeneous graph structure of questions and knowledge points. The nodes of the heterogeneous graph include knowledge points and questions, while the edges represent the hierarchical relationship, dependency relationship or association path between questions and knowledge points. This structured design helps the model to infer students' performance on new knowledge points based on the relationship between knowledge points during cold start, thereby making up for the lack of direct interaction data. Moreover, this approach of the present invention provides a dynamic and scalable knowledge network, which enables the model to adaptively update and expand new knowledge points and new questions; through the graph convolutional network, students' mastery of certain knowledge points can be smoothly propagated to other related knowledge points, thereby providing support for the dynamic update of knowledge status. Through this structured knowledge association, the model can better capture students' learning paths and mastery of knowledge points during the knowledge status update process. The heterogeneous graph construction and aggregation method adopted in the present invention not only enhances the understanding of knowledge association under cold start conditions, but also provides structured support for knowledge transfer, helping the model to comprehensively capture students' knowledge mastery.
[0030] 3) This invention combines automatic semantic annotation with structural information encoding to improve the semantic representation quality and relevance of exercises. The former generates high-quality semantic features for new exercises, and the latter enhances its contextual relevance by modeling graph structures. Together, they provide an effective means to alleviate the cold start problem. This method not only improves the adaptability of the knowledge tracking model, but also provides theoretical support and practical reference for achieving efficient prediction and recommendation in personalized learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 A schematic diagram of the steps of the cold start optimization method for knowledge tracing based on a large language model;
[0032] Figure 2 This is a schematic diagram of the construction process of the complete first prompt;
[0033] Figure 3 This is a schematic diagram of the construction process of the complete second prompt;
[0034] Figure 4 This is a schematic diagram of the construction process of the complete third prompt;
[0035] Figure 5 This is a module diagram of the knowledge tracking system based on the large language model;
[0036] Figure 6 This is a flowchart of the knowledge tracing cold start optimization in an embodiment. DETAILED DESCRIPTION
[0037] In order to make the above-mentioned purpose, features and advantages of the present invention more obvious and easy to understand, the specific implementation mode of the present invention is described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined accordingly without conflicting with each other.
[0038] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for the purpose of distinguishing descriptions, and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features.
[0039] like Figure 1As shown, in a preferred embodiment of the present invention, a knowledge tracking cold start optimization method based on a large language model is provided, and its specific steps are shown in S1 to S4. The specific implementation method of each step is described in detail below.
[0040] S1. For each question extracted from the online answer records of students on the online learning platform, the large language model first outputs a set of solution steps for the question text, and then the large language model annotates the set of knowledge concepts involved in the set of solution steps for the question text and the set of solution steps. Finally, the large language model outputs a set of binary pairing relationships between the solution steps and the relevant knowledge concepts for the question text, the set of solution steps and the set of knowledge concepts under constraints.
[0041] In an embodiment of the present invention, the large language model needs to generate different information in sequence through three links, and finally obtain a binary pairing relationship between the problem-solving steps and the relevant knowledge concepts. In different environments, different prompt templates need to be constructed through prompt word engineering. Since the present invention needs to process a series of different problems in batches, it is necessary to set placeholders for personalized fields in the prompt templates in the three links, and automatically generate complete prompts for actual input into the large language model by replacing the placeholders. The specific implementation of the above-mentioned S1 step of the present invention is described in detail below.
[0042] S11. Obtain students' online answer records on the online learning platform, and extract a question set consisting of all questions and the question text and question-related information corresponding to each question, wherein the question-related information includes the question type and the answers given by students to the questions.
[0043] It should be noted that the online learning platform in the present invention can be any online education platform that can obtain and store students' online answer records, such as a MOOC platform. The "question" mentioned in the present invention refers to the "exercise" for students to learn on the platform, which can be in the form of fill-in-the-blank questions, multiple-choice questions, etc. The "question text" mentioned in the present invention is the content of the question in text form, such as "What is 3+2?" "Question-related information" is some information related to this question. In addition to the question type and the answer given by the student to the question, it can also have a Chinese analysis of the question, or it can also associate other information according to different types of questions, such as multiple-choice questions can have optional options for the question.
[0044] S12. Construct a first prompt template to prompt the large language model to output problem-solving steps for the question text; then, for each question in the question set, combine the corresponding question text with the first prompt template and input them into the large language model, so that the large language model outputs a set of problem-solving steps for each question through chain reasoning.
[0045] It should be noted that the prompt template in the present invention is a template required for constructing a prompt for inputting a large language model (LLM). Since the question set contains a large number of questions, it is necessary to apply the template in batches to generate prompt texts that meet the conditions of each question. The construction of the prompt template can refer to the construction method of the prompt word engineering in the prior art. Specifically, it can be combined with the actual data situation, clarify the instructions and input data, and continuously optimize to ensure that the model can understand and generate outputs that meet expectations by selecting appropriate template types and template variables.
[0046] like Figure 2 As shown, in an embodiment of the present invention, the first prompt template contains a first placeholder for the question text. For each question in the question set, the corresponding question text replaces the first placeholder in the first prompt template to form a complete first prompt, which is input into the large language model. When there is a series of questions, the relevant information can be automatically extracted by the script to generate a complete first prompt corresponding to each question, and the output is obtained by inputting the large language model through the calling interface. In addition, in order to better enable the large model to output complete problem-solving steps, in addition to the question text, placeholders for one or more question-related information in the question type, question answer, and Chinese analysis of the question can be further set in the first prompt template. These question-related information can assist the large model to output more accurate problem-solving steps. Therefore, the question text in the first prompt must be entered, and other question-related information can be selectively entered according to actual conditions.
[0047] In the present invention, an exemplary complete first prompt is as follows (in order to distinguish the starting position of the prompt text, {} is used as the starting mark):
[0048] {Problem text: Find the solution to the equation x + 3 = 7.
[0049] Final answer to the question: 4.
[0050] Chinese solution to the problem: Subtract 3 from both sides of the equation and we get x = 4.
[0051] Please generate the problem-solving steps step by step:}
[0052] Based on the complete first prompt above, the output generated by the expected large model is:
[0053] 1. Subtract 3 from both sides of the equation to get x = 7 - 3.
[0054] 2. Calculate the result on the right and get x = 4.
[0055] Of course, the above-mentioned complete first prompt form is only an implementation and can be optimized according to actual conditions.
[0056] S13. Construct a second prompt template to prompt the large language model to mark the knowledge concepts involved in the problem-solving step set for the question text and the problem-solving step set; then, for each question in the problem set, combine the corresponding question text, the problem-solving step set and the second prompt template and input them into the large language model, so that the large language model outputs all the knowledge concepts contained in the problem-solving step set of each question through chain reasoning, and obtains the knowledge concept set corresponding to each question.
[0057] It should be noted that the construction principles of the second prompt template and the first prompt template are similar, and the only difference is that the variable placeholders and required instructions are different between the two.
[0058] like Figure 3 As shown, in an embodiment of the present invention, the second prompt template has a second placeholder corresponding to the question text and the set of problem-solving steps. For each question in the problem set, the corresponding question text and the set of problem-solving steps replace the corresponding second placeholders in the first prompt template to form a complete second prompt, which is then input into the large language model. Similarly, when there is a series of questions, the script can automatically extract relevant information to generate a complete second prompt corresponding to each question, and the output is obtained by inputting the large language model through the calling interface.
[0059] S14. Construct a third prompt template with constraints, which is used to prompt the large language model to output the binary pairing relationship between the problem-solving steps and the related knowledge concepts for the problem text, the problem-solving step set and the knowledge concept set, wherein the constraints are used to prompt the large language model to meet the pairing principles when generating the binary pairing relationship; then, for each problem in the problem set, the corresponding problem text, the problem-solving step set and the knowledge concept set are combined with the third prompt template and input into the large language model, so that the large language model outputs the binary pairing relationship set contained in all the problem-solving steps in each problem through chain reasoning.
[0060] like Figure 4 As shown, in an embodiment of the present invention, the third prompt template contains third placeholders corresponding to the question text, the set of problem-solving steps, and the set of knowledge concepts. For each question in the set of questions, the corresponding question text, the set of problem-solving steps, and the set of knowledge concepts replace the corresponding third placeholders in the first prompt template to form a complete third prompt, which is then input into the large language model. Similarly, when there is a series of questions, the script can automatically extract relevant information to generate a complete third prompt corresponding to each question, and the output can be obtained by inputting the large language model through the calling interface.
[0061] In addition, the third prompt template also needs to include constraints for generating pairing relationships. In an embodiment of the present invention, the constraints are set to simultaneously satisfy the following conditions: each problem-solving step is associated with at least one knowledge concept; each knowledge concept is associated with at least one problem-solving step; a many-to-many mapping is allowed between problem-solving steps and knowledge concepts; the large language model needs to combine historical records that record generated binary pairing relationships to sequentially generate new binary pairing relationships, and if the newly generated binary pairing relationship does not exist in the historical records, it will be updated to the historical records.
[0062] It should be noted that there are certain differences in the construction principles of the third prompt template and the second prompt template and the first prompt template. Stronger constraints need to be set in the internal instructions to inform the large language model of the need to generate pairing relationships.
[0063] In the present invention, an exemplary complete third prompt is as follows (in order to distinguish the starting position of the prompt text, {} is used as the starting mark):
[0064] {You are an expert in the field of mathematics education. You will receive a math problem, its solution steps, and knowledge concepts (which you have marked before). Your task is to associate the solution steps with the knowledge concepts and determine which knowledge concepts are needed for each solution step. Please note that all solution steps and knowledge concepts must be associated, and the association can be one-to-many or many-to-many. Each solution step and knowledge concept has been numbered. Your output should match the numbers of all solution steps and knowledge concepts, separated by commas, and output as one line.
[0065] Your output must meet the following conditions:
[0066] Each problem-solving step must be paired.
[0067] Every knowledge concept must be paired.
[0068] Only related problem-solving steps and knowledge concepts can be paired.
[0069] The output cannot contain non-existent solution step numbers. For example, if the solution has 4 steps, "5-2" is illegal.
[0070] The output cannot contain non-existent knowledge concept numbers. For example, if there are 3 knowledge concepts, "3-5" is illegal.
[0071] Your output should be in a comma-separated format on one line. For example, if there are 4 problem-solving steps and 5 knowledge concepts, possible output is: `1-1, 1-3, 1-5, 2-4, 3-2, 3-5, 4-2, 4-3, 4-5`.
[0072] Note that this output needs to satisfy all of the above conditions.
[0073] Now, please provide your mapping results based on the following questions, problem-solving steps, and knowledge concepts.
[0074] Question: The base of a triangle is 5 and the height is 10. Find its area.
[0075] Steps to solve the problem:
[0076] 1) Find the formula for the area of a triangle.
[0077] 2) Substitute the values for the base and height into the formula.
[0078] 3) Perform calculations to find the area
[0079] Knowledge concept:
[0080] 1) Triangle area formula
[0081] 2) Substitute the values
[0082] 3) Mathematical operations}
[0083] Based on the complete third prompt above, binary pairing of problem-solving steps and the IDs of relevant knowledge concepts is performed, and the output generated by the expected large model is: 1-1, 2-2, 3-3.
[0084] Of course, the above complete third prompt form is only an implementation and can be optimized according to actual conditions.
[0085] S2. Treat all questions and knowledge concepts on the online learning platform as two types of nodes in the heterogeneous graph, and establish edge connections between knowledge concepts and questions in the heterogeneous graph based on the binary pairing relationship set of all questions. Establish edge connections between knowledge concepts and knowledge concepts based on the correlation between knowledge concepts, and then generate semantic embedding vectors for each node in the heterogeneous graph through the encoder.
[0086] It should be noted that the correlation between knowledge concepts can be determined based on an external knowledge base, including hierarchical and subordinate relationships or correlation relationships between knowledge concepts.
[0087] It should be noted that the encoder used to generate the semantic embedding vector of each node can adopt a pre-trained language model, such as the GPT-4o and Bert models. When generating the semantic embedding vector of a node, if the node is a question node, the question text corresponding to the question node is input into the pre-trained language model and converted into a semantic embedding vector; if the node is a concept node, the knowledge concept text corresponding to the concept node is input into the pre-trained language model and converted into a semantic embedding vector. The semantic embedding vector generated by the pre-trained language model can be used as the original embedding information of the node in the heterogeneous graph and input into the heterogeneous graph encoder for further information aggregation.
[0088] S3. The heterogeneous graph is used as the input of the heterogeneous graph encoder. In each encoding layer of the heterogeneous graph encoder, the neighbor information of the graph nodes is encoded and aggregated through different graph attention network layers, and the jumping knowledge mechanism is combined to aggregate the neighbor information into the node embedding while retaining the original embedding information of the node. Finally, the last encoding layer of the heterogeneous graph encoder outputs the final node embedding of the knowledge concept and the question.
[0089] In an embodiment of the present invention, the above-mentioned heterogeneous graph encoder is composed of a series of encoding layers, and the specific execution process in each encoding layer is as follows: encoding and aggregation are performed on the knowledge concept-problem edges through the first graph attention network layer, and the neighbor problem nodes connected to the central knowledge concept by edges are embedded and weighted to aggregate into the neighbor information of the central knowledge concept; encoding and aggregation are performed on the knowledge concept-knowledge concept edges through the second graph attention network layer, and the neighbor knowledge concept nodes connected to the central knowledge concept by edges are embedded and weighted to aggregate into the neighbor information of the central knowledge concept; encoding and aggregation are performed on the problem-knowledge concept edges through the third graph attention network layer, and the neighbor knowledge concept nodes connected to the central problem by edges are embedded and weighted to aggregate into the neighbor information of the central problem; and then, for each problem node and knowledge concept node in the heterogeneous graph, the neighbor information is aggregated into its own node embedding through the skip knowledge mechanism.
[0090] It should be noted that the knowledge concept-problem edge refers to the edge between the knowledge concept node and the problem node in the heterogeneous graph. During aggregation, the knowledge concept node is used as the central node, and the information of the adjacent problems is aggregated to the central node. The knowledge concept-knowledge concept edge refers to the edge between the knowledge concept node and the knowledge concept node in the heterogeneous graph. During aggregation, the previous knowledge concept node is used as the central node, and the information of the adjacent knowledge concept nodes is aggregated to the central node. The problem-knowledge concept edge refers to the edge between the problem node and the knowledge concept node in the heterogeneous graph. During aggregation, the problem node is used as the central node, and the information of the adjacent knowledge concept nodes is aggregated to the central node.
[0091] S4. Train a question prediction model through contrastive learning to predict the target student's answer to the target question and complete the model cold start. The above-mentioned question prediction model is based on the heterogeneous graph output after encoding by the heterogeneous graph encoder. The student state encoder encodes the student learning representation and student answer representation corresponding to all questions in the target student's online answer record, and then inputs the gated recurrent unit to model the target student's learning history to obtain the target student's current knowledge state. Finally, the target student's current knowledge state, the semantic embedding vector of the target question and the student learning representation of the target question are input into the prediction head to obtain the student's answer prediction for the next question.
[0092] In an embodiment of the present invention, in the above-mentioned student state encoder, the process of obtaining the current knowledge state of the target student is: traverse the online answer records in turn to find the questions that have generated question-answering interactions with the target student, and at the same time determine the knowledge concept set contained in each interactive question, and average the final node embeddings of all knowledge concepts in the knowledge concept set as the student learning representation of the target student for this interactive question, and then combine the student learning representation and the correct or incorrect mark of the target student's answer to the interactive question to construct the student answer representation of the target student for this interactive question. When combining the two types of information, it is assumed that the student is at time of students responded that , then it can be represented by learning and answer correctness (1 means correct, 0 means wrong) , where [⋅,⋅] represents the concatenation operation and 0 is the zero vector. Finally, the student learning representations and student answer representations corresponding to all the interactive questions of the target student are sequentially input into the gated recurrent unit (GRU) for temporal modeling, thereby outputting the current knowledge state of the target student.
[0093] The answer prediction model of the above-mentioned step S4 needs to be trained through contrastive learning before being used for actual reasoning. In an embodiment of the present invention, when training the answer prediction model, it is necessary to construct a first positive-negative sample pair for calculating the contrast loss of the question and a second positive-negative sample pair for calculating the contrast loss of the problem-solving steps. Among them, in the first positive-negative sample pair, the positive sample pair is a single question and the knowledge concept related to this question, and the negative sample pair is a single question and the knowledge concept unrelated to this question; and in the second positive-negative sample pair, the positive sample pair is a single problem-solving step and a set of knowledge concepts related to this problem-solving step, and the negative sample pair is a single problem-solving step and a set of knowledge concepts unrelated to this problem-solving step.
[0094] In addition, in an embodiment of the present invention, when the answer prediction model is trained by contrastive learning, the total loss function is obtained by weighted calculation of three parts: the question contrast loss, the problem-solving step contrast loss, and the cross-entropy loss output by the prediction head in the answer prediction model.
[0095] The core of the above-mentioned knowledge tracking cold start optimization method based on a large language model in the present invention is achieved through improvements in two aspects: automatic semantic annotation and heterogeneous graph information encoding. Automated semantic annotation uses the powerful semantic understanding ability of the pre-trained large language model (LLM) to automatically extract high-quality semantic representations from the exercise text, covering multi-dimensional information such as the scope of knowledge points and logical complexity. This process not only reduces the dependence on manual annotation, but also provides additional initial semantic information for new exercises, making up for the lack caused by data sparsity. In addition, the heterogeneous graph information encoder models the explicit and implicit relationships between exercises through the graph neural network (GNN), embeds the exercises into the global graph structure, and enhances its representation through dynamic information aggregation. This joint modeling method enables new exercises to benefit from related exercises and knowledge points, and obtain reasonable initialization features even in the absence of interactive data.
[0096] Based on the same inventive concept, in another embodiment of the present invention, Figure 5 As shown, a knowledge tracking system based on a large language model is also provided, which includes an interaction module, a model cold start optimization module and a prediction module. The specific functions of each module are described in detail below.
[0097] An interactive module is used for users to specify target students and target questions for knowledge tracking;
[0098] A model cold start optimization module, used to obtain the heterogeneous graph output after encoding by the heterogeneous graph encoder and the trained answer prediction model according to the aforementioned knowledge tracking cold start optimization method based on the large language model, and store them for calling;
[0099] The prediction module is used to call the heterogeneous graph and the answer prediction model stored in the model cold start optimization module according to the specified information in the interaction module, output the target student's answer to the target question, and push or visualize it according to the preset logic.
[0100] It should be noted that the specific push or visualization can be set according to system requirements, for example, in various forms such as messages, pop-ups, charts, etc.
[0101] The following will use a specific embodiment to demonstrate the specific implementation and technical effect of the knowledge tracking cold start optimization method based on a large language model described in S1 to S4 above.
[0102] Example
[0103] In this embodiment, the GPT model is used as the large language model to implement the knowledge tracking cold start optimization method based on the large language model. The model cold start process in this method is as follows: Figure 6 The specific implementation of each step is described in detail below.
[0104] 1. Data Acquisition
[0105] In this embodiment, it is necessary to obtain the student online answer records on the online learning platform in order to extract the question set consisting of all questions and the question text and question related information corresponding to each question. These student online answer records can be recorded online, or the corresponding offline data set can be obtained. The question related information includes the question type, the answer given by the student to the question, etc.
[0106] 2. Automatic semantic annotation
[0107] 2.1. Problem-solving steps generation
[0108] 2.1.1. Question input and analysis
[0109] The purpose of this step is to receive students' online answer records and analyze the questions, extract the type of questions and other question-related information, and provide a basis for the subsequent problem-solving steps. The implementation process is as follows:
[0110] (1) Receive input question q.
[0111] (2) Parse question q and extract the question type and question text, specifically:
[0112] Question Type: Determine the type of question (e.g. fill-in-the-blank, multiple choice, etc.).
[0113] Question Text: Extract the content of the question itself.
[0114] 2.1.2. Prompt Generation
[0115] The purpose of this step is to prepare appropriate prompts for generating problem-solving steps and ensure that the generation of each step is consistent with the content and type of the problem. The implementation process is as follows:
[0116] (1) Based on the type and content of question q, construct a prompt template as the input of the GPT model.
[0117] If the question is a math problem, you can include relevant math concepts, formulas, or symbols.
[0118] If it is a multiple-choice question, you can include eliminating the options one by one.
[0119] If it is a fill-in-the-blank question, you can guide the model to deduce the answer through specific steps.
[0120] (2) Combine the question and related prompts to generate a complete prompt text.
[0121] (3) The output of this step is a complete prompt , which will be passed to the GPT model to generate problem-solving steps.
[0122] 2.1.3. Step-by-Step Solution Generation
[0123] The purpose of this step is to generate clear and coherent problem-solving steps to ensure that the model output meets the context and type requirements of the problem. The specific implementation process is as follows:
[0124] (1) Input prompts and questions
[0125] The built prompt The question content is passed to the large language model (GPT-4) as input for generating the problem-solving steps. Contains question text and question type.
[0126] (2) Chain reasoning
[0127] Each problem-solving step The generation of the problem depends on the prompt And the steps generated previously The large language model will be through conditional probability prompt Determine the output of the next step so that the problem-solving process unfolds step by step and in a coherent manner:
[0128]
[0129] where prompt It is a prompt message designed based on the question content. is a sequence of inferences generated by previous steps. Temperature parameter Adjust the randomness and diversity of the generation steps: lower temperatures (such as It tends to choose steps with high probability, and the resulting problem-solving process is more certain and accurate. ) increases the sampling of low-probability steps, and the generated problem-solving process is more exploratory and diverse. Through temperature sampling, LLM is allowed to generate diverse problem-solving paths to cover multiple potential solution strategies in the cold start phase when training is insufficient or the labeled data is insufficient, thereby improving the coverage and richness of the problem representation. The output of this step is a list of problem-solving steps consisting of multiple lines of text. The generated problem-solving steps are post-processed to ensure that they meet the requirements and have no redundant or erroneous parts.
[0130] 2.1.4. Final Output
[0131] When processing a specified number of problems, the present invention will automatically save the current progress to ensure that the processed data will not be lost even if the program is interrupted. The saved file contains all the problems and their step-by-step solutions, and the data is stored in JSON format.
[0132] When all questions are processed, the system will perform a final save, storing all generated problem-solving steps and related data in a specified JSON file, containing detailed information for each question, including the question description text, question type, answer, options (multiple-choice questions require a selection type), analysis, and generated problem-solving steps. This step ensures that all answers and solution processes are properly saved and can be used or checked later.
[0133] 2.2 Knowledge Concept Generation
[0134] 2.2.1. Input Preparation
[0135] In this embodiment, the generation of knowledge concepts is based on the following information:
[0136] (1) Problem description That is, the text content of the question itself, which is used to provide the background and core information of the question.
[0137] (2) Problem-solving steps A sequence of problem-solving steps generated by the model that is relevant to the problem. These steps reflect the problem-solving logic and the application process of knowledge concepts.
[0138] (3) Generate instructions: Explicitly require the model to label the knowledge concepts involved according to the problem and the steps to solve the problem. These instructions ensure that the generated results meet the task objectives. The above content is combined into prompt information (prompt) to form the context for generating knowledge concepts. Prompt is the direct input for LLM (large language model) to generate knowledge concepts.
[0139] 2.2.2 Calling Model
[0140] Use the constructed prompt to call LLM. LLM responds to the input content. , contextual reasoning about the problem and its solution process. The model extracts and generates a set of knowledge concepts from the problem-solving logic. . These knowledge concepts reflect the skills and knowledge concepts involved in solving the problem.
[0141] 2.2.3. Result analysis
[0142] The output returned by the model is usually presented in text form, listing the sequence of knowledge concepts These concepts may include mathematical skills, domain knowledge, or other knowledge concepts directly related to problem solving. The results are directly parsed into an independent list of concepts for subsequent teaching analysis or modeling tasks.
[0143]
[0144] Each knowledge concept The generation of depends not only on the question content, but also on the previously generated concept sequence. , to ensure the logical consistency between knowledge concepts. The present invention avoids the problems of high cost and incomplete coverage of manual annotation in the cold start phase through automatic annotation, especially when adding new problems or switching fields. The knowledge concepts automatically generated by LLM are based on the problems and their solution steps, which can ensure the logical consistency of the annotation results, thereby quickly establishing knowledge concept mapping for new problems and shortening the cold start time.
[0145] 2.3 Problem-solving steps and knowledge concept mapping
[0146] In order to achieve the mapping between problem-solving steps and knowledge concepts, the pre-generated problem-solving steps are used and knowledge concepts , combined with the content of the question , the model generates a pairing relationship between problem-solving steps and knowledge concepts The complete process of mapping a large model is as follows:
[0147] 2.3.1 Context Preparation
[0148] The content of the question , Problem Solving Steps Collection , knowledge concept set As input, construct prompt information The prompt information should include: Question content Clarify the context of problem-solving steps and knowledge concepts; Problem-solving steps Provide detailed information for each step; knowledge concepts Display all candidate knowledge concepts.
[0149] 2.3.2. Generate pairs step by step
[0150] Each pairing is a pairing relationship in the form of a binary, consisting of a knowledge concept and a problem-solving step. In the process of LLM generating pairs, it is necessary to set constraints when generating pairs: each problem-solving step is associated with at least one knowledge concept; each knowledge concept is associated with at least one problem-solving step; many-to-many mapping is allowed (for example, one step is associated with multiple knowledge concepts, and one knowledge concept is associated with multiple steps). The model generates each pairing in sequence. In the When generating the next step, based on contextual prompts and previous Pairing history , generating new pairings.
[0151] 2.3.3. Conditional Update
[0152] After each new pair is generated, it is added to the history record and updated with new condition constraints to generate the next pair. Pairs that already exist in the history record do not need to be added to the history record.
[0153] 2.3.4 Termination Conditions
[0154] If all problem-solving steps and related knowledge concepts have been generated, or the maximum limit of generated pairs has been reached The build process terminated.
[0155] The logic of generating the binary pairing relationship in this step can be expressed as:
[0156]
[0157] in: No. Pair the generated problem-solving steps with the knowledge concepts; The model's generation logic is used to generate new pairs based on input and historical pairs; prompt Contains question content , contextual prompts for problem-solving steps and knowledge concepts; forward The pairing history generated this time is used as a conditional constraint.
[0158] 2.3.5 Output results
[0159] The paired sequences generated in this step Represents the mapping of each problem-solving step to its related knowledge concept. This result can clarify the problem-solving steps The knowledge concept relied on Provide accurate association information for subsequent learning path optimization or cognitive analysis. The mapping relationship can accurately display the knowledge concepts corresponding to each problem-solving step, and automatically generate requirements that adapt to different problem characteristics. This method not only reduces the time cost of manually constructing mapping relationships, but also improves the adaptability to new problem scenarios, making the model more efficient in the cold start stage when data is scarce. During the mapping process, LLM will not only generate corresponding knowledge concepts based on the problem-solving steps, but also automatically adjust the strength of the association between each step and the knowledge concept according to the characteristics of the problem. Through the above steps, the automatic semantic annotation module realizes the efficient construction of problem-solving steps, knowledge concepts, and the mapping relationship between the two. The automatic characteristics of the module greatly alleviate the dependence on manual annotation in the cold start stage, provide support for the rapid generation of diversified problem-solving paths and knowledge concept annotations for new problems, and provide a complete and accurate semantic information foundation for problem representation.
[0160] 3. Contrastive Learning to Generate Embeddings
[0161] To alleviate the cold start problem and improve the model's adaptability to new knowledge concepts, this embodiment uses a customized contrastive learning (CL) method to generate embeddings of questions and problem-solving steps that are semantically related to knowledge concepts (KCs). Traditionally, pre-trained large language models (LLMs) are used to generate general embedding representations, but their performance in specific tasks in the education field is limited. To overcome this limitation, this embodiment designs a contrastive learning method that explicitly guides the encoder to align the embeddings of questions and problem-solving steps with the embeddings of their related knowledge concepts. The "enhanced" embeddings obtained through contrastive learning can not only improve the performance of downstream knowledge tracking (KT) models, but also alleviate the cold start problem of models that perform poorly on new questions or knowledge concepts.
[0162] In the framework of this embodiment, the core goal of contrastive learning loss is to generate a more generalized embedding representation by optimizing the semantic distance between positive sample pairs (such as questions and related knowledge concepts) and negative sample pairs (such as questions and irrelevant knowledge concepts). At the same time, in order to avoid learning degradation caused by false negative sample problems (i.e., semantically similar but unrelated knowledge concepts are incorrectly labeled as negative samples), this embodiment designs a mechanism in contrastive learning to pre-process knowledge concepts through clustering algorithms to ensure the accuracy of negative sample selection, thereby further improving the embedding quality and alleviating the cold start phenomenon.
[0163] 3.1. Generating Embedded Representations
[0164] This embodiment requires that the problem content, problem-solving steps, and knowledge concepts be represented as embedded vectors to facilitate model training and reasoning. , and its corresponding set of problem-solving steps and knowledge concept set , this embodiment uses an encoder model To generate their embedding representations:
[0165]
[0166] in are the embedding vectors of the problem, the problem-solving steps, and the knowledge concept, respectively. These embeddings extract semantic features from the original text through the encoder. To alleviate the cold start problem, these features enhance the model's representation of new problems, new steps, and new knowledge concepts through contrastive learning. By narrowing the embedding distance between problems and problem-solving steps and their related knowledge concepts, and pushing the embedding distance away from irrelevant knowledge concepts, the model can adapt to new data more efficiently.
[0167] 3.2 Optimization Objectives Based on Contrastive Learning
[0168] 3.2.1. Question Embedding Optimization
[0169] Each question Usually related to one or more knowledge concepts The goal is to embed the problem Embed as close to relevant knowledge concepts as possible , while staying away from irrelevant concept embedding The specific steps are as follows:
[0170] (1) Construct positive sample pairs: All related knowledge concepts Form a positive sample pair .
[0171] (2) Construct negative sample pairs: Randomly sample from irrelevant knowledge concept sets to form negative sample pairs .
[0172] (3) Calculate similarity: Use the cosine similarity function to calculate the similarity between the question and each knowledge concept:
[0173]
[0174] in is the temperature parameter, which is used to adjust the distribution range of similarity.
[0175] Each question is usually related to one or more knowledge concepts. The goal of this embodiment is to make the embedding vector of the question The embedding vectors of these knowledge concepts At the same time, this embodiment hopes that the similarity between the embedding of the problem and the embedding of irrelevant knowledge concepts is low. To achieve this, this embodiment designs the following contrastive learning loss function:
[0176]
[0177] Among them: indicator function It is used to eliminate semantically similar knowledge concepts to avoid the "false negative sample" problem. In the above formula, this embodiment assumes that the clustering algorithm The knowledge concepts are semantically grouped. The purpose of clustering is to classify semantically similar knowledge concepts into the same category, thereby avoiding mislabeling similar knowledge concepts as negative samples. For problems related to multiple knowledge concepts , this embodiment needs to consider the comparative learning objectives of all relevant knowledge concepts, and the specific formula is:
[0178]
[0179] 3.2.2. Comparative Learning of Problem Solving Steps
[0180] The embedding optimization of the problem-solving steps aims to align the semantic features of the step semantics with the semantic features of the knowledge concepts, so that the model can understand the knowledge concept expression in the "problem-solving process". The specific steps are as follows:
[0181] (1) Related knowledge concepts: For each problem-solving step , mark its related knowledge concept set .
[0182] (2) Positive and negative sample pairs: Related knowledge concepts Constitute positive sample pairs and irrelevant knowledge concepts constitute a negative sample pair.
[0183] The goal of the problem-solving steps is similar to the goal of the problem content, that is, we hope to embed the problem-solving steps into Embedding related knowledge concepts The loss function is expressed as follows:
[0184]
[0185] Unlike the problem, the problem-solving steps are not necessarily related to all knowledge concepts. This embodiment selects the set of knowledge concepts that are truly related to the problem-solving steps by marking the correspondence between the steps and the knowledge concepts. , and calculate the associated loss:
[0186]
[0187] 3.2.3. False Negative Sample Processing
[0188] In actual tasks, there may be "false negative samples", that is, concepts in irrelevant knowledge that are semantically close to the current problem or step. To prevent the model from degrading due to learning these samples, this example recommends the following strategy:
[0189] (1) Knowledge concept clustering: Use clustering algorithms (such as K-means) to divide knowledge concepts into multiple clusters, and knowledge concepts in the same cluster are considered to be semantically similar.
[0190] (2) Sample elimination: Knowledge concepts belonging to the same cluster are not considered negative samples and are not included in the loss calculation.
[0191] 3.3 Heterogeneous Graph Encoder
[0192] The heterogeneous graph encoder encodes the information in the graph structure through a multi-layer graph neural network (GNN) to capture the complex relationship between questions and knowledge concepts. The entire heterogeneous graph consists of two types of nodes: question nodes and and knowledge concept nodes ) and three edge types:
[0193] 1. Concept-Question edge: represents the relationship between a knowledge concept and a related question, for example, a knowledge concept is the core examination point of a question;
[0194] 2. Concept-Concept edge: represents the relationship between concepts, such as hierarchical relationship or correlation;
[0195] 3. Question-Concept edge: Indicates the knowledge concept involved in the question. The question-concept edge and the knowledge concept-question edge are actually the same edge in the graph, but for the subsequent graph network, the central node for aggregating information is different.
[0196] In the present invention, the initial representation of the question and knowledge concepts is first generated by the automatic semantic annotation module, and the semantic embedding of the question text and knowledge concepts is extracted using the pre-trained language model. These initial embedding vectors and It will be used as the input of the Graph Attention Network (GAT) to further encode the contextual information of the node.
[0197] To capture the multi-level relationships of nodes, this embodiment uses three different graph attention network (GAT) layers to encode and aggregate the information on the three types of edges respectively:
[0198] 3.3.1. Knowledge Concept-Problem GAT Layer
[0199] This layer is used to integrate target knowledge concepts The adjacency problem information.
[0200] 1. Calculate the attention coefficient:
[0201]
[0202] in, represents vector concatenation, is a learning parameter.
[0203] 2. Normalized attention weights:
[0204]
[0205] 3. Neighbor information aggregation:
[0206]
[0207] 3.3.2 Knowledge Concept-Knowledge Concept GAT Layer
[0208] This layer is used to connect adjacent knowledge concepts The information is integrated into the target knowledge concept
[0209] 1. Calculate the attention coefficient:
[0210]
[0211] 2. Normalized attention weights:
[0212]
[0213] 3. Neighbor information aggregation:
[0214]
[0215] 3.3.3. Problem-Knowledge Concept GAT Layer
[0216] This layer will be adjacent to the knowledge concept Aggregate information to the target problem .
[0217] 1. Calculate the attention coefficient:
[0218]
[0219] 2. Normalized attention weights:
[0220]
[0221] 3. Neighbor information aggregation:
[0222]
[0223] 3.3.4 Integration and Update
[0224] In order to alleviate the cold start problem, the model introduces a jumping knowledge mechanism, which not only aggregates neighbor information through GAT, but also retains the original semantic representation of concepts and questions. This mechanism can utilize the semantic embedding of questions and concepts in the cold start phase to avoid insufficient information caused by missing historical data.
[0225] Each layer retains the original information of the node through a skipping mechanism while aggregating the neighbor information into the embedding.
[0226] For the concept of knowledge
[0227]
[0228] For the problem
[0229]
[0230] The initial embedding is provided by LLM:
[0231]
[0232] go through After the layers, the final embedding of knowledge concepts and questions is:
[0233]
[0234] in, is the embedding dimension; is the number of layers, which is a hyperparameter.
[0235] 3.4 Student Status Encoder
[0236] The student state encoder is used to capture the student's learning history and current knowledge mastery. The present invention regards the question sequence in the student's online answer record as questions at different times, and the last question as the question at the current moment. By encoding the student's learning record (such as the correctness of the answer to the question), the model can be helped to better predict the student's answer to future questions.
[0237] (1) Student learning representation:
[0238] For each question , this embodiment constructs the student's learning representation through the representation of its associated concepts. Assume a problem Relating multiple concepts , then the student at time The learning representation is:
[0239]
[0240] in, It's a concept Representation.
[0241] (2) Students responded:
[0242] In order to consider the correctness of the student's answer, this embodiment defines the student at time The answer indicated , which is represented by the learned representation and answer correctness (1 means correct, 0 means wrong) is composed of
[0243]
[0244] Among them, [⋅,⋅] represents the concatenation operation and 0 is the zero vector.
[0245] (3) Student status update:
[0246] In order to capture students' learning behavior, this embodiment uses a gated recurrent unit (GRU) to model students' learning history.
[0247]
[0248]
[0249]
[0250]
[0251] in, is the sigmoid activation function, and is a trainable parameter, Students at the moment state of knowledge.
[0252] 3.5 Answer prediction
[0253] In the answer prediction module, this embodiment combines the student's current knowledge status , semantic representation of the problem and learning representations of concepts , predict the student's answer to the next question. Answers to questions
[0254]
[0255] in, is a trainable weight matrix, is the bias term, is the sigmoid activation function. Through the above formula, the model can use the rich representations generated by the historical state and the cold start phase to accurately predict the next step.
[0256] Therefore, in this embodiment, the constructed answer prediction model can be expressed as the following model structure: based on the heterogeneous graph output by the heterogeneous graph encoder, the student state encoder encodes the student learning representation and student answer representation corresponding to all questions in the target student's online answer record, and then inputs the gated recurrent unit to model the target student's learning history to obtain the target student's current knowledge state, and finally the target student's current knowledge state, the semantic embedding vector of the target question and the student learning representation of the target question are input into the prediction head to obtain the student's answer prediction to the next question.
[0257] During training, this embodiment uses the cross entropy loss function to optimize the prediction of the answer prediction model:
[0258]
[0259] in, is the number of training time steps, is a true answer (1 for true, 0 for false), is the predicted output of the model.
[0260] In addition, since the problem and the problem-solving steps are closely related in the actual task, this embodiment designs a joint training objective to integrate the contrastive learning loss of the two parts and the cross entropy loss in the prediction module into an overall optimization objective:
[0261]
[0262] in is a set of questions in a training batch, is the batch size; and Control the balance between losses.
[0263] By jointly training the answer prediction model, the semantic embeddings of the question and the solution steps are close to each other in the semantic space, while maintaining a high degree of distinction from irrelevant knowledge concepts.
[0264] 4. Validity Verification
[0265] To verify the effectiveness of the present invention, this embodiment uses the widely used online education datasets ASSIST2009, ASSIST2012 and Programming as student online answer record datasets on the online learning platform, and conducts experimental evaluations on these datasets to evaluate the prediction performance of the model. These datasets contain a large number of interactive records of students at different times, questions and skill points on the online learning platform, which can fully reflect the students' knowledge status and behavior patterns, and are suitable for testing knowledge tracking tasks. In the experiment, the performance of the model is measured by two indicators: accuracy (ACC) and area under the receiver operating characteristic curve (AUC). Among them, ACC represents the accuracy of the model's prediction of the correctness of the student's answer, that is, the proportion of correct answers predicted by the model. AUC represents the distinguishing ability of the model prediction, which reflects the performance of the model under different thresholds by calculating the distinguishing effect of the predicted positive and negative samples, and more comprehensively evaluates the performance of the model. The closer the AUC is to 1.0, the better the model is in distinguishing correct and incorrect answers, and has a higher discrimination power.
[0266] Table 1 shows the prediction results of the answer prediction model trained in this embodiment (referred to as this model) and the comparison model on the above three data sets, including ACC and AUC scores. The test results show that the performance of this model on different data sets is stable, and high ACC and AUC values are achieved, which proves the effectiveness and reliability of this model in the knowledge tracking task. This excellent performance shows that the present invention provides more precise concept and question representation, and more accurate estimation of students' knowledge status.
[0267] Table 1: Comparison results of knowledge tracking models
[0268]
[0269] The models used for comparison are: DKT: Deep knowledge tracing; DKVMN: Dynamic Key-Value Memory Networks for knowledge tracing; SAKT: A Self-Attentive model for Knowledge Tracing; AKT: Context-aware Attentive Knowledge Tracing; GKT: Graph-based Knowledge Tracing; DGEKT: Dual Graph Ensemble learning method for Knowledge Tracing.
[0270] In addition, the cold start capability is also verified in this embodiment. The Programming dataset is collected on a commercial code learning platform. The dataset contains plain text of concepts and problem descriptions, while the other two datasets do not contain text information of the problems. In order to study the performance of the present invention in dealing with newly added problems, this embodiment is set to randomly remove a quarter of the problems from the training set, and test the model's predictive ability for these problems on the validation set and the test set. This embodiment mainly compares the performance of the existing DGEKT and AKT models (best and second best) with the performance of the present invention on the Programming dataset. The results are shown in Table 2:
[0271] Table 2 Comparison of performance with and without new questions
[0272]
[0273] The results show that the present invention performs better in predicting unseen questions, while existing models perform poorly in this task. The present invention can better alleviate the cold start problem of knowledge tracking.
[0274] The above-described embodiments are only some preferred implementations of the present invention, but are not intended to limit the present invention. A person skilled in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.
Claims
1. A knowledge tracking cold start optimization method based on a large language model, characterized in that: include: S1. For each question extracted from the student online answer records on the online learning platform, the large language model first outputs a set of problem-solving steps for the question text, then the large language model annotates the set of knowledge concepts involved in the set of problem-solving steps for the question text and the set of problem-solving steps, and finally the large language model outputs a set of binary pairing relationships between the problem-solving steps and the related knowledge concepts for the question text, the set of problem-solving steps and the set of knowledge concepts under constraints; S2. All questions and knowledge concepts on the online learning platform are regarded as two types of nodes in the heterogeneous graph, and according to the binary pairing relationship set of all questions, edge connections between knowledge concepts and questions are established in the heterogeneous graph, and according to the correlation between knowledge concepts, edge connections between knowledge concepts are established, and then the semantic embedding vector of each node is generated; S3, taking the heterogeneous graph as the input of the heterogeneous graph encoder, encoding and aggregating the neighbor information of the graph nodes through different graph attention network layers in each encoding layer of the heterogeneous graph encoder, and aggregating the neighbor information into the node embedding while retaining the original embedding information of the node in combination with the skip knowledge mechanism, and finally outputting the final node embedding of the knowledge concept and the question by the last encoding layer of the heterogeneous graph encoder; S4. Train a question answer prediction model through comparative learning to predict the target student's answer to the target question and complete the model cold start; the question answer prediction model is based on the heterogeneous graph output after encoding by the heterogeneous graph encoder, and the student state encoder encodes the student learning representation and student answer representation corresponding to all questions in the target student's online answer record, and then inputs the gated recurrent unit to model the target student's learning history to obtain the target student's current knowledge state, and finally inputs the target student's current knowledge state, the target question's semantic embedding vector and the student learning representation of the target question into the prediction head to obtain the student's answer prediction to the next question.
2. The knowledge tracking cold start optimization method based on a large language model as claimed in claim 1, characterized in that: The specific implementation steps of S1 are as follows: S11, obtaining the student's online answer record on the online learning platform, and extracting a question set consisting of all questions and the question text and question-related information corresponding to each question; S12, constructing a first prompt template for prompting the large language model to output a solution step for the question text; then for each question in the question set, combining the corresponding question text with the first prompt template and inputting them into the large language model, so that the large language model outputs a solution step set for each question through chain reasoning; S13, constructing a second prompt template, used to prompt the large language model to mark the knowledge concepts involved in the problem-solving step set for the question text and the problem-solving step set; then for each question in the problem set, combining the corresponding question text, the problem-solving step set and the second prompt template and inputting them into the large language model, so that the large language model outputs all the knowledge concepts included in the problem-solving step set of each question through chain reasoning, and obtains the knowledge concept set corresponding to each question; S14. Construct a third prompt template with constraints, which is used to prompt the large language model to output the binary pairing relationship between the problem-solving steps and the related knowledge concepts for the problem text, the problem-solving step set and the knowledge concept set, wherein the constraints are used to prompt the large language model to meet the pairing principles when generating the binary pairing relationship; then, for each problem in the problem set, the corresponding problem text, the problem-solving step set and the knowledge concept set are combined with the third prompt template and input into the large language model, so that the large language model outputs the binary pairing relationship set contained in all the problem-solving steps in each problem through chain reasoning.
3. The knowledge tracking cold start optimization method based on a large language model as claimed in claim 2, characterized in that: The question-related information includes the question type and the student's answer to the question.
4. The knowledge tracking cold start optimization method based on a large language model as claimed in claim 2, characterized in that: It is necessary to set placeholders for personalized fields in the first prompt template, the second prompt template, and the third prompt template, and automatically generate complete prompts that are actually input into the large language model by replacing the placeholders.
5. The knowledge tracking cold start optimization method based on a large language model as claimed in claim 2, characterized in that: The constraints in the third prompt template are set to meet the following conditions at the same time: each problem-solving step is associated with at least one knowledge concept; each knowledge concept is associated with at least one problem-solving step; a many-to-many mapping is allowed between problem-solving steps and knowledge concepts; the large language model needs to combine the historical records that record the generated binary pairing relationships to sequentially generate new binary pairing relationships, and if the newly generated binary pairing relationship does not exist in the historical records, it will be updated to the historical records.
6. The knowledge tracking cold start optimization method based on a large language model as claimed in claim 1, characterized in that: In each encoding layer of the heterogeneous graph encoder, encoding and aggregation are performed on the knowledge concept-problem edges through the first graph attention network layer, and the neighbor problem nodes connected to the central knowledge concept by edges are embedded in the weighted aggregation into the neighbor information of the central knowledge concept; encoding and aggregation are performed on the knowledge concept-knowledge concept edges through the second graph attention network layer, and the neighbor knowledge concept nodes connected to the central knowledge concept by edges are embedded in the weighted aggregation into the neighbor information of the central knowledge concept; encoding and aggregation are performed on the problem-knowledge concept edges through the third graph attention network layer, and the neighbor knowledge concept nodes connected to the central problem by edges are embedded in the weighted aggregation into the neighbor information of the central problem; and for each problem node and knowledge concept node in the heterogeneous graph, the neighbor information is aggregated into its own node embedding through the skip knowledge mechanism.
7. The knowledge tracking cold start optimization method based on a large language model as claimed in claim 1, characterized in that: In the student state encoder, the process of obtaining the current knowledge state of the target student is as follows: from the online answer records, traverse in sequence to find the questions that have generated question-answer interactions with the target student, and at the same time determine the set of knowledge concepts contained in each interacted question, and average the final node embeddings of all knowledge concepts in the knowledge concept set as the student learning representation of the target student for this interacted question, and then combine the student learning representation and the correct or incorrect mark of the target student's answer to the interacted question to construct the student answer representation of the target student for this interacted question, and input the student learning representation and student answer representation corresponding to all the interacted questions of the target student into the gated recurrent unit in sequence for timing modeling, thereby outputting the current knowledge state of the target student.
8. The knowledge tracking cold start optimization method based on a large language model as claimed in claim 1, characterized in that: When the answer prediction model is trained by contrastive learning, it is necessary to construct a first positive-negative sample pair for calculating the contrast loss of the question and a second positive-negative sample pair for calculating the contrast loss of the problem-solving steps; in the first positive-negative sample pair, the positive sample pair is a single question and knowledge concepts related to the question, and the negative sample pair is a single question and knowledge concepts unrelated to the question; in the second positive-negative sample pair, the positive sample pair is a single problem-solving step and a set of knowledge concepts related to the problem-solving step, and the negative sample pair is a single problem-solving step and a set of knowledge concepts unrelated to the problem-solving step.
9. The knowledge tracking cold start optimization method based on a large language model as claimed in claim 8, characterized in that: When the answer prediction model is trained by contrastive learning, the total loss function is obtained by weighted calculation of three parts: the question contrast loss, the problem-solving step contrast loss, and the cross-entropy loss output by the prediction head in the answer prediction model.
10. A knowledge tracking system based on a large language model, characterized in that: include: An interactive module is used for users to specify target students and target questions for knowledge tracking; A model cold start optimization module, used to obtain the heterogeneous graph output after encoding by the heterogeneous graph encoder and the trained answer prediction model according to the knowledge tracking cold start optimization method based on the large language model as described in any one of claims 1 to 9, and store them for calling; The prediction module is used to call the heterogeneous graph and the answer prediction model stored in the model cold start optimization module according to the specified information in the interaction module, output the target student's answer to the target question, and push or visualize it according to the preset logic.
Citation Information
Patent Citations
Deep heterogeneous graph embedding model based on feature fusion
CN114565053A
Medical field question and answer algorithm based on graph attention mechanism
CN115757717A