A knowledge tracing method based on dual paths
Patent Information
- Application Number
- CN202411191786.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-08-28
AI Technical Summary
目前已有的知识追踪模型要么只关注习题难度等特征对未来答题表现的影响,要么只构建学习者知识水平等特征来增强预测效果,忽略其中一种容易导致预测结果准确度的降低
本申请实施例提供的一种基于对偶路径的知识追踪方法,通过构建知识追踪模型KTM学习关于学习者和习题的参数,来进一步地理解学习者知识熟练程度和知识掌握程度,以及习题难度和区分度等要素对学生未来答题中的表现所起的作用。通过融入学习者和习题特征的学生行为数据,能够较为全面的表示学习者的学习情况,从而能够更精确地对学生群体进行分类,对学习者的学习情况进行进一步地划分,将学习者的答题情况和学习行为信息和学习风格、学习习惯等方面进行对应,从而有利于教师更明确的了解每个学生的学习状态,并根据学生的实际学习状态对教学方案作出调整,也可以提示每个学习者了解自身学习的不足之处,使得学习者可以根据实际学习情况更好的规划学习计划。
Smart Images

Figure CN119066450B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of knowledge tracing technology, and in particular to a knowledge tracing method based on dual paths. Background Technology
[0002] With the advancement of information technology, traditional learning methods are undergoing profound changes. A significant indicator of this transformation is the rise of online tutoring systems. These systems collect massive amounts of user learning data, accumulating rich information on learning behavior, which provides solid data support for knowledge tracking technology. Based on this data, knowledge tracking technology can accurately analyze students' learning progress, knowledge mastery, and learning difficulties. The task of knowledge tracking is to track students' mastery of knowledge concepts based on their historical answer sequences and interaction records from past learning activities, thereby predicting their performance in future questions.
[0003] Current knowledge tracing models can be mainly divided into two types: classic knowledge tracing models and deep learning-based knowledge tracing models. Existing knowledge tracing models either focus solely on the impact of features such as question difficulty on future performance, or they only construct features such as learner knowledge level to enhance predictive effectiveness. Ignoring either of these approaches can easily lead to a decrease in prediction accuracy. Meanwhile, although other learning behaviors of students cannot be directly encoded into the answer sequence, analysis of student groups reveals that behavioral factors such as learners' learning habits and styles may have a certain degree of influence on answering questions.
[0004] Therefore, when conducting attribution analysis on the knowledge and abilities of learning groups, it is necessary to collect as comprehensive an information collection as possible on learners' learning behaviors on online education platforms and to deeply explore the representations of students' internal cognitive structures. Summary of the Invention
[0005] This application provides a knowledge tracing method based on dual paths to address the shortcomings of the aforementioned related technologies. This application aims to construct parameters about the exercises and learners that the model needs to learn by collecting students' answer information and as comprehensive a set of behavioral information as possible on a learning platform. This allows for a deeper attribution analysis of students' knowledge mastery. The technical solution is as follows: In a first aspect, embodiments of this application provide a knowledge tracing method based on dual paths, including: Obtain learning record data for each student on the online learning platform; The learning record data of each student is input into the trained knowledge tracking model, so that the trained knowledge tracking model can extract features based on the learning record data and output the learning behavior data of each student; wherein, the learning behavior data includes the student feature data and exercise feature data of each student; The learning behavior data and learning record data of each student are input into the clustering model so that the clustering model can perform cluster analysis on the learning behavior data and learning data of all students, and output the knowledge tracking result label corresponding to each cluster center according to the clustering results; The learning record data of each student corresponding to each knowledge tracking result tag is analyzed, the learning data analysis results corresponding to the knowledge tracking result tag are output, and the corresponding learning guidance information is generated based on the learning data analysis results and fed back to the corresponding students.
[0006] In one alternative embodiment of the first aspect, after obtaining the learning record data of each student on the online learning platform, the method further includes: Based on each student's learning data, answer record data and learning log record data for each student are obtained. After processing the answer record data and learning log record data, the answer sequence for each student is obtained, and the answer data for each student is output. Based on the answer record data and the learning log record data, the corresponding knowledge point set and exercise set are extracted. Based on the association relationship between each exercise and at least one knowledge point, the exercise-knowledge point association matrix is output.
[0007] In one alternative to the first aspect, the association between each exercise and at least one knowledge point is determined based on the following steps: Based on the knowledge graph, entities of preset data types are extracted from the answer record data and the learning log record data; Using exercises as the main entity and knowledge points as the secondary entities, identify secondary entities that are related to each main entity, and add the related relationships between each main entity and secondary entity to obtain multiple triples <exercise, relationship, knowledge point> constructed based on exercises and knowledge points.
[0008] In one alternative embodiment of the first aspect, inputting the learning record data of each student into a trained knowledge tracking model, so that the trained knowledge tracking model performs feature extraction based on the learning record data, includes: The answer data and the exercise-knowledge point association matrix corresponding to each student are input into the trained knowledge tracking model. The trained knowledge tracking model performs feature extraction based on the answer data and the exercise-knowledge point association matrix to extract the student feature data and the exercise feature data for each student. The student characteristic data includes the student's proficiency in knowledge points, the student's mastery of knowledge points, the probability of guessing answers, and the probability of making mistakes in answers. The exercise characteristic data includes the difficulty of the exercises and the discrimination of the exercises.
[0009] In one alternative to the first aspect, the proficiency level of the knowledge point is calculated based on the following steps: Based on the answer data for each student and the exercise-knowledge point association matrix, the information space of all exercises answered by all students at all times is obtained. ; Vector of answer results The proficiency level of each student in the knowledge points is calculated by multiplying the result with the problem-knowledge point association matrix. in, Students At any moment t On the issue The vector of the answer results, This indicates that the student At any moment t On the issue Make the correct answer, This indicates that the student At any moment t On the issue Making a wrong answer, This indicates that the student At any moment t On the issue No response was given; the total number of students was n The total number of exercises is m .
[0010] In one alternative to the first aspect, the level of mastery of the knowledge points is calculated based on the following steps: Get students At any moment t To assess students' mastery of knowledge points At any moment t +1 indicates proficiency in all knowledge points; Based on students At any moment t The degree of mastery of knowledge points and the time t +1 Calculates student's proficiency level across all knowledge points. exist t Time's up t The change in the mastery of knowledge points during the learning process at time +1; Calculate the students' performance t The level of understanding of the knowledge points at time +1; The student's initial knowledge level at the initial moment is determined based on the student's historical answer data.
[0011] In one alternative to the first aspect, the exercise feature data is calculated based on the following steps: Obtain the answer sequence of each student for each exercise, calculate the number of students who gave the correct answer to each exercise, and determine the difficulty of the exercise based on the number of students who gave the correct answer; Obtain the answers from multiple students to the same exercise, obtain the distribution of the answers and the knowledge points involved in the same exercise, and determine the discrimination index of the exercise based on the distribution of the answers, the knowledge points involved in the same exercise, and the difficulty of the same exercise.
[0012] Secondly, embodiments of this application also provide a knowledge tracing device based on dual paths, comprising: The first data acquisition unit is used to acquire the learning record data of each student on the online learning platform; The second data acquisition unit is used to input the learning record data of each student into the trained knowledge tracking model, so that the trained knowledge tracking model can perform feature extraction based on the learning record data and output the learning behavior data of each student; wherein, the learning behavior data includes the student feature data and exercise feature data of each student; The data analysis unit is used to input the learning behavior data and learning record data of each student into the clustering model, so that the clustering model can perform cluster analysis on the learning behavior data and learning data of all students, and output the knowledge tracking result label corresponding to each cluster center according to the clustering results; The data analysis unit is also used to analyze the learning record data of each student corresponding to each knowledge tracking result tag, output the learning data analysis results corresponding to the knowledge tracking result tag, and generate corresponding learning guidance information based on the learning data analysis results and feed it back to the corresponding students.
[0013] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method provided by the first aspect or any implementation thereof of the embodiments of this application.
[0014] Fourthly, this application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method provided by the first aspect of the embodiments of this application or any implementation thereof.
[0015] The beneficial effects of the technical solutions provided in some embodiments of this application include at least the following: This application provides a knowledge tracing method based on dual paths. By constructing a Knowledge Tracing Model (KTM) to learn parameters about learners and exercises, it further understands the role of learners' knowledge proficiency and mastery, as well as the difficulty and discrimination of exercises, in students' future performance. By incorporating student behavior data based on learner and exercise characteristics, it can comprehensively represent learners' learning situations, thereby enabling more accurate classification of student groups and further segmentation of learners' learning situations. It correlates learners' answer performance and learning behavior information with learning styles and habits, which helps teachers better understand each student's learning status and adjust teaching plans accordingly. It also helps each learner recognize their own learning weaknesses, allowing them to better plan their learning based on their actual situation. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a knowledge tracing method based on dual paths according to an embodiment of this application. Figure 2 This is a schematic diagram of the structure of a knowledge tracking device based on dual paths provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or modules is not limited to the steps or modules listed, but may optionally include steps or modules not listed, or may optionally include other steps or modules inherent to such process, method, product, or apparatus.
[0020] It should be noted that the terms "first" and "second" used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that "first" and "second" can be interchanged in a specific order or sequence where permitted. It should be understood that the objects distinguished by "first" and "second" can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in an order other than those described or illustrated herein.
[0021] The present application will now be described in detail with reference to specific embodiments.
[0022] Next, combine Figure 1 Taking the calculation method of application usage time on the terminal as an example, this application introduces a knowledge tracking method based on dual paths provided in its embodiments. For details, please refer to... Figure 1 , Figure 1 This paper illustrates a flowchart of a knowledge tracing method based on dual paths provided in an embodiment of this application. The method includes the following steps: S101. Obtain the learning data of each student on the online learning platform and construct an exercise-knowledge point association matrix for each student. Obtain students' answer data D Resp .
[0023] S102, student answer data D Resp The exercise-knowledge point association matrix is input into the trained knowledge tracking model, and the learning behavior data of each student is obtained through the trained knowledge tracking model.
[0024] S103. Based on the students' learning record data and the characteristics of the students and exercises output by the model, perform cluster analysis on all students and output the knowledge tracking result label corresponding to each student.
[0025] S104: Analyze the student learning records corresponding to various knowledge tracking result tags, determine the influence weight of various indicators on learning outcomes, and generate corresponding learning guidance results based on the weights.
[0026] Specifically, in S101, learning data for each student can be obtained through an online learning platform. This learning data includes, but is not limited to, students' reading records of learning materials, viewing duration and progress of learning videos, number of interactions and content of interactions with teachers in online classes, time spent completing exercises, and answer records. The online learning platform can be any network platform used to provide practice questions, video learning materials, or text learning materials, so that students can learn and complete the content given in the course on the relevant platform. This application embodiment does not limit this.
[0027] Specifically, the learning data may include each student's answer record data and learning log record data. The answer record data may include the student's answer content and answer result for each exercise. The learning log record is used to record the student's learning process data, including but not limited to the learning time of various learning materials, the learning progress of class hours, the order of answering questions, and the time of answering questions. For example, the learning platform can record each student's mouse click operation and the learning time of learning materials (such as text reading time) during the process of answering specific exercises to record the student's learning behavior. This application embodiment does not limit this.
[0028] In some embodiments, after S101, the answer record data of each student can be obtained based on each student's learning data. R Resp and learning log data R Log Based on answer record data R Resp and learning log data R Log After data preprocessing, the corresponding student's answer sequence is obtained: in, For students exist t The exercises done at all times. For students exist t Practice at all times The answer, =1 indicates that the student answered correctly. =0 indicates that the student answered incorrectly.
[0029] For example, as shown in Table 1 below, students The answer sequence can be represented as: Table 1 Answer Sequence In some embodiments, considering that each exercise may correspond to multiple knowledge points, it is difficult to represent the relationship between each knowledge point and the exercise through a knowledge graph using a knowledge tracing model. Therefore, with exercises as the main entity and knowledge points as secondary entities, secondary entities that are associated with each main entity are identified, and <exercise> are constructed accordingly. Includes knowledge points The triples of > represent the correspondence between exercises and knowledge points.
[0030] Specifically, the knowledge graph method of exercises and knowledge points is used to process all knowledge points related to all the exercises answered by students, and output a set of knowledge points. ,in , Represents a set of knowledge points The first in k There are 1 knowledge point; the set Q of all exercises can be represented as This allows us to construct a corresponding knowledge matrix based on the relationship between each exercise and knowledge point. Through the exercise-knowledge point association matrix, that is, the knowledge matrix It can characterize the relevance of all knowledge points and exercises involved in all the exercises that learners need to answer, specifically including: ; in, Each element in the matrix Then it represents the exercise. With knowledge points Relationship, Representing exercises Knowledge points involved ,and Then it represents the exercise. Knowledge points not covered .
[0031] It should be noted that a single exercise may involve multiple knowledge points; therefore, a knowledge matrix is required. A single row in the array may contain more than one non-zero element.
[0032] Furthermore, in S102, the students' answer data... D Resp And exercises—knowledge point association matrix The data is input into the trained knowledge tracing model KTM, which extracts the learning behavior data of each student.
[0033] It should be noted that when processing students' answer data... D Resp And exercises—knowledge point association matrix Before inputting the answer data into the trained knowledge tracking model KTM, it is necessary to process the answer data. D Resp And exercises—knowledge point association matrix Preprocessing is performed to conform to the format requirements of the input data of the trained knowledge tracking model.
[0034] Specifically, student learning behavior data can include student characteristic data and exercise characteristic data. Student characteristic data includes the student's proficiency in knowledge points, their level of mastery of those knowledge points, the probability of guessing answers, and the probability of making mistakes in answering questions. , The problem feature data includes problem difficulty and problem discrimination.
[0035] Among them, the proficiency level of knowledge points is used to represent the students' proficiency in various knowledge points. It is mainly based on the analysis of students' answers to exercises related to these knowledge points. If the percentage of correct answers is high, it indicates a higher level of proficiency in the knowledge points. The specific calculation process includes: First, obtain the information space of all exercises answered by all students at all times. , Students At any moment t On the issue The vector of the answer results, This indicates that the student At any moment t On the issue Make the correct answer, This indicates that the student At any moment t On the issue Making a wrong answer, This indicates that the student At any moment t On the issue No response was given; the total number of students was n The total number of exercises is m .
[0036] Vector of answer results With knowledge matrix Multiplication is used to calculate students' answers to exercises related to specific knowledge points, thus determining their level of proficiency in those knowledge points. .
[0037] Furthermore, it can be based on each student's proficiency in the knowledge points. Calculate each student's mastery of all knowledge points. ,in, The value ranges from 0 to 1. A higher value indicates a better grasp of all knowledge points by the student. The specific steps include: Get students At any moment t Knowledge points mastery , obtain students At any moment t +1 Proficiency level in all knowledge points .
[0038] Based on students At any moment t Knowledge points mastery And at the moment t +1 Calculates student's proficiency level across all knowledge points. exist t Time's up t Change in the mastery of knowledge points during the learning process at time +1 Apply the formula: ; in, Matrix representing students The sequence of answer results for each knowledge point. Matrix representing each exercise The sequence of knowledge points covered; Understandably, the change in the degree of mastery of knowledge points is used to characterize students'... exist t Time's up t The student's answers to exercises at time +1 can indicate their learning progress. Changes in the level of mastery of relevant knowledge points.
[0039] Specifically, the formula can be used to calculate the student's performance. t Mastery of the knowledge points at time +1 Apply the formula: ; Understandably, Students The initial level of knowledge mastery can be determined based on the student's historical answer data or a preset value. This application does not limit this.
[0040] In some embodiments, even if a student answers a practice question incorrectly during the test-taking process, it should be considered that the student's mastery of the knowledge point has changed during this process. Therefore, the change in the mastery of the knowledge point is considered. A non-negativity constraint is required; apply the formula: in, This indicates a non-negative activation function, ensuring that the resulting change in the level of knowledge mastery better reflects the changes in students' knowledge mastery during the learning process.
[0041] Furthermore, it can be based on each student's understanding of each exercise. The answer result determines the corresponding exercise. The difficulty and discrimination of exercises are considered. Difficulty can be calculated based on the number of students who answered the same exercise correctly. A higher number of correct answers indicates lower difficulty, and vice versa. Discrimination can be calculated based on difficulty, the distribution of scores among students on the same exercise, and the knowledge points covered. If most students' scores are similar, discrimination is poor; if all students' scores follow a normal distribution, discrimination is good. Discrimination is used to characterize the ability of exercises to assess students' mastery of knowledge points, distinguishing between knowledge points with good and poor mastery.
[0042] Specifically, it is possible to obtain student information. exist t Exercises to be answered at any time All the knowledge points involved are calculated using the formula: in, This represents one-hot encoding, exercises. After undergoing single heat treatment, the following is obtained , This indicates that the matrix is transposed. , representing exercises Relevance to the corresponding knowledge points.
[0043] Furthermore, regarding the exercises Difficulty and discrimination It can be calculated using the following formula: in, and All of these are trainable parameters.
[0044] It should be noted that during the training process of the knowledge tracing model, the knowledge tracing model can be used to monitor students' performance. t Predict the probability based on the answer at time +1: in, This indicates that students are in the lead. t Based on the students' responses at each moment, predict their performance. t +1 time for the problem The performance on the test represents a probability. .
[0045] Understandably, the knowledge tracking model needs to combine the learner's proficiency in the knowledge points, mastery of the knowledge points, difficulty of the questions, and discrimination of the questions to predict relevant parameters. Specifically, a parameterized Rasch model can be used to estimate the students' answer performance, thereby calculating the probability of guessing and the probability of making mistakes in answering questions during the training process.
[0046] Specifically, the formula is applied as follows: ; in, A factor vector representing a combination of students' proficiency in knowledge points, their mastery of those knowledge points, the difficulty of practice problems, and the discrimination effect of those problems can be used to characterize students. Answer the exercise at time t. The complex interaction process, This represents element-wise multiplication between vectors.
[0047] Furthermore, based on the above... A parametric Rasch model can be used to estimate students' initial performance on questions: ; in, Students Answer the exercises The initial probability of answering the question. is a constant associated with the given dataset, which can usually be preset to 0 or 0.25. d is a constant in the Rasch model, which can be set to d=1.702.
[0048] Furthermore, during the student's answering process, it is usually necessary to consider guessing and error factors in order to more accurately predict the student's probability of answering correctly, using the formula: ; in, Students Answer the exercises The final probability of answering the question. The overall probability of students making mistakes when answering all exercises. For students The overall probability of guessing the answer. Answering questions for multiple students The overall probability of guessing, the probability of error, and the probability of guessing can be calculated in the process of the model predicting the probability of students answering questions.
[0049] Specifically, the above probability is... Based on the characteristics of the binary classification task in the knowledge tracing model KTM, the loss function is designed as follows: ; in, Indicates that students are t The true value of the answer at any given moment, and This indicates that the knowledge tracing model outputs the student's understanding of the exercises. The goal of the knowledge tracing model is to minimize the difference between the actual and predicted answers in order to more accurately predict students' future performance, thereby calculating the probability of error. And guessing probability .
[0050] Therefore, the data obtained through this knowledge tracking model includes: proficiency level of knowledge points. Mastery of knowledge points Difficulty Discrimination Error probability And guessing probability Combined with the learning log data obtained in S101 R Log After data processing, the output is the learning behavior data corresponding to each student. D Act .
[0051] Execute S103, based on the level of proficiency in the knowledge points. Mastery of knowledge points Difficulty Discrimination And guessing probability g Error probability s Data processing is performed to obtain the feature vector of each student, which is then input into a clustering model. Based on the clustering results of the model, the knowledge tracking result label corresponding to each student is output, specifically including: The encoder part learns parameters for each layer and parameters To obtain the output parameters hThe decoder part of each layer learns parameters. and parameters To obtain the output parameters y Its goal is to minimize the variance. .
[0052] The parameters of the previous layer h As input to the next layer, and using the ReLU activation function in all encoder-decoder pairs (except for the parameters of the first pair). And the parameters of the last pair The parameters of the model are .
[0053] In the parameter optimization phase, optimization is performed through iterative computation of soft assignments and minimization of KL divergence. The depth mapping is updated using an auxiliary target distribution learned from the current high-confidence assignments. And refine the cluster centroids.
[0054] Furthermore, the soft assignment of embedded features to cluster centroids is calculated, including: Preset k For each cluster center, the distance between the student's embedded features and the centroid satisfies... t If the distribution is such that the soft allocation result is calculated using the following formula: ; Furthermore, the auxiliary distribution calculation of the target includes: soft allocation The auxiliary distribution is calculated by raising the power to the square and then normalizing it according to the frequency of each cluster. : ; Specifically, the objective function of the deep clustering model DEC is defined as soft allocation. and auxiliary distribution KL divergence: ; Furthermore, the clustering model DEC uses stochastic gradient descent to jointly optimize cluster centers. and encoder parameters , L Relative to each data point and the centroid of each cluster The gradient calculation for the feature space embedding is as follows: ; Updated via backpropagation This allows for the updating of certain encoder parameters. The purpose.
[0055] For clustering algorithms, the clustering model is preset with... k The number of cluster centroids determines the number of cluster centers, which can be determined using kL line plots, the elbow principle, and other indicators. k The value.
[0056] Therefore, we can use the learning behavior data of different students D Act Clustering yields multiple clusters with the same knowledge tracking result label, and the knowledge tracking result label corresponding to each student is determined based on each student's learning behavior data.
[0057] Specifically, the knowledge tracking result labels can be as shown in Table 2 below. The corresponding labels are determined based on the specific error probability, guess probability, and the learning data of each student obtained from the online learning platform. The learning data may include knowledge points. Viewing duration of related videos, Text reading time, The amount of practice problems Knowledge mastery level (proficiency and understanding of knowledge points) and accuracy rate in answering questions, etc.
[0058] Specifically, the actual error probability, guess probability, viewing time, text reading time, amount of practice questions, knowledge mastery (proficiency and mastery of knowledge points), and answer accuracy can be compared with the average values of all students. Alternatively, a corresponding threshold can be preset to determine the difference between the actual value and the threshold. Based on the comparison results with the threshold and / or the difference, the corresponding knowledge tracking result label can be determined.
[0059] Table 2 Knowledge Tracking Results Tags Here, the superscript avg indicates the average level of the corresponding data. The average of the accuracy rate. This indicates the average duration of video viewing. This indicates the average time spent reading the text. An average value representing the degree of knowledge mastery. This indicates the average number of exercises practiced; "-" indicates that this item is not compared. This indicates that the value is lower than the corresponding value and the difference between the two values is small. ">>" indicates a smaller absolute value of the difference from the corresponding value; ">>" indicates a value greater than the corresponding value and a larger difference from the corresponding value; and "<<" indicates a value less than the corresponding value and a larger difference from the corresponding value.
[0060] Therefore, the tracking results for each student can be determined separately, allowing for an analysis of the specific learning situation based on the student's learning outcomes. For example, in the first type of tracking results, it is clearly observed that students' time investment in learning materials is significantly insufficient, and the amount of practice exercises is also significantly inadequate. In the second type of tracking results, it is clearly observed that students' time investment in learning materials is somewhat insufficient, the amount of practice exercises is also inadequate, their knowledge mastery is slightly below average, and their practice exercise results are below average. In the third type of tracking results, students' error rate when doing exercises is higher than average, their time investment in learning materials is significantly higher than average, their amount of practice exercises is slightly below average, their knowledge mastery is higher than or equal to average, and their exercise accuracy rate is slightly below average. The tracking results for the fourth and fifth types will not be elaborated here.
[0061] Furthermore, in step S104, the obtained knowledge tracking result labels can be used to output learning data analysis results for the corresponding students based on their learning behavior data, thereby providing learning guidance suggestions for the students.
[0062] Specifically, based on the knowledge tracking result labels determined in S103, common problems of a certain type of students can be analyzed and identified, thereby providing learning guidance for that type of students. This approach can reflect the common problems of most students and improve the efficiency of generating learning guidance.
[0063] Specifically, the learning data analysis results can be in text form. They can be combined with a large language model to build prompt templates using the knowledge tracking result tags output by the knowledge tracking model, and the large language model can be used to generate corresponding analysis results and guidance.
[0064] For example, for the first type of tracking results, feedback can be provided such as increasing study time and increasing the amount of practice questions to strengthen learning behavior; for the third type of tracking results, feedback can be provided such as increasing the amount of practice questions and adjusting the knowledge points corresponding to the practice questions based on the amount of practice questions.
[0065] Specifically, learning data analysis results can be provided for each student based on their learning behavior data, as shown in Table 3 below: Table 3 Student Groups Learning data analysis Specifically, the data analysis projects mentioned above can determine the impact of students' learning behavior data on learning outcomes. Then, based on the various data points in students' learning behavior data, weighted calculations can be performed to generate targeted learning guidance results. This is more conducive to pointing out students' shortcomings in specific projects and promoting students' learning.
[0066] Specifically, based on the learning data analysis results of each student, the learning records of students with good learning performance can be obtained from a data perspective, such as learning time, total number of exercises, number of exercises for various knowledge points, etc. The learning records of students with good learning performance can be used to plan the learning progress of other students. This application embodiment does not limit this.
[0067] Optionally, online learning platforms can provide a visual interface to help teachers and students understand the overall learning status.
[0068] The following are apparatus embodiments of this application, which can be used to execute the method embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments of this application.
[0069] Please see below. Figure 2 This is a schematic diagram of a knowledge tracing device based on dual paths, provided as an exemplary embodiment of this application. This device can be implemented as all or part of a terminal through software, hardware, or a combination of both, or it can be integrated as an independent module on a server. The knowledge tracing device based on dual paths in this embodiment can be applied to a terminal or the cloud. The device 20 includes a first data acquisition unit 201, a second data acquisition unit 202, and a data analysis unit 203, wherein: The first data acquisition unit 201 is used to acquire the learning record data of each student on the online learning platform; The second data acquisition unit 202 is used to input the learning record data of each student into the trained knowledge tracking model, so that the trained knowledge tracking model can perform feature extraction based on the learning record data and output the learning behavior data of each student; wherein, the learning behavior data includes the student feature data and exercise feature data of each student; The data analysis unit 203 is used to input the learning behavior data and learning record data of each student into the clustering model, so that the clustering model can perform cluster analysis on the learning behavior data and learning data of all students, and output the knowledge tracking result label corresponding to each cluster center according to the clustering result; The data analysis unit 203 is also used to analyze the learning record data of each student corresponding to each knowledge tracking result tag, output the learning data analysis results corresponding to the knowledge tracking result tag, and generate corresponding learning guidance information based on the learning data analysis results and feed it back to the corresponding students respectively.
[0070] It should be noted that the device 20 provided in the above embodiments is only illustrated by the division of the above functional modules when executing the knowledge tracing method based on dual paths. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the knowledge tracing method embodiments based on dual paths belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.
[0071] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the methods described above.
[0072] Please see Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of this application.
[0073] like Figure 3 As shown, the electronic device 300 includes a processor 301 and a memory 302.
[0074] In this embodiment, the processor 301 is the control center of the computer system, and can be a processor of a physical machine or a processor of a virtual machine. The processor 301 may include one or more processing cores, such as a 4-core processor or an 8-core processor. The processor 301 can be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array).
[0075] Processor 301 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake-up state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor used to process data in the standby state.
[0076] Memory 302 may include one or more computer-readable storage media, which may be non-transitory. Memory 302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments of this application, the non-transitory computer-readable storage media in memory 302 is used to store at least one instruction, which is executed by processor 301 to implement the method in the embodiments of this application.
[0077] In some embodiments, the electronic device 300 further includes a peripheral device interface 303 and at least one peripheral device 304. The processor 301, memory 302, and peripheral device interface 303 can be connected via a bus or signal line. Each peripheral device 304 can be connected to the peripheral device interface 303 via a bus, signal line, or circuit board. Specifically, the peripheral device 304 includes: a display screen, a camera, and audio circuitry. The peripheral device interface 303 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 301 and memory 302.
[0078] In some embodiments of this application, the processor 301, memory 302, and peripheral device interface 303 are integrated on the same chip or circuit board; in other embodiments of this application, any one or two of the processor 301, memory 302, and peripheral device interface 303 can be implemented on separate chips or circuit boards. This application does not specifically limit the implementation in this regard.
[0079] The electronic device structural block diagram shown in the embodiments of this application does not constitute a limitation on the electronic device 300. The electronic device 300 may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0080] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the methods in any of the foregoing embodiments. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, as well as magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.
[0081] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A knowledge tracing method based on dual paths, characterized in that, include: The learning record data of each student on the online learning platform is obtained. Based on the learning record data of each student, the answer record data and learning log record data of each student are obtained. After processing the answer record data and the learning log record data, the answer sequence of each student is obtained, and the answer data of each student is output. Based on the answer record data and the learning log record data, the corresponding knowledge point set and exercise set are extracted. Based on the association relationship between each exercise and at least one knowledge point, the exercise-knowledge point association matrix is output. The exercise-knowledge point association matrix is represented as follows: The set of knowledge points is represented by C. Let K represent the k-th knowledge point in the set of knowledge points, where K represents the total number of knowledge points. The set of exercises is represented as... , Let represent the j-th exercise in the exercise set, M represent the total number of exercises, and the element in the j-th row and k-th column of the exercise-knowledge point association matrix is represented as . , This represents the exercises in row j. Knowledge points related to column k Relationship, Representing exercises Knowledge points involved , Then it represents the exercise. Knowledge points not covered Each row of the exercise-knowledge point association matrix includes at least one element that is 1; The steps to determine the relationship between exercises and knowledge points include: Based on the knowledge graph, entities of preset data types are extracted from the answer record data and the learning log record data; Using exercises as the main entity and knowledge points as the secondary entities, each secondary entity is associated with a specific main entity. The associated relationships are then added between each associated main entity and secondary entity, resulting in multiple triples <exercise, relationship, knowledge point> constructed based on exercises and knowledge points. The learning record data of each student is input into a trained knowledge tracing model, enabling the model to extract features based on the learning record data and output each student's learning behavior data. This learning behavior data includes each student's student characteristic data and exercise characteristic data. The learning behavior data and learning record data of each student are input into the clustering model so that the clustering model can perform cluster analysis on the learning behavior data and learning data of all students. Based on the clustering results, the knowledge tracking result label corresponding to each cluster center is output. The parameter optimization process of the clustering model includes: The encoder part of each layer of the clustering model learns parameters. and parameters To obtain the output parameters h The decoder part of each layer learns parameters. and parameters To obtain the output parameters y ; The parameters of the previous encoder layer h It serves as the input to the next layer encoder and is used with the ReLU activation function in all encoder-decoder pairs; The parameters of the clustering model are optimized through iterative computation of soft assignments and minimization of KL divergence. The depth mapping is updated by learning from the current high-confidence assignments using an auxiliary target distribution. And refine the cluster centroids; Calculating the soft assignment of embedded features to cluster centroids includes: Preset k The centroids of each cluster, and the distances from the embedded features of students to the centroids satisfy the following conditions: t The formula for calculating the distribution and soft allocation is as follows: ; The calculation process for the auxiliary distribution includes: soft allocation The power is increased to the square, and then normalized according to the frequency of each cluster to calculate the auxiliary distribution. The calculation formula is as follows: ; in, Indicate the embedding features of the i-th student and the i-th student. Soft allocation of cluster centroids, This represents the embedding feature of the i-th student. Indicates the first The centroid of each cluster, express t Parameters of the distribution; Denotes the auxiliary distribution. Indicates the first The frequency of each cluster, the pre-defined value of the clustering model. k The centroid of each cluster determines the number of clusters. The objective function of the clustering model DEC is defined as the KL divergence of the soft assignment and auxiliary distribution. The gradient of the KL divergence with respect to the feature space embedding of each embedded feature and each cluster centroid is calculated. Stochastic gradient descent is used to jointly optimize the cluster centroid and encoder parameters. The learning record data of each student corresponding to each knowledge tracking result tag is analyzed, the learning data analysis results corresponding to the knowledge tracking result tag are output, and the corresponding learning guidance information is generated based on the learning data analysis results and fed back to the corresponding students.
2. The knowledge tracing method based on dual paths according to claim 1, characterized in that, The step of inputting the learning record data of each student into the trained knowledge tracking model, so that the trained knowledge tracking model can perform feature extraction based on the learning record data, includes: The answer data and the exercise-knowledge point association matrix corresponding to each student are input into the trained knowledge tracking model. The trained knowledge tracking model performs feature extraction based on the answer data and the exercise-knowledge point association matrix to extract the student feature data and the exercise feature data for each student. The student characteristic data includes the student's proficiency in knowledge points, the student's mastery of knowledge points, the probability of guessing answers, and the probability of making mistakes in answers. The exercise characteristic data includes the difficulty of the exercises and the discrimination of the exercises.
3. The knowledge tracing method based on dual paths according to claim 2, characterized in that, The level of proficiency in the knowledge points is calculated based on the following steps: Based on the answer data for each student and the exercise-knowledge point association matrix, the information space of all exercises answered by all students at all times is obtained. ; Vector of answer results The proficiency level of each student in the knowledge points is calculated by multiplying the result with the problem-knowledge point association matrix. in, Students At any moment t On the issue The vector of the answer results, This indicates that the student At any moment t On the issue Make the correct answer, This indicates that the student At any moment t On the issue Making a wrong answer This indicates that the student At any moment t On the issue No response was given; the total number of students was n The total number of exercises is m .
4. The knowledge tracing method based on dual paths according to claim 3, characterized in that, The level of mastery of the knowledge points is calculated based on the following steps: Get students At any moment t To assess students' mastery of knowledge points At any moment t +1 indicates proficiency in all knowledge points; Based on students At any moment t The degree of mastery of knowledge points and the time at which t +1 Calculates student's proficiency level across all knowledge points. exist t Time's up t The change in the mastery of knowledge points during the learning process at time +1; Calculate the students' performance t The level of understanding of the knowledge points at time +1; The student's initial knowledge level at the initial moment is determined based on the student's historical answer data.
5. The knowledge tracing method based on dual paths according to claim 3, characterized in that, The exercise feature data is calculated based on the following steps: Obtain the answer sequence of each student for each exercise, calculate the number of students who gave the correct answer to each exercise, and determine the difficulty of the exercise based on the number of students who gave the correct answer; Obtain the answers from multiple students to the same exercise, obtain the distribution of the answers and the knowledge points involved in the same exercise, and determine the discrimination index of the exercise based on the distribution of the answers, the knowledge points involved in the same exercise, and the difficulty of the same exercise.
6. An apparatus based on the knowledge tracing method based on dual paths as described in any one of claims 1-5, characterized in that, The device includes: The first data acquisition unit is used to acquire the learning record data of each student on the online learning platform; The second data acquisition unit is used to input the learning record data of each student into the trained knowledge tracking model, so that the trained knowledge tracking model can perform feature extraction based on the learning record data and output the learning behavior data of each student; wherein, the learning behavior data includes the student feature data and exercise feature data of each student; The data analysis unit is used to input the learning behavior data and learning record data of each student into the clustering model, so that the clustering model can perform cluster analysis on the learning behavior data and learning data of all students, and output the knowledge tracking result label corresponding to each cluster center according to the clustering results; The data analysis unit is also used to analyze the learning record data of each student corresponding to each knowledge tracking result tag, output the learning data analysis results corresponding to the knowledge tracking result tag, and generate corresponding learning guidance information based on the learning data analysis results and feed it back to the corresponding students.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.