Knowledge tracking method, system and device based on attention mechanism and hypergraph convolutional network, and medium

By introducing attention mechanism and hypergraph convolutional network into the knowledge tracking method, combining feature enhancement and historical interaction modules, the problems of long sequence dependency capture and data quality dependence are solved, and higher accuracy and robustness are achieved.

CN119990296APending Publication Date: 2025-05-13XIDIAN UNIV

Patent Information

Application Number
CN202510087355.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing deep learning-based knowledge tracking methods have difficulties in long-sequence dependency capture and interpretability, and have high requirements for data quality, making computing power and algorithm design difficult.

Method used

The knowledge tracking method based on attention mechanism and hypergraph convolutional network is adopted, combined with feature enhancement module, feature fusion module, historical interaction module and difficulty constraint module, the many-to-many relationship between the hypergraph convolutional network modeling questions and skills is used to capture long sequence dependencies using channel attention mechanism and sparse multi-head self-attention mechanism, and the model robustness is improved by fuzzing modules and difficulty penalty terms.

Benefits of technology

It improves the accuracy and robustness of the knowledge tracking task, can more effectively capture long sequence dependencies, enhance feature representation ability, and adapt to the learning intensity of questions of different difficulty levels, reducing the dependence on data quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990296A_ABST
    Figure CN119990296A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge tracking method, system and equipment based on an attention mechanism and a hypergraph convolutional network, and a medium. The method comprises the following steps: acquiring online record data of a to-be-trained student and preprocessing the online record data; a hypergraph convolutional network module is used for modeling the many-to-many relationship between the questions and the skills, and question embedding fusing complex associated information between the questions and the skills is constructed; the method comprises the following steps: preprocessing different behavior characteristics of a student, fusing the preprocessed different behavior characteristics of the student through a channel attention mechanism, improving tracking precision by using the fused behavior characteristics, capturing long sequence dependence of a student online record data sequence by using a sparse multi-head self-attention mechanism, and providing support for state updating by using a student historical practice record as assistance; a question difficulty feature and a fuzzification module are introduced, the difficulty feature subjected to fuzzification processing is used for assisting training to improve the performance of the model, meanwhile, a difficulty penalty term is added to a loss function, the learning strength of the model for different difficulty questions is adjusted to improve the robustness, and an optimized overall model is obtained; and finally realizing the knowledge tracking method based on the attention mechanism and the hypergraph convolutional network. The system, the equipment and the medium can perform knowledge tracking based on the method; the method has the advantages of high accuracy and good robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of educational big data mining based on online data, and in particular to a knowledge tracking method, system, device and medium based on an attention mechanism and a hypergraph convolutional network. Background Art

[0002] With the rapid development of information technology and the continuous deepening of educational informatization, the application of educational big data has become an important means to promote educational reform and improve the quality of education. In modern educational scenarios, a large number of students' online learning behaviors, academic performance and daily activity data are widely collected and stored. How to provide precise and intelligent decision-making support for education and teaching through the analysis and mining of these massive data has become one of the hot issues in the field of education research. Knowledge tracking, as an important tool for evaluating students' learning status, is one of the core tasks in educational intelligence. It aims to predict students' mastery of future questions through their historical learning behaviors, thereby providing support for personalized teaching and helping students improve their learning efficiency. The results of knowledge tracking can help teachers understand students' knowledge mastery, optimize teaching strategies, and provide a scientific basis for educational decision-making.

[0003] However, although deep learning-based knowledge tracking has certain advantages over traditional knowledge tracking methods (such as BKT and IRT) in terms of feature modeling, data processing, scalability, and accuracy, it still has problems such as difficulty in capturing long sequence dependencies and insufficient interpretability. It also depends on the quality of the data, which not only requires high computing power, but also puts higher requirements on algorithm design.

[0004] Knowledge tracking methods can be mainly divided into three stages, including statistical methods, matrix decomposition methods, and deep learning methods. The first two traditional methods usually rely on artificial feature design and statistics, and can only handle single knowledge points, cannot handle complex data relationships, and have difficulty capturing the dynamic changes of students' knowledge status. Therefore, the accuracy and efficiency of these methods are limited. With the continuous development and improvement of deep learning technology, knowledge tracking methods based on this technology are also constantly being optimized and improved.

[0005] However, the long sequences and large batches of data in knowledge tracking tasks need to be processed, such as other state-related features appearing in the sequence, complex factors affecting the tracking process, etc. In addition, due to the existence of the black box model of deep learning technology itself, it is still challenging to achieve reliable, high-performance, and explainable knowledge tracking tasks.

[0006] The patent application is a knowledge tracking method based on the item response principle and GRU-Attention (CN118627615A). First, the gated recurrent unit neural network is used to extract features from the concept summary vector in the information reading and writing module, which contains both the student's mastery level of the exercises and the difficulty level of the exercises, to capture long-term dependencies. Then, the attention mechanism is used to assign weights to the extracted feature information to obtain the final knowledge concept feature vector. The neural network is used to infer the project difficulty and the student's ability characteristics, and the correctness of the answer is predicted through the item response principle. This application uses a knowledge tracking method based on the item response principle and GRU-Attention, but in long sequence scenarios, the model is difficult to avoid the influence of noise and invalid information, resulting in inaccurate allocation of attention weights, and the difficulty characteristics are not further utilized in the model training loss, which is not conducive to improving the performance of the model in dynamically capturing the student's status. Summary of the invention

[0007] In order to overcome the problems of the above-mentioned prior art, the purpose of the present invention is to propose a knowledge tracking method, system, device and medium based on attention mechanism and hypergraph convolutional network, combining feature enhancement module FAM, feature fusion module FEMFFM, history interaction module HIM and difficulty constraint module DCM, using hypergraph convolutional network HGNN as encoder, constructing many-to-many relationship between questions and skills through hypergraph, and using hypergraph convolution to enhance question features, using channel attention mechanism to weight different behavior features and fuse them to obtain new feature vectors, using history interaction module HIM to review students' historical practice records and associate historical practice records at different time steps to capture long sequence dependencies, using fuzzification module to convert difficulty coefficient features into membership feature vectors, using difficulty features to add penalty terms to loss function BCELoss to improve model robustness; it has the advantages of high accuracy and good robustness of knowledge tracking tasks.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] A knowledge tracking method based on attention mechanism and hypergraph convolutional network. First, the online recorded data of the students to be trained is obtained and preprocessed. Secondly, the many-to-many relationship between questions and skills is modeled using the hypergraph convolutional network module, and the question embedding that integrates the complex correlation information between questions and skills is constructed. In addition, the different behavioral characteristics of students after preprocessing are fused through the channel attention mechanism, and the fused behavioral characteristics are used to improve the tracking accuracy. The sparse multi-head self-attention mechanism is used to capture the long sequence dependency of the students' online recorded data sequence, and the students' historical practice records are used as an auxiliary to provide support for state updates. Finally, the question difficulty feature and fuzzification module are introduced, and the difficulty feature after fuzzification is used to assist training to improve model performance. At the same time, a difficulty penalty term is added to the loss function, and the learning intensity of the model for questions of different difficulty levels is adjusted to improve robustness, and the optimized overall model is obtained. Finally, a knowledge tracking method based on attention mechanism and hypergraph convolutional network is realized.

[0010] A knowledge tracking method based on attention mechanism and hypergraph convolutional network specifically includes the following steps:

[0011] Step 1, obtain the online record data of the students to be trained: use public network resources to collect the student practice record data set published on the Internet, perform preprocessing operations on the student practice record data, including invalid field cleaning, feature dimension screening, feature value conversion, and sequence data construction, and obtain the student practice sequence data set, which includes the question features, score features, and difficulty features required for the practice, and constructs state features to represent the different behavioral characteristics of students during the practice process;

[0012] Step 2: The dataset obtained by preprocessing in step 1 is used as input for training, and the question features are enhanced by the feature enhancement module FAM, including: using a hypergraph convolutional network to model the many-to-many relationship between questions and skills, constructing a hypergraph, and constructing a question embedding that integrates the complex correlation information between questions and skills through the hypergraph convolutional network, thereby enriching the information contained in the embedding;

[0013] Step 3: Use the channel attention mechanism to assign weights to and fuse the different behavioral features of students preprocessed in step 1; use the sparse multi-head self-attention mechanism SMHSA to associate the historical practice records of students at different time steps, capture the long sequence dependency of the students' online record data sequence, update the student status, and improve the model performance;

[0014] Step 4: By introducing question difficulty features and fuzzification modules, the fuzzification module is used to process the difficulty features to assist training and improve model performance. A difficulty penalty term is added to the loss function BCELoss, and the model's learning intensity for questions of different difficulty levels is adjusted, so that the model tends to have a lower accuracy rate for high-difficulty questions, while conforming to the actual relationship between difficulty and accuracy. The robustness of the model is improved by integrating question difficulty into the training process to assist training.

[0015] The step 2 comprises:

[0016] Step 2.1 Use the hypergraph convolutional network to model the many-to-many relationship between questions and skills and construct a hypergraph:

[0017] First, use the question-skill matrix qs_table in the preprocessed data set in step 1 to calculate the Euclidean distance and obtain the distance matrix dis_mat between questions:

[0018]

[0019] A=diag(X·X T ),

[0020] B=X·X T ,

[0021] in, represents the question-skill matrix, N represents the number of questions, and D represents the characteristic dimension of each question;

[0022] Use the k-nearest neighbor algorithm to select the k nearest neighbors of each node in the hypergraph through the distance matrix to construct the vertex-edge matrix H, and then calculate the association matrix G of the hypergraph through the nearest neighbor constructed vertex-edge matrix H:

[0023]

[0024] The hypergraph is defined as G = (V, E, W), which includes the vertex set V and the hyperedge set E, which can be represented by the incidence matrix H; D is the feature dimension, x id and x jd Respectively represent the values ​​of point i and point j in the dth dimension, d ij represents the calculated Euclidean distance between each training vector, DV is the node degree vector, DE is the hyperedge degree vector, and H is the vertex-edge matrix of the hypergraph;

[0025] Step 2.2 uses the hypergraph convolutional network HGNN as the encoder, and constructs the question embedding that contains the complex correlation information between questions and skills through the linear layer and the nonlinear activation layer ReLU combined with the hypergraph G:

[0026]

[0027] Among them, W is the hyperedge weight matrix, and the generated question' represents the question embedding vector obtained by the encoder.

[0028] The specific method of step 3 is:

[0029] Step 3.1: Concatenate the different student behavior feature vectors obtained through preprocessing in step 1 in the channel dimension through Concat, input them into the channel attention mechanism CAM, and assign weights to different student behavior feature vectors:

[0030]

[0031] Among them, Concat(·) represents the feature concatenation of the channel dimension, CAM(·) represents the channel attention mechanism, and B h ,B w ,B c Represent different behavioral feature vectors, B f is the high-level feature vector after the channel attention mechanism;

[0032] The student's behavior feature vector obtained by Concat splicing passes through the global average pooling layer and one-dimensional convolution layer of the channel attention mechanism CAM to calculate the importance weight of each channel. Finally, the behavior feature vector is obtained after weighting. The adaptive convolution kernel k is used for convolution in the one-dimensional convolution layer, and the receptive field can adapt to different numbers of channels:

[0033]

[0034] Among them, GAP(·) represents global average pooling, Conv1D(·) represents one-dimensional convolution with adaptive convolution kernel k;

[0035] Step 3.2 uses the sparse multi-head self-attention mechanism SMHSA to associate students' historical practice records at different time steps, obtain the long-term sequence dependency of students' online record data sequences, update the current student state features, and obtain a more accurate state feature representation.

[0036] The specific method of step 3.2 is as follows: the question embedding vector question' passed through the hypergraph convolutional network module in step 2.2 is constructed into a key vector K (Key) and a query vector Q (Query) through a sparse multi-head self-attention mechanism SMHSA, and the embedding vector cognition representing the student status characteristics in training is constructed into a value vector V (Value), and a multi-head self-attention operation is performed:

[0037]

[0038] Among them, Softmax(·) represents normalization, SMHSA(·) represents sparse multi-head self-attention mechanism, cognition' represents the latest student state feature vector after sparse multi-head self-attention mechanism; I represents attention distribution, I t,i represents the attention weight of the t-th query on the i-th key, represents the scaling factor, v i Represents interaction x i The corresponding value vector.

[0039] The sparse multi-head self-attention mechanism SMHSA in step 3.2 uses the k-sparse sparse attention mechanism to select the top k interactions with the largest attention scores in I as the most relevant historical interactions. These interactions are used as valid input auxiliary state updates when updating the state. The remaining interactions are regarded as irrelevant or minor information and ignored. The k-sparse sparse attention mechanism adopts top-K sparse attention:

[0040]

[0041] Among them, I i Represents interaction x i The obtained attention score s is the kth largest attention score in the attention score I.

[0042] The specific method of step 4 is:

[0043] Step 4.1 The difficulty feature of the question obtained through preprocessing in step 1 is passed through the fuzzification module to obtain the fuzzy membership feature vector, and the continuous fraction feature in the student exercise sequence is processed into a fuzzy fraction (expressed as ) Gaussian fuzzy logic system is used as the membership function to describe the fuzzy score:

[0044]

[0045] here, represents the continued fraction input, Represented by the fuzzification module Fuzzy members of

[0046] Step 4.2 introduces the difficulty feature obtained through preprocessing in step 1 into the loss function BCELoss to add a difficulty penalty term, adjust the learning intensity of the model for questions of different difficulty levels, and make the prediction consistent with the relationship between actual difficulty and accuracy:

[0047] Loss=base_loss+λ·penalty, (8)

[0048]

[0049] Among them, base_loss is the basic loss of model prediction, penalty is the penalty loss using the difficulty coefficient feature, and λ is the penalty coefficient, which is a hyperparameter.

[0050] A system for knowledge tracking method based on attention mechanism and hypergraph convolutional network, comprising:

[0051] Feature enhancement module FAM, used in step 2, constructs a hypergraph G representing the many-to-many relationship between questions and skills, and uses the hypergraph convolutional network HGNN as an encoder to model the high-order relationship of each question feature, thereby obtaining a question embedding that incorporates the complex correlation information between questions and skills;

[0052] The feature fusion module FEM and the historical interaction module HIM are used in step 3. In the feature fusion module FEM, the channel attention mechanism is used to assign weights to different learning behavior features of students, prioritize different behavior features, and fuse them to obtain a new behavior feature vector. In the historical interaction module HIM, the sparse multi-head self-attention mechanism SMHSA is used to review students' historical practice records, associate historical practice records at different time steps, capture long sequence dependencies, and update the state more accurately.

[0053] The difficulty constraint module DCM is used in step 4 to introduce the fuzzification module. The difficulty coefficient features preprocessed in step 1 are obtained through the fuzzification module to obtain the fuzzy membership feature vector for feature enhancement. By introducing the difficulty feature, a penalty term is added to the loss function BCELoss, and the learning intensity of the model for questions of different difficulty levels is adjusted to improve the robustness of the model.

[0054] An electronic device for a knowledge tracking method based on an attention mechanism and a hypergraph convolutional network, comprising:

[0055] Memory for storing computer programs;

[0056] A processor is used to implement the knowledge tracking method based on the attention mechanism and the hypergraph convolutional network described in any one of steps 1 to 4 when executing the computer program.

[0057] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the computer program can perform knowledge tracking based on the method described in any one of steps 1 to 4.

[0058] Compared with the prior art, the present invention has the following beneficial effects:

[0059] 1. Effective solution to the long sequence dependency problem: The present invention first focuses on the long sequence dependency problem existing in the RNN and LSTM architectures used in the current traditional deep learning-based knowledge tracking. The present invention adopts a multi-head self-attention mechanism to enhance the ability to capture multiple dependencies at different time steps in the long sequence of the model, and adopts a sparse attention mechanism to effectively filter out noise information and irrelevant dependencies, focusing on key related dependencies, so that the model can efficiently process long sequences and avoid the problem of long sequence dependencies not being able to be captured.

[0060] 2. Application of feature enhancement module: The present invention adopts channel attention mechanism, hypergraph convolutional network module and fuzzification module to enhance the feature embedding to be used in the knowledge tracking process and incorporate more information into the constructed embedding. The channel attention mechanism assigns weights to different behavioral features and fuses them, the hypergraph convolutional network module associates the relationship between questions and skills and integrates the associated information into the question embedding vector through the convolutional network, and the fuzzification module converts the difficulty feature into the difficulty affiliation feature, which has fuzziness and can accommodate more numerical errors. This enables the present invention to obtain more information in the features.

[0061] 3. A more reasonable loss function for adapting difficulty: The present invention introduces difficulty features into the loss function, and adopts a difficulty penalty term combined with the loss function BCELoss to adjust the model's learning intensity for questions of different difficulty levels, so that the model tends to have a lower accuracy rate for high-difficulty questions, while conforming to the actual relationship between difficulty and accuracy. The robustness of the model is improved by integrating the difficulty of the questions into the training process to assist in training.

[0062] 4. Potential for wide application: In view of the good robustness and generalization ability of the present invention, it has wide application potential, including smart education, data mining, performance prediction, knowledge tracking, big data analysis and other fields, and has broad application prospects.

[0063] The present invention focuses on optimizing sequence features and strengthening state tracking capabilities. First, the hypergraph convolutional network module is used to obtain enhanced question features, the channel attention mechanism is used to fuse behavioral features, and the fuzzification module is used to fuzzify difficulty features, so that the features in the sequence can contain more contextual information and are highly relevant to the knowledge tracking task; for state tracking capabilities, a sparse multi-head self-attention mechanism and a loss function that introduces difficulty are used for optimization, which strengthens the ability to capture long sequence historical dependencies while eliminating the influence of noise and irrelevant information, and can adapt the difficulty of the questions to prevent overfitting in extreme cases, and finally achieve high robustness of the model.

[0064] In general, the present invention combines the feature enhancement module FAM, the feature fusion module FEM, the history interaction module HIM and the difficulty constraint module DCM to accurately capture the state using effective sequence features, and can show higher accuracy and robustness when facing complex long sequence data. This method is not only universal in theory, but also shows significant performance advantages in practical applications.

[0065] In summary, the present invention shows universality in different data sets and does not require personalized adjustments for each data set. At the same time, it can better represent sequence characteristics, improve state tracking capabilities in long sequence and complex data scenarios, and achieve high-precision knowledge tracking prediction and classification results, providing efficient and reliable solutions for applications in multiple fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 Flowchart of the knowledge tracking method based on attention mechanism and hypergraph convolutional network.

[0067] Figure 2 This is the overall architecture diagram of the knowledge tracking model based on the attention mechanism and hypergraph convolutional network.

[0068] Figure 3 This is the framework diagram of the feature enhancement module FAM.

[0069] Figure 4 This is the FEM framework diagram of the feature fusion module.

[0070] Figure 5 This is the framework diagram of the historical interaction module HIM. DETAILED DESCRIPTION

[0071] This example presents a knowledge tracking method based on attention mechanism and hypergraph convolutional network. The process is shown in Figure 1 The specific model structure framework is as follows: Figure 2 As shown:

[0072] Step 1, obtain the online record data of the students to be trained: use public network resources to collect the student practice record data set published on the Internet, perform preprocessing operations on the student practice record data, including invalid field cleaning, feature dimension screening, feature value conversion, and sequence data construction, and obtain the student practice sequence data set, which includes the question features, score features, and difficulty features required for the practice, and constructs state features to represent the different behavioral characteristics of students during the practice process;

[0073] Step 2: Use the dataset obtained from the preprocessing in step 1 as input for training, and enhance the question features through the feature enhancement module FAM, including: using a hypergraph convolutional network to model the many-to-many relationship between questions and skills, constructing a hypergraph, and constructing a question embedding that integrates the complex correlation information between questions and skills through the hypergraph convolutional network to enrich the information contained in the embedding. Figure 3 As shown, the specific method is:

[0074] Step 2.1 Use the hypergraph convolutional network to model the many-to-many relationship between questions and skills and construct a hypergraph:

[0075] First, use the question-skill matrix qs_table in the preprocessed data set in step 1 to calculate the Euclidean distance to obtain the distance matrix dis_mat between questions:

[0076]

[0077] A=diag(X·X T ),

[0078] B=X·X T ,

[0079] in, represents the question-skill matrix, N represents the number of questions, and D represents the characteristic dimension of each question.

[0080] Use the k-nearest neighbor algorithm to select the k nearest neighbors of each node in the hypergraph through the distance matrix to construct the vertex-edge matrix H, and then calculate the association matrix G of the hypergraph through the nearest neighbor constructed vertex-edge matrix H:

[0081]

[0082] The hypergraph is defined as G = (V, E, W), which includes the vertex set V and the hyperedge set E, which can be represented by the incidence matrix H; D is the feature dimension, x id and x jd Respectively represent the values ​​of point i and point j in the dth dimension, d ij represents the calculated Euclidean distance between each training vector, DV is the node degree vector, DE is the hyperedge degree vector, and H is the vertex-edge matrix of the hypergraph;

[0083] Step 2.2 uses the hypergraph convolutional network HGNN as the encoder, and constructs the question embedding that contains the complex correlation information between questions and skills through two linear layers and one nonlinear activation layer ReLU combined with the hypergraph G:

[0084]

[0085] Among them, W is the hyperedge weight matrix, and the generated question' represents the question embedding vector obtained by the encoder.

[0086] Step 3: Take the training result of step 2 as input, and use the channel attention mechanism to process the different behavioral characteristics of students, assign weights to different feature vectors and fuse them to assist training, such as Figure 4 As shown in the feature fusion module FEM, a sparse multi-head self-attention mechanism is used to associate students' historical practice records at different time steps, capture long sequence dependencies and update student status more accurately, thus improving model performance. Figure 5 The historical interaction module HIM is shown.

[0087] The specific method is:

[0088] Step 3.1 The behavior feature fusion module FEM concatenates the student behavior feature vectors obtained through preprocessing in step 1 in the channel dimension through Concat, inputs them into the channel attention mechanism CAM, and assigns weights to different student behavior feature vectors:

[0089]

[0090] Among them, Concat(·) represents the feature concatenation of the channel dimension, CAM(·) represents the channel attention mechanism, and B h ,B w ,B c Represent different behavioral feature vectors, B f is the high-level feature vector after the channel attention mechanism;

[0091] The concatenated features are passed through the global average pooling layer and one-dimensional convolution layer of the channel attention mechanism CAM in the behavior feature vector fusion module to calculate the importance weight of each channel. Finally, the feature vector is obtained after weighting. The adaptive convolution kernel k is used for convolution in the one-dimensional convolution layer, and the receptive field can adapt to different numbers of channels:

[0092]

[0093] Here, GAP(·) represents global average pooling, and Conv1D(·) represents one-dimensional convolution with adaptive kernel size k.

[0094] Step 3.2 obtains the long-term dependency of the sequence through the sparse multi-head self-attention mechanism SMHSA, updates the current student status by associating the historical practice records of students at different time steps, obtains the long-term dependency of the sequence of the student's online record data sequence, updates the current student status features, and obtains a more accurate status feature representation:

[0095] The question embedding vector question' passed through the hypergraph convolutional network module in step 2.2 is constructed into a key vector K (Key) and a query vector Q (Query) through the sparse multi-head self-attention mechanism SMHSA, and the embedding vector cognition representing the student status in training is constructed into a value vector V (Value), and the multi-head self-attention operation is performed:

[0096]

[0097]

[0098] Among them, Softmax(·) represents normalization, SMHSA(·) represents sparse multi-head self-attention mechanism, cognition' represents the latest student state feature vector after sparse multi-head self-attention mechanism; I represents attention distribution, I t,i represents the attention weight of the t-th query on the i-th key, represents the scaling factor, v i Represents interaction x i The corresponding value vector;

[0099] The sparse multi-head self-attention mechanism SMHSA in step 3.2 uses the k-sparse sparse attention mechanism to select the interactions with the top k largest attention scores in I as the most relevant historical interactions. These interactions are used as valid input auxiliary state updates when updating the state. The remaining interactions are regarded as irrelevant or minor information and are ignored. The k-sparse sparse attention mechanism adopts top-K sparse attention:

[0100]

[0101] Among them, I i Represents interaction x i The obtained attention score s is the kth largest attention score in the attention score I.

[0102] Step 4: By introducing the difficulty feature of the question and the fuzzification module, the difficulty feature enhanced by fuzzification processing is used to assist training to improve model performance. A difficulty penalty term is added to the loss function BCELoss, and the learning intensity of the model for questions of different difficulty levels is adjusted, so that the model tends to have a lower accuracy rate for high-difficulty questions, while conforming to the actual relationship between difficulty and accuracy. The robustness of the model is improved by integrating the difficulty of the question into the training process to assist training. The specific method is:

[0103] Step 4.1 The difficulty feature of the question obtained through preprocessing in step 1 is passed through the fuzzification module to obtain the fuzzy membership feature vector, and the continuous fraction feature in the student exercise sequence is processed into a fuzzy fraction (expressed as ) Gaussian fuzzy logic system is used as the membership function to describe the fuzzy score:

[0104]

[0105] here, represents the continued fraction input, Represented by the fuzzification module Fuzzy members of

[0106] Step 4.2 introduces the difficulty feature obtained through preprocessing in step 1 into the loss function BCELoss to add a difficulty penalty term, adjust the learning intensity of the model for questions of different difficulty levels, and make the prediction consistent with the relationship between actual difficulty and accuracy:

[0107]

[0108] Among them, base_loss is the basic loss of model prediction, penalty is the penalty loss using the difficulty coefficient feature, and λ is the penalty coefficient, which is a hyperparameter.

[0109] A system for knowledge tracking method based on attention mechanism and hypergraph convolutional network, comprising:

[0110] Feature Enhancement Module FAM, used in step 2, constructs a hypergraph G representing the many-to-many relationship between questions and skills, uses the hypergraph convolutional network HGNN as an encoder to model the high-order relationship of each question feature, obtains the question embedding that integrates the complex correlation information between questions and skills, and enhances the question embedding vector;

[0111] The feature fusion module FEM and the historical interaction module HIM are used in step 3. In the feature fusion module FEM, weights are assigned to different learning behavior features of students through the channel attention mechanism, and the priorities of different behavior features are divided. The new behavior feature vector is obtained by fusion. The mutual interference and influence between behaviors are eliminated on the premise of distinguishing the priorities of different behaviors, and the model accuracy is enhanced by using new features. In the historical interaction module HIM, the sparse multi-head self-attention mechanism SMHSA is used to review the historical practice records of students, associate the historical practice records of different time steps, capture long sequence dependencies, and combine the multi-head self-attention and sparse attention mechanisms to eliminate the influence of noise and irrelevant interactions when obtaining attention weights, so as to update the state more accurately.

[0112] The difficulty constraint module DCM is used in step 4. The fuzzification module is introduced to obtain the fuzzy membership feature vector through the fuzzification module for feature enhancement. By introducing the difficulty feature, a penalty term is added to the loss function BCELoss, and the learning intensity of the model for questions of different difficulty levels is adjusted to improve the model robustness. The fuzzification module can improve the model's tolerance to numerical errors, reduce the interference of noise on the prediction results, and further improve the model performance in combination with the updated loss function.

[0113] The present invention introduces an innovative knowledge tracing method, which improves the accuracy and robustness of the knowledge tracing task by strengthening the knowledge tracing method guided by feature vector construction and fusion and attention mechanism.

[0114] Through sufficient ablation experiments, the experimental and performance advantages are verified: the effectiveness of the feature enhancement module FAM, the feature fusion module FEM, the historical interaction module HIM, and the difficulty constraint module DCM are proved, which strengthens the feasibility and performance of the invention. Through comparative experiments with existing LPKT-based methods, DKT-based methods, AKT-based methods, and DKVMN-based methods, the model hyperparameters are tuned through experimental analysis. Using three large-scale public datasets, assist2009, assist2017, and assist2012; the performance advantages on the assist2009, assist2012, and assist2017 datasets are demonstrated, which means that this model achieves better results in knowledge tracking tasks.

[0115] The processed data set is passed into the knowledge tracking model based on the attention mechanism and hypergraph convolutional network. In order to track the student status more accurately, after preprocessing the data, the hypergraph convolutional network is used to obtain enhanced question features, the channel attention mechanism is used to fuse the behavioral features, and the fuzzification module is used to fuzzy the difficulty features, so that the features in the sequence can contain more contextual information and are highly relevant to the knowledge tracking task; for the state tracking ability, the sparse multi-head self-attention mechanism and the loss function of the difficulty are used for optimization, which strengthens the ability to capture long sequence historical dependencies while eliminating the influence of noise and irrelevant information, and can adapt the difficulty of the questions to prevent overfitting in extreme cases, and finally achieve high robustness of the model.

[0116] The operating system used in the experiment is Windows 11, and the deep learning framework used is PyTorch. The specific configurations involved in the experiment are shown in Table 1.

[0117] Table 1 Experimental configuration table

[0118]

[0119] Through sufficient ablation experiments, the effectiveness of the feature enhancement module FAM, feature fusion module FEM, historical interaction module HIM, and difficulty constraint module DCM of the present invention is proved. This strengthens the feasibility and performance of the present invention. Through comparative experiments with existing knowledge tracking models such as DKT, LPKT, AKT and DKVMN, the present invention demonstrates performance advantages on three different data sets. This means that the model of this paper has achieved better results in action classification. Table 2 shows the knowledge tracking performance of the model of this paper relative to other models on public data sets.

[0120] Table 2 Model comparison experimental results

[0121]

[0122] Table 3 lists the specific impact of each module added in step 2, step 3 and step 4 on the model performance. Through comprehensive ablation experimental analysis, the loss function integrating channel attention mechanism, sparse multi-head self-attention mechanism, hypergraph convolutional network, fuzzification module and difficulty improvement all achieved the best performance in terms of two key indicators, AUC and ACC, which proves the effectiveness and necessity of these modules in improving model performance. In this section, the assist2009 public dataset is selected for demonstration.

[0123] The channel attention mechanism of the present invention allocates the importance of different behaviors of students and performs feature fusion to obtain high-level feature vectors; the sparse multi-head self-attention mechanism is combined to improve the accuracy of the model's state update, and the AUC is improved by 1.74% and the ACC is improved by 1.43% compared with the performance of the basic model. The hypergraph convolutional network constructs the embedding of questions containing complex relationships between questions and skills, and enhances the expression ability of the embedding vector; using the difficulty feature, the enhanced fuzzy membership feature vector is obtained through the fuzzification module, and a penalty term is added to the loss function BCELoss in the difficulty improvement loss function to adjust the model's learning intensity for questions of different difficulty levels. Compared with the performance of the basic model, the AUC is improved by 2.94% and the ACC is improved by 1.67%.

[0124] Finally, the knowledge tracking model based on the attention mechanism and hypergraph convolutional network of the present invention has a good performance compared with the basic model, with AUC improved by 3.06% and ACC improved by 2.00%.

[0125] Table 3 Ablation experiment results in the knowledge tracking module

[0126]

Claims

1. A knowledge tracking method based on attention mechanism and hypergraph convolutional network, characterized in that: First, the online recorded data of the students to be trained is obtained and preprocessed; secondly, the hypergraph convolutional network module is used to model the many-to-many relationship between questions and skills, and to construct a question embedding that integrates the complex correlation information between questions and skills; in addition, the different behavioral characteristics of students after preprocessing are fused through the channel attention mechanism, and the fused behavioral characteristics are used to improve the tracking accuracy, and the sparse multi-head self-attention mechanism is used to capture the long sequence dependency of the students' online recorded data sequence, and the students' historical practice records are used as an auxiliary to provide support for state updates; finally, the question difficulty feature and fuzzification module are introduced, and the difficulty feature after fuzzification is used to assist training to improve model performance. At the same time, a difficulty penalty term is added to the loss function, and the learning intensity of the model for questions of different difficulty levels is adjusted to improve robustness, and the optimized overall model is obtained; finally, a knowledge tracking method based on attention mechanism and hypergraph convolutional network is realized.

2. The knowledge tracking method based on attention mechanism and hypergraph convolutional network according to claim 1, characterized in that: The specific steps include: Step 1, obtain the online record data of the students to be trained: use public network resources to collect the student practice record data set published on the Internet, perform preprocessing operations on the student practice record data, including invalid field cleaning, feature dimension screening, feature value conversion, and sequence data construction, and obtain the student practice sequence data set, which includes the question features, score features, and difficulty features required for the practice, and constructs state features to represent the different behavioral characteristics of students during the practice process; Step 2: The dataset obtained by preprocessing in step 1 is used as input for training, and the question features are enhanced by the feature enhancement module FAM, including: using a hypergraph convolutional network to model the many-to-many relationship between questions and skills, constructing a hypergraph, and constructing a question embedding that integrates the complex correlation information between questions and skills through the hypergraph convolutional network, thereby enriching the information contained in the embedding; Step 3: Use the channel attention mechanism to assign weights to and fuse the different behavioral features of students preprocessed in step 1; use the sparse multi-head self-attention mechanism SMHSA to associate the historical practice records of students at different time steps, capture the long sequence dependency of the students' online record data sequence, update the student status, and improve the model performance; Step 4: By introducing question difficulty features and fuzzification modules, the fuzzification module is used to process the difficulty features to assist training and improve model performance. A difficulty penalty term is added to the loss function BCELoss, and the model's learning intensity for questions of different difficulty levels is adjusted, so that the model tends to have a lower accuracy rate for high-difficulty questions, while conforming to the actual relationship between difficulty and accuracy. The robustness of the model is improved by integrating question difficulty into the training process to assist training.

3. The knowledge tracking method based on attention mechanism and hypergraph convolutional network according to claim 2 is characterized in that: The step 2 comprises: Step 2.1 Use the hypergraph convolutional network to model the many-to-many relationship between questions and skills and construct a hypergraph: First, use the question-skill matrix qs_table in the preprocessed data set in step 1 to calculate the Euclidean distance and obtain the distance matrix dis_mat between questions: A=diag(X·X T ), B=X·X T , in, represents the question-skill matrix, N represents the number of questions, and D represents the characteristic dimension of each question; Use the k-nearest neighbor algorithm to select the k nearest neighbors of each node in the hypergraph through the distance matrix to construct the vertex-edge matrix H, and then calculate the association matrix G of the hypergraph through the nearest neighbor constructed vertex-edge matrix H: The hypergraph is defined as G = (V, E, W), which includes the vertex set V and the hyperedge set E, which can be represented by the incidence matrix H; D is the feature dimension, x id and x jd Respectively represent the values ​​of point i and point j in the dth dimension, d ij represents the calculated Euclidean distance between each training vector, DV is the node degree vector, DE is the hyperedge degree vector, and H is the vertex-edge matrix of the hypergraph; Step 2.2 uses the hypergraph convolutional network HGNN as the encoder, and constructs the question embedding that contains the complex correlation information between questions and skills through the linear layer and the nonlinear activation layer ReLU combined with the hypergraph G: Among them, W is the hyperedge weight matrix, and the generated question′ represents the question embedding vector obtained by the encoder.

4. The knowledge tracking method based on attention mechanism and hypergraph convolutional network according to claim 2, characterized in that: The specific method of step 3 is: Step 3.1: Concatenate the different student behavior feature vectors obtained through preprocessing in step 1 in the channel dimension through Concat, input them into the channel attention mechanism CAM, and assign weights to different student behavior feature vectors: Among them, Concat(·) represents the feature concatenation of the channel dimension, CAM(·) represents the channel attention mechanism, and B h , B w , B c Represent different behavioral feature vectors, B f is the high-level feature vector after the channel attention mechanism; The student's behavior feature vector obtained by Concat splicing passes through the global average pooling layer and one-dimensional convolution layer of the channel attention mechanism CAM to calculate the importance weight of each channel. Finally, the behavior feature vector is obtained after weighting. The adaptive convolution kernel k is used for convolution in the one-dimensional convolution layer, and the receptive field can adapt to different numbers of channels: Among them, GAP(·) represents global average pooling, Conv1D(·) represents one-dimensional convolution with adaptive convolution kernel k; Step 3.2 uses the sparse multi-head self-attention mechanism SMHSA to associate students' historical practice records at different time steps, obtain the long-term sequence dependency of students' online record data sequences, update the current student state features, and obtain a more accurate state feature representation.

5. The knowledge tracking method based on attention mechanism and hypergraph convolutional network according to claim 4, characterized in that: The specific method of step 3.2 is as follows: the question embedding vector question' passed through the hypergraph convolutional network module in step 2.2 is constructed into a key vector K (Key) and a query vector Q (Query) through a sparse multi-head self-attention mechanism SMHSA, and the embedding vector cognition representing the student status characteristics in training is constructed into a value vector V (Value), and a multi-head self-attention operation is performed: Among them, Softmax(·) represents normalization, SMHSA(·) represents sparse multi-head self-attention mechanism, cognition' represents the latest student state feature vector after sparse multi-head self-attention mechanism; I represents attention distribution, I t,i represents the attention weight of the t-th query on the i-th key, represents the scaling factor, v i Represents interaction x i The corresponding value vector.

6. The knowledge tracking method based on attention mechanism and hypergraph convolutional network according to claim 4, characterized in that: The sparse multi-head self-attention mechanism SMHSA in step 3.2 uses the k-sparse sparse attention mechanism to select the top k interactions with the largest attention scores in I as the most relevant historical interactions. These interactions are used as valid input auxiliary state updates when updating the state. The remaining interactions are regarded as irrelevant or minor information and ignored. The k-sparse sparse attention mechanism adopts top-K sparse attention: Among them, I i Represents interaction x i The obtained attention score s is the kth largest attention score in the attention score I.

7. The knowledge tracking method based on attention mechanism and hypergraph convolutional network according to claim 2, characterized in that: The specific method of step 4 is: Step 4.1 The difficulty feature of the question obtained through preprocessing in step 1 is passed through the fuzzification module to obtain the fuzzy membership feature vector, and the continuous fraction feature in the student exercise sequence is processed into a fuzzy fraction (expressed as ) Gaussian fuzzy logic system is used as the membership function to describe the fuzzy score: here, represents the continued fraction input, Represented by the fuzzification module Fuzzy members of Step 4.2 introduces the difficulty feature obtained through preprocessing in step 1 into the loss function BCELoss to add a difficulty penalty term, adjust the learning intensity of the model for questions of different difficulty levels, and make the prediction consistent with the relationship between actual difficulty and accuracy: Among them, base_loss is the basic loss of model prediction, penalty is the penalty loss using the difficulty coefficient feature, and λ is the penalty coefficient, which is a hyperparameter.

8. A system based on the attention mechanism and the knowledge tracking method of the hypergraph convolutional network according to any one of claims 2 to 7, characterized in that: include: Feature enhancement module FAM, used in step 2, constructs a hypergraph G representing the many-to-many relationship between questions and skills, and uses the hypergraph convolutional network HGNN as an encoder to model the high-order relationship of each question feature, thereby obtaining a question embedding that incorporates the complex correlation information between questions and skills; The feature fusion module FEM and the historical interaction module HIM are used in step 3. In the feature fusion module FEM, the channel attention mechanism is used to assign weights to different learning behavior features of students, prioritize different behavior features, and fuse them to obtain a new behavior feature vector. In the historical interaction module HIM, the sparse multi-head self-attention mechanism SMHSA is used to review students' historical practice records, associate historical practice records at different time steps, capture long sequence dependencies, and update the state more accurately. The difficulty constraint module DCM is used in step 4 to introduce the fuzzification module. The difficulty coefficient features preprocessed in step 1 are obtained through the fuzzification module to obtain the fuzzy membership feature vector for feature enhancement. By introducing the difficulty feature, a penalty term is added to the loss function BCELoss, and the learning intensity of the model for questions of different difficulty levels is adjusted to improve the robustness of the model.

9. An electronic device based on the attention mechanism and the knowledge tracking method of the hypergraph convolutional network according to any one of claims 1 to 7, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the knowledge tracking method based on the attention mechanism and the hypergraph convolutional network as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it is capable of performing knowledge tracking based on the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Knowledge tracking method based on project reaction principle and GRU-Attention

    CN118627615A

Cited By

  • Intelligent monitoring image recognition processing system for community emergency disposal

    CN121982656A

  • Intelligent monitoring image recognition processing system for community emergency disposal

    CN121982656B