A Hierarchical Attention Cognitive Diagnosis Method Based on the Mixture of Knowledge Association and Abnormal Behavior
By combining graph neural networks and long-term memory networks, a hierarchical attention mechanism is built, which solves the problem of knowledge correlation and timing characteristics neglected in the dynamic cognitive diagnosis model, realizes a comprehensive and consistent diagnosis of learners' cognitive state, and improves prediction accuracy.
Patent Information
- Application Number
- CN202411674133.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The existing dynamic cognitive diagnostic model ignores the learner's historical timing characteristics and knowledge correlation, resulting in inconsistent implicit cognition and learning behavior, making it difficult to fully portray the learner's cognitive process.
A hierarchical attention cognitive diagnosis method based on knowledge association and abnormal behavior is adopted. Through the graph neural network, aggregating test attributes, combining long and short-term memory networks and click flow characteristics, guessing and carelessness gates are set to realize the triple-second-order coupling of knowledge, timing and behavior, and predict the future performance of learners.
The accuracy and consistency of cognitive diagnosis are improved, and the correspondence between learners' implicit cognition and explicit behavior is corrected through multi-task learning methods, which enhances the comprehensiveness and predictive ability of the learning process.
Smart Images

Figure CN119514591B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of educational big data mining, graph neural networks, and student behavior modeling, and particularly relates to a hierarchical attention cognitive diagnosis method based on a mixture of knowledge association and abnormal behavior. Background Art
[0002] Cognitive diagnosis aims to model the data in a single assessment scenario of students. To address the problem of one-sidedness in diagnostic results caused by ignoring learners' historical priors in static cognitive diagnosis, Corbett and Anderson proposed Knowledge Tracing in 1994, which achieved the coupling of knowledge and time series, also known as "dynamic cognitive diagnosis". Since then, dynamic cognitive diagnosis models have not been limited to single learning scenarios but are applicable to multiple consecutive learning scenarios, realizing the expansion of cognitive diagnosis in terms of time granularity and the continuous analysis of the learning process.
[0003] (1) Traditional dynamic cognitive diagnosis models predict learners' future performance by tracking the knowledge fluctuations of learners at the time granularity. Bayesian Knowledge Tracing (BKT) is a classic dynamic cognitive diagnosis model based on probabilistic graphs. It abstracts the knowledge state as a set of binary variables, takes the real-time interactions of learners as input, and simulates the changes in learners' knowledge mastery during the learning process through a hidden Markov model. There are many variants of the BKT model, such as three-parameter BKT, incorporating factors of guessing and slip, time factors, and problem difficulty estimation. Such models have good interpretability.
[0004] (2) The dynamic cognitive diagnosis model based on deep learning was proposed in the context of the rapid development of deep learning technology. It solves the problem of complex parameter estimation in the BKT series of models, and its performance is significantly better than that of probabilistic graphical models. The dynamic cognitive diagnosis model based on deep learning can perform more complex modeling of knowledge states and learning processes, track the evolution process of knowledge through neural networks, and has stronger generalization ability. ① The dynamic cognitive diagnosis model based on recurrent neural networks uses recurrent neural networks to track the knowledge states of learners at different times. For example, DKT is the earliest and most popular dynamic cognitive diagnosis model. The EKT model makes full use of the multi-modal features of test questions to enhance the representation of test questions in the dynamic cognitive diagnosis model. The IEKT model models the knowledge acquisition sensitivity by estimating learners' cognition of test questions and uses reinforcement learning to search for the best parameters. The DIMKT model introduces test question difficulty into the recurrent neural network to explore the impact of test question difficulty on learning performance. ② The dynamic cognitive diagnosis model based on memory networks uses memory networks to store the implicit knowledge states of learners. Each answer will dynamically update the memory network through key-value pairs. DKVMN is a classic model in this category. The KIEDLKD model combines IRT theory to improve the interpretability of the model in predicting knowledge proficiency. The dynamic cognitive diagnosis model based on graph neural networks first proposed the GKT model, which uses graph neural networks to model learners' mastery of knowledge. The further proposed ABKT model introduces ability factors into learning feedback attribution and uses graph neural networks and ensemble learning to model the learning process, etc. The SGKT model uses session graphs to model learners' answering processes and uses relational graphs to simulate the relationship between practice and skills. The performance of the deep dynamic cognitive diagnosis series of models is significantly better than that of traditional probabilistic models and is the current mainstream dynamic cognitive diagnosis models. However, these models assume that the sequential data of learners' answers are equidistant, ignoring the temporal information hidden in historical data and making it difficult to explain the process of their knowledge evolution based on cognitive laws.
[0005] (3) Temporally Enhanced Dynamic Cognitive Diagnosis Model. In multiple consecutive learning scenarios, there are rich temporal features, which are important factors affecting the learner's cognitive process and can enhance the representation of the learning process. According to cognitive processing theory, there are forgetting laws, learning laws, short-term memory laws, etc. in the human brain's memory of knowledge. First of all, the forgetting law is the most common memory law, which describes the phenomenon that an individual's mastery of knowledge decreases over time. Nagatani et al. combined three forgetting characteristics with the representation of knowledge states and introduced forgetting into the learner's learning process. The AKT model, RKT model, and HawkesKT model use exponential decay functions and response time intervals to model the forgetting curve. The DGMN model combines the forgetting gating mechanism with the attention structure to dynamically capture forgetting behavior. Secondly, the learning law reveals that repeated learning of knowledge can improve an individual's mastery of knowledge, and it is often jointly modeled with the forgetting law. For example, EKPT uses learning curves and forgetting curves to simulate the change of concept proficiency over time. The LPKT model takes the temporal features in the learning process as learning and forgetting factors. The KTM-DLF model uses matrix factorization technology to integrate learning, forgetting factors, learner ability, and test question difficulty. The GFLDKT model simulates the explicit forgetting and learning behaviors in the learner's learning process through a gating mechanism. Finally, the short-term memory law highlights the response records in the most recent time period. For example, the MF-DAKT model records the learner's recent attempts at related concepts based on the recency effect to highlight the influence of the most recent responses. The CKT model uses a convolutional neural network to simulate the enhancement of the current knowledge state by the records in the adjacent time period. The MC-DRER model and the HMN model design a hierarchical memory network that fuses working memory and long-term memory to simulate the change process of the learner's memory of knowledge. The temporally enhanced dynamic cognitive diagnosis model, due to following the memory laws in educational psychology, improves the interpretability of deep learning models and enhances the prediction performance to a certain extent. However, such models focus on knowledge and temporal features and are difficult to capture the hidden knowledge correlation fluctuations behind different behaviors, and there are still inconsistencies in their diagnostic results.
[0006] Therefore, the dynamic cognitive diagnosis model can diagnose the knowledge state of learners in the continuous learning process, making up for the defect of cognitive diagnosis that ignores historical temporal features and achieving different degrees of coupling between knowledge and time series. Its essence is the binary coupling of knowledge and time series. However, the current dynamic cognitive diagnosis models either do not consider the association of multiple knowledge or do not consider the influence of answering behavior factors on the learner's knowledge state, which is likely to cause inconsistencies between implicit cognition and learning behavior and weaken the binary coupling of process behavior, resulting in limitations and ultimately the problem of incomplete learning characterization. Summary of the Invention
[0007] To overcome the deficiencies of the above-mentioned existing technologies, the present invention provides a hierarchical attention cognitive diagnosis method based on a mixture of knowledge association and abnormal behavior. The purpose of the present invention is to address the problems of information fragmentation and cognitive inconsistency caused by the neglect of complex knowledge associations in the temporal sequence in existing dynamic cognitive diagnosis models, and to propose a hierarchical attention cognitive diagnosis method based on a mixture of knowledge association and abnormal behavior, so as to achieve dynamic cognitive diagnosis with knowledge association perception in the time series and consistent internal cognition and external behavior performance.
[0008] The purpose of the present invention is achieved through the following technical measures: A hierarchical attention cognitive diagnosis method based on a mixture of knowledge association and abnormal behavior, comprising the following steps:
[0009] Step 1, obtain the learning unit data of the student during the assessment, including the knowledge characteristics of the questions answered, the behavioral characteristics of the student during answering, and the temporal characteristics formed by the answering. Arrange several learning units in a learning process sequence data in chronological order, and thus construct a temporal dataset. Based on the constructed temporal dataset, obtain the knowledge association representation of the questions;
[0010] Step 2, based on the obtained knowledge representation, combine the answering responses of the student to the questions, and construct a long-short-term hierarchical attention network to dynamically update the knowledge state of the student, realizing the first-order coupling of knowledge and temporal characteristics;
[0011] Step 3, based on the click stream behavioral characteristics generated by the learner during the learning process due to clicks, calculate the guessing index and the carelessness index of the learner when doing questions, realizing the first-order coupling of behavior and temporal characteristics;
[0012] Step 4, based on the knowledge state obtained after the coupling of knowledge and time series, combined with the guessing index and the carelessness index obtained after the coupling of behavior and time series, set the guessing gate and the carelessness gate, realizing the second-order coupling of knowledge-time series-behavior, and updating the comprehensive ability state of the learner;
[0013] Step 5, set the anomaly index and the future performance prediction task, where the anomaly index includes the guessing index and the carelessness index, both of which are predicted using a feedforward neural network and a multi-sensory perceptron, and set two hyperparameters to balance the three prediction tasks, and use the anomaly index prediction to enhance the prediction of the learner's final performance.
[0014] Further, the specific content of Step 1 includes:
[0015] (1-1) Embedding and encoding of question attributes;
[0016] Among them, the attributes of the questions include concepts, content, and difficulty;
[0017] (1-2) Construction of the heterogeneous graph of question characteristics;
[0018] First, based on different information of test questions, three heterogeneous nodes of concept, difficulty, and content are set up to construct a heterogeneous graph of test questions;
[0019] Secondly, the edge relationships in the heterogeneous graph are constructed. Given that there are three edge relationships among the three heterogeneous nodes, namely content-concept, content-difficulty, and concept-concept, the first two are obtained through the test question-concept association matrix and the test question-difficulty association matrix respectively, while the edge relationship between concepts is obtained through similarity calculation and threshold setting;
[0020] (1-3) Enhanced representation of test questions by aggregating neighbor node information;
[0021] First, for the same type of nodes of test questions, a bidirectional LSTM is set up to aggregate the same type of nodes, and the average pooling method is used to obtain the aggregated vector of the corresponding type. Its formula can be expressed as:
[0022]
[0023] where n i (v) The set of neighbor nodes of type i of the current node v, emb(ω) represents the embedding representation of the neighbor node w, is the vector after aggregating the neighbor nodes of type i of node v;
[0024] Secondly, for different types of nodes, three different types of vectors are aggregated through the attention mechanism and combined with the embedding of node v itself. The formula is expressed as:
[0025]
[0026] where α (v,v) and α (v,i) represent weight coefficients, is the vector representation of the aggregated node v; thus, the multi-dimensional heterogeneous features of test questions are aggregated through the heterogeneous graph, and the knowledge association representation of test questions is realized.
[0027] Furthermore, the embedding and encoding of test question attributes specifically include:
[0028] First, the concept features of test questions belong to structured data, and one-hot encoding is used for the initialization of embedding;
[0029] Secondly, the difficulty features of test questions adopt one-hot encoding;
[0030] Again, the content nodes of the test questions, including the test question content, belong to unstructured data. For formula information, the ontology replacement technology in the subject field is adopted to maintain the standardization of the content. Then, it is concatenated with the text information, and the BERT model is used to perform a unified semantic representation of the test questions. For picture information, CNN is used for picture recognition and representation. Finally, the embedded representations of the text, formula, and picture are aggregated in a cascaded manner to obtain the initial embedded representation of the test question content.
[0031] Further, step 2 specifically includes:
[0032] First, perform an embedded representation on the triple features of knowledge, time sequence, and behavior, and use the long short-term memory neural network LSTM to encode it to obtain the current knowledge state;
[0033] Then, according to the principle that the higher the student's ability, the greater their working memory is likely to be, use the Rash model to calculate the latent traits of the student, and thus calculate the working memory capacity of the student;
[0034] The memory capacity of the student directly affects their short-term memory. Use the short-term attention mechanism with a triple gate function to enhance the student's current knowledge state;
[0035] Finally, the knowledge in short-term memory needs to be converted into long-term memory to be truly mastered. Therefore, use the long-term attention based on the laws of learning and forgetting to further update their knowledge state.
[0036] Further, the specific implementation of step 2 includes the following steps:
[0037] (2-1) Embedding and encoding of triple features
[0038] First, define a learning unit composed of the triple elements of knowledge, time sequence, and behavior, denoted as LC n =(q n , s n , b n ), which is used to represent the learning event that occurs within a certain time unit n; where, q n represents the knowledge feature under investigation; s n represents the time sequence feature in the time sequence data, including the time unit interval Δs n , the same concept learning interval Δr n and the repetition times Δc n ; b n represents the click stream behavior feature generated by the learner during the learning process; the embedded representation of the learning unit is
[0039] X n =ε v +emb(s n )+emb(r n ),
[0040] Among them, ε v is the knowledge association representation based on the heterogeneous graph neural network, and emb(s n ) is the embedding of the response time. r n ∈b n is the correctness of the last click in the click stream, and emb(r n ) is the embedding of the response result;
[0041] Subsequently, the influence of the previous answer on the subsequent answer is simulated through a recurrent neural network to initially obtain the potential knowledge state. The encoding process can be expressed as:
[0042] H n = LSTM(X n ),
[0043] where LSTM is the long short-term memory neural network, represents the potential knowledge state at the nth question;
[0044] (2-2) Working memory capacity calculation
[0045] First, according to the learner's answering situation, the Rasch model is used to calculate the learner's initial potential trait. The formula is as follows:
[0046] P(r in = 1|θ i ,β n ) = exp(θ i -β n )[1 + exp(θ i -β n )] -1 ,
[0047] where r in is the binary variable of whether the learner answers correctly or wrongly, 1 represents a correct answer, 0 represents a wrong answer, θ i is the potential trait parameter of learner i, and β n is the difficulty parameter of question n;
[0048] Subsequently, the potential trait value obtained by further using the Rasch model is used to determine the size of the learner's working memory capacity. Specifically, the working memory capacity is adjusted using the learner's trait and question difficulty, and the calculation is as follows:
[0049] WMC n = ΔC × σ(α·θ i -β·difff n ) + ∈
[0050] where WMC nThe working memory capacity when answering the nth question, difff n is the difficulty of the nth question, and α and β are weight factors used to adjust the influence of latent traits and question difficulty; σ is the Sigmoid activation function used to limit the output value between 0 and 1; ΔC is the capacity floating range, and ∈ is the minimum value of the capacity to ensure that the capacity is not zero;
[0051] (2 - 3) Short-term attention based on triple masking
[0052] To mask irrelevant or excessive information and enhance short-term dependencies in the time series, a triple mask is set for the attention mechanism, including a short-term mask, a capacity mask, and a temporal mask, to converge the answering effects within a limited capacity and a defined time, thereby simulating the working memory operation mechanism;
[0053] First, the short-term mask M s is used to mask learning units with a long time interval, simulating the characteristic that humans have more active memories of recent events; by calculating the time interval Δs between the current learning unit n and the previous unit; if the time interval exceeds a certain short-term window threshold δ s , the attention weight of this unit is set to 0, thus avoiding over-focus on long-term information;
[0054] Secondly, the capacity mask M w , based on the limitation of the working memory capacity, masks learning units outside the capacity range to ensure that the model only focuses on important information within the current capacity; according to the working memory capacity WMC when answering the nth question obtained in the previous step n , the mask is set so that only the information of the first WMC n learning units is retained, and the weights of learning units outside this capacity range are set to 0;
[0055] Finally, the temporal mask M c is used to obscure subsequent learning units to ensure temporal consistency and prevent the influence of future information on the current state, that is, to avoid information leakage; for the current learning unit n, all learning units after it are masked, making the attention weight of future time steps to the current state 0;
[0056] Combining the triple mask, the formula of the attention mechanism is expressed as:
[0057] H′ n = Attention(H n , M s , M w , M c ),
[0058] where, Represents the aggregated knowledge state after triple masking, and Attention refers to the multi-head self-attention network layer. M s 、M w 、M c are the short-term mask, the capacity mask, and the temporal mask respectively, which are used to regulate the information screening from different sources;
[0059] (2-4) Long-term Attention Based on Learning and Forgetting
[0060] On the basis of enhanced short-term memory, a long-term attention mechanism based on forgetting and learning curves is set up to capture the behavioral associations between long-cycle learning units;
[0061] First of all, the memory decay function F s (Δs n ) is used to simulate the process of human memory gradually weakening over time. We use the polynomial decay model, so that the memory decays rapidly at the beginning and then gradually slows down:
[0062] F s (Δs n )=1 / (1+γ·Δs n ) p
[0063] Among them, Δs n is the sequence time distance; γ is the parameter to adjust the decay speed, the larger the value, the faster the decay; p is the order of the polynomial, which is used to control the smoothness of the decay;
[0064] Secondly, the learning enhancement function F t (Δr n ,Δc n ) is used to describe the enhancement effect of repeated learning on memory, which is expressed as follows:
[0065] F t (Δr n ,Δc n )=(1+α·Δc n )·exp(-β·Δr n )
[0066] Among them, Δr n represents the time interval between the current learning unit and the last learning of the same concept, and Δc n represents the number of repeated learning; α and β are the set hyperparameters, α is the enhancement coefficient of the number of repetitions, which is used to control the memory growth amplitude of each repeated learning; β is the decay rate of the learning interval, which is used to control the influence of the interval time on memory enhancement;
[0067] Finally, after memory decay and learning enhancement, the knowledge state of the learner evolves into:
[0068]
[0069] Among them, H″ n+1 represents the predicted future knowledge state after learning and forgetting adjustment; is the query vector in the attention mechanism, q n+1 is the representation of the test question to be answered at the (n + 1)-th moment; represents the knowledge state of the past learning unit and is used for the key matrix in the attention mechanism; then serves as the value matrix in the attention mechanism; are all weight matrices in the attention mechanism and are used for linear transformation of the query, key, and value vectors.
[0070] Furthermore, the click - stream - based guessing index in step 3 includes:
[0071] First, based on the click - stream data formed when the learner answers a certain question, the accumulative method is used to calculate the number of times of changing answers;
[0072] Second, based on the click - stream data, the counting method is used to calculate the option coverage rate;
[0073] Then, based on the click - stream data, the residence - time ratio of the correct option is obtained by the accumulative method;
[0074] Finally, using these behavioral characteristics in the answering process, the guessing index is derived. The calculation formula of the guessing index is as follows:
[0075]
[0076] Among them, guess ∈ [0, 1] is the guessing index, f c (count) represents the enhancement function of the number of times of changing answers, f v (cover) represents the enhancement function of the option coverage rate, f t (time) is the weakening function representing the residence - time ratio of the correct option; specifically, f c (count) = log(1 + α·count). The more times of replacement, the stronger the learner's uncertainty about the answer and the higher the degree of guessing; f v (cover) = β·cover. The higher the option coverage rate, it means that the learner has tried more options when answering, and may guess without a clear answer; f t (time) = 1 / (1 + γ·time). The higher the residence - time ratio, it indicates that the learner is more confident about the answer and the lower the possibility of guessing; α, β, and γ are all trainable parameters used to adjust these functions.
[0077] Further, the carelessness index based on clickstream in step 3 includes:
[0078] First, calculate the ratio of the total time spent in the clickstream to the standard time, that is, the spending time ratio spend;
[0079] Then, by analyzing the clickstream characteristics and time ratio of the learner, calculate its carelessness index, and the formula calculation is as follows:
[0080] skip = sigmoid(f′ c (count)·f′ v (cover)·f s (spend)) -1
[0081] Among them, skip ∈ [0, 1] is the carelessness index, f s (spend) = 1 / (1 + ζ·spend) is the weakening function of the spending time ratio. The smaller the spending time ratio, it indicates that the learner may be careless due to insufficient time; ζ is also a trainable parameter, f′ c (count) and f′ v (cover) and f c (count) and f v (cover) are calculated in the same way as the enhancement function, but only the trainable parameters are replaced, that is, α′ and β′.
[0082] Further, the specific implementation method of step 4 is as follows:
[0083] First, define a gating function GLU(), which is used to control the influence of guessing and carelessness. The function can be expressed as:
[0084]
[0085] Among them, the GLU function is used to control the influence of abnormal behavior on the knowledge state. The input value X is the student's knowledge state H″ n+1 , W1, W2 are weight matrices, b1, b2 are bias terms, represents the element-wise product;
[0086] Subsequently, if a guessing behavior is detected, reduce the error caused by guessing through the guessing gate, update the knowledge state, and subtract the error caused by guessing from the knowledge mastery level. The formula is:
[0087] K n+1 = H″ n+1 - GLU g (guess n+1 ·W g + b g )
[0088] Among them, K n+1 represents the updated knowledge state after passing through the guess gate, and H″ n+1 is the future knowledge state, and guess n+1 is the guess exponent at time step n + 1, and W g and b g are the weight and bias parameters of the guess gate;
[0089] Finally, if careless behavior is detected, the knowledge loss caused by carelessness is enhanced through the carelessness gate, and the knowledge state is further updated by adding the knowledge enhancement ignored by carelessness to the knowledge mastery level. The formula is:
[0090] θ n+1 = K n+1 + GLU s (skip n+1 ·W s + b s )
[0091] Among them, θ n+1 is the comprehensive ability state of the learner after passing through the carelessness gate, K n+1 is the knowledge state after passing through the guess gate, skip n+1 is the carelessness exponent at time step n + 1, and W s and b s are the weight and bias parameters of the proficiency gate.
[0092] Furthermore, in step 5, first, corresponding pseudo-labels are generated according to the behavior characteristics, which are respectively represented as and Specifically, the "guess" behavior is marked by setting thresholds for the replacement times, option coverage rate, and dwell time ratio, and the samples that meet the conditions are assigned as otherwise 0; similarly, the "carelessness" behavior is marked by combining the time ratio and replacement times to set a threshold, and the samples that meet the conditions are assigned as otherwise 0;
[0093] Secondly, a guess exponent prediction task is set. In the prediction task, the feed-forward network layer is used to transform the comprehensive ability state into the prediction space of the guess exponent, and a multi-layer perceptron is set to complete the prediction, that is
[0094]
[0095] where FNN and MLP are the feed-forward network layer and the multi-layer perceptron;
[0096] To optimize this prediction task, a loss function is defined to measure the difference between the predicted guess exponent of the model and the expected distribution. The loss function is defined as:
[0097]
[0098] Among them, is the predicted guess index, is the target value for training;
[0099] Finally, set the careless index prediction task. In the prediction task, use the feedforward network layer to transform the comprehensive ability state into the prediction space of the careless index, and set up a multi-layer perceptron to complete the prediction, that is
[0100]
[0101] Correspondingly, the loss function is:
[0102]
[0103] Among them, is the predicted careless index, is the target value for training.
[0104] Furthermore, in order to predict the future performance of the learner, the knowledge state of the learner is concatenated with the predicted anomaly index to obtain the current comprehensive state, as follows:
[0105]
[0106] Secondly, set a similar multi-layer perceptron to complete the prediction of the learner's response result, that is
[0107]
[0108] Among them is the predicted correct / incorrect result; its loss function is:
[0109]
[0110] Finally, in order to balance the three prediction tasks, the loss functions are weighted and summed to obtain the global loss function of the multi-task, that is:
[0111]
[0112] Among them is the global loss function, and the hyperparameters λ1 and λ2 are used to adjust the weights of each task to achieve the balance of different tasks.
[0113] Compared with the prior art, the beneficial effects of the present invention are as follows: Existing dynamic cognitive diagnosis is essentially a binary coupling of "time series - knowledge". The binary coupling that weakens the process behavior has limitations and will ultimately lead to the problem of incomplete learning characterization. The present invention introduces knowledge association and abnormal behavior into time series modeling, sets a ternary second-order coupling framework, analyzes the comprehensive ability of students through two-stage coupling of the ternary features of knowledge, time series, and behavior, and uses abnormal behavior analysis and prediction to correct the diagnosis result. This method first uses graph neural network technology to aggregate heterogeneous test question attributes such as test question text, concepts, and difficulty, and obtains knowledge association through neighbor node update; subsequently, a short-term attention mechanism based on working memory capacity is used to aggregate the pre-order knowledge information associated in the time series to enhance the current cognitive state representation of the learner, and a long-term attention based on the laws of learning and forgetting is established to update the weight of the knowledge association sequence to extract long-distance dependence relationships, realizing the first-order coupling of "knowledge - time series"; again, key features reflecting students' guessing and carelessness are extracted from click stream features, and guessing and carelessness functions are set to calculate the guessing index and carelessness index, realizing the first-order coupling of "behavior - time series"; based on the knowledge state and abnormal behavior index obtained in the first two steps, guessing gates and carelessness gates are set to correct the cognitive state of students and obtain the comprehensive ability state of students, realizing the second-order coupling of "knowledge - time series - behavior"; finally, a multi-task learning method is adopted, and an abnormal index prediction task is introduced to correct the prediction of the learner's future answering performance and ensure the consistency diagnosis of implicit cognition and learning behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0114] Figure 1 It is a principle block diagram of the method of the embodiment of the present invention.
[0115] Figure 2 It is a flow block diagram of the method of the embodiment of the present invention.
[0116] Figure 3 It is a structural example diagram of the learning unit in the embodiment of the present invention.
[0117] Figure 4 It is a comparison diagram of the coupling mechanism between the embodiment of the invention and the traditional cognitive diagnosis method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0118] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0119] A hierarchical attention cognitive diagnosis method based on a mixture of knowledge association and abnormal behavior, the principle of which is as Figure 1As shown, its essence is a ternary second-order cognitive diagnosis architecture. At the bottom layer, "learning units" are used to describe the educational context. A learning unit includes three types of features: knowledge, time series, and behavior. For the knowledge and behavior features, first-order coupling analysis is performed with the time series feature respectively to obtain the cognitive state. Specifically, by performing first-order coupling analysis on knowledge-time series, the evolution of the learner's knowledge state can be obtained; by performing first-order coupling analysis on behavior-time series, the behavior changes of the learner can be analyzed. Then, the knowledge state and the behavior state are second-order coupled at the time granularity to obtain the learner's comprehensive ability state. According to the ability state, dual tasks of behavior and response are set to enhance the prediction of the learner's performance.
[0120] As Figure 2 shown, a hierarchical attention cognitive diagnosis method based on a mixture of knowledge association and abnormal behavior provided by an embodiment of the present invention includes the following steps:
[0121] (1) Obtain the time series data set when the student participates in the assessment. Completing the answer to a test question at a certain moment is called a "learning unit", including the knowledge features of the test question answered, the behavior features of the student when answering the question, and the time series features formed by the answer. In chronological order, several learning units can form a learning process sequence data, and based on this, a time series data set is constructed. Based on the constructed time series data set, for the knowledge representation of the test question, cross-modal test question data is collected, such as the text, pictures, and formulas of the test question, and the unified representation of the cross-modal test question is realized through different methods. Then, according to the relevant information such as the test question itself, the test question difficulty, and the test question concept, a heterogeneous graph is constructed, and by aggregating the feature information of the nodes of the test question heterogeneous graph, the associated features of the knowledge are mined, and the knowledge representation with associated information is obtained;
[0122] (2) Based on the obtained knowledge representation, combined with the student's answer response to the test question, a long-short-term hierarchical attention network is constructed to dynamically update the student's knowledge state and achieve the first-order coupling of knowledge and time series features. Specifically, first, the ternary features of knowledge, time series, and behavior are embedded and represented, and the long short-term memory neural network LSTM is used to encode them to obtain the current knowledge state; the Rasch model is used to calculate the potential initial traits of each scholar, and based on this, the working memory capacity of the learner is calculated; the self-attention mechanism is used to set triple masks of short-term mask, capacity mask, and time series mask to converge the answer effects within a limited capacity and a limited time, and enhance the current knowledge state; at the same time, based on the learning and forgetting rules, forgetting decay and learning enhancement functions are set to capture long-term attention and infer the learner's knowledge state;
[0123] (3) Calculate the guessing index and carelessness index of the learner when doing questions based on the click-stream behavior characteristics generated by the learner during the learning process due to clicks, and achieve the first-order coupling of behavior and time-series characteristics; specifically, based on the click-stream data, calculate features such as the number of option replacements, option coverage rate, residence time ratio, and time spent ratio, and set various anomaly enhancement and weakening functions based on the behavior characteristics, and output the guessing index and carelessness index of the learner when answering this question.
[0124] (4) Based on the knowledge state obtained after the coupling of knowledge and time-series in (2), combined with the guessing index and carelessness index obtained after the coupling of behavior and time-series in (3), set the guessing gate and carelessness gate to achieve the second-order coupling of knowledge-time-series-behavior, and update the comprehensive ability state of the learner. Specifically, based on the GLU gating function, set the guessing gate, and subtract the pseudo-knowledge growth obtained due to guessing from the original knowledge state; at the same time, set the carelessness gate, and add the mastered knowledge lost due to carelessness to the current knowledge state, and output the comprehensive ability state of the student.
[0125] (5) Set three prediction tasks of guessing index, carelessness index, and performance, all of which are predicted by a feed-forward neural network and a multi-sensory perceptron, and set two hyperparameters to balance the three prediction tasks, and use anomaly index prediction to enhance the prediction of the learner's final performance.
[0126] The specific implementation methods of each step are as follows:
[0127] (1) Cross-modal question representation and knowledge association mining
[0128] (1-1) Embedding and encoding of question attributes
[0129] The attributes of the questions include concepts, content, and difficulty, etc. The steps to achieve the embedding and encoding of question attributes are as follows:
[0130] First, the concept features of the questions belong to structured data, with a fixed data structure and few categories, and one-hot encoding is used for the initialization of embedding;
[0131] Second, the difficulty features of the questions can be divided into 3-5 levels from simple to difficult, and one-hot encoding is used;
[0132] Third, the content nodes of the questions, including the question content, carry a large amount of heterogeneous information such as text, formulas, and pictures, and belong to unstructured data. For formula information, the subject domain ontology replacement technology is used to keep the content standardized; then it is spliced with the text information, and the BERT model is used to perform a unified semantic representation of the questions; for picture information, CNN is used for picture recognition and representation. Finally, the embedding representations of text, formula, and picture are aggregated in a cascaded manner to obtain the initial embedding representation of the question content.
[0133] (1-2) Construction of the Heterogeneous Graph of Test Question Features
[0134] First, based on the different information of the above test questions, three types of heterogeneous nodes, namely concepts, difficulty levels, and content, are set up to construct a heterogeneous graph of test questions.
[0135] Secondly, the edge relationships in the heterogeneous graph are constructed. Given that there are three edge relationships among the three types of heterogeneous nodes, namely content - concept, content - difficulty level, and concept - concept, the first two can be obtained through the test question - concept association matrix and the test question - difficulty level association matrix respectively, while the edge relationship between concepts is obtained by similarity calculation and threshold setting.
[0136] (1-3) Enhanced Representation of Test Questions by Aggregating Neighbor Node Information
[0137] First, for the same - type nodes of test questions, a bidirectional LSTM is set up to aggregate the same - type nodes, and the average pooling method is used to obtain the aggregated vector of the corresponding type. Its formula can be expressed as
[0138]
[0139] where n i (v) is the set of neighbor nodes of type i of the current node v, emb(ω) represents the embedding representation of neighbor node w, is the vector after aggregating the neighbor nodes of type i of node v.
[0140] Secondly, for different - type nodes, the attention mechanism is used to aggregate three different types of vectors, and they are combined with the embedding of node v itself. The formula is expressed as
[0141]
[0142] where α (v,v) and α (v,i) represent weight coefficients, is the vector representation of the aggregated node v. In this way, the multi - dimensional heterogeneous features of test questions are aggregated through the heterogeneous graph, and the knowledge - related representation of test questions is realized.
[0143] (2) Knowledge State Update Based on Long - Short - Term Hierarchical Attention (Knowledge - Time Series First - Order Coupling)
[0144] As Figure 2For the "long - short - term hierarchical attention" part, based on the obtained knowledge association representation, combined with the student's response to the test questions, a long - short - term hierarchical attention network is constructed to dynamically update the student's knowledge state. Specifically, first, the triple features of knowledge, time series, and behavior are embedded and represented, and the long - short - term memory neural network (LSTM) is used to encode them to obtain the current knowledge state. According to the principle that "the higher the student's ability, the larger their working memory is likely to be", the Rash model is used to calculate the student's latent trait, and thus the student's working memory capacity is calculated. The student's memory capacity directly affects their short - term memory, and a short - term attention mechanism with a triple gate function is used to enhance the student's current knowledge state. Finally, the knowledge in short - term memory needs to be converted into long - term memory to be truly mastered, and the knowledge state is further updated based on the learning and forgetting laws.
[0145] (2 - 1) Embedding and Encoding of Triple Features
[0146] First, as Figure 3 shown, a learning unit composed of triple elements of knowledge, time series, and behavior is defined, denoted as LC n =(q n , s n , b n ), which is used to represent the learning event occurring within a certain time unit n. Among them, q n represents the knowledge feature under investigation, such as the test question and the concepts it considers; s n represents the time - series feature in the time - series data, including the time interval Δs n between each time unit, the learning interval Δr n of the same concept, and the repetition times Δc n , etc.; b n represents the click - stream behavior feature generated by the learner due to clicks during the learning process. The embedding representation of the learning unit is
[0147] X n =ε v +emb(s n )+emb(r n ),
[0148] where ε v is the knowledge association representation based on the heterogeneous graph neural network, emb(s n ) is the embedding of the reaction time, r n ∈b n is the correctness of the last click in the click - stream, and emb(r n ) is the embedding of the reaction result.
[0149] Subsequently, the influence of previous responses on subsequent responses is simulated through a recurrent neural network to initially obtain the potential knowledge state. The encoding process can be expressed as
[0150] H n = LSTM(X n ),
[0151] where LSTM is a long short - term memory neural network, representing the latent knowledge state at the n - th question.
[0152] (2 - 2) Working memory capacity calculation
[0153] First, according to the learner's answering situation (right or wrong), the Rasch model is used to calculate the learner's initial latent trait. The formula is as follows:
[0154] P(r in = 1|θ i ,β n ) = exp(θ i - β n )[1 + exp(θ i - β n )] -1 ,
[0155] where r in is a binary variable indicating whether the learner answers correctly or wrongly (1 represents a correct answer, 0 represents a wrong answer), θ i is the latent trait parameter of learner i, and β n is the difficulty parameter of question n.
[0156] Subsequently, we further use the latent trait value obtained from the above - mentioned Rasch model to determine the size of the learner's working memory capacity. The working memory capacity reflects the amount of knowledge and information that the learner can process during answering. The higher the latent trait, the larger the learner's working memory capacity; conversely, it is smaller. Specifically, we adjust the working memory capacity using the learner's trait and question difficulty, and the calculation is as follows:
[0157] WMC n = ΔC×σ(α·θ i - β·difff n ) + ∈
[0158] where WMC n is the working memory capacity when answering the n - th question, difff n is the difficulty of the n - th question, α and β are weight factors used to adjust the influence of the latent trait and question difficulty. σ is the Sigmoid activation function used to limit the output value between 0 and 1. ΔC is the capacity floating range, and ∈ is the minimum value of the capacity to ensure that the capacity is not zero. The working memory theory holds that a person's memory capacity is 5 ± 2. Therefore, we set ΔC to 4 and ∈ to 3.
[0159] (2-3) Triple-Masked Short-Term Attention
[0160] To mask irrelevant or excessive information and enhance short-term dependencies in time series, we set up triple masks for the attention mechanism, including a short-term mask, a capacity mask, and a temporal mask, to converge the answering effects within limited capacity and a defined time, thereby simulating the working memory operation mechanism.
[0161] First, the short-term mask M s is used to mask learning units with long time intervals, simulating the characteristic that human memory of recent events is more active. By calculating the time interval Δs between the current learning unit n and the previous unit. If the time interval exceeds a certain short-term window threshold δ s , the attention weight of this unit is set to 0, thus avoiding over-focus on long-term information.
[0162] Second, the capacity mask M w , based on the limitation of working memory capacity, masks learning units beyond the capacity range to ensure that the model only focuses on important information within the current capacity. According to the working memory capacity WMC at the nth question obtained in the previous step n , the mask is set so that only the information of the first WMC n learning units is retained, and the weights of learning units beyond this capacity range are set to 0.
[0163] Finally, the temporal mask M c is used to obscure subsequent learning units to ensure temporal consistency and prevent the influence of future information on the current state, that is, to avoid information leakage. For the current learning unit n, all subsequent learning units are masked, making the attention weight of future time steps on the current state 0.
[0164] Combining the triple masks, the formula of the attention mechanism is expressed as:
[0165] H′ n =Attention(H n ,M s ,M w ,M c ),
[0166] where represents the aggregated knowledge state after triple masking, and Attention refers to the multi-head self-attention network layer. M s , M w , M c are the short-term mask, the capacity mask, and the temporal mask respectively, used to regulate information screening from different sources.
[0167] (2-4) Long-Term Attention Based on Learning and Forgetting
[0168] Based on the enhancement of short-term memory, a long-term attention mechanism based on forgetting and learning curves is set up to capture the behavioral associations between long-cycle learning units.
[0169] First, the memory decay function F s (Δs n ) is used to simulate the process of human memory gradually weakening over time. We use a polynomial decay model, so that the memory decays rapidly at the beginning and then gradually slows down:
[0170] F s (Δs n ) = 1 / (1 + γ·Δs n ) p
[0171] where Δs n is the sequence time distance. γ is a parameter that adjusts the decay speed, and the larger the value, the faster the decay. p is the order of the polynomial, which we set to 2 to control the smoothness of the decay.
[0172] Second, the learning enhancement function F t (Δr n , Δc n ) is used to describe the enhancement effect of repeated learning on memory. The shorter the learning interval and the more the number of repetitions, the more obvious the memory strengthening effect. It is expressed as follows:
[0173] F t (Δr n , Δc n ) = (1 + α·Δc n )·exp(-β·Δr n )
[0174] where Δr n represents the time interval between the current learning unit and the last learning of the same concept, and Δc n represents the number of repeated learning. α and β are set hyperparameters. α is the enhancement coefficient of the number of repetitions, used to control the memory growth amplitude of each repeated learning. β is the decay rate of the learning interval, used to control the impact of the interval time on memory enhancement.
[0175] Finally, after memory decay and learning enhancement, the learner's knowledge state evolves into:
[0176]
[0177] where, H″ n+1 represents the predicted future knowledge state after learning and forgetting adjustment. is the query vector in the attention mechanism, q n+1is the test question representation to be answered at the (n + 1)-th moment. represents the knowledge state of past learning units and is used for the key matrix in the attention mechanism. serves as the value matrix in the attention mechanism. Both are weight matrices in the attention mechanism and are used for linear transformation of query, key, and value vectors. We use the knowledge state of future test questions to be answered in past learning units, combined with the memory decay function and learning enhancement function, to capture the long-term impact of past learning units on the current unit, making the model more conform to the laws of human memory and forgetting.
[0178] (3) Abnormal behavior analysis based on clickstream behavior (behavior-temporal first-order coupling)
[0179] Such as Figure 2 In the "abnormal index analysis" section, guessing and carelessness are phenomena that are likely to occur during the learner's answering process. According to clickstream behavior, the degree of guessing and carelessness can be obtained. Guessing means that the student does not master the knowledge but guesses correctly, while carelessness means that the student masters the knowledge but answers incorrectly due to carelessness.
[0180] (3-1) Guessing index analysis based on clickstream
[0181] First, based on the clickstream data formed by the learner when answering a certain question, the accumulative method is used to calculate the number of times of changing answers. For example, if the clickstream of a learner answering a certain test question is cs = {'A', 'B', 'A', 'C', 'A', 'B', 'A'}, then the number of changes is 6, denoted as count = Φ(cs) = 6.
[0182] Second, based on the clickstream data, the counting method is used to calculate the option coverage rate. For example, if the clickstream of a learner answering a certain test question is cs = {'A', 'B', 'A', 'C', 'A', 'B', 'A'}, and this question has 4 options, then the option coverage rate is cover = Ψ(cs) = 3 / 4 = 0.75.
[0183] Then, based on the clickstream data, the cumulative method is used to obtain the residence time ratio of the correct option. For example, in the above clickstream, the learner's cumulative residence time on A is 40s, and the total residence time is 70s, then the residence time ratio is time = Ω(cs) = 40 / 70 = 0.57.
[0184] Finally, we use these behavioral characteristics during the answering process to derive the guessing index. The more times of changing options, the larger the option coverage rate, and the less the residence time ratio of the correct answer, the greater the degree of guessing. The calculation formula of the guessing index is as follows:
[0185]
[0186] where guess ∈ [0, 1] is the guessing exponent, and f c (count) is an enhancement function representing the number of times the answer is changed, and f v (cover) is an enhancement function representing the option coverage rate, and f t (time) is a weakening function representing the ratio of the stay time of the correct option. Specifically, f c (count) = log(1 + α·count). The more times the answer is changed, the stronger the learner's uncertainty about the answer and the higher the degree of guessing. f v (cover) = β·cover. The higher the option coverage rate, it means that the learner has tried more options when answering questions and may guess without a clear answer. f t (time) = 1 / (1 + γ·time). The higher the ratio of the stay time, it usually indicates that the learner is more confident about the answer and the possibility of guessing is lower. α, β, and γ are all trainable parameters used to adjust these functions.
[0187] (3 - 2) Carelessness Index Analysis Based on Clickstream
[0188] First, calculate the ratio of the total time spent in the clickstream to the standard time. For example, if the time spent on clicking all options for answering a question is 3s and the average time for answering this question is 120s, then the ratio of the time spent is spend = Θ(cs) = 3 / 120.
[0189] Then, by analyzing the clickstream characteristics of the learner and the ratio of the time spent, calculate its carelessness index. The smaller the ratio of the time spent, the fewer the number of times of changing options, and the smaller the option coverage rate, the greater the probability of carelessness. The formula calculation is as follows:
[0190] skip = sigmoid(f′ c (count)·f′ v (cover)·f s (spend)) -1
[0191] where skip ∈ [0, 1] is the carelessness index, and f s (spend) = 1 / (1 + ζ·spend) is a weakening function of the ratio of the time spent. The smaller the ratio of the time spent, it indicates that the learner may be careless due to lack of time. ζ is also a trainable parameter. f′ c (count) and f′ v (cover) have the same calculation method as the enhancement functions in (3 - 1), except that the trainable parameters are replaced, namely α′ and β′.
[0192] (4) Ability status update of the fusion anomaly index (second-order coupling of knowledge - time series - behavior)
[0193] As Figure 2 in the "ability update" section, a guess gate and a carelessness gate are set to obtain the comprehensive ability status of the learner.
[0194] First, we define a gating function GLU() to control the influence of guessing and carelessness, and the function can be expressed as:
[0195]
[0196] Among them, the GLU function is used to control the influence of abnormal behavior on the knowledge state, and the input value X is the knowledge state H″ of the student n+1 , W1 and W2 are weight matrices, and b1 and b2 are bias terms, represents the element-wise product.
[0197] Subsequently, if guessing behavior is detected, the error caused by guessing is reduced through the guess gate, and the knowledge state is updated. We subtract the error caused by guessing from the knowledge mastery level, and the formula is:
[0198] K n+1 =H″ n+1 -GLU g (guess n+1 ·W g +b g )
[0199] Among them, K n+1 represents the updated knowledge state after passing through the guess gate, H″ n+1 is the future knowledge state obtained in (2 - 4), guess n+1 is the guessing index at time step n + 1, W g and b g are the weight and bias parameters of the guess gate.
[0200] Finally, if carelessness behavior is detected, the knowledge loss caused by carelessness is enhanced through the carelessness gate, and the knowledge state is further updated. We add the knowledge enhancement ignored by carelessness to the knowledge mastery level. The formula is:
[0201] θ n+1 =K n+1 +GLU s (skip n+1 ·W s +b s )
[0202] Among them, θ n+1 is the comprehensive ability status of the learner after passing through the carelessness gate, K n+1is the knowledge state after passing through the guessing gate, skip n+1 is the carelessness index at time step n+1, W s and b s are the weight and bias parameters of the proficiency gate.
[0203] (5) Prediction of enhanced performance with anomaly index
[0204] As Figure 2 in the "Behavior Prediction" section, we set the guessing index prediction, carelessness index prediction, and future performance prediction respectively, and enhance the final performance prediction through the anomaly behavior prediction of the first two.
[0205] (5-1) Anomaly index prediction
[0206] First, we set the prediction tasks for the guessing index and carelessness index to capture the anomaly behavior tendency of learners. Since the guessing index and carelessness index are difficult to observe directly, we generate corresponding pseudo-labels according to the behavior characteristics, denoted as and Specifically, we set thresholds through features such as the number of replacements, option coverage rate, and dwell time ratio to label "guessing" behavior, and assign the samples that meet the conditions to otherwise 0. Similarly, we set thresholds by combining features such as time ratio and number of replacements to label "careless" behavior, and assign the samples that meet the conditions to otherwise 0;
[0207] Second, we set the guessing index prediction task. In the prediction task, we use the feedforward network layer to transform the comprehensive ability state into the prediction space of the guessing index, and set a multi-layer perceptron to complete the prediction, that is
[0208]
[0209] where FNN and MLP are the feedforward network layer and multi-layer perceptron.
[0210] To optimize this prediction task, we define a loss function to measure the difference between the predicted guessing index of the model and the expected distribution. The loss function is defined as:
[0211]
[0212] where, is the predicted guessing index, is the target value for training.
[0213] Finally, similarly, we set the carelessness index prediction task. In the prediction task, we use the feedforward network layer to transform the comprehensive ability state into the prediction space of the carelessness index, and set a multi-layer perceptron to complete the prediction, that is
[0214]
[0215] Accordingly, the loss function is as follows:
[0216]
[0217] where is the predicted carelessness index, is the target value for training.
[0218] (5-2) Prediction of future performance
[0219] To predict the future performance of the learner, the knowledge state of the learner is concatenated with the predicted anomaly index to obtain the current comprehensive state, as follows:
[0220]
[0221] Secondly, a similar multi-layer perceptron is set up to complete the prediction of the learner's response result, that is:
[0222]
[0223] where is the predicted correct / incorrect result. Its loss function is:
[0224]
[0225] Finally, to balance the three prediction tasks, the loss functions are weighted and summed to obtain the global loss function for multi-tasks, that is:
[0226]
[0227] where is the global loss function, and the hyperparameters λ1 and λ2 are used to adjust the weights of each task to achieve the balance of different tasks.
[0228] The proposed method of knowledge association enhanced temporal hierarchical attention cognitive diagnosis in the present invention is essentially a ternary second-order coupling framework. The learning units in the learning process are divided into a "knowledge, time series, behavior" triple, and a variety of deep learning technologies are introduced to achieve second-order modeling of cognition and ability. Its comparison with the existing static cognitive diagnosis and dynamic cognitive diagnosis methods is shown in Table 1. As shown in Table 1, the static cognitive diagnosis method only uses the knowledge features in the educational context and does not consider the time series and behavior features in the real educational context, and is applicable to single learning scenarios. Its essence is the unary coupling of knowledge; the dynamic cognitive diagnosis method considers the knowledge features and the sequence features of the learning process, is applicable to multiple learning scenarios, and ignores the time interval attribute and behavior process features in the learning process. Its essence is the weak coupling of knowledge and time series features, such as Figure 4As shown in (a) therein, the binary coupling of the weakening process behavior has limitations, which will eventually lead to the problem of incomplete learning characterization; while the cognitive diagnosis method proposed by the present invention considers the second-order coupling of three features, as Figure 4 shown in (b) therein, it can scientifically and comprehensively model the cognitive process of learners and achieve accurate prediction of learners' future performance.
[0229] Table 1 Comparison table of coupling mechanisms between the method of this example and traditional cognitive diagnosis methods
[0230]
[0231] The content not described in detail in this specification belongs to the prior art known to those skilled in the art.
[0232] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A hierarchical attention cognitive diagnosis method based on the mixture of knowledge association and abnormal behavior, characterized in that, It includes the following steps: Step 1: Obtain the learning unit data of students during the assessment, including the knowledge characteristics of the questions answered, the behavioral characteristics of students during answering, and the temporal characteristics formed by the answers. Arrange several learning units in chronological order to form learning process sequence data, and thus construct a temporal dataset. Based on the constructed temporal dataset, obtain the knowledge association representation of the questions; Step 2: Based on the obtained knowledge representation, combined with the answering responses of students to the questions, construct a long short-term hierarchical attention network to dynamically update the knowledge state of students, realizing the first-order coupling of knowledge and temporal characteristics; Step 3: Based on the click stream behavioral characteristics generated by learners during the learning process, calculate the guessing index and carelessness index of learners when doing questions, realizing the first-order coupling of behavior and temporal characteristics; Step 4: Based on the knowledge state obtained after the coupling of knowledge and time series, combined with the guessing index and carelessness index obtained after the coupling of behavior and time series, set the guessing gate and carelessness gate to realize the second-order coupling of knowledge-time series-behavior, and update the comprehensive ability state of learners; Step 5: Set the anomaly index and future performance prediction task. The anomaly index includes the guessing index and the carelessness index, both of which are predicted using a feedforward neural network and a multi-sensory perceptron, and two hyperparameters are set to balance the three prediction tasks. The anomaly index prediction is used to enhance the prediction of learners' final performance.
2. The hierarchical attention cognitive diagnosis method based on the mixture of knowledge association and abnormal behavior according to claim 1, characterized in that: The specific content of Step 1 includes: (1-1) Embedding and encoding of question attributes; Among them, the attributes of the questions include concepts, content, and difficulty; (1-2) Construction of a question feature heterogeneous graph; First, based on different information of the questions, set three heterogeneous nodes of concepts, difficulty, and content to construct a question heterogeneous graph; Secondly, construct the edge relationships in the heterogeneous graph. Given that there are three edge relationships among the three heterogeneous nodes, namely content-concept, content-difficulty, and concept-concept, the first two are obtained through the question-concept association matrix and the question-difficulty association matrix respectively, and the edge relationship between concepts is obtained through similarity calculation and threshold setting; (1-3) Enhanced representation of questions by aggregating neighbor node information; First, for the same type of nodes of the questions, set a bidirectional LSTM to aggregate the same type of nodes, and use the average pooling method to obtain the aggregated vector of the corresponding type, and its formula is expressed as: where n i (v) the set of neighbor nodes of type i of the current node v, and emb(ω) represents the embedding representation of neighbor node w, is the vector after aggregating the neighbor nodes of type i of node v; Secondly, for different types of nodes, aggregate three different types of vectors through the attention mechanism, and combine them with the embedding of node v itself, and the formula is expressed as: Among them, α (v,v) and α (v,i) represent weight coefficients, is the vector representation after the aggregation of node v; thus, the multi-dimensional heterogeneous features of the test questions are aggregated through the heterogeneous graph, realizing the knowledge association representation of the test questions.
3. The hierarchical attention cognitive diagnosis method based on the mixture of knowledge association and abnormal behavior according to claim 2, characterized in that: The embedding and encoding of question attributes specifically include: First, the concept features of the questions belong to structured data, and one-hot encoding is used for the initialization of the embedding; Secondly, the difficulty features of the questions use one-hot encoding; Thirdly, the content nodes of the questions, including the question content, belong to unstructured data; for formula information, the subject domain ontology replacement technology is used to keep the content standardized; then it is concatenated with the text information, and the BERT model is used to perform a unified semantic representation of the questions; for picture information, CNN is used for picture recognition and representation; finally, the embedding representations of text, formula, and picture are aggregated in a cascaded manner to obtain the initial embedding representation of the question content.
4. A hierarchical attention cognitive diagnosis method based on a mixture of knowledge association and abnormal behavior, characterized in that: Step 2 specifically includes: First, perform embedding representation on the triple features of knowledge, time sequence, and behavior, and use a long short-term memory neural network (LSTM) to encode them to obtain the current knowledge state; Then, according to the principle that the higher the student's ability, the larger their working memory is likely to be, use the Rasch model to calculate the latent trait of the student, and thus calculate the working memory capacity of the student; The memory capacity of the student directly affects their short-term memory. Use a short-term attention mechanism with a triple gate function to enhance the student's current knowledge state; Finally, the knowledge in short-term memory needs to be converted into long-term memory to be truly mastered. Therefore, use long-term attention based on the laws of learning and forgetting to further update their knowledge state.
5. A hierarchical attention cognitive diagnosis method based on a mixture of knowledge association and abnormal behavior, as described in claim 1 or 4, characterized in that: The specific implementation of Step 2 includes the following steps: (2-1) Triple Feature Embedding and Encoding First, define a learning unit composed of three elements: knowledge, time series, and behavior, denoted as LC n =(q n , s n , b n ), which is used to represent the learning event that occurs within a certain time unit n; among them, q n represents the knowledge feature under investigation; s n represents the time series feature in the time series data, including the time unit interval Δs n , the same concept learning interval Δr n and the repetition times Δc n ; b n represents the click stream behavior feature generated by the learner due to clicks during the learning process; the embedding of the learning unit is expressed as X n = ε v + emb(s n ) + emb(r n ), Among them, ε v is the knowledge association representation based on the heterogeneous graph neural network, emb(s n ) is the embedding of the reaction time, r n ∈b n is the correctness of the last click in the click stream, emb(r n ) is the embedding of the reaction result; Subsequently, use a recurrent neural network to simulate the influence of previous answers on subsequent answers, and initially obtain the latent knowledge state; the encoding process is expressed as: H n = LSTM(X n ) Among them, LSTM is a long short-term memory neural network, represents the latent knowledge state at the nth question; (2-2) Working Memory Capacity Calculation First, according to the answering situation of the learner, use the Rasch model to calculate the initial latent trait of the learner. The formula is as follows: P(r in = 1|θ i , β n ) = exp(θ i - β n )[1 + exp(θ i - β n )] -1 , where r in is a binary variable indicating whether the learner answers correctly or incorrectly, with 1 representing a correct answer and 0 representing an incorrect answer, and θ i is the latent trait parameter of learner i, and β n is the difficulty parameter of question n; Subsequently, further use the latent trait value obtained by the Rasch model to determine the size of the learner's working memory capacity; specifically, adjust the working memory capacity using the learner's trait and the difficulty of the test questions. The calculation is as follows: WMC n = ΔC × σ(α·θ i - β·difff n ) + ∈ Among them, WMC n is the working memory capacity when answering the nth question, difff n is the difficulty of the nth question, α and β are weight factors used to adjust the influence of latent traits and question difficulty; σ is the Sigmoid activation function used to limit the output value between 0 and 1; ΔC is the capacity floating range, and ∈ is the minimum value of the capacity to ensure that the capacity is not zero; (2-3) Short-Term Attention Based on Triple Mask To mask irrelevant or excessive information and enhance short-term dependencies in the time series, set a triple mask for the attention mechanism, including a short-term mask, a capacity mask, and a time-sequence mask, to converge the answering influence within a limited capacity and a limited time, thereby simulating the working memory operation mechanism; First, the short-term mask M s is used to mask learning units with a longer time interval, simulating the characteristic that human memory of recent events is more active; by calculating the time interval Δs between the current learning unit n and the previous unit; if the time interval exceeds a certain short-term window threshold δ s , then the attention weight of this unit is set to 0, thus avoiding over-focus on long-term information; Secondly, the capacity mask M w Based on the limitation of working memory capacity, learning units beyond the capacity range are masked to ensure that the model only focuses on important information within the current capacity; according to the working memory capacity WMC at the nth question obtained in the previous step n , set the mask so that only the information of the first WMC n learning units is retained, and the weights of learning units beyond this capacity range are set to 0; Finally, the timing mask M c is used to mask subsequent learning units, ensuring temporal consistency and preventing the influence of future information on the current state; for the current learning unit n, all subsequent learning units are masked, making the attention weights of future time steps to the current state zero; Combined with the triple mask, the formula for the attention mechanism is expressed as: H n ′ = Attention(H n , M s , M w , M c ), Among them, represents the aggregated knowledge state after triple masking, and Attention refers to the multi-head self-attention network layer; M s , M w , M c are the short-term mask, the capacity mask, and the temporal mask respectively, which are used to regulate the information screening from different sources; (2-4) Long-Term Attention Based on Learning and Forgetting On the basis of enhancing short-term memory, set a long-term attention mechanism based on the forgetting and learning curves to capture the behavioral associations between long-cycle learning units; First, the memory decay function F s (Δs n ) is used to simulate the process of human memory gradually weakening over time. Using a polynomial decay model, the memory decays rapidly in the initial stage and then gradually slows down: F s (Δs n ) = 1 / (1 + γ·Δs n ) p where Δs n is the sequence time distance; γ is a parameter for adjusting the attenuation rate, and the larger the value, the faster the attenuation; p is the order of the polynomial, which is used to control the smoothness of the attenuation; Secondly, the learning enhancement function F t (Δr n , Δc n ) is used to describe the enhancement effect of repeated learning on memory and is expressed as follows: F t (Δr n , Δc n ) = (1 + α·Δc n )·exp(-β·Δr n ) where Δr n represents the time interval between the current learning unit and the last learning of the same concept, and Δc n represents the number of repeated learning; α and β are set hyperparameters. α is the enhancement coefficient of the number of repetitions, which is used to control the memory growth amplitude of each repeated learning; β is the decay rate of the learning interval, which is used to control the impact of the interval time on memory enhancement; Finally, after memory decay and learning enhancement, the knowledge state of the learner evolves into: where, H″ n+1 represents the speculated future knowledge state after learning and forgetting adjustment; is the query vector in the attention mechanism, q n+1 is the representation of the test question to be answered at the (n + 1)-th moment; represents the knowledge state of the past learning unit and is used for the key matrix in the attention mechanism; is used as the value matrix in the attention mechanism; W Q , W K , are all weight matrices in the attention mechanism and are used for linearly transforming the query, key, and value vectors.
6. The hierarchical attention cognitive diagnosis method based on the mixture of knowledge association and abnormal behavior as claimed in claim 1, wherein: The clickstream-based guessing index in Step 3 includes: First, based on the clickstream data formed when the learner answers a certain question, use an accumulation method to calculate the number of times they change their answers; Second, based on the clickstream data, use a counting method to calculate the option coverage rate; Then, based on the clickstream data, obtain the residence time ratio of the correct option through an accumulation method; Finally, use these behavioral characteristics during the answering process to derive the guessing index. The formula for the guessing index is as follows: where guess ∈ [0, 1] is the guessing exponent, f c (count) is an enhancement function for the number of answer replacements, f v (cover) is an enhancement function for option coverage, f t (time) is a weakening function for the ratio of the stay time of the correct option; specifically, f c (count) = log(1 + α·count), the more replacements are made, the stronger the learner's uncertainty about the answer and the higher the degree of guessing; f v (cover) = β·cover, the higher the option coverage, it means that the learner has tried more options when answering and may guess without a clear answer; f t (time) = 1 / (1 + γ·time), the higher the ratio of the stay time, it indicates that the learner is more confident about the answer and the possibility of guessing is lower; α, β, and γ are all trainable parameters used to adjust these functions.
7. The hierarchical attention cognitive diagnosis method based on the mixture of knowledge association and abnormal behavior according to claim 6, wherein: The clickstream-based carelessness index in Step 3 includes: First, calculate the ratio of the total time spent in the clickstream to the standard time, that is, the time spent ratio spend; Then, by analyzing the clickstream characteristics and time ratio of the learner, calculate their carelessness index. The formula calculation is as follows: skip=sigmoid(f′ c (count)·f′ v (cover)·f s (spend)) -1 where skip ∈ [0, 1] is the carelessness index, and f s (spend) = 1 / (1 + ζ·spend) is the attenuation function of the proportion of time spent. The smaller the proportion of time spent, the more likely it indicates that the learner is careless due to insufficient time; ζ is also a trainable parameter, and f′ c (count) and f′ v (cover) are calculated in the same way as the enhancement functions of f c (count) and f v (cover), except that the trainable parameters are replaced with α′ and β′ respectively.
8. The hierarchical attention cognitive diagnosis method based on the mixture of knowledge association and abnormal behavior according to claim 1, wherein: The specific implementation method of Step 4 is as follows: First, define a gated function GLU() to control the influence of guessing and carelessness. The function is expressed as: Among them, the GLU function is used to control the impact of abnormal behavior on the knowledge state, and the input value X is the student's knowledge state H″ n+1 , W1 and W2 are weight matrices, and b1 and b2 are bias terms, represents the element-wise product; Subsequently, if a guessing behavior is detected, reduce the error caused by guessing through the guessing gate, update the knowledge state, and subtract the error caused by guessing from the knowledge mastery level. The formula is: K n+1 = H″ n+1 - GLU g (guess n+1 ·W g + b g ) Among them, K n+1 represents the updated knowledge state after passing through the guess gate, H″ n+1 is the future knowledge state, guess n+1 is the guess exponent at time step n+1, W g and b g are the weight and bias parameters of the guess gate; Finally, if careless behavior is detected, the knowledge gaps caused by carelessness are enhanced through the carelessness gate, and the knowledge state is further updated by adding the knowledge enhancement ignored by carelessness to the knowledge mastery level. The formula is as follows: θ n+1 = K n+1 + GLU s (skip n+1 · W s + b s ) Among them, θ n+1 is the comprehensive ability state of the learner after passing through the carelessness gate, K n+1 is the knowledge state after passing through the guessing gate, skip n+1 is the carelessness index at time step n + 1, W s and b s are the weight and bias parameters of the proficiency gate.
9. The hierarchical attention cognitive diagnosis method based on the mixture of knowledge association and abnormal behavior according to claim 1, wherein: In step 5, first, corresponding pseudo-labels are generated according to behavioral characteristics, denoted as and Specifically, thresholds are set by the number of replacements, option coverage rate, and residence time ratio to mark "guessing" behavior, and samples that meet the conditions are assigned as Otherwise, it is 0; similarly, thresholds are set by combining the time ratio and the number of replacements to mark "careless" behavior, and samples that meet the conditions are assigned as Otherwise, it is 0; Secondly, a guessing index prediction task is set up. In the prediction task, the feed-forward network layer is used to transform the comprehensive ability state into the prediction space of the guessing index, and a multi-layer perceptron is set up to complete the prediction, that is where FNN and MLP are the feed-forward network layer and the multi-layer perceptron; To optimize this prediction task, a loss function is defined to measure the difference between the predicted guessing index of the model and the expected distribution. The loss function is defined as: Among them, is the predicted guess index, is the target value for training; Finally, a carelessness index prediction task is set up. In the prediction task, the feed-forward network layer is used to transform the comprehensive ability state into the prediction space of the carelessness index, and a multi-layer perceptron is set up to complete the prediction, that is Accordingly, the loss function is: Among them, is the predicted carelessness index, is the target value for training.
10. A hierarchical attention cognitive diagnosis method based on a mixture of knowledge association and abnormal behavior, characterized in that: To predict the future performance of the learner, the knowledge state of the learner is concatenated with the predicted anomaly index to obtain the current comprehensive state, as follows: Secondly, a similar multi-layer perceptron is set up to complete the prediction of the learner's response result, that is wherein is the predicted correct / incorrect result; its loss function is as follows: Finally, to balance the three prediction tasks, the loss functions are weighted and summed to obtain the global loss function of the multi-task, that is: Among them is the global loss function, and the hyperparameters λ1 and λ2 are used to adjust the weights of each task to achieve the balance of different tasks.
Citation Information
Patent Citations
Cognitive diagnosis method based on learner cognitive response model
CN112765830A
Test question resource recommendation method and system
CN114155124A