A Dynamic Cognitive Diagnosis Method with Second-Order Coupling of Knowledge-Time Series-Behavior Ternary
By introducing knowledge-time-order-behavior ternary second-order coupling technology into the dynamic cognitive diagnostic model, comprehensive modeling of learners' cognitive process and accurate prediction of future performance, the limitations of existing models in reflecting the association of ‘knowledge, timing and behavior’ is solved.
Patent Information
- Application Number
- CN202411674130.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The existing dynamic cognitive diagnostic model has limitations in reflecting the association between ‘knowledge, timing and behavior’, resulting in the incomplete characterization of learners’ performance and inconsistent cognitive status and performance.
A dynamic cognitive diagnosis method of knowledge-time-order-behavior ternary second-order coupling is proposed. Through neural network model and behavior feature extraction technology, the first-order coupling analysis of ‘knowledge-time-order’ and ‘behavior-time-order’ is realized. Combining click flow and activity flow behavior, GLU gating technology and time distance attention mechanism are used to perform second-order coupling analysis to obtain the learner’s comprehensive ability state.
It realizes scientific and comprehensive modeling of learners' cognitive process, can accurately predict learners' future performance and assist teachers in accurate teaching.
Smart Images

Figure CN119514590B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of educational big data mining, attention networks, and learner modeling, and particularly to a dynamic cognitive diagnosis method with a ternary second-order coupling of knowledge, time series, and behavior. Background Art
[0002] Cognitive diagnosis is the key to characterizing the learning subject and carrying out personalized learning in an intelligent educational environment. Cognitive diagnosis aims to obtain the most intuitive response data based on the interaction between the learner and the test questions during the test-taking process, construct a cognitive diagnosis model to model the learner's internal cognitive processing process, and then diagnose the learner's knowledge mastery status, and evaluate the effectiveness of the model by predicting the learner's future performance. The results of cognitive diagnosis can help teachers and students quickly locate the learner's weak knowledge. Teachers can design corresponding intervention strategies according to the group learning situation to carry out differentiated teaching, and learners can carry out targeted remedial exercises according to their personal learning situation to achieve large-scale personalized learning. Therefore, it is urgent to use intelligent technology to implement accurate diagnosis, comprehensively characterize the learner's cognitive level, reduce the burden on students and improve efficiency by providing accurate educational services, and achieve high-quality personalized learning.
[0003] The modeling technology of cognitive diagnosis has evolved from the traditional cognitive diagnosis model with a manually constructed probabilistic graph structure to an intelligent cognitive diagnosis model based on deep learning for automatic learning, so as to improve the fitting ability of learners and item characteristics. According to the different data sources of cognitive diagnosis, the current cognitive diagnosis can be roughly divided into two categories: static cognitive diagnosis and dynamic cognitive diagnosis.
[0004] Static cognitive diagnosis aims to infer the learner's knowledge mastery status through the diagnosis and evaluation of individual cognitive processes, processing skills, or knowledge structures, mainly for single learning scenarios, such as final exams. The most classic cognitive diagnosis models, such as the DINA model, use the error and guessing parameters in the learning process to construct a diagnostic function to infer the learner's knowledge mastery status. With the development of artificial intelligence technology, many studies have introduced deep learning methods into static cognitive diagnosis modeling, that is, deep cognitive diagnosis models. For example, Cheng et al. proposed the DIRT model, which uses deep learning algorithms to mine the test question text and the relationship between test questions and concepts to enhance the test question representation and strengthen the diagnosis process; Wang et al. proposed the NeuralCD framework, which captures the complex interaction relationship between learners and test questions through neural networks; Gao et al. proposed the DeepCDM model, which uses deep learning to simulate the learner's mastery of skills and test questions. However, due to the lack of attention to the historical prior data in the learning process, the current results of static cognitive diagnosis are one-sided.
[0005] In order to solve the one-sided problem of static cognitive diagnosis, dynamic cognitive diagnosis came into being. Dynamic cognitive diagnosis accumulates multiple assessment data to form time series data, and obtains the changes in the learner's knowledge mastery status by analyzing the time series data. For example, the classic BKT model abstracts the cognitive state into a set of binary variables, takes the learner's real-time interaction as input, and simulates the changes in the learner's knowledge mastery during the learning process through the hidden Markov model. In fact, due to the good performance of deep learning algorithms, dynamic cognitive diagnosis based on deep learning has become a research hotspot in the field of cognitive diagnosis. For example, the DKT model uses a recurrent neural network to track the learner's cognitive state at different times; the DKVMN model uses a memory network to store the learner's implicit cognitive state, and dynamically updates the network through key-value pairs each time the answer is given; and the EKT model makes full use of natural language processing, domain ontology replacement, and convolutional neural networks to characterize multimodal test features to enhance cognitive diagnosis performance.
[0006] However, most of the current dynamic cognitive diagnosis models based on deep learning take into account the learner's historical answer records, and preliminarily realize the "knowledge-time sequence" binary coupling through time series modeling, which can dynamically track the learner's cognitive level. However, the existing models are still mainly based on result-oriented behavior, and the integration of scattered process behavior characteristics still cannot reflect the relationship between "knowledge, time sequence and behavior". This binary coupling that weakens process behavior has limitations and easily leads to incomplete characterization of learners' performance and inconsistency between their performance and their cognitive state. Summary of the invention
[0007] In order to overcome the above-mentioned deficiencies of the prior art, the present invention provides a knowledge-sequence-behavior three-element second-order coupled dynamic cognitive diagnosis method, comprising the following steps:
[0008] Step 1: Take a certain moment in the learning process as a unit to obtain the learning unit data when students participate in the assessment, which consists of knowledge features, behavior features and time series features. According to the chronological order, several learning units can form the learning process sequence data, so as to construct a time series data set. Knowledge features include the questions answered by students at a certain moment and the concepts related to the questions. Behavior features include the click stream formed by students repeatedly modifying their answers during the answering process, and the activity stream formed by students viewing the analysis and video resources after the answering. Time series features refer to the time interval between each time unit, the time and number of repetitions of the same concept learning interval, etc. Based on the constructed time series data set, a ternary standardized representation of the learning unit is carried out. For knowledge features, each concept related to the current question is first embedded, and then the average pooling method is used to obtain the concept representation associated with the question, which is added to the embedded question to obtain the knowledge representation of the current answer question; for time series features, the embedding method is used to achieve representation.
[0009] Step 2: For knowledge-temporal features and behavior-temporal features, respectively adopt a neural network model and behavior feature extraction technology to achieve the first-order coupling analysis of the "knowledge-temporal" and "behavior-temporal" cognitive states. Specifically, it includes: for knowledge-temporal coupling, input knowledge and temporal representations into an LSTM neural network, and set forgetting attenuation and learning enhancement functions for dynamic adjustment; for behavior-temporal coupling, according to the timestamp information in the click stream and activity stream, adopt behavior feature extraction technology to extract five click feature vectors and four activity features, and encode them through addition and self-attention network with causal mask respectively.
[0010] Step 3: Based on the behavior analysis and encoding of the click stream and activity stream, combined with the knowledge state analysis, achieve the second-order coupling analysis of the ability state of knowledge-behavior-temporal coupling. Specifically, first, based on the output result of the first-order coupling, set the GLU gating function to obtain the initial cognitive state of the student; then, set the "guess gate" and "proficiency gate" through click stream encoding to weaken and enhance the current cognitive state respectively; subsequently, set the "learning gate" through average pooling of the activity stream to further enhance the current cognitive state; finally, use the time distance attention mechanism to calculate the comprehensive ability state representation of the student.
[0011] Step 4: Based on the obtained comprehensive ability state, enhance the final performance prediction through the dual prediction tasks of behavior and reaction. Specifically, set three behavior prediction tasks of reaction time, click stream length, and activity stream length, transform the ability state into the corresponding behavior space, and then complete the behavior prediction task through a multi-layer perceptron; aggregate learning behavior prediction and predict the final reaction result at future moments through a multi-layer perceptron.
[0012] Furthermore, the triple normalization representation of the learning unit specifically includes:
[0013] (1-1) Knowledge feature embedding
[0014] All test questions and concepts together constitute a knowledge space, represented as a triple (Q, C, τ), where Q is a non-empty set of test questions, C is a non-empty set of concepts, and τ is a mapping from Q to concept C; for any test question q n ∈Q, τ(q) represents the set of concepts required to solve test question q, and a test question can be associated with multiple concepts; specifically, it includes three steps: concept embedding, test question embedding, and knowledge comprehensive representation;
[0015] First, concept embedding: establish the association between test questions and multiple concepts through concept embedding; specifically, first embed each concept, and then obtain the comprehensive knowledge representation of the test question through average pooling, that is Among them, is the overall embedding representation of the concepts involved in the nth test question, is the embedded representation of the concept involved in the nth test question, d k is the dimension of the embedding, c n,i is derived from q n 's knowledge concept set τ(q n );
[0016] Secondly, test question embedding: The test questions are encoded in an embedded way. The embedding of the test question index involved in the nth learning unit is represented as
[0017] Finally, the knowledge comprehensive representation of the test question: It is obtained by adding the index embedding of the test question and the concept embedding , that is, k n = q n + c n , where represents the knowledge comprehensive representation of the test question;
[0018] (1 - 2) Temporal feature embedding
[0019] Each learning unit in the learning process occurs at a certain time, which is recorded as the timestamp when the learning event occurs. The timestamps in the learning trajectory are used to extract temporal features, including timestamp distance, time interval, and knowledge review times;
[0020] First, the timestamp distance is used to measure the real distance between learning units. By calculating the timestamp distance between adjacent answerings before and after, the sequence time interval Δs n ;
[0021] Secondly, the time interval refers to the time interval between the current knowledge answering time and the time of the last answering of the same knowledge point, denoted as the review time interval Δr n ;
[0022] Then, the knowledge review times refer to the cumulative answering times of the same knowledge point in the learning process, denoted as the knowledge review times Δc n ; Thus, the temporal feature T n =(Δs n , Δr n , Δc n );
[0023] (1 - 3) Behavioral feature extraction
[0024] Use clickstream and activity stream to describe the multi - learning behavioral features in the learning process;
[0025] First, extract the clickstream cs n from the learning trajectory; When answering a certain test question q nWhen learners may modify their answers multiple times due to uncertainty, several click-and-answer operations are thus formed. Extracting them in chronological order results in a click stream. Among them, represents the result of the i-th click-and-answer operation when answering question q n ;
[0026] Secondly, extract the activity stream as n from the learning trajectory; after each completion of answering a question, learners will conduct thinking and further learning through activities such as viewing answer analysis and watching video lectures; extract the post-answer activity behavior of learners to form an activity stream, denoted as where represents the i-th activity after the student answers question q n , including viewing the analysis explanation or watching the video lecture, that is
[0027] Furthermore, the specific implementation process of knowledge-temporal coupling is as follows:
[0028] First, embed the temporal feature T n =(Δs n , Δr n , Δc n ) into the same dimension d k and then add them to obtain the comprehensive representation of the temporal feature Subsequently, concatenate it with the knowledge feature embedding to form the final input vector x n , and the comprehensive representation of the learning unit is as follows:
[0029]
[0030] Input the knowledge feature k n and the temporal feature T n together into the LSTM neural network to capture the dynamic relationship between knowledge points and learning time factors;
[0031] Subsequently, based on the original forgetting gate, use the timestamp distance to adjust the forgetting degree, and the calculation method is as follows:
[0032] f n =σ(W f ·[h n-1 , x n +b f )·exp(-α·Δs n )
[0033] where is the output of the adjusted forgetting gate, is the hidden state of the previous learning unit; σ is the sigmoid activation function, and exp(·) is the exponential function. and are the trainable weight matrix and bias term; Δs n is the timestamp distance between adjacent learning units, and α is a tuning parameter used to control the impact of the timestamp distance on the forget gate.
[0034] Next, to enhance the learning process, a review time interval Δr n and the knowledge review count Δc n are introduced to adjust the output of the input gate, and the calculation method is as follows:
[0035]
[0036] where is the output of the input gate, and are the trainable parameters; γ is a set hyperparameter used to control the impact of the time interval and review count on the input gate.
[0037] Finally, based on the adjusted forget gate f n and the input gate i n , the cell state of the LSTM is updated, and the hidden state is updated, and the calculation method is as follows:
[0038]
[0039] where is the cell state of the LSTM, is the candidate cell state, and are the trainable parameters; the hidden state is calculated through the output gate o n :
[0040] o n = σ(W o · [h n-1 , x n + b o ), h n = o n · tanh(C n )
[0041] where is the output of the output gate, represents the hidden state of the current learning unit.
[0042] Furthermore, behavior-temporal coupling includes clickstream behavior-temporal coupling and activity stream behavior-temporal coupling;
[0043] The specific implementation process of clickstream behavior - temporal coupling is as follows:
[0044] First, according to the clickstream Based on its timestamp information, five behavior change features are extracted: ① Read whether the learner's answer at time t = 0 matches the standard answer, that is, whether the answer is correct or not, and record it as "initial answer correctness "; ② Read the correctness information at the last modified time t = i in the same way, and record it as "final answer correctness "; ③ Calculate the time difference between the last time and the initial time, and record it as "reaction time consumed t n ", where t n = t i - t0; ④ Calculate the length of the entire clickstream, and record it as "clickstream length l n = |cs n |; ⑤ Analyze the behavior changes of the learner in the clickstream. The learner may finally modify the answer, record it as the "modified" type, or may change it back to the initial answer after repeated changes, record it as "hesitation", and call this variable "clickstream type p n ";
[0045] Secondly, combine these five embedding vectors into a vector and map it to a unified vector space where is the embedding representation of the activity stream cs n at the nth answer, D k is the dimension of the embedding, and emb represents the corresponding embedding layer function;
[0046] The specific implementation process of activity stream behavior - temporal coupling is as follows:
[0047] First, according to the activity stream Based on its timestamp information, four behavior change features are extracted: ① Read the type of the current activity, and record it as "activity type ", where ② Calculate the duration from the start to the end of the activity, and record it as "total activity time "; ③ Obtain the number of times the activity is to view the analysis, that is , and record it as "number of times of viewing analysis "; ④ Obtain the number of times the activity is to view the video, that is , and record it as "number of video operations ";
[0048] Subsequently, an embedding matrix is used to represent the 4 activity stream features. Each row of the matrix corresponds to one activity, that is where is the activity stream as nThe embedded representation of the i-th activity; combining the corresponding activities in the activity stream to form the activity stream where la n is the length of the activity stream.
[0049] Furthermore, the specific implementation of step 3 includes:
[0050] (3-1) Based on the output result of the first-order coupling, set the GLU gating function to obtain the initial cognitive state of the student, and set the guess gate and proficiency gate through the click stream encoding to weaken and enhance the current cognitive state respectively;
[0051] (3-2) Based on the activity stream, set the GLU-based learning gate to enhance the current cognitive state;
[0052] (3-3) Through the time distance attention mechanism, obtain the comprehensive ability state representation of the student.
[0053] Furthermore, the specific implementation method of (3-1) is as follows:
[0054] First, the knowledge state of the learner is controlled by the gating function GLU(), and its function can be expressed as:
[0055]
[0056] where, is the initial cognitive state, is the output result of the first-order coupling of the cognitive state in the previous step, and represent the corresponding weight and bias parameters;
[0057] Second, according to the modification operations of the student in the click stream, calculate the guessing degree during the learner's answering process; adjust the cognitive state by setting the guess gate, and the function is expressed as:
[0058] K′ n =K n -GLU g (CS n ·W g +b g ),
[0059] where, is the knowledge state after passing through the guess gate; represents the embedded representation of the click stream, and are the weight and bias parameters of the guess gate;
[0060] Third, according to the time attribute in the click stream, infer the proficiency degree of the learner's knowledge mastery; enhance the knowledge state by setting the proficiency gate, and the function is expressed as:
[0061] K″ n = K′ n + GLU p (CS n ·W p + b p )
[0062] Among them, is the knowledge state after passing through the proficiency gate; and are the weight and bias parameters of the proficiency gate.
[0063] Furthermore, the specific implementation of (3 - 2) is as follows:
[0064] For unified representation learning activities, the average pooling method is used to obtain the average value of each activity embedding in the activity stream, thereby aggregating all activity information. The calculation method is as follows:
[0065]
[0066] Among them, is the aggregated representation of the activity stream of the nth response, is the activity stream as n the length of, that is, the number of activities, represents the embedding representation of the i-th activity in the nth response;
[0067] Secondly, through the learning activity behavior after answering, the learner's knowledge mastery can be enhanced, and a learning gate is set to simulate the improvement degree of its high-order cognitive state; the function is expressed as:
[0068] K″′ n = K″ n + GLU l (AS n )
[0069] Among them, represents the knowledge state enhanced by the learning gate, is the knowledge state before entering the learning gate, and GLU l is the learning gate function, which is used to enhance the influence of the activity stream on the knowledge state.
[0070] Furthermore, the specific implementation of (3 - 3) is as follows:
[0071] Based on the obtained current cognitive state h n and (3 - 2) to obtain the knowledge state change K″′ n caused by the behavior, and further infer the potential comprehensive ability state θ n+1 of the student through the time distance attention mechanism; To introduce the influence of time factors, the query vector is set to the representation k of the test question at the (n + 1)-th moment n+1 , and an exponential decay function is used to weight the historical information at different moments, and the time interval between the answering unit at the (n + 1)-th moment and each historical unit is introduced during attention calculation;
[0072] First, define the query, key, and value vectors: The query vector Q n+1 is generated using the embedded representation k of the test question at the (n + 1)-th moment n+1 , while the key and value vectors come from the current and all historical learning units; the calculations are as follows:
[0073] Q n+1 = W Q ·k n+1
[0074] K 1:n = [W K ·h1, W K ·h2,..., W K ·h n
[0075] V 1:n = [W V ·h1, W V ·h2,..., W V ·h n
[0076] where represents the embedded test question at the (n + 1)-th moment, is the mapping matrix used to map the test question and the hidden state to the query, key, and value spaces; is the query vector, is the key vector matrix, is the value vector matrix;
[0077] Subsequently, calculate the similarity through the query vector Q n+1 and the key vector K 1:n , and adjust the attention weights by combining with the time decay function exp(-βΔt n+1,i ); the calculation formula for the attention weights is:
[0078]
[0079] where α n+1,i represents the attention weight of the (n + 1)-th unit to the i-th historical unit, Δt n+1,i is the time interval between the (n + 1)-th unit and the i-th historical unit, and β is the hyperparameter of time decay, used to control the influence of the time interval on the attention weights;
[0080] Finally, use the time-weighted attention weights αn+1,i Perform a weighted sum with the value vector V 1:n to generate the potential comprehensive ability state θ of the (n + 1)-th learning unit n+1 :
[0081]
[0082] where represents the speculated potential comprehensive ability state on the (n + 1)-th learning unit
[0083] Furthermore, in step 4, the reaction time prediction formula is expressed as: The loss function is denoted as where t n+1 is the true reaction time at time n + 1;
[0084] The click stream length prediction formula is expressed as: The loss function is denoted as where l n+1 is the true click length at time n + 1;
[0085] The activity stream length prediction formula is expressed as: The loss function is denoted as where a n+1 is the true activity length at time n + 1;
[0086] where is the predicted reaction time, is the predicted click stream length, is the predicted activity stream length, and θ n+1 is the potential ability state of the (n + 1)-th one. In each behavior prediction, a feed-forward neural network layer FNN is respectively set to transform the ability state into the corresponding behavior space, and then the prediction is completed through the respective multi-layer perceptron MLP
[0087] Furthermore, the prediction of the final reaction result in step 4 includes:
[0088] First, obtain the behavior aggregation representation by aggregating the three vectors used to predict the multi-dimensional learning behaviors and the current potential ability state θ n+1 to form a comprehensive behavior representation:
[0089]
[0090] where is the behavior representation of the (n + 1)-th unit, which combines the average information of each behavior feature; FNN(·) represents the feed-forward neural network layer that maps θ n+1 to their respective behavior spaces
[0091] Secondly, based on the aggregated behavior representation R n+1 , a multi-layer perceptron is set to predict the final response result of the learner, and the calculation process is expressed as:
[0092]
[0093] wherein, is the predicted response result, and MLP is the multi-layer perceptron; the loss function is denoted as where r n+1 is the true response at the n+1th moment;
[0094] The final loss function is obtained by weighted summation, that is where λ1, λ2 and λ3 are hyperparameter weights used to adjust the influence of the latter two loss functions.
[0095] Compared with the prior art, the beneficial effects of the present invention are as follows: The existing dynamic cognitive diagnosis is essentially a "time series - knowledge" binary coupling. The binary coupling that weakens the process behavior has limitations and will ultimately lead to the problem of incomplete learning characterization. The method of the present invention proposes a dynamic cognitive diagnosis method of ternary second-order coupling for the learning process, divides the learning units in the learning process into a "knowledge, time series, behavior" triple, obtains the first-order coupling of "knowledge - time series" and "behavior - time series" through a neural network model and behavior feature extraction technology to analyze the changes in knowledge states and behavior states. At the same time, the GLU gating technology is used to set a guess gate, a proficiency gate and a learning gate to realize the complex modeling of the cognitive state changes of the learner during the learning process and obtain the potential ability state of the learner; finally, through the multi-task prediction of behavior features and response results, the diagnosis of the learner's cognitive state and the prediction of the learner's future performance are realized. The present invention can scientifically and comprehensively predict the learning situation of the learner and achieve the purpose of assisting teachers in precise teaching. Description of the Drawings
[0096] Figure 1 is the framework diagram of the method of this example.
[0097] Figure 2 is the flow chart of the method of this example.
[0098] Figure 3 is the structural example diagram of the learning unit.
[0099] Figure 4 is the coupling mechanism comparison diagram between the method of this example and the traditional cognitive diagnosis method. Detailed Embodiments
[0100] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0101] A dynamic cognitive diagnosis method with ternary second-order coupling of knowledge-time-behavior, as Figure 1 shown, is a ternary second-order cognitive diagnosis architecture. The bottom layer uses "learning units" to describe the educational context. The learning units include three types of features: knowledge, time sequence, and behavior. For the knowledge and behavior features, first-order coupling analysis is respectively performed with the time sequence feature to obtain the cognitive state. That is, by performing first-order coupling analysis on knowledge-time sequence, the evolution of the learner's knowledge state can be obtained. By performing first-order coupling analysis on behavior-time sequence, the behavior changes of the learner can be analyzed. Then, the knowledge state and the behavior state are second-order coupled at the time granularity to obtain the comprehensive ability state of the learner. According to the ability state, dual tasks of behavior and response are set to enhance the performance prediction of the learner.
[0102] A dynamic cognitive diagnosis method with ternary second-order coupling of knowledge-time-behavior provided by an embodiment of the present invention includes the following steps:
[0103] Step 1: Taking a certain moment in the learning process as a unit, obtain the learning unit data of the student when taking the assessment, which is composed of knowledge features, behavior features, and time sequence features. According to the chronological order, several learning units can form the learning process sequence data, and based on this, a time sequence data set is constructed. The knowledge features include the questions answered by the student at a certain moment and the concepts related to the questions. The behavior features include the click stream formed by the student repeatedly modifying the answers during the answering process, and the activity stream formed by the student viewing the analysis and video resources after answering. The time sequence features refer to the time interval of each time unit, the repeated time and the number of repetitions of the same concept learning interval, etc. Based on the constructed time sequence data set, perform ternary standardized representation of the learning unit. For the knowledge features, first, each concept related to the current question is embedded, and then the average pooling method is used to obtain the concept representation associated with the question, which is added to the embedded question to obtain the knowledge representation of the current answered question. For the time sequence features, the representation is realized by embedding.
[0104] Step 2: For the knowledge-time sequence features and the behavior-time sequence features, respectively use a neural network model and behavior feature extraction technology to realize the first-order coupling analysis of the "knowledge-time sequence" and "behavior-time sequence" cognitive states. Specifically, for the knowledge-time sequence coupling, input the knowledge and time sequence representations into the LSTM neural network, and set forgetting attenuation and learning enhancement functions to dynamically adjust. For the behavior-time sequence coupling, according to the timestamp information in the click stream and the activity stream, use behavior feature extraction technology to extract five click feature vectors and four activity features, and encode them respectively through addition and the self-attention network with causal masking.
[0105] Step 3: Based on clickstream and activity stream behavior analysis and coding, combined with knowledge state analysis, conduct second-order coupling analysis of the ability state with knowledge-behavior-temporal coupling. Specifically, first, based on the output results of first-order coupling, set the GLU gating function to obtain the initial cognitive state of the student; then, based on clickstream coding, set the guess gate and proficiency gate to weaken and enhance the current cognitive state respectively; subsequently, based on the activity stream, set the GLU-based learning gate to enhance the current cognitive state; finally, through the time-distance attention mechanism, obtain the comprehensive ability state representation of the student.
[0106] Step 4: Based on the obtained comprehensive ability state, enhance the final performance prediction through double prediction tasks of behavior and response. Specifically, set three behavior prediction tasks of reaction time, clickstream length, and activity stream length, transform the ability state into the corresponding behavior space, and then complete the behavior prediction task through a multi-layer perceptron; aggregate the learning behavior prediction and predict the final reaction result at the future moment through a multi-layer perceptron.
[0107] The specific implementation methods of each step are as follows:
[0108] (1) Ternary standardized representation of the learning unit
[0109] As in Figure 2 the "learning unit" layer in, it describes the ternary standardized representation of the learning unit. First, define a learning unit composed of three elements of knowledge, time sequence, and behavior, denoted as LC i =(K i , T i , B i ), which is used to represent the learning event occurring within a certain time unit i. Among them, K i represents the knowledge characteristics under investigation, such as the test questions and the concepts considered, T i represents the time sequence characteristics in the time sequence data, and B i represents the behavior characteristics generated by the learner during the learning process. And the learning process X of the learner consists of N + 1 learning units LC i , denoted as X = {LC1, LC2,..., LC n ,..., LC N+1}, where LC n represents the nth learning unit.
[0110] Example: The structural example of the learning process is as shown in Figure 3 , the learning process consists of multiple learning units, LC i =(K i , T i , B i ), where K iRefers to the knowledge formed by integrating the test questions done by the student at this moment and their associated concepts. A test question q includes multiple concepts c; T i Refers to the temporal features at the i-th moment, such as sequence time interval, review time interval, and knowledge review times; B i Refers to the temporal features at the i-th moment, including click stream and activity stream.
[0111] The calculation process of the standardized representation of the learning unit includes: knowledge feature embedding, temporal feature embedding, and behavior feature embedding.
[0112] (1-1) Knowledge feature embedding
[0113] All test questions and concepts together constitute a knowledge space, represented as a triple (Q, C, τ). Among them, Q is a non-empty set of test questions, C is a non-empty set of concepts, and τ is a mapping from Q to the concept C. For any test question q n ∈Q, τ(q) represents the set of concepts required to solve the test question q. A test question can be associated with multiple concepts. Specifically, it includes three steps: concept embedding, test question embedding, and knowledge comprehensive representation.
[0114] First, concept embedding. Establish the association between test questions and multiple concepts through concept embedding. Specifically, first embed each concept, and then obtain the comprehensive knowledge representation of the test question through average pooling, that is Among them, is the overall embedding representation of the concepts involved in the n-th test question, is the embedding representation of the concepts involved in the n-th test question, d k is the dimension of the embedding, c n,i comes from n the knowledge concept set τ(q n ).
[0115] Secondly, test question embedding. Encode the test questions in an embedding manner. The embedding of the test question index involved in the n-th learning unit is represented as
[0116] Finally, the knowledge comprehensive representation of the test question. Obtained by adding the index embedding of the test question and the concept embedding , that is k n = q n + c n , where represents the knowledge comprehensive representation of the test question.
[0117] (1-2) Temporal feature embedding
[0118] Each learning unit in the learning process occurs at a certain moment and is recorded as the timestamp when the learning event occurs. We use the timestamps in the learning trajectory to extract temporal features, including timestamp distance, time interval, and knowledge review times.
[0119] First, we use the timestamp distance to measure the real distance between learning units. By calculating the timestamp distance between adjacent answers before and after, we denote the sequence time interval as Δs n ;
[0120] Second, the time interval refers to the time interval between the current knowledge answer and the time of the last answer to the same knowledge point, which can be used to reflect the learner's familiarity with a certain knowledge point. We denote the review time interval as Δr n ;
[0121] Then, the knowledge review times refer to the cumulative number of answers to the same knowledge point during the learning process. Similarly, to a certain extent, it can reflect the learner's familiarity with this knowledge point. We denote the knowledge review times as Δc n ; Thus, the temporal feature T is constituted n =(Δs n , Δr n , Δc n ).
[0122] (1 - 3) Extraction of behavioral features
[0123] To achieve the comprehensiveness and diversity of learning behavior characterization, we propose to use clickstream and activity stream to describe the multiple learning behavior features in the learning process.
[0124] First, we extract the clickstream cs n from the learning trajectory. When answering a certain question q n , the learner may modify the answer multiple times due to uncertainty, thus forming several click answers. We extract them in chronological order to form the clickstream where represents the result of the i-th click answer when answering question q n . For example, when answering a multiple-choice question, the learner may first select A, then select B, and finally select C again. Then the clickstream of this learner is cs = {'A', 'B', 'C'}.
[0125] Second, we extract the activity stream as n from the learning trajectory. After each question answer, the learner will conduct thinking and further learning through activity behaviors such as viewing answer analysis, watching video lectures, etc. We extract the learner's post-answer behavior activities to form the activity stream, denoted as where represents the student's answer to question q nThe i-th activity after that, including viewing the explanation or watching the lecture, i.e.,
[0126] (2) First-order coupling analysis of cognitive states
[0127] As Figure 2 In the "first-order coupling" layer in, it describes the first-order coupling analysis of cognitive states. For knowledge, temporal features and behaviors, and temporal features, neural network models and behavior feature extraction techniques are respectively used to achieve the first-order coupling analysis of "knowledge-temporal" and "behavior-temporal" cognitive states. Specifically, for knowledge-temporal coupling, knowledge and temporal representations are input into the LSTM neural network, and forgetting attenuation and learning enhancement functions are set to dynamically adjust; for behavior-temporal coupling, according to the timestamp information in the click stream and activity stream, behavior feature extraction techniques are used to extract five click feature vectors and four activity features, which are respectively encoded by addition and self-attention networks with causal masks.
[0128] (2-1) Change analysis of cognitive states (knowledge-temporal coupling)
[0129] First, we embed the temporal feature T n =(Δs n ,Δr n ,Δc n ) into the same dimension d k After that, the combined representation of the temporal feature is obtained by addition Subsequently, it is concatenated with the knowledge feature embedding to form the final input vector x n , and the combined representation of the learning unit is as follows:
[0130]
[0131] We input the knowledge feature k n and the temporal feature T n into the LSTM neural network together to capture the dynamic relationship between knowledge points and learning time factors.
[0132] Subsequently, the basic update process of the traditional LSTM unit involves a forget gate, an input gate, and an output gate. Based on the original forget gate, we use the timestamp distance to adjust the forgetting degree, and the calculation method is as follows:
[0133] f n =σ(W f ·[h n-1 ,x n +b f )·exp(-α·Δs n )
[0134] Among them, is the output of the adjusted forget gate, is the hidden state of the previous learning unit. σ is the sigmoid activation function, exp(·) is the exponential function, and are the trainable weight matrix and bias term. Δs n is the timestamp distance between adjacent learning units, and α is a tuning parameter used to control the impact of the timestamp distance on the forget gate. Through such a forget gate design, the greater the time interval between learning units, the greater the possibility of forgetting, and the deeper the model forgets previous knowledge.
[0135] Next, to enhance the learning process, we introduce the review time interval Δr n and the knowledge review times Δc n to adjust the output of the input gate, and the calculation method is as follows:
[0136]
[0137] Among them, is the output of the input gate, and are the trainable parameters. γ is a set hyperparameter used to control the impact of the time interval and review times on the input gate. In the input gate, if the learner reviews frequently (i.e., Δc n is large) and has reviewed recently (i.e., Δr n is small), then the learning effect is better, and the weight of the memory input is increased.
[0138] Finally, based on the adjusted forget gate f n and the input gate i n , we update the cell state of the LSTM and update the hidden state, and the calculation method is as follows:
[0139]
[0140] Among them, is the cell state of the LSTM, is the candidate cell state, and are the trainable parameters. We calculate the hidden state through the output gate o n :
[0141] o n = σ(W o · [h n-1 , x n + b o ), h n = o n · tanh(Cn )
[0142] Among them, is the output of the output gate, representing the hidden state of the current learning unit. In summary, we use temporal features to adjust the forget gate and the input gate, and obtain the hidden state of the learning unit based on the LSTM architecture, initially realizing the knowledge-temporal coupling.
[0143] (2-2) Click behavior change analysis (click stream behavior - temporal coupling)
[0144] First, according to the click stream Based on its timestamp information, five behavior change features are extracted: ① Read whether the learner's answer at time t = 0 matches the standard answer, that is, whether the answer is correct, denoted as "initial answer correctness "; ② Read the correctness information at the last modification time t = i in the same way, denoted as "final answer correctness "; ③ Calculate the time difference between the last time and the initial time, denoted as "response time t consumed n ", where, t n = t i - t0; ④ Calculate the length of the entire click stream, denoted as "click stream length l n = |cs n |; ⑤ Analyze the behavior changes of the learner in the click stream. The learner may finally modify the answer, denoted as the "modified" type, or may change back to the initial answer after repeated changes, denoted as "hesitation". This variable is called "click stream type p n ".
[0145] Secondly, combine these five embedding vectors into a vector and map it to a unified vector space Among them is the embedding representation of the activity stream cs n at the nth answer, D k is the dimension of the embedding, and emb represents the corresponding embedding layer function.
[0146] (2-3) Activity behavior change analysis (activity stream behavior - temporal coupling)
[0147] First, according to the activity stream Based on its timestamp information, four behavior change features are extracted: ① Read the type of the current activity, denoted as "activity type ", where, ② Calculate the duration from the start to the end of the activity, denoted as "total activity time "; ③ Obtain the number of times the activity is to view the analysis (that is ), denoted as "number of times of viewing the analysis ”; ④ Obtain the number of times the activity is to view the video (i.e., ), and record the “number of video operations ”.
[0148] Subsequently, an embedding matrix is used to represent the four activity flow features. Each row of the matrix corresponds to an activity, i.e., where is the embedding representation of the i-th activity in the activity flow as n at the n-th answer. Combine the corresponding activities in the activity flow to form the activity flow where la n is the length of the activity flow.
[0149] (3) Second-order coupling analysis of ability state
[0150] Figure 2 The “second-order coupling” layer in
[0151] describes the second-order coupling analysis of the ability state. Based on the click stream and activity flow behavior analysis and coding, combined with the knowledge state analysis, the second-order coupling analysis of the ability state of knowledge-behavior-temporal coupling is realized. Specifically, first, based on the output result of the first-order coupling, set the GLU gating function to obtain the initial cognitive state of the student; then, set the guess gate and proficiency gate through the click stream coding to weaken and enhance the current cognitive state respectively; subsequently, based on the activity flow, set the GLU-based learning gate to enhance the current cognitive state; finally, through the time distance attention mechanism, obtain the comprehensive ability state representation of the student.
[0151] (3-1) Cognitive state update based on click stream
[0152] First, the knowledge state of the learner can be controlled by the gating function GLU(), and its function can be expressed as:
[0153]
[0154] where is the initial cognitive state, is the output result of the first-order coupling of the cognitive state in the previous step, and represent the corresponding weight and bias parameters.
[0155] Second, according to the modification operations of the student in the click stream, calculate the guessing degree during the learner's answering process. If there is a guess, the knowledge mastery level should subtract the error caused by the guess. Adjust the cognitive state by setting the “guess gate”, and the function is expressed as:
[0156] K′ n = K n - GLU g (CS n ·Wg +b g ),
[0157] Among them, is the knowledge state after passing through the guess gate, and GLU g represents the gating function of the guess gate; represents the embedded representation of the click stream, and are the weight and bias parameters of the guess gate.
[0158] Thirdly, according to the time attribute in the click stream, infer the proficiency of the learner in mastering knowledge. A reasonable allocation of the learner's problem-solving time can enhance the student's proficiency, and then the knowledge mastery level can be added with the knowledge enhancement brought by the proficiency. Set the "proficiency gate" to enhance the knowledge state, and the function is expressed as:
[0159] K″ n =K′ n +GLU p (CS n ·W p +b p ),
[0160] Among them, is the knowledge state after passing through the proficiency gate, and GLU p represents the gating function of the proficiency gate; and are the weight and bias parameters of the proficiency gate.
[0161] (3 - 2) Cognitive state enhancement based on the activity stream
[0162] First of all, learners have various learning activities, such as viewing analysis and watching videos. To uniformly represent learning activities, we use the method of average pooling to obtain the average value of the embeddings of each activity in the activity stream, so as to aggregate all activity information. The calculation method is as follows:
[0163]
[0164] Among them, is the aggregated representation of the activity stream of the nth answer, is the length of the activity stream as n , that is, the number of activities, represents the embedded representation of the i-th activity in the nth answer.
[0165] Secondly, through the learning activity behavior after answering, the learner's knowledge mastery level can be enhanced. That is, the learner's knowledge level needs to be enhanced after learning activities. Set the "learning gate" to simulate the improvement degree of its high-order cognitive state. The function is expressed as:
[0166] K″′ n = K″ n + GLU l (AS n )
[0167] where represents the knowledge state enhanced by the learning gate, is the knowledge state before entering the learning gate, and GLU l is the learning gate function, which is used to enhance the influence of the activity stream on the knowledge state.
[0168] (3-3) Inference of potential ability based on time-distance attention mechanism
[0169] Based on the current cognitive state h n obtained from (2-1) and the knowledge state change K″′ n caused by the behavior in (3-2), we further infer the potential comprehensive ability state θ n+1 of the student through the time-distance attention mechanism. To introduce the influence of time factors, we set the query vector as the test question representation k n+1 at the n+1th moment, and use the exponential decay function to weight the historical information at different moments, introducing the time interval between the answer unit at n+1 and each historical unit during the attention calculation.
[0170] First, we define the query, key, and value vectors. The query vector Q n+1 is generated using the test question embedding representation k n+1 at the n+1th moment, while the key and value vectors come from the current and all historical learning units. The calculation is as follows:
[0171] Q n+1 = W Q · k n+1
[0172] K 1:n = [W K · h1, W K · h2,..., W K · h n
[0173] V 1:n = [W V · h1, W V · h2,..., W V · h n
[0174] where represents the test question embedding at the n+1th moment, is the mapping matrix, which is used to map the test questions and hidden states to the query, key, and value spaces. is the query vector, is the key vector matrix, is the value vector matrix.
[0175] Subsequently, the similarity is calculated through the query vector Q n+1 and the key vector K 1:n and the attention weights are adjusted by combining with the time decay function exp(-βΔt n+1,i ). The calculation formula of the attention weights is:
[0176]
[0177] where α n+1,i represents the attention weight of the (n + 1)-th unit to the i-th historical unit, Δt n+1,i is the time interval between the (n + 1)-th unit and the i-th historical unit, and β is the hyperparameter of time decay, which is used to control the influence of the time interval on the attention weights.
[0178] Finally, we use the time-weighted attention weights α n+1,i and the value vector V 1:n to perform weighted summation to generate the potential comprehensive ability state θ n+1 of the (n + 1)-th learning unit:
[0179]
[0180] where represents the potential comprehensive ability state on the speculated (n + 1)-th learning unit.
[0181] (4) Dual-task prediction of behavior and response
[0182] As in Figure 2 the performance prediction part, based on the obtained comprehensive ability state, the final performance prediction is enhanced through the dual prediction tasks of behavior and response. Specifically, first, three behavior prediction tasks of reaction time, click stream length, and activity stream length are set, the ability state is transformed into the corresponding behavior space, and then the behavior prediction tasks are completed through a multi-layer perceptron; finally, the learning behavior prediction is aggregated, and the prediction of the final reaction result at the future moment is realized through a multi-layer perceptron.
[0183] (4-1) Multivariate learning behavior recognition and prediction
[0184] Predict the learner's behavior performance according to the learner's ability state, including reaction time, click stream length, and activity stream length.
[0185] The reaction time prediction formula is expressed as: The loss function is denoted as where tn+1 is the true reaction time at the (n + 1)-th moment.
[0186] The clickstream length prediction formula is expressed as: The loss function is denoted as where l n+1 is the true click length at the (n + 1)-th moment.
[0187] The activity stream length prediction formula is expressed as: The loss function is denoted as where a n+1 is the true activity length at the (n + 1)-th moment.
[0188] where is the predicted reaction time, is the predicted clickstream length, is the predicted activity stream length, and θ n+1 is the potential ability state at the (n + 1)-th. In each behavior prediction, a feedforward neural network layer FNN is respectively set to transform the ability state into the corresponding behavior space, and then the prediction is completed through their respective multi-layer perceptrons MLP.
[0189] (4 - 2) Prediction of Reaction Results Enhanced by Behavior Analysis
[0190] First, we obtain the behavior aggregation representation by aggregating the three vectors used to predict the multi-variate learning behavior in the previous step, as well as the current potential ability state θ n+1 , to form a comprehensive behavior representation:
[0191]
[0192] where, is the behavior representation of the (n + 1)-th unit, which combines the average information of each behavior feature. FNN(·) represents the feedforward neural network layer that maps θ n+1 to their respective behavior spaces.
[0193] Second, based on the aggregated behavior representation R n+1 , a multi-layer perceptron is set to predict the final reaction result of the learner, and the calculation process is expressed as:
[0194]
[0195] where, is the predicted reaction result, and MLP is the multi-layer perceptron. The loss function is denoted as where r n+1 is the true reaction at the (n + 1)-th moment.
[0196] The final loss function is obtained by weighted summation, that is where λ1, λ2, and λ3 are hyperparameter weights used to adjust the influence of the latter two loss functions.
[0197] A ternary second-order coupled dynamic cognitive diagnosis method for the learning process proposed by the present invention divides the learning units in the learning process into a "knowledge, time sequence, behavior" triple, and introduces a variety of deep learning technologies to achieve second-order modeling of cognition and ability. Its comparison with existing static cognitive diagnosis and dynamic cognitive diagnosis methods is shown in Table 1. As shown in Table 1, the static cognitive diagnosis method only uses the knowledge features in the educational context and does not consider the time sequence and behavior features in the real educational context. It is applicable to single learning situations and is essentially a unary coupling of knowledge; the dynamic cognitive diagnosis method considers knowledge features and learning process sequence features and is applicable to multiple learning situations, but ignores the time interval attribute and behavior process features in the learning process. It is essentially a weak coupling of knowledge and time sequence features, as shown in (a) of Figure 4 There are limitations in weakening the binary coupling of process behaviors, which will ultimately lead to the problem of incomplete learning characterization; while the ternary second-order coupled cognitive diagnosis method proposed by the present invention considers the second-order coupling of three features, as shown in (b) of Figure 4 It can scientifically and comprehensively model the cognitive process of learners and achieve accurate prediction of learners' future performance.
[0198] Table 1 Comparison table of coupling mechanisms between the method of this example and traditional cognitive diagnosis methods
[0199]
[0200] The content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0201] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A knowledge-sequence-behavior three-element second-order coupling dynamic cognitive diagnosis method, characterized by: The steps include: Step 1: Obtain the learning unit data of students when they take the assessment, including the three elements of knowledge, behavior, and time sequence. According to the chronological order, several learning units are combined into learning process sequence data to construct a time series data set. Based on the constructed time series data set, obtain the ternary standardized representation of the learning unit; Among them, knowledge features include the test questions answered by students at a certain moment and the concepts related to the test questions; behavioral features include the click stream formed by students repeatedly modifying their answers during the answering process, and the activity stream formed by students viewing analysis and video resources after answering; temporal features refer to the time interval between each time unit, the time and number of repetitions between learning the same concept; Step 2: For the represented knowledge and behavior characteristics, conduct first-order coupling analysis with the time series characteristics to obtain the cognitive state, that is, conduct first-order coupling analysis on knowledge-time series to obtain the evolution of the learner's knowledge state, and conduct first-order coupling analysis on behavior-time series to analyze the learner's behavior changes; Specifically, for knowledge-time series coupling, knowledge and time series representations are input into the LSTM neural network, and forgetting reduction and learning enhancement functions are set to dynamically adjust; for behavior-time series coupling, according to the timestamp information in the click stream and activity stream, the behavior feature extraction technology is used to extract the click feature vector and activity feature, and they are encoded by addition and self-attention network with causal mask respectively; Step 3: perform second-order coupling of knowledge state and behavior state at the time granularity to obtain the learner's comprehensive ability state; Step 4: Based on the acquired comprehensive ability state representation, the final performance prediction is enhanced through the dual prediction task of behavior and reaction; Specifically, it includes setting three behavior prediction tasks: reaction time, click stream length, and activity stream length, converting the capability state to the corresponding behavior space, and then completing the behavior prediction task through a multi-layer perceptron; aggregate learning behavior prediction, and realizing the prediction of the final reaction result at future moments through a multi-layer perceptron.
2. The knowledge-sequence-behavior three-element second-order coupled dynamic cognitive diagnosis method as claimed in claim 1, characterized in that: The three-element standardized representation of the learning unit specifically includes: (1-1) Knowledge Feature Embedding All test questions and concepts together constitute the knowledge space, which is represented by a triple (Q, C, τ), where Q is a non-empty set of test questions, C is a non-empty set of concepts, and τ is a mapping from Q to concept C. For any test question q n ∈Q, τ(q) represents the set of concepts required to solve question q. One question can be associated with multiple concepts. Specifically, it includes three steps: concept embedding, question embedding, and comprehensive knowledge representation. First, concept embedding: concept embedding is used to establish the association between test questions and multiple concepts. Specifically, each concept is first embedded, and then the comprehensive knowledge representation of the test question is obtained through average pooling, that is, in, is the overall embedding representation of the concepts involved in the nth test question, is the embedding representation of the concepts involved in the nth test question, d k is the dimension of the embedding, c n,i From q n The knowledge concept set τ(q n ); Secondly, question embedding: the questions are encoded in an embedded manner, and the embedding of the question index involved in the nth learning unit is represented as Finally, the comprehensive knowledge representation of the test question: embedded by the index of the test question and concept embedding Add them together to get k n =q n +c n ,in A comprehensive representation of the knowledge of the test questions; (1-2) Temporal feature embedding Each learning unit in the learning process occurs at a certain moment and is recorded as the timestamp when the learning event occurs. The timestamps in the learning trajectory are used to extract temporal features, including timestamp distance, time interval, and number of knowledge reviews. First, the timestamp distance is used to measure the real distance between learning units. By calculating the timestamp distance between the adjacent answers, the time interval Δs is recorded. n ; Secondly, the time interval refers to the time between the current knowledge answer and the last answer of the same knowledge point, which is recorded as the review time interval Δr n ; Then, the number of knowledge reviews refers to the cumulative number of answers to the same knowledge point during the learning process, and the number of knowledge reviews is Δc n ; Thus, the timing feature T n =(Δs n ,Δr n ,Δc n ); (1-3) Behavioral feature extraction Use clickstream and activity stream to describe the multi-faceted learning behavior characteristics in the learning process; First, extract the clickstream cs from the learning trajectory n ;When answering a test question n When learning, learners may modify their answers many times due to uncertainty, thus forming several click answers, which are extracted in chronological order to form a click stream. in, Indicates answering the test question q n The result of the i-th click answer; Second, extract the activity flow from the learning trajectory as n ; After completing each test, learners will think and learn further by looking at the answer analysis and watching the video lecture. The learners' behavior after answering the questions is extracted to form an activity flow, recorded as in Indicates that the student answers the test question q n The i-th activity after that includes viewing the explanation or watching the video lecture, i.e.
3. The knowledge-sequence-behavior three-element second-order coupled dynamic cognitive diagnosis method as claimed in claim 2, characterized in that: The specific implementation process of knowledge-timing coupling is as follows: First, the time series feature T n =(Δs n ,Δr n ,Δc n ) are embedded into the same dimension d k After that, the comprehensive representation of the time series features is obtained by adding Then embed it with knowledge features Concatenate to form the final input vector x n , the comprehensive representation of the learning unit is as follows: The knowledge feature k n and the time series characteristics T n They are input into the LSTM neural network together to capture the dynamic relationship between knowledge points and learning time factors; Then, based on the original forget gate, the timestamp distance is used to adjust the forgetting degree. The calculation method is as follows: f n =σ(W f ·[h n-1 ,x n ]+b f )·exp(-α·Δs n ) in, is the adjusted forget gate output, is the hidden state of the previous learning unit; σ is the sigmoid activation function, exp(·) is the exponential function, and is the trainable weight matrix and bias term; Δs n is the timestamp distance between adjacent learning units, and α is an adjustment parameter used to control the impact of timestamp distance on the forget gate; Next, in order to enhance the learning process, the review interval Δr is introduced n and knowledge review times Δc n To adjust the output of the input gate, the calculation method is as follows: in, is the input gate output, and is a trainable parameter; γ is a set hyperparameter used to control the impact of time interval and number of reviews on the input gate; Finally, based on the adjusted forget gate f n and input gate i n , update the cell state of LSTM, and update the hidden state. The calculation method is as follows: in, is the cell state of LSTM, is a candidate cell state, and is a trainable parameter; through the output gate o n Compute the hidden state: o n =σ(W o ·[h n-1 ,x n ]+b o ),h n =o n ·tanh(C n ) in, is the output of the output gate, Represents the hidden state of the current learning unit.
4. The knowledge-sequence-behavior three-element second-order coupled dynamic cognitive diagnosis method as claimed in claim 2, characterized in that: Behavior-timing coupling includes clickstream behavior-timing coupling and activitystream behavior-timing coupling; The specific implementation process of click flow behavior-sequence coupling is as follows: First, according to the click stream According to its timestamp information, five behavioral change characteristics are extracted: ① Read whether the learner's answer at time t=0 matches the standard answer, that is, whether the answer is correct or not, and record "the correct or incorrect first answer" ”; ② Using the same method, read the correctness information of the last modification at time t=i, and record the correctness of the final answer ”; ③ Calculate the difference between the final time and the initial time, and record the reaction time t n ”, where t n =t i -t0; ④ Calculate the length of the entire click stream, record "click stream length l n =|cs n ⑤ Analyze the behavioral changes of learners in the clickstream. The learner may eventually modify the answer, which is recorded as "modification" type, or may change back to the initial answer after repeated changes, which is recorded as "hesitation". This variable is called "clickstream type p n ”; Secondly, these five embedding vectors are combined into one vector and mapped to a unified vector space in is the activity flow cs when the nth answer is given n The embedding representation of k is the embedding dimension, and emb represents the corresponding embedding layer function; The specific implementation process of activity flow behavior-timing coupling is as follows: First, according to the activity flow According to its timestamp information, four behavior change characteristics are extracted: ① Read the type of the current activity and record the "activity type ",in, ② Calculate the time from the start to the end of the activity and record the total activity time "; ③ Get the activity is to view the analysis The number of times, record "View the number of parsing "; ④ The acquisition activity is to watch the video The number of times the video is operated is recorded as " ” Then, an embedding matrix is used to represent the four activity flow features, where each row of the matrix corresponds to an activity, namely in is the activity flow when answering the nth question n The embedding representation of the i-th activity of is; the corresponding activities in the activity stream are combined to form an activity stream Among them n is the length of the active stream.
5. The knowledge-sequence-behavior three-element second-order coupled dynamic cognitive diagnosis method as claimed in claim 1, characterized in that: The specific implementation of step 3 includes: (3-1) Based on the output results of the first-order coupling, the GLU gating function is set to obtain the student's initial cognitive state, and the guessing gate and the proficiency gate are set through click stream coding to weaken and enhance the current cognitive state respectively; (3-2) Setting GLU-based learning gates based on activity flow to enhance the current cognitive state; (3-3) Through the time-distance attention mechanism, the comprehensive ability status representation of students is obtained.
6. The knowledge-sequence-behavior three-element second-order coupled dynamic cognitive diagnosis method as claimed in claim 5, characterized in that: The specific implementation of (3-1) is as follows: First, the learner's knowledge state is controlled by the gate function GLU(), which can be expressed as: in, is the initial cognitive state, is the output result of the first-order coupling of the cognitive state in the previous step, and Represents the corresponding weight and bias parameters; Second, based on the students’ modification operations in the clickstream, the degree of guessing during the learner’s answering process is calculated; the cognitive state is adjusted by setting a guessing gate, and the function is expressed as: K′ n =K n -GLU g (CS n ·W g +b g ), in, is the state of knowledge after passing through the guessing gate; represents the embedding representation of the click stream, and To guess the weight and bias parameters of the gate; Third, based on the time attributes in the clickstream, the learner’s proficiency in knowledge is inferred; the knowledge state is enhanced by setting a proficiency gate, and the function is expressed as: in, It is the state of knowledge after passing the proficiency gate; and are the weight and bias parameters of the proficiency gate.
7. The knowledge-sequence-behavior three-element second-order coupled dynamic cognitive diagnosis method according to claim 6, characterized in that: The specific implementation of (3-2) is as follows: In order to unify the representation learning activities, the average pooling method is used to obtain the average value of each activity embedding in the activity stream, thereby aggregating all activity information. The calculation method is as follows: in, is the aggregate representation of the activity flow of the nth answer, For activity streams n The length of , that is, the number of activities, represents the embedding representation of the i-th activity in the n-th answer; Secondly, the learning activities after answering the questions can enhance the learner's knowledge mastery, and the learning gate is set to simulate the improvement of their high-order cognitive state; the function is expressed as: K″′ n =K″ n +GLU l (AS n ), in, represents the knowledge state after being enhanced by the learning gate, GLU is the knowledge state before entering the learning gate. l It is a learning gate function used to enhance the influence of activity flow on knowledge state.
8. The knowledge-sequence-behavior three-element second-order coupled dynamic cognitive diagnosis method according to claim 7, characterized in that: The specific implementation of (3-3) is as follows: Based on the current cognitive state h n And (3-2) get the knowledge state change K″′ caused by the behavior n , and further infer the potential comprehensive ability state of students θ through the time distance attention mechanism n+1 ; In order to introduce the influence of time factors, the query vector is set to the question representation k at time n+1 n+1 , and use an exponential decay function to weight the historical information at different times, and introduce the time interval between the n+1 answer unit and each historical unit in the attention calculation; First, define the query, key, and value vectors: n+1 Use the test question embedding representation k at time n+1 n+1 Generate, and the key and value vectors come from the current and all historical learning units; the calculation is as follows: Q n+1 =W Q ·k n+1 K 1:n =[W K ·h1,W K ·h2,...,W K ·h n ] V 1:n =[W V ·h1,W V ·h2,...,W V ·h n ] in, represents the test embedding at the n+1th moment, is the mapping matrix, used to map questions and hidden states to query, key and value spaces; is the query vector, is the key vector matrix, is a matrix of value vectors; Then, through the query vector Q n+1 and the key vector K 1:n Calculate the similarity and combine it with the time decay function exp(-βΔt n+1,i ) adjusts the attention weight; the calculation formula of attention weight is: Among them, α n+1,i represents the attention weight of the n+1th unit to the i-th historical unit, Δt n+1,i The time interval between the n+1th unit and the i-th historical unit. β is a hyperparameter of time decay, which is used to control the effect of the time interval on the attention weight. Finally, the time-weighted attention weight α is used n+1,i With the value vector V 1:n Perform weighted summation to generate the potential comprehensive ability state θ of the n+1th learning unit n+1 : in, Represents the potential comprehensive ability state of the inferred n+1th learning unit.
9. The knowledge-sequence-behavior three-element second-order coupled dynamic cognitive diagnosis method as claimed in claim 1, characterized in that: In step 4, the reaction time prediction formula is expressed as: The loss function is denoted as where t n+1 is the actual reaction time at time n+1; The click stream length prediction formula is expressed as: The loss function is denoted as Among them l n+1 is the actual click length at time n+1; The activity flow length prediction formula is expressed as: The loss function is denoted as where a n+1 is the actual activity length at time n+1; in is the predicted reaction time, is the predicted clickstream length, is the predicted activity flow length, θ n+1 It is the potential ability state of the n+1th one. In each behavior prediction, a feedforward network layer FNN is set to convert the ability state into the corresponding behavior space, and then the prediction is completed through the respective multi-layer perceptron MLP.
10. The knowledge-sequence-behavior three-element second-order coupled dynamic cognitive diagnosis method according to claim 9, characterized in that: The prediction of the final reaction results in step 4 includes: First, we obtain the behavior aggregation representation by aggregating the three vectors used to predict the multivariate learning behavior and the current potential ability state θ n+1 , forming a comprehensive behavioral representation: in, is the behavior representation of the n+1th unit, combining the average information of each behavior feature; FNN(·) represents the n+1 Feed-forward network layers that map to their respective behavior spaces; Secondly, based on the aggregated behavior representation R n+1 , set up a multilayer perceptron to predict the learner's final response result, and the calculation process is expressed as: in, is the predicted reaction result, MLP is a multi-layer perceptron; the loss function is recorded as where r n+1 is the true response at time n+1; The final loss function is obtained by weighted summation, that is, Among them, λ1, λ2 and λ3 are hyperparameter weights used to adjust the impact of the latter two loss functions.
Citation Information
Patent Citations
Public resource demand prediction method based on coupling relation learning
CN116362383A
Methods and systems for state navigation
US20220245109A1