Cognitive diagnosis method for heterogeneous graph learning based on unbalanced perception
By introducing a heterogeneous graph learning method of imbalance perception into the cognitive diagnostic model, using graph neural network and Transformer architecture to extract the characteristics of students, exercises and knowledge concepts, the problem of existing models failing to fully consider synergistic relationships and data imbalances is solved, and higher diagnostic accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510198137.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-22
- Publication Date
- 2025-06-13
AI Technical Summary
The existing cognitive diagnostic model fails to fully consider the synergistic relationship between students and the imbalance of data distribution, resulting in inaccurate diagnostic results.
Using a heterogeneous graph learning method based on imbalance perception, by constructing heterogeneous graphs between students, exercises and knowledge concepts, using graph neural networks and Transformer architecture to extract local and global features of nodes, establish a diagnostic prediction model, and optimize model parameters to improve diagnostic accuracy.
It effectively alleviates the problem of unbalanced data distribution, improves the quality of characterization learning, and significantly improves the accuracy and reliability of cognitive diagnostic models.
Smart Images

Figure CN120144950A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent education cognitive diagnosis, and specifically to a cognitive diagnosis method based on heterogeneous graph learning with unbalanced perception. Background Art
[0002] In recent years, intelligent education has gradually become an important research field, and cognitive diagnosis technology plays a key role in it. By analyzing students' answering records, cognitive diagnosis can evaluate students' mastery of specific knowledge concepts and provide data support for personalized learning. These diagnostic results not only help to improve teaching strategies but also provide services such as course recommendation, exercise push, and student adaptive testing for online education platforms.
[0003] With the successful application of neural networks, cognitive diagnosis technology has also achieved new breakthroughs. More and more researchers have started to use neural networks to construct new cognitive diagnosis architectures or improve the performance of models by enhancing the representation of input data. In particular, methods based on graph neural networks have shown excellent performance because they can fully capture the complex relationships between students and exercises, and between exercises and concepts. The advantage of graph neural networks is that they can learn the relationships between different types of nodes, even if these relationships are distant, so as to comprehensively capture multi-dimensional information that affects students' performance.
[0004] However, existing cognitive diagnosis models still face some challenges. The main problem is that the models fail to fully consider the influence of some implicit relationships, such as the collaborative relationship between students. This may lead to information propagation obstacles, thus affecting the accurate assessment of students' abilities by the models. In addition, existing models often ignore the imbalance in the quantity and relationships among students, exercises, and concepts, and this imbalance may lead to inaccurate diagnostic results. Therefore, how to solve these problems and further improve the accuracy and robustness of the models remains a difficult problem to be overcome in the field of cognitive diagnosis. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a cognitive diagnosis method based on heterogeneous graph learning with unbalanced perception, which can solve the existing problems.
[0006] To achieve the above purpose, the technical solution of the present invention is as follows:
[0007] The present invention is realized through the following technical solutions: A cognitive diagnosis method based on heterogeneous graph learning with unbalanced perception, comprising the following steps:
[0008] S1. Obtain data and perform preprocessing to respectively form multiple data sets regarding students, exercises, knowledge concepts, and answering records;
[0009] S2. Obtain the explicit relationships and implicit relationships of the dataset, and construct a heterogeneous graph based on the dataset, explicit relationships, and implicit relationships;
[0010] S3. Use two different graph neural network models to process the imbalance-aware local feature extraction module, and learn the local structural features and local semantic features of the nodes in the heterogeneous graph;
[0011] S4. Use an improved Transformer architecture to process the global representation learning module, and learn global context features from the local structural features and local semantic features;
[0012] S5. Establish a diagnostic prediction model based on the global context features, and optimize the diagnostic prediction model based on the constructed loss function;
[0013] S6. Use the optimized diagnostic prediction model to predict the cognitive diagnosis results.
[0014] Further, in S1, the multiple datasets include:
[0015] The student set S = {s 1 , s 2 , …, s N}, where N is the total number of students;
[0016] The exercise set E = {e 1 , e 2 , …, e M}, where M is the total number of exercises;
[0017] The knowledge concept set C = {c 1 , c 2 , …, c K}, where K is the total number of knowledge concepts;
[0018] The answer record set R = {(s i , e j , r ij ) | s i ∈ S, e j ∈ E, r ij ∈ {0, 1}}, where r ij represents the answering situation of student s i to exercise e j .
[0019] Further, the obtaining of the explicit relationships and implicit relationships of the dataset, and the construction of the heterogeneous graph based on the dataset, explicit relationships, and implicit relationships specifically include:
[0020] The explicit relationships include the interaction relationship between students and exercises, the inclusion relationship between exercises and knowledge concepts, and the dependency relationship between knowledge concepts;
[0021] The implicit relationship includes the collaborative relationship between students;
[0022] The construction of the heterogeneous graph includes defining the heterogeneous graph = {V, E}, where the node set V = S ∪ E ∪ C, and the edge set E = R se ∪R ec ∪R cc ∪R ss , S represents the set of students; E represents the set of exercises; C represents the set of knowledge concepts; R se represents the interaction relationship between students and exercises, R ec represents the inclusion relationship between exercises and knowledge concepts, R cc represents the dependency relationship between knowledge concepts, R ss represents the collaborative relationship between students.
[0023] Furthermore, the calculation of the implicit relationship, that is, the collaborative relationship between students, is as follows:
[0024] Calculate the relationship between students s i and s l The calculation formula is shown in the following formula (1): The calculation formula is shown in the following formula (1):
[0025]
[0026] In formula (1), T i represents the set of exercises correctly answered by student s i , F i represents the set of exercises wrongly answered by student s i , T l and F l are the same in this regard; J(·) represents the Jaccard similarity of two sets; rt il represents the similarity of the answering time of student s i and student s l on the set of exercises U il correctly answered in common; α is the weight controlling the influence of correct and wrong answers; β is the weight controlling the influence of time;
[0027] where the calculation process of rt il is as follows:
[0028]
[0029] In formula (2), t ij represents the time taken by student s i to answer exercise e jThe answering time, represents the average answering time of student s i for answering questions;
[0030] represents the average answering time of exercise e j that has been answered; represents the answering time of the later student s i for exercise e j ;
[0031] Using formula (2), calculate the debiased time of student s i and student s l respectively on the set U of exercises that are both answered correctly il to obtain the answering time similarity rt between the two students il ;
[0032]
[0033] In formula (3), EucliD(·) represents the Euclidean distance between two vectors, represents the debiased time vector of student s i on the set of exercises U il , and t l Similarly;
[0034] Finally, if is greater than a set threshold η, then let represent that student s i and student s l are similar; for student s i , if student s l is the last P students similar to it, then let represent that student s i and student s l are the least similar;
[0035] Thus, the collaborative relationship between students is obtained
[0036] Furthermore, the use of two different graph neural network models to process the unbalanced perception local feature extraction module and learn the local structural features and local semantic features of the nodes in the heterogeneous graph specifically includes:
[0037] S3.1. Obtain the initial embedding vectors of the local structural features and local semantic features, which are calculated and determined by the following formula (4):
[0038]
[0039] In formula (4), represents student si Practice e j and knowledge concept c k Initial structure embedding vector of the node; W S ∈R N×d W E ∈R M×d W C ∈R K×d is a trainable matrix mapped to the d-dimensional space; Represents student s i Practice e j and knowledge concept c k One-hot encoding; Represents the initial semantic embedding vector of the node; Represents the initial semantic embedding vector of the relationship; Em∈R 8×8 Represents the one-hot encoding matrix;
[0040] S3.2. Generate virtual knowledge concept nodes based on the initial embedding vectors; Calculate through the following formula (5):
[0041]
[0042] In formula (5), is the initial embedding of knowledge concept nodes u and k, N k is the knowledge concept node closest to k in the embedding space. Based on the knowledge point N k and k, generate a new virtual knowledge point node k v , where δ is a random value following a uniform distribution;
[0043] Use formula (6) to set the inclusion relationship v and the dependency relationship of the virtual knowledge point k as the union of the relationships of knowledge point k and N k ;
[0044]
[0045] In formula (6), r = 1 indicates the existence of the corresponding relationship;
[0046] S3.3. Calculate the local structure feature and the local semantic feature; Determine through the following formulas (7) and (8);
[0047]
[0048] In formula (7), 1 ≤ l ≤ L, representing the number of aggregation layers; Agg(·) represents the aggregation function in the graph neural network; v represents a student, practice, or knowledge concept node, N v represents the set of neighbor nodes of v; σ represents the activation function, Wl-1 ∈R d×d represents the trainable matrix of the (l - 1)-th layer;
[0049]
[0050] In formula (8), φ(·) represents a mapping function. If it contains a single node, it maps to the type of that node.
[0051] Furthermore, the improved Transformer architecture is used to process the global representation learning module, and global context features are learned from the local structural features and local semantic features; it includes
[0052] S4.1. Apply the personalized PageRank technique to the constructed heterogeneous graph to obtain the PPR matrix P ∈ R |V|×|V| , P ij represents the importance of node j to i;
[0053] For each node v ∈ V, according to the matrix P, extract the set of the top T student nodes with the highest values Similarly, obtain the set of exercise nodes and the set of knowledge point concept nodes which together form the global relevant node set SP of node v v ={s v , e v , c v};
[0054] S4.2. For each node v and its global relevant node set SP v , process and calculate its local structural features and local semantic features respectively through an improved Transformer architecture;
[0055] Specifically, use formula (9) and formula (10) to calculate the attention coefficient matrices of the structural features and semantic features respectively;
[0056]
[0057] In formula (9), is the attention parameter matrix, is the trainable vector; H g ∈R |SPv+1|×d represents the local structural features of node v and its global relevant node set SP v ; ⊕ represents tensor addition;
[0058]
[0059] In formula (10), is the attention parameter matrix;
[0060] Using formula (11), the fused attention weight matrix Att is obtained l ;
[0061]
[0062] In formula (11), λ represents the weight controlling semantic information;
[0063] Using formula (12), the fused features are obtained
[0064]
[0065] In formula (12), LN(·) represents layer normalization, MHA(·) represents multi-head attention, and W f l is the trainable parameter matrix, and FFN(·) represents the feed-forward network layer;
[0066] After L-layer propagation, is the final representation of node v. All nodes in the heterogeneous graph are operated through the above steps to obtain the final representation.
[0067] Furthermore, the diagnostic prediction model is established according to the global context features; and the diagnostic prediction model is optimized based on the constructed loss function; including:
[0068] S5.1. Determine the ability of student s i using formula (13) of exercise e j difficulty and exercise discrimination
[0069]
[0070] Predict the answering score of student s i for exercise j using formula (14);
[0071]
[0072] In formula (14), represents element-wise multiplication of vectors, and Q j represents the j-th row of the Q-matrix, representing the knowledge concepts included in exercise e j ;
[0073] S5.2. Construct the loss function L; specifically as shown in the following formula (15):
[0074]
[0075] In formula (15), r i represents the true value i answered by student s and represents the predicted value;
[0076] S5.3. Use the gradient descent method to perform backpropagation training on the loss function L to update the model parameters until the loss function L converges and stops, and finally obtain an optimized training model.
[0077] Furthermore, in S1, the obtaining data and performing preprocessing are specifically as follows: obtain the total number of students, the total number of exercises, the total number of knowledge concepts, and the corresponding answer record situation; perform data cleaning, data transformation, feature engineering, and data integration on the obtained data.
[0078] A cognitive diagnosis device based on unbalanced perception of heterogeneous graph learning, characterized in that: it includes a processor and a memory; the memory is used to store programs; the processor executes the programs to implement the method described in any one of the above.
[0079] A computer-readable storage medium, characterized in that: the storage medium stores a program, and the program is executed by a processor to implement the method described in any one of the above.
[0080] Compared with the prior art, the beneficial effects of the present invention include:
[0081] The cognitive diagnosis method based on unbalanced perception of heterogeneous graph learning of the present invention realizes more accurate representation learning by introducing an improved heterogeneous graph neural network and a Transformer architecture. First of all, according to the historical answer records and answer times of students, the collaborative relationship between students is constructed, so as to extract more comprehensive student feature representations; secondly, combined with the unbalanced perception of heterogeneous graph learning strategy and the improved Transformer model, the problem of data distribution imbalance is effectively alleviated in the heterogeneous graph. Through the method of the present invention, the local information and global information of the nodes in the graph can be captured simultaneously, significantly improving the quality of representation learning, thereby effectively improving the accuracy and reliability of the cognitive diagnosis model. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] The disclosure of the present invention will be described with reference to the accompanying drawings. It should be understood that the drawings are only for illustrative purposes and are not intended to limit the protection scope of the present invention. In the drawings, the same reference numerals are used to refer to the same components. Among them:
[0083] Figure 1 is a schematic flow chart of a cognitive diagnosis method based on unbalanced perception of heterogeneous graph learning of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0084] It is easy to understand that according to the technical solution of the present invention, without changing the essence of the present invention, those of ordinary skill in the art can propose various interchangeable structural forms and implementation methods. Therefore, the following specific embodiments and drawings are only exemplary descriptions of the technical solution of the present invention, and should not be regarded as all of the present invention or as a limitation or restriction on the technical solution of the present invention.
[0085] The present invention provides a cognitive diagnosis method based on unbalanced perception of heterogeneous graph learning, and its process is as Figure 1 shown, including the following steps:
[0086] S1. Obtain data and perform preprocessing to form multiple data sets about students, exercises, knowledge concepts, and answer records respectively;
[0087] Specifically, obtaining data includes obtaining the total number of students, the total number of exercises, the total number of knowledge concepts, and the corresponding answer record situation; among which, performing data preprocessing includes but is not limited to:
[0088] 1.1. Data cleaning, specifically including: removing duplicate records, for example, checking and deleting duplicate answer records. Handling missing values, filling or deleting question information, etc. Outlier detection, for example, identifying and handling abnormal answer times or handling the situation of unanswered questions.
[0089] 1.2. Data conversion, including encoding conversion and normalization / standardization processing.
[0090] 1.3. Feature engineering, for example, exercise features, including: question difficulty score: calculating the difficulty score of each question based on historical answer results. Question type distribution: counting the proportion of different types of questions answered by students. Knowledge point coverage rate: counting the number of knowledge points involved in the questions and the coverage rate.
[0091] Knowledge concept features, including knowledge point importance: calculating the importance of knowledge points according to the question difficulty and the number of questions covering the knowledge points. Knowledge point relationship graph: constructing a relationship graph between knowledge points, including parent-child relationships and correlations.
[0092] Answer record features, including answer accuracy rate: calculating the accuracy rate of students on different questions. Answer speed: calculating the average answer speed of students on different questions. Time series feature: serializing the answer time for subsequent time series analysis.
[0093] 1.4. Data integration, including integrating the processed data into a unified data set to ensure the integrity and consistency of the data. Formatting the data into a format suitable for subsequent analysis and modeling.
[0094] Among them, a data set regarding students, exercises, knowledge concepts, and answering records is formed; Exemplarily:
[0095] Define the student set S = {s 1 , s 2 , …, s N}, where N is the total number of students;
[0096] Define the exercise set E = {e 1 , e 2 , …, e M}, where M is the total number of exercises;
[0097] Define the knowledge concept set C = {c 1 , c 2 , …, c K}, where K is the total number of knowledge concepts;
[0098] Let the answering record set R = {(s i , e j , r ij ) | s i ∈ S, e j ∈ E, r ij ∈ {0, 1}}, where r ij represents the answering situation of student s i to exercise e j .
[0099] S2. Obtain the explicit relationships and implicit relationships of the data set, and construct a heterogeneous graph based on the data set, explicit relationships, and implicit relationships;
[0100] Exemplarily, the constructed explicit relationships include constructing the interaction relationship between students and exercises; the inclusion relationship between exercises and knowledge concepts; the dependency relationship between knowledge concepts;
[0101] The implicit relationships include the collaborative relationship between students; the calculation process is as follows:
[0102] Calculate the relationship between student s i and s l Specifically, it is calculated and determined using formula (1):
[0103]
[0104] In formula (1), T i represents the set of exercises correctly answered by student s i , F i represents the set of exercises wrongly answered by student s i , T l and F lSimilarly, J(·) represents the Jaccard similarity between two sets; rt il represents student s i and student s l on the set U of exercises answered correctly together il is the similarity of answering time; α is the weight controlling the influence of correct and incorrect answers; β is the weight controlling the influence of time;
[0105] where rt il The calculation process of rt il is as follows:
[0106]
[0107] In formula (2), t ij represents the answering time of student s i for exercise e j , represents the average answering time of student s i answering, represents the average answering time of exercise e j being answered; represents the shifted answering time of student s i for exercise e j ;
[0108] Using formula (2) to find the debiased time of student s i and student s l respectively on the set U of exercises answered correctly together il to obtain the answering time similarity rt il between the two students;
[0109]
[0110] In formula (3), EucliD(·) represents the Euclidean distance between two vectors, represents student s i on the set U of exercises il is the debiased time vector, t l Similarly;
[0111] Finally, if is greater than a set threshold η, then let represent that student s i and student s l are similar; for student s i , if student s l is the last P students similar to it, then let represent that student s i and student s l are the least similar;
[0112] Thus, the collaborative relationship between students is obtained.
[0113] Constructing a heterogeneous graph includes defining a heterogeneous graph = {V, E}, where the node set V = S ∪ E ∪ C, and the edge set E = R se ∪R ec ∪R cc ∪R ss , S represents the set of students; E represents the set of exercises; C represents the set of knowledge concepts; R se represents the interaction relationship between students and exercises, R ec represents the inclusion relationship between exercises and knowledge concepts, R cc represents the dependency relationship between knowledge concepts, R ss represents the collaborative relationship between students.
[0114] S3. Use two different graph neural network models to process the imbalance-aware local feature extraction module, and learn the local structural features and local semantic features of the nodes in the heterogeneous graph;
[0115] In a heterogeneous graph, the types of nodes and edges are diverse, which poses challenges to the extraction of local features; in order to effectively capture the local structural features and local semantic features of nodes, two different graph neural network models can be used to process the imbalance-aware local feature extraction module.
[0116] First, preprocess the heterogeneous graph data to solve the data imbalance problem, then use two different graph neural network models respectively to extract the local features of nodes from different angles, and finally integrate these features so that they can fully reflect the local structure and semantic information of nodes, providing strong support for subsequent analysis or task applications; specifically:
[0117] Data preprocessing and imbalance processing: Data cleaning and formatting, organize the node and edge information in the heterogeneous graph, remove invalid or incorrect data records, unify the data formats such as node attributes and edge types, and ensure the standardization and consistency of the data for subsequent model processing; Imbalance data processing, analyze the distribution of different types of nodes and edges in the heterogeneous graph, determine the existing data imbalance phenomenon, and sampling strategies and weight adjustment strategies can be used;
[0118] Specifically, the structural feature extraction module uses a Structural-aware GNN; the semantic feature extraction module uses a Semantic-aware GNN. The specific steps are as follows:
[0119] S3.1. Obtain the initial embedding vectors of local structural features and local semantic features. Exemplarily, calculate and determine them through the following formula (4):
[0120]
[0121] In formula (4), represents the initial structural embedding vector of student s i , exercise e j and knowledge concept c k nodes; W S ∈R N×d , W E ∈R M×d , W C ∈R K×d is a trainable matrix mapped to the d-dimensional space; represents the one-hot encoding of student s i , exercise e j and knowledge concept c k ; represents the initial semantic embedding vector of the node; represents the initial semantic embedding vector of the relationship; Em ∈ R 8×8 represents the one-hot encoding matrix.
[0122] S3.2. Generate virtual knowledge concept nodes based on the initial embedding vectors. Exemplarily, calculate through the following formula (5):
[0123]
[0124] In formula (5), is the initial embedding of knowledge concept nodes u and k, N k is the knowledge concept node closest to k in the embedding space. Based on knowledge point N k and k, generate a new virtual knowledge point node k v , where δ is a random value following a uniform distribution;
[0125] Use formula (6) to set the inclusion relationship v and dependency relationship of virtual knowledge point k as the union of the relationships between knowledge point k and N k ;
[0126]
[0127] In formula (6), r = 1 indicates the existence of a corresponding relationship;
[0128] S3.3. Calculate and obtain local structural features and local semantic features. Exemplarily, calculate and determine them through the following formulas (7) and (8); specifically include:
[0129]
[0130] In formula (7), 1 ≤ l ≤ L, representing the number of aggregation layers; Agg(·) represents the aggregation function in the graph neural network; v represents a student, exercise, or knowledge concept node, and N v represents the set of neighbor nodes of v; σ represents the activation function, and W l-1 ∈R d×d represents the trainable matrix of the (l - 1)-th layer;
[0131]
[0132] In formula (8), φ(·) represents a mapping function. If it contains a single node, it maps to the type of that node. Exemplarily, s i —>S. If it contains an edge, it maps to the corresponding relationship, such as (s i , e j )—>r se ; w represents a trainable weight; N si represents the set of neighbor nodes of the student node s i , and N ej , N ck is the same.
[0133] S4. Use the improved Transformer architecture to process the global representation learning module and learn global context features from the local structural features and local semantic features;
[0134] S4.1. Apply the personalized PageRank technique to the constructed heterogeneous graph to obtain the PPR matrix P ∈ R |V|×|V| , where P ij represents the importance of node j to i;
[0135] For each node v ∈ V, according to the matrix P, extract the set of the top T student nodes with the highest values Similarly, obtain the set of exercise nodes and the set of knowledge concept nodes to form the global relevant node set SP v ={s v , e v , c v}.
[0136] S4.2. For each node v and its global relevant node set SP v , process and calculate its local structural features and local semantic features respectively through an improved Transformer architecture;
[0137] Specifically, the attention coefficient matrices of the structural features and semantic features are calculated using formulas (9) and (10) respectively;
[0138]
[0139] In formula (9), is the attention parameter matrix, is a trainable vector; H g ∈R |SPv+1|×d represents the local structural features of node v and its set of globally relevant nodes SP v such as ⊕ represents tensor addition;
[0140]
[0141] In formula (10), is the attention parameter matrix;
[0142] Using formula (11), the fused attention weight matrix Att l ;
[0143]
[0144] In formula (11), λ represents the weight controlling semantic information;
[0145] Using formula (12), the fused feature H g l ;
[0146]
[0147] In formula (12), LN(·) represents layer normalization, MHA(·) represents multi-head attention, and W f l is a trainable parameter matrix, and FFN(·) represents the feed-forward network layer;
[0148] After L layers of propagation, is the final representation of node v, and all nodes in the heterogeneous graph obtain the final representation through the operations of S4.
[0149] S5. Establish a diagnostic prediction model based on the global context features and optimize the diagnostic prediction model based on the constructed loss function.
[0150] S5.1. Obtain the ability i of student s and the difficulty j of exercise e and the exercise discrimination
[0151]
[0152] Predict student s using formula (14). i The response score for exercise j;
[0153]
[0154] In formula (14), denotes element-wise multiplication of vectors, and Q j denotes the j-th row of the Q-matrix and represents the knowledge concepts included in exercise e j ;
[0155] S5.2. Construct the loss function L; specifically, it is shown as formula (15) below:
[0156]
[0157] In formula (15), r i denotes the true value of the response of student s i , denotes the predicted value;
[0158] S5.3. Use the gradient descent method to perform backpropagation training on the loss function L to update the model parameters until the loss function L converges and stops, and finally obtain the optimized training model.
[0159] Specifically, the gradient descent method is a commonly used optimization algorithm for minimizing the loss function (L) to update the model parameters and improve the prediction performance of the model; the specific steps include:
[0160] S5.31. Initialize the model parameters
[0161] Steps: Randomly initialize the weights and biases of the model. Reasonable initialization can accelerate convergence and avoid local minima.
[0162] S5.32. Forward propagation
[0163] Steps: The input data is calculated through the various layers of the model; the output of each node is calculated; and finally the prediction result is output. It is used to generate the predicted value of the model and calculate the loss.
[0164] S5.33. Calculate the loss
[0165] Steps: Use the loss function (L) to evaluate the difference between the predicted value and the true label. It is used to quantify the accuracy of the model prediction.
[0166] S5.34. Backpropagation
[0167] Step: Calculate the gradient of the loss function with respect to each parameter.
[0168] Starting from the output layer, calculate the gradient layer by layer backward.
[0169] Use the chain rule to pass the gradient to the parameters of each layer.
[0170] Used to determine the direction of the impact of the loss function on each parameter.
[0171] S5.35. Update parameters
[0172] Step: Update each parameter using the gradient descent method.
[0173] The choice of learning rate is very important. Too large may lead to divergence, and too small may result in slow convergence.
[0174] Adaptive learning rate methods (such as AdaGrad, RMSprop, Adam, etc.) can be used to adjust the learning rate.
[0175] S5.36. Iterative training
[0176] Step: Repeat the above process (forward propagation, calculate loss, backward propagation, update parameters) several times until the loss function converges.
[0177] The maximum number of iterations or the change threshold of the loss function can be set as the stopping condition.
[0178] Used to gradually approximate the minimum value of the loss function through continuous iteration.
[0179] S5.37. Model evaluation
[0180] Step:
[0181] Evaluate the performance of the model using the validation set.
[0182] If the model performs well on the validation set, it is considered that the training is completed.
[0183] If the model is overfitting, regularization techniques (such as L1, L2 regularization) or increasing the amount of data can be considered.
[0184] Used to ensure that the model has good generalization ability.
[0185] Implementation case
[0186] To verify the effectiveness of the method of the present invention, the present invention selects the real-world commonly used intelligent education datasets Assistments2017 (ASSIST2017) and EdNet. For these two datasets, preprocessing is first performed, mainly by deleting student records with less than 15 response records. For the dataset ASSIST2017, we only select the records of students' first response to the exercises; for the dataset EdNet, we select the response records of 20,000 students.
[0187] For the student response prediction task, this example uses AUC (Area Under the Curve), ACC (Accuracy), and RMSE (Root Mean Square Error) as evaluation indicators.
[0188] In this embodiment, ten methods including DINA (Deterministic Input NoisyAnd Gate, a cognitive diagnosis model), MIRT (Multidimensional Item Response Theory), NCDM (Neural Cognitive Diagnosis Model), HierNCD (hierarchical cognitive diagnosis model), KaNCD (Knowledge-Aware Neural Cognitive Diagnosis), GCN (Graph Convolutional Network), HGT (Heterogeneous Graph Transformer), Hinormer (Hierarchical Information Normalization Transformer), RCD (Regularized Cognitive Diagnosis), and SCD (Sparse Cognitive Diagnosis) are selected to compare with the method of the present invention, ImCHCD (hierarchical cognitive diagnosis model). As shown in Table 1:
[0189] Table 1 Experimental Results of the Present Invention
[0190]
[0191] Specifically, as can be seen from Table 1, compared with all methods, the ImCHCD of the present invention has achieved the best results on the datasets Assistments2017 and EdNet, and the experiment proves the effectiveness of the method of the present invention in the prediction task.
[0192] The present invention realizes more accurate representation learning by introducing an improved heterogeneous graph neural network and Transformer architecture. First, based on the historical answering records and answering times of students, the collaborative relationships between students are constructed to extract more comprehensive student feature representations. Secondly, by combining an imbalance-aware heterogeneous graph learning strategy with an improved Transformer model, the problem of data distribution imbalance is effectively alleviated in the heterogeneous graph. Through this method, the present invention can capture both the local information and global information of the nodes in the graph, significantly improving the quality of representation learning, thereby effectively enhancing the accuracy and reliability of the cognitive diagnosis model.
[0193] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this article.
[0194] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0195] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be in electrical, mechanical, or other forms of connection.
[0196] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments herein.
[0197] In addition, in each embodiment of this document, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0198] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this document, in essence, or the part that contributes to the prior art, or all or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this document. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0199] In this document, specific embodiments are used to elaborate on the principles and implementation manners of this document. The descriptions of the above embodiments are only used to help understand the methods and their core ideas of this document; at the same time, for those of ordinary skill in the art, according to the ideas of this document, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this document.
Claims
1. A cognitive diagnosis method based on imbalance-aware heterogeneous graph learning, characterized by: The following steps are involved: S1, obtain data and preprocess them to form multiple data sets about students, exercises, knowledge concepts and answer records; S2, obtaining explicit relations and implicit relations of the data set, and constructing a heterogeneous graph based on the data set, explicit relations and implicit relations; S3, using two different graph neural network models to process the imbalance-aware local feature extraction module to learn the local structural features and local semantic features of the nodes in the heterogeneous graph; S4, using an improved Transformer architecture to process a global representation learning module, and learning global context features from the local structural features and local semantic features; S5. Establishing a diagnosis prediction model according to the global context features, and optimizing the diagnosis prediction model based on the constructed loss function; S6. Predict cognitive diagnosis results using the optimized diagnosis prediction model.
2. The cognitive diagnosis method based on imbalance-aware heterogeneous graph learning according to claim 1, characterized in that: In S1, the plurality of data sets include: Student set S = {s1, s2, ..., s N }, where N is the total number of students; Training set E = {e1, e2, ..., e M }, where M is the total number of exercises; Knowledge concept set C = {c1, c2, ..., c K }, where K is the total number of knowledge concepts; Answer record set R = {(s i , e j , r ij )|s i ∈S,e j ∈E,r ij ∈{0,1}}, where r ij Indicates students i Exercise e j The answer situation.
3. The cognitive diagnosis method based on imbalance-aware heterogeneous graph learning according to claim 2, characterized in that: The obtaining of explicit relations and implicit relations of the data set, and constructing a heterogeneous graph based on the data set, the explicit relations and the implicit relations specifically includes: The explicit relationships include the interactive relationship between students and exercises, the inclusion relationship between exercises and knowledge concepts, and the dependency relationship between knowledge concepts; The implicit relationships include collaborative relationships between students; The construction of the heterogeneous graph includes defining a heterogeneous graph = {V, E}, where the node set V = S∪E∪C, the edge set E = R se ∪R ec ∪R cc ∪R ss , S represents the student set; E represents the exercise set; C represents the knowledge concept set; R se represents the interaction between students and exercises, R ec Represents the inclusion relationship between practice and knowledge concepts, R cc represents the dependency relationship between knowledge concepts, R ss Represents the collaborative relationship between students.
4. The cognitive diagnosis method based on imbalance-aware heterogeneous graph learning according to claim 3 is characterized by: The implicit relationship, i.e. the collaborative relationship between students, is calculated as follows: Calculate students i and l The relationship between The calculation formula is shown in the following formula (1): In formula (1), T i Indicates students i Correct answer set, F i Indicates students i A collection of incorrect answers, T l and F l Similarly; J(·) represents the Jaccard similarity of two sets; rt il Indicates students i and students l In the commonly answered correct exercise set U il Similarity of answering time on α is the weight controlling the impact of correct and incorrect answers; β is the weight to control the effect of time; where rt il The calculation process is as follows: In formula (2), t ij Indicates students i Exercise e j The answering time, Indicates students i The average time to answer the questions, Indicates practice j The average time it takes to answer the questions; Indicates the students who are behind i Exercise e j Time to answer questions; Using formula (2), we can find the student s i and students l In the common correct answer set U il The debiasing time on the , we get the similarity rt between the answering time of the two students il ; In formula (3), EucliD(·) represents the Euclidean distance between two vectors. Indicates students i In the practice set U il The debiasing time vector on t l Similarly; Finally, if If it is greater than a set threshold η, then Indicates students i and students l are similar; for students i , if student s l is similar to P students, then let Indicates students i and students l is the least similar; Thus, we can obtain the collaborative relationship between students.
5. The cognitive diagnosis method based on imbalance-aware heterogeneous graph learning according to claim 1, characterized in that: The imbalance-aware local feature extraction module is processed by using two different graph neural network models to learn the local structural features and local semantic features of the nodes in the heterogeneous graph, specifically including: S3.
1. Obtain the initial embedding vector of the local structural features and local semantic features, and calculate and determine it by the following formula (4): In formula (4), Indicates students i , Exercise e j and knowledge concept c k The initial structural embedding vector of the node; W S ∈R N×d , W E ∈R M×d , W C ∈R K×d is a trainable matrix mapped to d-dimensional space; Indicates students i , Exercise e j and knowledge concept c k One-hot encoding of ; Represents the initial semantic embedding vector of the node; The initial semantic embedding vector representing the relationship; Em∈R 8×8 represents the one-hot encoding matrix; S3.
2. Generate a virtual knowledge concept node based on the initial embedding vector; calculate it through the following formula (5): In formula (5), is the initial embedding of knowledge concept nodes u and k, N k is the knowledge concept node closest to k in the embedding space, based on knowledge point N k and k, generate a new virtual knowledge point node k v , δ is a random value that follows a uniform distribution; Using formula (6), the virtual knowledge point k v The inclusion relationship and dependencies Set as knowledge points k and N k The union of relations; In formula (6), r = 1 indicates that there is a corresponding relationship; S3.3, calculate and obtain local structural features and local semantic features; calculate and determine by the following formulas (7) and (8); In formula (7), 1≤l≤L represents the number of aggregation layers; Agg(·) represents the aggregation function in the graph neural network; v represents a student, exercise, or knowledge concept node, and N v represents the set of neighbor nodes of v; σ represents the activation function, W l-1 ∈R d×d Represents the trainable matrix of the l-1th layer; In formula (8), φ(·) represents a mapping function, and if a node is included, it is mapped to the type of the node.
6. The cognitive diagnosis method based on imbalance-aware heterogeneous graph learning according to claim 1, characterized in that: The improved Transformer architecture is used to process the global representation learning module to learn global context features from the local structural features and local semantic features; include S4.
1. Use Pagerank technology to construct heterogeneous graphs and obtain the PPR matrix P∈R |V|×|V| , P ij Indicates the importance of node j to i; For each node v∈V, extract the top T student nodes with the highest value according to the matrix P. Similarly, we get the set of practice nodes and knowledge point concept node set The set of global related nodes SP that make up node v v ={s v , e v , c v }; S4.
2. For each node v and its globally related node set SP v , its local structural features and local semantic features are processed and calculated respectively through an improved Transformer architecture; Specifically, the attention coefficient matrices of structural features and semantic features are calculated using formula (9) and formula (10) respectively; In formula (9), is the attention parameter matrix, is a trainable vector; H g ∈R |SPv+1|×d Represents node v and its globally related node set SP v The local structural characteristics of represents tensor addition; In formula (10), is the attention parameter matrix; Using formula (11), we get the fused attention weight matrix Att l ; In formula (11), λ represents the weight of controlling semantic information; Using formula (12), we get the fused features In formula (12), LN(·) represents layer normalization, MHA(·) represents multi-head attention, is a trainable parameter matrix, FFN(·) represents a feed-forward network layer; After propagation through L layers, It is the final representation of node v. All nodes in the heterogeneous graph undergo the above steps to obtain the final representation.
7. The cognitive diagnosis method based on imbalance-aware heterogeneous graph learning according to claim 1, characterized in that: The step of establishing a diagnosis prediction model according to the global context features and optimizing the diagnosis prediction model based on the constructed loss function comprises: S5.
1. Determine student s using formula (13) i Capabilities Exercise e j Difficulty and practice discrimination Using formula (14) to predict student s i Score your answers to exercise j; In formula (14), Represents vector element-by-element multiplication, Q j represents the jth row of the Q-matrix, representing exercise e j The knowledge concepts included; S5.
2. Construct a loss function L; specifically, it is shown in the following formula (15): In formula (15), r i Indicates students i The true value of the answer, represents the predicted value; S5.
3. Use the gradient descent method to perform back-propagation training on the loss function L to update the model parameters until the loss function L converges and stops, and finally obtains the optimized training model.
8. The cognitive diagnosis method based on imbalance-aware heterogeneous graph learning according to claim 1, characterized in that: In S1, the data acquisition and preprocessing specifically include: acquiring the total number of students, the total number of exercises, the total number of knowledge concepts and the corresponding answer records; and performing data cleaning, data conversion, feature engineering and data integration on the acquired data.
9. A cognitive diagnosis device based on imbalance-aware heterogeneous graph learning, characterized in that: It comprises a processor and a memory; the memory is used to store a program; the processor executes the program to implement the method as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 8.