Cognitive diagnosis method for fine-grained knowledge level constraint perception
By combining student answer records and Q matrix, random grouping and multi-scale relational learning methods are used to dynamically adjust the student similarity relationship network, which solves the deviation problem of fine-grained knowledge and ability reasoning in cognitive diagnosis, improves the accuracy and interpretability of the diagnosis, and supports the development of personalized education.
Patent Information
- Application Number
- CN202510501062.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-22
AI Technical Summary
The existing cognitive diagnostic framework has significant deviations from students' real abilities in the process of fine-grained knowledge and ability reasoning, and lacks systematic exploration of the representation of knowledge proficiency, resulting in insufficient diagnostic accuracy and interpretability.
Through the knowledge proficiency assessment module combined with student answer records and Q matrix, a student similarity construction method based on random grouping is adopted, and a multi-scale relationship learning strategy and a graph network mechanism with Top-k attention enhancement is used to dynamically adjust the student similarity relationship network, accurately model the complex learning relationship between students, and finally optimize the output of each module through the joint training mechanism.
It significantly improves the rationality, interpretability and accuracy of cognitive diagnosis, and provides important support for personalized education and intelligent education systems.
Smart Images

Figure CN120354084A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cognitive diagnosis, and more specifically, to a cognitive diagnosis method with fine-grained knowledge level constraint perception. Background Art
[0002] Cognitive diagnosis (CD) plays a key role in educational data mining, aiming to reveal students' mastery of specific knowledge concepts by deeply analyzing their answer records. With the increasing popularity of intelligent education technology, the demand for cognitive diagnosis in improving individual learning effects and supporting personalized learning is growing. By analyzing the interaction data between students and tasks, cognitive diagnosis can not only effectively model students' knowledge levels but also evaluate their current learning states, thus helping students better understand their learning processes and improve learning efficiency.
[0003] Although existing cognitive diagnosis frameworks have improved the accuracy and interpretability of diagnosis to a certain extent through students' explicit interaction records, exercise texts, and exercise relationships, these methods mainly focus on the automatic inference of knowledge proficiency based on diagnosis models and lack a systematic exploration of optimizing the representation of knowledge proficiency. This deficiency may lead to significant deviations from students' true abilities in the process of fine-grained knowledge ability reasoning. Therefore, how to effectively combine the patterned information of students' past learning behaviors to construct a more accurate knowledge proficiency representation model to accurately infer students' true knowledge mastery levels remains a key problem to be solved. Solving this challenge can not only further promote the research depth in the field of cognitive diagnosis but also significantly improve the fairness and interpretability of modern educational evaluation. Summary of the Invention
[0004] The purpose of the present invention is to provide a cognitive diagnosis method with fine-grained knowledge level constraint perception, which statically evaluates the knowledge mastery level through a knowledge proficiency evaluation module combined with students' answer records and Q matrix; adopts a method for constructing student similarity based on random grouping to reveal potential learning mode associations; and uses a multi-scale relationship learning strategy and a graph network mechanism enhanced by Top-k attention to dynamically adjust the student similarity relationship network, accurately modeling the complex learning relationships between students. Finally, through a joint training mechanism, the outputs of each module are comprehensively optimized, significantly improving the rationality, interpretability, and accuracy of cognitive diagnosis, providing important support for the further development of personalized education and intelligent education systems.
[0005] To achieve the above purpose, the technical solution provided by an embodiment of the present invention is as follows:
[0006] A cognitive diagnosis method with fine-grained knowledge level constraint perception, comprising the following steps:
[0007] S1. First, the knowledge proficiency evaluation module combines students' answering records and the Q matrix to statically evaluate the knowledge mastery level;
[0008] S2. Second, a student similarity construction method based on random grouping is adopted to reveal potential learning pattern associations, and a multi-scale relationship learning strategy and a graph network mechanism enhanced by Top-k attention are used to dynamically adjust the student similarity relationship network, accurately modeling the complex learning relationships among students;
[0009] S3. Finally, through the joint training mechanism, the outputs of each module are comprehensively optimized, significantly improving the rationality, interpretability, and accuracy of cognitive diagnosis.
[0010] As a further improvement of the present invention, multiple core sets are defined, specifically including:
[0011] Student set S = {s1, s2,..., s n}, which contains n students;
[0012] Exercise set E = {e1, e2,..., e m}, which includes m different exercises;
[0013] Knowledge point set KC = {kc1, kc2,..., kc k}, which contains k different knowledge points;
[0014] Knowledge concept scoring matrix KS = KS ij n×k , where
[0015] KS ij represents the average score of the i-th student on the j-th knowledge point using binary scoring, and the answering response set R = {0, 1} records the answering situations of students, where 1 means completely correct and 0 means other situations.
[0016] As a further improvement of the present invention, the test records of students are modeled as a triple set (s, e, ks), where s ∈ S represents a certain student in the student set, e ∈ E represents a certain exercise in the exercise set, and ks ∈ KS is the average score of the binary scoring of knowledge concepts based on prior statistics. A pre-determined Q matrix is introduced, denoted as Q = Q ij m×k , in the matrix, if exercise e i is associated with knowledge point kc j , then Q ij = 1, otherwise it is 0.
[0017] As a further improvement of the present invention, in the model for evaluating students' abilities, first, the diagnostic model automatically infers the knowledge mastery vector θ cdBy multiplying the one-hot encoded vector \(x_s\) of the student with a trainable knowledge proficiency matrix \(S\), it is specifically expressed as follows:
[0018] \(\theta\) cd \(=\sigma(x\) s \(\times S),\)
[0019] where \(\theta\) cd \(\in(0,1)\) 1×k represents the proficiency of the student in \(k\) knowledge points, \(x\) s \(\in\{0,1\}\) 1×n is the one-hot encoding of the student, used to identify the specific student, \(S\in\mathbb{R}\) n×k is the trainable proficiency matrix, and the function \(\sigma(\cdot)\) represents the sigmoid activation function. The fine-grained knowledge constraint perception vector is calculated through the scoring matrix \(k_s\) ij of the student and the knowledge point weight vector \(k\):
[0020] \(\theta\) f \(=\sigma(k_s\) ij \(\times k),\)
[0021] where, \(\theta\) f \(\in(0,1)\) 1×k represents the fine-grained knowledge level of the student in each knowledge point, \(k_s\) s \(\in\{0,1\}\) 1×n is the scoring encoding, \(k\in\mathbb{R}\) n×k is the trainable proficiency matrix;
[0022] For exercise \(e\), we extract its knowledge point relevance vector \(Q\) e from the \(Q\) matrix, and the formula is as follows:
[0023] \(Q\) e \(=x\) e \(\times Q,\)
[0024] where, \(x_e\in\{0,1\}\) 1×m is the one-hot encoding of the exercise ID, \(Q\) e \(\in\{0,1\}\) 1×k represents the knowledge point relevance vector of this exercise, capturing the relevance between each exercise and its related knowledge points.
[0025] As a further improvement of the present invention, the performance of the student when answering questions can generally be described by the following formula:
[0026] \(r = CDM(\theta,\omega\) e ),
[0027] where, \(r\) represents the student's answering response (such as score or correct / incorrect), and \(\theta\) represents the student's knowledge mastery, \(\omega\) eIt includes various parameters of the exercise questions;
[0028] The knowledge proficiency θ of the student is divided into two components, so as to ensure that the evaluation result of the student's ability is more consistent with their true knowledge level;
[0029] θ = φ(θ cd , θ f ), where θ cd represents the knowledge mastery vector automatically inferred by the diagnostic model, and θ f represents the ability characteristics constrained by the fine-grained knowledge level. The function φ is used to fuse these two types of characteristics and comprehensively reflect their overall impact on the student's knowledge proficiency.
[0030] As a further improvement of the present invention, the prior knowledge mastery information of the student is constructed based on the binary scoring knowledge concept. Specifically, is an indicator function used to determine whether the score of student i on question j meets the binary scoring standard. When it means that student i gets a score of 1 on question j, and assigns 1 point to each knowledge point involved in this question; when it means that student i does not get a score on question j, so no points are assigned to all the knowledge points involved in this question;
[0031] Next, calculate the score accumulation A i,k of student i on knowledge point k and calculate the number of answering times C i,k of student i on knowledge point k. The formula is:
[0032]
[0033] where Qi represents the set of questions answered by student i; k ∈ K j means that question j contains knowledge point k. On this basis, the mastery degree θ i,k of student i on knowledge point k is defined as the ratio of the score accumulation to the number of answering times:
[0034]
[0035] where max(C i,k , 1) avoids the denominator being zero;
[0036] Then, calculate the score accumulation G k of all students on each knowledge point k and the cumulative occurrence times H k :
[0037]
[0038] where S represents the set of all students. Based on this, the overall mastery degree θ of knowledge point kk It is defined as the ratio of the cumulative scores of all students to the cumulative number of occurrences of knowledge points:
[0039]
[0040] Using the difference Δθ i,k As a metric, it can effectively reflect the deviation between the mastery level of students on knowledge point k and the overall mastery level θ k of the group, thereby more precisely revealing the learning progress of students relative to the average level. Its calculation formula is:
[0041] Δθ i,k = θ i,k - θ k ,
[0042] KS = [Δθ i,k n×k,
[0043] where the element Δθ of matrix KS i,k represents the difference in the mastery level of student i on knowledge point k.
[0044] As a further improvement of the present invention, based on the binary scoring knowledge state matrix KS of students, a similarity relationship graph between students is constructed to explore complex relationships. The student set S is randomly divided into m groups. Through random grouping, the local aggregation characteristics of the knowledge state of students can be effectively utilized, thereby improving the calculation efficiency and reducing the influence of noise and sparsity that may be introduced in global modeling.
[0045]
[0046] KS′ = σ(KS·W k + b),
[0047] M t = KS′[G t , :],
[0048] where G t represents the t-th student subset, W k ∈R k×k is a trainable weight matrix, KS′ ∈R n×k is the transformed knowledge state matrix, Mt ∈R |Gt|×k represents the knowledge state submatrix corresponding to the t-th student subset G t extracted from the knowledge state matrix KS′ of all students, and b ∈R k is a trainable bias vector;
[0049] Using the multi-head attention mechanism F attn to calculate the intra-group feature matrix A t and the similarity weight matrix Wt :
[0050] A t , W t = F attn (M t ),
[0051] where F attn represents the multi-head attention mechanism, and M t is the input feature matrix of the t-th subset;
[0052] Calculate the mean μt and standard deviation σ of the statistical similarity matrix t :
[0053]
[0054] Next, filter the significant relationships by calculating the dynamic threshold τ t :
[0055] τ t = μ t + α · σ t ,
[0056] E t = {(i, j)|W t,ij > τ t},
[0057] where α is a hyperparameter used to adjust the strictness of the filtering condition. By constructing the edge index E t , extract the significant student similarity relationships to better model the intra-group collaborative learning pattern;
[0058] Apply a multi-layer perceptron (MLP) to the feature matrix A t to perform non-linear mapping and generate the intra-group relationship representation F t :
[0059] F t = f MLP (A t ; Θ),
[0060] where f MLP (·) represents the MLP model, which extracts key relationship features by gradually compressing the feature dimensions. Its structure is as follows:
[0061] H1 = ρ(A t W1 + b1),
[0062] H2 = ρ(H1W2 + b2),
[0063] F t = H2W3 + b3,
[0064] Among them, Θ = {W1, b1, W2, b2, W3, b3} are the model parameters, and ρ(·) represents the ReLU activation function.
[0065] As a further improvement of the present invention, through multi-scale feature interaction, the association of students is made more flexible and dynamically adaptable, thereby enhancing the model's ability to accurately depict the knowledge state of students. First, the Top-k-AEGN single-scale update algorithm is adopted, and its core formula is:
[0066] Z t (l) = F t (l-1) W l ,
[0067] Among them, W l ∈R k×k is the linear transformation matrix;
[0068] For nodes i and j, the similarity S ij (l) is calculated as follows:
[0069]
[0070] Among them, S t (l) ∈R |Gt|×|Gt| ,
[0071] For each node i, the top k = t (l) nodes with the highest similarity are selected from its similarity distribution S to form its neighbor set N i (l) , where r is the proportional hyperparameter, and |G t | is the total number of nodes within the group:
[0072]
[0073] Using the neighbor set N i (l) to construct the adjacency matrix A t (l) ∈R |Gt|×|Gt| , which is defined as follows:
[0074]
[0075] Combined with the dynamic adjacency matrix A t (l) , calculate the attention coefficient α ij (l) between node i and node j:
[0076]
[0077] Among them, ⊕ represents the splicing operation, and φ(·) is the LeakyReLU activation function;
[0078] Through the recursive update mechanism, multi-scale interaction of features between nodes is realized. Specifically, the node features of each layer are weighted and aggregated by the features of neighbor nodes, and the feature representation of the next layer is generated through non-linear transformation. The update formula is as follows:
[0079]
[0080] where, z i (l) represents the input feature of node i at the l-th layer, and N i (l) represents the neighbor set of node i. The final feature representation: After L-layer recursive update, the final feature representation of node i is:
[0081]
[0082] where, z i (0) is the initial input feature of the node, and Θ = {Θ0, Θ1,..., Θ L-1} are the trainable parameters of all layers.
[0083] As a further improvement of the present invention, an adaptive optimization personalized weighting mechanism is adopted to dynamically adjust the fusion method of the two features, so as to more accurately reflect the personalized knowledge mastery trajectory of students:
[0084]
[0085] where, w θi is the personalized weighting coefficient of student i, which is optimized and adjusted during the model training according to the learning characteristics of the student, and the response prediction is based on the following formula:
[0086]
[0087] As a further improvement of the present invention, the cross-entropy loss function is adopted to measure the difference between the predicted value ŷ and the actual answer label r of the student. The specific form of the loss function is as follows:
[0088]
[0089] where, r i represents the actual answer label of the student, and ŷ i represents the predicted probability of the model.
[0090] Compared with the prior art, the advantages of the present invention are:
[0091] Through the knowledge proficiency evaluation module, the present invention combines students' answering records and Q matrices to statically evaluate the knowledge mastery level; adopts a method for constructing student similarity based on random grouping to reveal potential learning pattern associations; and uses a multi-scale relationship learning strategy and a graph network mechanism enhanced by Top-k attention to dynamically adjust the student similarity relationship network, accurately modeling the complex learning relationships among students. Finally, through the joint training mechanism, the outputs of each module are comprehensively optimized, significantly improving the rationality, interpretability, and accuracy of cognitive diagnosis, providing important support for the further development of personalized education and intelligent education systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 It is a cognitive diagnosis framework diagram of the present invention;
[0093] Figure 2 It is a kernel density diagram of the application effect of the fine-grained knowledge level constraint of all students of the present invention on the Math dataset;
[0094] Figure 3 It is a comparison diagram of the fine-grained level constraint situations of randomly selecting 3 groups from the Science dataset of the present invention;
[0095] Figure 4 It is a dynamic evolution diagram of the knowledge state that the CEKT model of the present invention can track students during the process of solving three programming problems;
[0096] Figure 5 It is the fine-grained knowledge level constraint ability θ f and the collaborative effect and influence degree of the diagnostic model automatically inferring the knowledge mastery vector θ cd in students' abilities;
[0097] Figure 6 It is a diagnostic result diagram of two randomly selected students from the Science and Math datasets of the present invention;
[0098] Figure 7 It is a data table of the Programme for International Student Assessment (PISA2015) in 2015 adopted in the embodiments of the present invention;
[0099] Figure 8 It is a result table of performance evaluation on the cognitive diagnosis model combined with FCD of the present invention;
[0100] Figure 9 It is a result table of the prediction of students' performance in different regions of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0101] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention; obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0102] Embodiment 1:
[0103] The data of the Programme for International Student Assessment 2015 (PISA2015) is used to verify the effectiveness of the fine-grained knowledge level constraint-aware cognitive diagnosis system. This dataset covers more than 500,000 students in 73 countries and regions, focusing on the subjects of science, reading, and mathematics. In the data processing stage, the questions marked as "Fullcredit" are regarded as correct (recorded as 1), while "Partialcredit" and "Nocredit" are regarded as incorrect (recorded as 0); the unattempted questions are marked as not answered. To ensure the accuracy of the analysis, the student data with less than 30 records is excluded, and the students with less than 20 records in the mathematics part are excluded. After processing, the data is randomly divided into a training set (70%), a validation set (10%), and a test set (20%).
[0104] In the experimental part, the Xavier initialization method is used to initialize the model parameters, and 5-fold cross-validation is used to train the model. The final result is obtained from the average value of the 5-fold cross-validation. The experiment is carried out on a server configured with a 64-bit Ubuntu 20.04.5 LTS system, which is equipped with an Intel Xeon Gold 5218 CPU with 2.30 GHz and a Tesla V100 GPU with 32 GB. The calculation process is implemented based on the PyTorch framework.
[0105] To study the impact of fine-grained knowledge level constraints on the model performance, four models based on the FCD framework are developed: FCD-IRT, FCD-MIRT, FCD-DINA, and FCD-NCDM. These models combine classical cognitive diagnosis methods. These models are reproduced, and the parameters are carefully adjusted to ensure that each model runs at its best state, thus ensuring the fairness of the comparison. The area under the curve (AUC), accuracy (ACC), and root mean square error (RMSE) are used as the main performance indicators to comprehensively evaluate the performance of each model.
[0106] IRT (Lord, 2012): This model measures students' mastery of specific knowledge points by evaluating their performance on different test questions and estimates students' abilities and question difficulties.
[0107] MIRT (Ackerman, 2014): As an extension of IRT, the MIRT model is used to analyze students' performance on multiple ability dimensions and can simultaneously estimate multiple ability dimensions and item parameters.
[0108] DINA (DeLaTorre, 2009): The model assumes that each question involves multiple knowledge points and infers students' mastery status of these knowledge points by analyzing their answering situations, and uses a binary classification method for description.
[0109] NCDM (Wang et al., 2022): The NCDM model uses neural networks to capture the complex relationships between students and questions, achieving accurate and interpretable cognitive diagnosis and replacing the traditional manually designed prediction functions.
[0110] HCD: The HCD framework integrates students' prior statistical information to construct a hierarchical structure and further introduces a hierarchical constraint-aware neural network to form a flexible plug-and-play solution. This framework deeply mines the complex relationships between different levels of students, thus more accurately simulating the students' knowledge states in real educational scenarios.
[0111] Since the true knowledge state of students is difficult to directly measure, the performance evaluation of cognitive diagnosis models always faces challenges. Although improving prediction accuracy is not the primary goal of the research, by analyzing the performance of the model in the student achievement prediction task, its rationality can be verified indirectly. This method also demonstrates the core value of fine-grained knowledge level constraint awareness (FCD) in improving the effect of cognitive diagnosis.
[0112] Figure 8 Shows the results of performance evaluation on the cognitive diagnosis model combined with FCD. Through analysis, it can be found that the FCD method (such as FCD-IRT) has significant advantages over traditional models (such as IRT) in various indicators, which indicates that the introduction of fine-grained knowledge constraints plays a key role in cognitive diagnosis. In addition, after integrating FCD into the NCDM model, it achieves the best performance in indicators such as AUC, ACC, and RMSE, fully demonstrating that neural networks can more efficiently capture the characteristics of students' knowledge states. Further comparison of the performance of the three datasets shows that FCD shows strong adaptability under different data scales. Whether in the data-rich Science or the data-scarce Read, satisfactory results are obtained, and it also performs outstandingly on Math with scarce interaction data. This also indicates that compared with HCD that focuses on the overall hierarchical constraints of students, FCD can more accurately reflect students' knowledge mastery by refining the constraints to the knowledge level.
[0113] To deeply analyze the roles of the three core modules, namely MSRL, BSS, and KPA, in the FCD framework, ablation experiments were designed. Specifically: - MSRL means removing the MSRL module and only retaining the other parts of the framework; - BSS replaces the BSS module with randomly generated similar edges within the student group while keeping the rest of the structure unchanged; - KPA replaces the KPA module with randomly initialized student prior knowledge proficiency, and the other components remain the same. The experimental results are as Figure 9 shown and the specific analysis is as follows: First, removing any one of the modules will lead to a performance decline, indicating that each module plays an important role in the overall framework and verifying its effectiveness in modeling different hierarchical relationships. Second, since there are no natural edge relationships in the dataset, reasonably constructing edges and conducting efficient learning have a significant impact on the model performance. Finally, the performance decline caused by removing the KPA module is the most significant, indicating that when modeling the fine-grained knowledge level, the student knowledge state based on prior statistics plays a key role in the constraints during the learning process.
[0114] Introducing the fine-grained knowledge level constraint aims to enhance the existing cognitive diagnosis model's accurate description of students' relative ability performance, thereby more reasonably inferring students' ability levels on various knowledge concepts. This constraint ensures that the evaluation results are more in line with the actual situation by embedding prior knowledge in the model. Figure 2 The application effect of the fine-grained knowledge level constraint of all students on the Math dataset is shown in the form of a kernel density map.
[0115] Several important conclusions can be observed from the figure. First, both the original NCDM model and the FCD-NCDM model with the fine-grained knowledge level constraint can better distinguish the ability distribution differences of students on the same knowledge point. Second, by comparing Figure a and Figure c, it can be found that the distribution fitting of the FCD-NCDM model for a single knowledge point is highly consistent with the prior statistics, indicating that the proposed fine-grained knowledge level constraint plays a key role in guiding the model to learn a reasonable distribution. In contrast, due to the lack of additional constraints, the NCDM model only relies on the optimization of the objective function, resulting in a large deviation between the learned distribution and the prior statistics. The overall results show that the FCD-NCDM model's assessment of students' knowledge states is more in line with the actual trend.
[0116] Further analyzing the distribution density of students with ability values close to 0.5 can more clearly verify the role of the fine-grained knowledge level constraint. For example, on the knowledge point KC11, the prior distribution shows that the proportion of students with an ability value of 0.5 is about 1.1%, while the result of the NCDM model is as high as 18.2%. In contrast, the proportion of the FCD-NCDM model at this stage is 1.2%, which is more consistent with the prior distribution. This result indicates that the fine-grained knowledge level constraint significantly improves the model's ability to infer students' knowledge states.
[0117] Accurately evaluating the knowledge state of students is crucial for cognitive diagnosis. Figure 2 It shows the situation of introducing fine-grained knowledge level constraints under the group effect, and the specific performance of the individual after being subject to this constraint still needs to be deeply analyzed. Three groups were randomly selected from the Science dataset, and the similarity between a certain student in each group and the students within his own group was calculated through Eq. (19). The student knowledge mastery vectors were calculated based on the prior statistical KPA method and the FCD framework proposed in this paper. The comparison results between the selected student and the top three students with the highest similarity were mainly analyzed.
[0118] Figure 3 Summarizes the relevant key observations. First, whether based on KPA or FCD, both methods can find other students similar to the target student within the group. Second, it is observed that the selection of similar students does not change much under the two methods. For example, in group 3, both KPA and FCD find the same similar students. However, in other groups, there will be certain deviations in the selection of student similarity. For example, in group 1, based on the KPA method, student s1 1 is similar to student s1 2 , while under the FCD method, its similar student is replaced by student s1 9 . Similarly, in group 2, the similar student of student s1 2 changes from s4 2 under the KPA method to s 25 2 under the FCD method.
[0119] Although the similarity results are not completely consistent, these findings indicate that the fine-grained knowledge level constraints play a positive role in the assessment of individual knowledge state, thus verifying its practicality in cognitive diagnosis. In cognitive diagnosis, the mastery level of students on knowledge points significantly affects their answering performance, and students with high mastery have a higher success rate (Chen et al., 2017). To evaluate the interpretability of the model, the DOA (Degree of Agreement) index (Fouss et al., 2007) is introduced. Traditional models (such as IRT and MIRT) describe students' abilities through latent trait vectors, but it is difficult to refine to specific knowledge points. This study reveals the differences in the model's portrayal of students' knowledge states by comparing the DOA performance of FCD with benchmark models (Random, DINA, NCDM) (see Figure 4 ).
[0120] The results show that although the DINA model lacks constraint awareness, its DOA value still performs outstandingly on all datasets, indicating the advantage of its theoretical framework in interpretability. However, the performance of this model in actual prediction is limited. In contrast, the DOA value of the model incorporating fine-grained knowledge constraints is higher, significantly improving interpretability, and the fine-grained constraint effect is better than the hierarchical constraint (such as HCD-NCDM), indicating that the fine analysis of knowledge points is more valuable. It is worth noting that the DOA value of the Random model is stable at around 0.5, while the DOA of the FCD-Random model with fine-grained knowledge constraints is significantly improved and is close to that of NCDM and HCD-NCDM. This further proves the importance of fine-grained knowledge awareness in improving model interpretability.
[0121] Figure 5 Shows the fine-grained knowledge level constraint ability θ f And the diagnostic model automatically infers the knowledge mastery vector θ cd The synergy and its impact degree in students' abilities. As can be seen from the figure, θ cd And θ f Show obvious differences among different students, which highlights the necessity of separately focusing on these two types of abilities in the model. Further analysis reveals that θ cd And θ f Do not contribute equally to students' abilities. For example, for student 1, the value of θ cd Is better than θ f , but θ f Has a more significant impact on their final ability θ; on the contrary, although student 23 scored higher in θ f , θ cd Has a greater impact. This shows that the abilities of different students are affected by these two characteristics in different ways. 610
[0122] Finally, from the regression lines, the regression line of the fine-grained knowledge level ability is higher than that of the diagnostic model's ability to automatically infer the knowledge mastery representation method, indicating that in the performance of these 30 students, the fine-grained knowledge level ability has a more significant impact on the final ability. The relationship between students' cognitive diagnosis results, their abilities, and the difficulty of the questions can be quantitatively analyzed through formula Eq. (26). Randomly select two students and six questions they answered from the Science and Math datasets respectively, Figure 6 The upper part of the figure (a) shows the diagnostic results of two randomly selected students in the Science dataset, and the lower part of the figure (b) shows the diagnostic results of two randomly selected students in the Math dataset.
[0123] As can be seen from the figure, when the students' knowledge mastery and the requirements of the questions are clear, the prediction results of the model are highly consistent with the students' actual answers. For example, inFigure 6 (a) In the case of student S1, the comparison between the student's final ability value θ and the question difficulty Difficulty shows that only for question e4, the student's ability is greater than the question difficulty. At this moment, the model predicts that the student can answer correctly, but the actual answer result does not match. This situation may be affected by various factors, such as the student's accidental mistakes, etc. For student S2, on question e3, the student's ability value is 0.21 and the question difficulty is 0.77. The model predicts that the student will answer wrongly, but the actual answer result is correct. This phenomenon may stem from the guessing factor in the student's answer, or it may indicate that there is still room for further optimization in the model when evaluating the student's ability. In contrast, Figure 6 (b) The prediction result of student S3 in (b) is completely consistent with the actual answer, which indicates that after introducing the fine-grained knowledge level constraint, the model can accurately evaluate the student's knowledge level and make reasonable predictions in most cases.
[0124] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be construed as limiting the claimed invention.
[0125] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other implementation manners that can be understood by those skilled in the art.
Claims
1. A cognitive diagnosis method with fine-grained knowledge level constraint awareness, characterized in that: Including the following steps: S1. First, the knowledge proficiency evaluation module combines the student's answer records and the Q matrix to statically evaluate the knowledge mastery level; S2. Second, a student similarity construction method based on random grouping is adopted to reveal potential learning pattern associations, and a multi-scale relationship learning strategy and a graph network mechanism enhanced by Top-k attention are used to dynamically adjust the student similarity relationship network, accurately modeling the complex learning relationships among students; S3. Finally, through the joint training mechanism, the outputs of each module are comprehensively optimized, significantly improving the rationality, interpretability, and accuracy of cognitive diagnosis.
2. The cognitive diagnosis method with fine-grained knowledge level constraint perception according to claim 1, characterized in that: Define multiple core sets, specifically including: The set of students S = {s1, s2,..., s n}, which contains n students; Practice set \(E = \{e_1, e_2, \ldots, e m \}\) contains \(m\) different practice problems; Knowledge point set KC = {kc1, kc2,..., kc k} contains k different knowledge points; Knowledge concept scoring matrix KS = KS ij n×k , where, KS ij It represents the average score of the i-th student on the j-th knowledge point using binary scoring. The response set R = {0, 1} records the answering situation of the students, where 1 indicates a complete correct answer and 0 indicates other situations.
3. The cognitive diagnosis method with fine-grained knowledge level constraint perception according to claim 2, characterized in that: Model the test records of students as a set of triples (s, e, ks), where s ∈ S represents a student in the set of students, e ∈ E represents an exercise in the set of exercises, and ks ∈ KS is the average score of the binary scoring of knowledge concepts based on prior statistics. Introduce a pre-determined Q matrix, denoted as Q = Q ij m×k , in the matrix, if the exercise e i is associated with the knowledge point kc j , then Q ij = 1, otherwise 0.
4. A cognitive diagnosis method with fine-grained knowledge level constraint perception according to claim 3, characterized in that: In the model for student ability assessment, first, the diagnostic model automatically infers the knowledge mastery vector θ cd by multiplying the one-hot encoded vector xs of the student with a trainable knowledge proficiency matrix S, which is specifically represented as follows: θ cd = σ(x s × S), where θ cd ∈(0,1) 1×k represents the proficiency of the student on k knowledge points, x s ∈{0,1} 1×n is the one - hot encoding of the student, used to identify the specific student, S∈R n×k is the trainable proficiency matrix, the function σ(·) represents the sigmoid activation function, the fine - grained knowledge - constraint awareness vector is calculated through the student's scoring matrix ks ij and the knowledge - point weight vector k: θ f = σ(ks ij × k), where, θ f ∈(0,1) 1×k represents the fine-grained knowledge level of students on each knowledge point, ks s ∈{0,1} 1×n is the scoring code, k ∈ R n×k is the trainable proficiency matrix; For exercise e, we extract the knowledge point relevance vector Q from the Q matrix e , and the formula is as follows: Q e = x e × Q, where \(x_e\in\{0,1\}\) 1×m is the one - hot encoding of the exercise ID, \(Q\) e \(\in\{0,1\}\) 1×k represents the knowledge - point relevance vector of the exercise, capturing the relevance between each exercise and its related knowledge points.
5. A cognitive diagnosis method with fine-grained knowledge level constraint awareness according to claim 4, characterized in that: The performance of students when answering questions can usually be described by the following formula: r = CDM(θ, ω e ), where r represents the student's response to the question (such as score or correct / incorrect), and θ represents the student's knowledge mastery, while ω e includes various parameters of the practice questions; The knowledge proficiency θ of students is divided into two components to ensure that the student ability assessment results are more consistent with their true knowledge level; θ = φ(θ cd , θ f ), where θ cd represents the knowledge mastery vector automatically inferred by the diagnostic model, and θ f represents the ability characteristics constrained by the fine-grained knowledge level. The function φ is used to fuse these two types of characteristics to comprehensively reflect their overall impact on the student's knowledge proficiency.
6. The cognitive diagnosis method with fine-grained knowledge level constraint perception according to claim 5, characterized in that: Construct the prior knowledge mastery information of students based on binary scoring for knowledge concepts. Specifically, is an indicator function used to determine whether the score of student i on question j meets the binary scoring criteria. When , it means that student i gets a score of 1 on question j, and assigns 1 point to each knowledge point involved in this question; when , it means that student i does not get a score on question j, so no points are assigned to all the knowledge points involved in this question; Next, calculate the cumulative score A of each student i on knowledge point k i,k and calculate the number of times student i has answered questions on knowledge point k, C i,k , with the formula: Among them, $Q_i$ represents the set of questions answered by student $i$; $k\in K$ j indicates that question $j$ covers knowledge point $k$. Based on this, the mastery level $\theta_{ik}$ of student $i$ on knowledge point $k$ i,k is defined as the ratio of the cumulative score to the number of questions answered: where max(C i,k , 1) avoids a zero denominator; Then, calculate the cumulative score G of all students for each knowledge point k k and the cumulative occurrence times H k : Among them, \(S\) represents the set of all students. Based on this, the overall mastery degree \(\theta\) of knowledge point \(k\) k is defined as the ratio of the cumulative scores of all students to the cumulative number of occurrences of the knowledge point: Use the difference Δθ i,k as a metric, which can effectively reflect the deviation between the mastery level of students on knowledge point k and the overall mastery level θ k of the group, thereby more accurately revealing the learning progress of students relative to the average level. Its calculation formula is: Δθ i,k = θ i,k - θ k , KS = [Δθ i,k n×k, Among them, the element Δθ of the matrix KS i,k represents the difference in the mastery level of student i on knowledge point k.
7. A cognitive diagnosis method with fine-grained knowledge level constraint perception according to claim 6, characterized in that: Based on the binary scoring knowledge state matrix KS of students, a similarity relationship graph among students is constructed to explore complex relationships. The student set S is randomly divided into m groups. Through random grouping, the local aggregation characteristics of the student knowledge state can be effectively utilized, thereby improving the calculation efficiency and reducing the noise and sparsity effects that may be introduced in global modeling. KS′ = σ(KS·W k + b), M t = KS′[G t , :], Among them, G t represents the t-th student subset, W k ∈R k×k is a trainable weight matrix, KS′ ∈ R n×k is the transformed knowledge state matrix, Mt ∈ R |Gt|×k represents the knowledge state sub-matrix extracted from the knowledge state matrix KS′ of all students corresponding to the t-th student subset G t and b ∈ R k is a trainable bias vector; Utilize the multi-head attention mechanism F attn Calculate the intra-group feature matrix A t and the similarity weight matrix W t : A t ,W t =F attn (M t ), Among them, F attn represents the multi-head attention mechanism, M t is the input feature matrix of the t-th subset; Calculate the mean μt and standard deviation σ of the statistical similarity matrix t : Next, calculate the dynamic threshold τ t Filter significant relationships: τ t = μ t + α·σ t , E t = {(i, j)|W t,ij > τ t}, where α is a hyperparameter used to adjust the strictness of the screening criteria, and by constructing the edge index E t , significant student similarity relationships are refined to better model the within-group collaborative learning pattern; For the feature matrix A t Apply a multi-layer perceptron (MLP) for non-linear mapping to generate the intra-group relationship representation F t : F t = f MLP (A t ; Θ), Among them, f MLP (·) represents the MLP model, which extracts key relationship features by gradually compressing the feature dimension, and its structure is as follows: H1 = ρ(A t W1 + b1), H2 = ρ(H1W2 + b2), F t = H2W3 + b3, where Θ = {W1, b1, W2, b2, W3, b3} are model parameters, and ρ(·) represents the ReLU activation function.
8. A fine-grained knowledge level constraint-aware cognitive diagnosis method according to claim 7, characterized in that: Through multi-scale feature interaction, the student association is made more flexible and dynamically adaptable, thereby improving the model's ability to accurately depict the student knowledge state. First, the Top-k-AEGN single-scale update algorithm is adopted, and its core formula is: Z t (l) = F t (l-1) W l , where, W l ∈R k×k is a linear transformation matrix; For nodes i and j, the similarity S ij (l) is calculated as follows: Among them, S t (l) ∈R |Gt|×|Gt| , For each node i, select the top t (l) nodes with the highest similarity from its similarity distribution S [i, :] to form its neighbor set N i (l) , where r is the proportional hyperparameter and |G t | is the total number of nodes within the group: Utilize the neighbor set N i (l) Construct the adjacency matrix A t (l) ∈R |Gt|×|Gt| , which is defined as follows: Combined with the dynamic adjacency matrix A t (l) , calculate the attention coefficient α between node i and node j ij (l) : Among them, represents the splicing operation, and φ(·) is the LeakyReLU activation function; Through the recursive update mechanism, multi-scale interaction of features between nodes is realized. Specifically, the node features of each layer are weighted and aggregated by the features of neighbor nodes, and the feature representation of the next layer is generated through a non-linear transformation. The update formula is as follows: where, z i (l) represents the input feature of node i at the l-th layer, and N i (l) represents the neighbor set of node i. The final feature representation: After L-layer recursive update, the final feature representation of node i is: where, z i (0) is the initial input feature of the node, and Θ = {Θ0, Θ1,..., Θ L-1} are the trainable parameters of all layers.
9. A cognitive diagnosis method with fine-grained knowledge level constraint perception according to claim 8, characterized in that: An adaptive optimization personalized weighting mechanism is adopted to dynamically adjust the fusion method of the two features, so as to more accurately reflect the personalized knowledge mastery trajectory of students: where, w θi is the personalized weighting coefficient of student i, which is optimized and adjusted during model training according to the learning characteristics of the student, and the reaction prediction is carried out based on the following formula:
10. A cognitive diagnosis method with fine-grained knowledge level constraint awareness according to claim 9, characterized in that: The cross-entropy loss function is used to measure the difference between the predicted value y ~ and the actual answer label r of the student. The specific form of the loss function is as follows: Among them, r i represents the actual answer label of the student, and y ~ i represents the predicted probability of the model.