Neuro-cognitive diagnostic method based on problem difficulty perception enhancement
By acquiring semantic and structurally perceptual representations of exercises and combining Bloom vectors and contrastive learning strategies, a comprehensive difficulty level for exercises is constructed. This addresses the shortcomings of existing cognitive diagnostic methods in terms of knowledge structure and text perception, enabling more accurate student ability assessment and exercise difficulty simulation.
Patent Information
- Application Number
- CN202610316760.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-16
- Publication Date
- 2026-06-16
AI Technical Summary
Existing cognitive diagnostic methods lack sufficient perception of knowledge structure and exercise text in modeling exercise difficulty, and fail to distinguish between the subjective and objective attributes of exercise difficulty, thus limiting their application in smart education.
We employ a neurocognitive diagnostic method based on enhanced perception of exercise difficulty. By acquiring semantic perception representations, structural perception representations, and Bloom vectors of exercises, and combining them with a contrastive learning strategy, we construct a comprehensive difficulty level for exercises, simulate the interaction process between students and exercises, and optimize the generation of subjective difficulty.
It effectively improves the performance and interpretability of the cognitive diagnostic model, enabling a more refined simulation of the interaction between students and exercises, and enhancing the interpretability and accuracy of diagnostic results.
Smart Images

Figure CN122222778A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of smart education and educational data mining technology, and in particular to a neurocognitive diagnostic method based on enhanced perception of exercise difficulty. Background Technology
[0002] Cognitive diagnostics is a core technology in smart education. Its goal is to accurately assess students' mastery of various knowledge concepts and predict their future performance based on their historical answer records and the relationships between exercises and knowledge concepts (Q-matrix). Although the introduction of deep learning technology has greatly improved the performance of cognitive diagnostic methods, the information about the perceived difficulty of exercises contained in the knowledge structure and exercise text has not been fully explored. Using only answer records and Q-matrix is often insufficient to characterize the complexity of real-world educational scenarios.
[0003] In recent years, cognitive diagnostic methods (CD) have made significant progress, exploring various strategies, including neural network-based, knowledge association-based, and knowledge enhancement-based methods. These methods utilize neural networks to integrate auxiliary information such as knowledge structures and exercise texts, thereby more realistically simulating the interaction between students and exercises. The paper "Fei Wang, Qi Liu, Enhong Chen, Zhenya Huang, Yuying Chen, YuYin, Zai Huang, and Shijin Wang. Neural cognitive diagnosis for intelligent education systems. In Proceedings of the 34th AAAI Conference on Artificial Intelligence, pages 6153–6161, 2020." proposes a neural cognitive diagnostic method, the NeuralCD model, which uses a three-layer neural network to simulate the interaction between students and exercises. The paper "Weibo Gao, Qi Liu, Zhenya Huang, Yu Yin, Haoyang Bi, Mu-Chun Wang, Jianhui Ma, Shijin Wang, and Yu Su. RCD: relation map driven cognitive diagnosis for intelligent education systems. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 501–510, 2021." proposes a relation map-driven cognitive diagnosis method and introduces the RCD model. This method incorporates hierarchical interactions into the cognitive diagnosis framework by constructing a three-layer heterogeneous graph of student-question-concept.The paper "Zhiang Dong, Jingyuan Chen, and Fei Wu. Knowledge is power: harnessing large language models for enhanced cognitive diagnosis. In Proceedings of the 39th AAAIConference on Artificial Intelligence, pages 164–172, 2025." proposes a knowledge-enhanced cognitive diagnosis model, namely the KCD model. This method integrates large language models (LLM) with cognitive diagnosis models and uses exercise texts to improve cognitive diagnosis performance.
[0004] However, current cognitive diagnostic methods face three fundamental limitations: first, problem difficulty modeling lacks sufficient perception of knowledge structure; second, problem difficulty modeling lacks cognitive depth in representing problem texts; and third, it fails to distinguish between the subjective and objective attributes of problem difficulty. All three limit their practical application in smart education.
[0005] The first limitation is that the modeling of exercise difficulty does not adequately consider the knowledge structure. This is manifested in the fact that real-world knowledge concepts often contain complex relationships such as pre-dependencies and hierarchical inclusions. However, many mainstream cognitive diagnostic datasets lack fine-grained relationship annotations, and the cost of manual annotation by experts is enormous. As a result, cognitive diagnostics does not adequately consider knowledge structure information when modeling exercise difficulty.
[0006] The second limitation is that the modeling of exercise difficulty lacks cognitive depth in representing the exercise text. Although existing methods attempt to extract semantic information from exercise text, their understanding of "difficulty" remains at a superficial level, such as lexical complexity or syntactic length, and fails to combine educational psychology theories to deeply analyze the cognitive requirements of the exercises.
[0007] The third limitation is that existing models fail to distinguish between the objective and subjective attributes of exercise difficulty. Exercise difficulty not only has an objective attribute determined by the exercise itself, but also a subjective attribute influenced by learners' prior knowledge and self-efficacy. Ignoring this distinction will prevent accurate simulation of the interaction between students' potential abilities and the potential difficulty of exercises. Summary of the Invention
[0008] The purpose of this invention is to provide a neurocognitive diagnostic method based on enhanced perception of exercise difficulty. This method integrates knowledge graph structure perception information and exercise text semantic perception information to enhance exercise difficulty modeling. It overcomes the problem that existing models lack sufficient perception of knowledge structure and exercise text when modeling exercise difficulty, and fail to distinguish between the subjective and objective attributes of exercise difficulty.
[0009] To achieve the above objectives, this invention provides a neurocognitive diagnostic method based on enhanced perception of exercise difficulty, comprising the following steps: S1. Obtain the semantic-aware representation of the exercise, the structural-aware representation of the exercise, the Bloom vector of the exercise, and the enhanced continuous Q matrix, respectively; S2. Using the Bloom vector obtained in S1 as a priori, the semantic perception representation and the structural perception representation of the exercises are adaptively fused to obtain the final fused perception representation of the exercises. S3. The potential objective difficulty of the exercises is obtained by combining the sensory representation projection of the exercises with the basic objective difficulty of the exercises. The subjective difficulty of the exercises is modeled based on the objective difficulty of the exercises and the cognitive state generated by the student ID. A contrastive learning strategy is used to optimize the generation of subjective difficulty. The objective difficulty and subjective difficulty are then combined to obtain the comprehensive difficulty of the exercises. S4. A three-layer neural network is used to simulate the interaction process between students and exercises. The comprehensive difficulty of the exercises obtained in S3, the student's cognitive state, the enhanced continuous Q matrix obtained in S1, and the exercise discrimination generated by the exercise ID are input into the three-layer neural network for training to obtain a trained cognitive diagnostic model. S5. Input the test data into the trained cognitive diagnostic model to obtain the probability of students answering the questions correctly, and then predict their performance and acquire their cognitive status.
[0010] Preferably, in S1, the number of students in the dataset is defined as... The number of exercises is The number of knowledge concepts is The semantic-aware representation, structure-aware representation, Bloom vector, and enhanced continuous Q-matrix of exercises are obtained from the dataset, including the following steps: S11. Use BGE to embed the exercise text to obtain the semantic-aware representation of a specific exercise. ; S12. Based on a large language model, design relation proposer and relation discriminator templates. Mine semantic associations from exercise texts, knowledge concepts, and Q-matrices to construct an enhanced exercise-knowledge concept graph. Then, use TransE to perform embedding learning on the graph to obtain a structure-aware representation of a specific exercise. ; S13. Bloom cognitive hierarchy of exercises labeled using mind chain technology, transforming the output of the large model into multi-hot vectors. It employs a recursive hierarchical weighting strategy to model the dependencies between cognitive levels, obtaining Bloom vectors for specific exercises. ; S14. Based on the Q matrix, an enhanced continuous Q matrix is obtained using a contrastive learning fine-tuning strategy of BGE. .
[0011] Preferably, in S13, six basic cognitive vectors are defined. ,in A hierarchical weighting strategy is adopted, and the calculation process is as follows: ; ; in, , Representing dimension, This represents the specific dimension of each defined basic cognitive vector, while These are learnable scalar parameters used to control from lower-order dimensions. To the current dimension That is, the first Information inheritance rate of each cognitive vector; , Formula Each of them The vector formed represents the attention coefficient of each basic cognitive vector; This is a learnable weight matrix.
[0012] Preferably, S2 is as follows: Bloom vector obtained using S1 As a priori, the semantic perception representation of the exercises is... and the perceptual representation of exercise structure Projecting onto a unified metric space specifically involves: ; in, To unify the feature dimensions of the metric space, , , The projection parameter matrix is learnable; Then, attention is performed based on metrics. and Two-way alignment, specifically: ; ; ; in, This indicates two fully connected layers; The gating coefficients are used to perform bidirectional feature correction. and These are the corrected semantic-aware vector and the structural-aware vector, respectively. Complete the semantic perception representation of the exercises and the perceptual representation of exercise structure The adaptive fusion yields the final problem fusion perceptual representation. Specifically: ; ; in, The symbol [.;.] represents a concatenation operation along the feature dimension, mapping the hidden layer features to two scalars, corresponding to the text-aware weights. and structure-aware weights and satisfy ; The weight matrix is a learnable matrix; For bias terms, This is a layer normalization operation.
[0013] Preferably, S3 is as follows: The exercises obtained in S2 are integrated with perceptual representations. After MLP projection, it serves as the potential objective difficulty of the exercises. The objective difficulty level of the exercises is generated by combining them with the exercise ID. By piecing together the data, we can obtain the objective difficulty level of the exercises. The calculation process is as follows: ; ; ; in, Let be the one-hot vector of the exercise. , For a trainable weight matrix, and For bias terms; Based on the objective difficulty of the exercises and cognitive states generated by student ID Subjective difficulty of modeling exercises The calculation process is as follows: ; ; ; ; ; in, For the student's one-hot vector, the gate value This indicates the student's level of acceptance of objective difficulty; offset value. This indicates the direction of fluctuation in the perceived difficulty level. To assess the overall difficulty of the exercises, we need to consider both the objective difficulty level and the overall difficulty level of the exercises. Subjective difficulty of exercises It was pieced together.
[0014] Preferably, in S3, a contrastive learning strategy is used to optimize the generation of subjective difficulty of exercises. The specific calculation process is as follows: Introducing group cognitive consensus and stratified loss as monitoring signals for subjective difficulty, for target exercises... The set of students who answered correctly is defined as the positive consensus group. The set of students who answered incorrectly is defined as the negative consensus group. ; Computing a positive consensus group and negative consensus groups The centroid in the subjective difficulty semantic space is used to represent the average cognitive state of the two groups. The calculation process of the centroid is as follows: ; in, The model represents students In the exercises The personalized subjective difficulty vector generated above; and The number of students in the positive consensus group and the negative consensus group, respectively; and These represent the centroids of the positive and negative consensus groups in the subjective difficulty semantic space, respectively. Based on the aforementioned centroid, cognitive consensus loss is adopted. Minimize the perceptual variance within the group: ; Layered loss Expand the positive consensus group and negative consensus groups To establish clear capability boundaries by addressing the perceived gap between them: ; in, This represents the L2 norm, used to quantify the overall magnitude of the subjective difficulty vector.
[0015] Preferably, S4 is as follows: A three-layer neural network is used to simulate the interaction between students and exercises, and the overall difficulty of the exercises obtained from S3 is analyzed. Students' cognitive status The enhanced continuous Q matrix obtained from S1 And the discrimination of exercises generated by exercise IDs The three-layer neural network is input for training; The main loss function is the binary cross-entropy loss function. And combined with the contrastive learning loss function of S3 and the policy loss function for the enhancement of the continuous Q matrix. The overall loss function is obtained. ; Based on the overall loss function The parameters of the neural network are optimized, and a well-trained cognitive diagnostic model is obtained through multiple iterations.
[0016] Preferably, in S4, the exercise discrimination index From the one-hot vector of the exercise The specific training process for acquiring cognitive diagnostics is as follows: ; in, For trainable Q-matrix One row represents the association vector between a specific exercise and a knowledge concept; Then, The three-layer neural network is trained using the following calculation process: ; ; ; in, The students output by the model For exercises The predicted score, obtained after training, is the student's score. Cognitive state vector .
[0017] Preferably, in S4, the binary cross-entropy loss function The expression is: ; in, For students For exercises The true score; For enhanced continuity Matrix-corrected policy loss function The expression is: ; in, For trainable augmented continuous matrix; By combining subjective difficulty comparison learning strategies, the overall loss function of the final cognitive diagnostic model is obtained. The expression is as follows: ; in, The weighting of the loss in the subjective difficulty comparison of the exercises. The weights of the strategy are adjusted for the Q matrix.
[0018] Therefore, the present invention employs the above-mentioned neurocognitive diagnostic method based on enhanced perception of exercise difficulty, and the beneficial effects are as follows: (1) Guided by educational theory, this invention explores deeper difficulty factors that affect the cognitive diagnosis process, decouples the difficulty of exercises into objective difficulty and subjective difficulty, thereby expanding the original cognitive diagnosis paradigm.
[0019] (2) Based on the large language model, this invention designs a relation proposer and relation discriminator strategy, constructs an exercise-knowledge concept graph, and annotates the Bloom cognitive level of the exercise based on the thinking chain technology, which alleviates the problem of huge manual annotation costs.
[0020] (3) The present invention designs a cognitive adaptive semantic-structural fusion mechanism, using Bloom vectors as a “cognitive regulator” to solve the problem of information dependence differences in exercises at different cognitive levels: for low-level cognitive exercises, the model is guided to focus on semantic entity information within the text; for high-level cognitive exercises, the model is guided to focus on topological structure information in the knowledge graph; this mechanism realizes the dynamic fusion of semantic perception information and structural perception information of exercises, generating a more realistic cognitive complexity fusion perception representation of exercises, effectively improving the performance and interpretability of the cognitive diagnostic model.
[0021] (4) The present invention designs a subjective difficulty modeling method for exercises, which extends the interaction between students’ cognitive state and exercise difficulty in the existing method to the interaction between students’ potential cognitive state and exercise potential difficulty, and simulates the interaction process between students and exercises in a more refined manner; in order to alleviate the problem of missing subjective difficulty labels, a contrastive learning strategy is adopted to supervise the generation of subjective difficulty.
[0022] (5) By decoupling the subjective and objective difficulty of the exercises, this invention effectively improves the interpretability of the results output by the cognitive diagnostic model.
[0023] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of the overall process of the neurocognitive diagnostic method based on enhanced perception of exercise difficulty in this invention; Figure 2 This is the relationship proposer template for the present invention based on the enhanced perception of exercise difficulty in neurocognitive diagnostic method; Figure 3 This invention provides a relation discriminator template for a neurocognitive diagnostic method based on enhanced perception of exercise difficulty. Figure 4 This is the Bloom annotation template for the neurocognitive diagnostic method based on enhanced perception of exercise difficulty in this invention; Figure 5 This is the hyperparameter of the atlas embedding dimension in Embodiment 1 of the present invention, which is based on the neurocognitive diagnostic method of enhanced problem difficulty perception. Figures showing the results of parameter sensitivity analysis on Math1 and ASSIST; Figure 6 This is the hyperparameter of the atlas embedding dimension in Embodiment 1 of the present invention, which is based on the neurocognitive diagnostic method of enhanced problem difficulty perception. Graphs showing the results of parameter sensitivity analysis on MOOPer and DS2023; Figure 7 This is the Q-matrix correction hyperparameter of Embodiment 1 of the present invention, which is based on the neurocognitive diagnostic method of enhanced problem difficulty perception. Figures showing the results of parameter sensitivity analysis on Math1 and ASSIST; Figure 8 This is the Q-matrix correction hyperparameter of Embodiment 1 of the present invention, which is based on the neurocognitive diagnostic method of enhanced problem difficulty perception. Graphs showing the results of parameter sensitivity analysis on MOOPer and DS2023; Figure 9 This is the contrastive learning weight hyperparameter of Implementation Example 1 of the Neurocognitive Diagnostic Method Based on Enhanced Perception of Exercise Difficulty in this Invention. Figures showing the results of parameter sensitivity analysis on Math1 and ASSIST; Figure 10 This invention relates to a comparative learning weight hyperparameter in an embodiment of a neurocognitive diagnostic method based on enhanced problem difficulty perception. Graphs showing the results of parameter sensitivity analysis on MOOPer and DS2023; Figure 11 This is a comparison chart of the interpretability index of Example 1 of the neurocognitive diagnostic method based on enhanced perception of exercise difficulty in this invention on MOOPer and DS2023. Figure 12 This is a visualization bar chart on Math1 of the diagnostic cases of Embodiment 1 of the neurocognitive diagnostic method based on enhanced perception of exercise difficulty of the present invention; Figure 13 This is a radar image on Math1 of the diagnostic case of the neurocognitive diagnostic method based on enhanced perception of exercise difficulty of the present invention. Detailed Implementation
[0025] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0026] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.
[0027] The process of this invention is as follows: Constructing a question-knowledge concept graph using a large language model; labeling question Bloom levels; obtaining an enhanced continuous Q matrix using BGE; extracting semantic-aware representations of questions using BGE; extracting structural-aware representations of questions using TransE; fusing the semantic-aware and structural-aware representations of questions using Bloom as prior guidance; projecting the fused question-aware representations as the potential objective difficulty of the questions, concatenating them with the basic objective difficulty to obtain the objective difficulty of the questions, and using a contrastive learning strategy to optimize the generation of subjective difficulty; finally, concatenating the objective and subjective difficulty to obtain the comprehensive difficulty of the questions; inputting the student's cognitive state, the comprehensive difficulty of the questions, the question's discrimination index, and the enhanced continuous Q matrix into a three-layer neural network for training, obtaining the student's probability of answering questions correctly, and performing performance prediction and cognitive state acquisition.
[0028] like Figure 1 As shown, the neurocognitive diagnostic method based on enhanced perception of exercise difficulty includes the following steps: S1. Define the number of students in the dataset as... The number of exercises is The number of knowledge concepts is The semantic-aware representation, structure-aware representation, Bloom vector, and enhanced continuous Q-matrix of exercises are obtained from the dataset, including the following steps: S11. Use BGE to embed the exercise text to obtain the semantic-aware representation of a specific exercise. .
[0029] S12, Design based on large language model, such as Figure 2 and Figure 3 The relation proposer and relation discriminator templates shown are used to mine semantic associations from exercise texts, knowledge concepts, and Q-matrices to construct an enhanced exercise-knowledge concept graph. Then, TransE is used to embed the graph into the graph to obtain a structure-aware representation of a specific exercise. .
[0030] S13. Bloom cognitive levels of exercises are annotated using mind chain technology. A template for annotating Bloom levels is as follows: Figure 4 As shown, the large model output labels are converted into multi-hot vectors. It employs a recursive hierarchical weighting strategy to model the dependencies between cognitive levels, obtaining Bloom vectors for specific exercises. .
[0031] This invention first defines six basic cognitive vectors. ,in A hierarchical weighting strategy is adopted, and the calculation process is as follows: ; ; in, , Representing dimension, This represents the specific dimension of each defined basic cognitive vector, while These are learnable scalar parameters used to control from lower-order dimensions. To the current dimension That is, the first Information inheritance rate of each cognitive vector; , Formula Each of them The vector formed represents the attention coefficient of each basic cognitive vector; This is a learnable weight matrix.
[0032] S14. Based on the Q matrix, an enhanced continuous Q matrix is obtained using a contrastive learning fine-tuning strategy of BGE. .
[0033] S2. Using the Bloom vector obtained in S1 as a priori, the semantic-aware representation and the structural-aware representation of the exercises are adaptively fused to obtain the final fused perceptual representation of the exercises, specifically: Bloom vector obtained using S1 As a priori, the semantic perception representation of the exercises is... and the perceptual representation of exercise structure Projecting onto a unified metric space specifically involves: ; in, To unify the feature dimensions of the metric space, , , is the learnable projection parameter matrix.
[0034] Then, attention is performed based on metrics. and Two-way alignment, specifically: ; ; ; in, This indicates two fully connected layers; The gating coefficients are used to perform bidirectional feature correction. and These are the corrected semantic-aware vector and the structural-aware vector, respectively.
[0035] Further improve the semantic perception representation of exercises and the perceptual representation of exercise structure The adaptive fusion yields the final problem fusion perceptual representation. Specifically: ; ; in, The symbol [.;.] represents a concatenation operation along the feature dimension, mapping the hidden layer features to two scalars, corresponding to the text-aware weights. and structure-aware weights and satisfy ; The weight matrix is a learnable matrix; For bias terms, This is a layer normalization operation.
[0036] S3. The projected perceptual representation of the exercises is used as the potential objective difficulty of the exercises. This is then concatenated with the basic objective difficulty of the exercises to obtain the objective difficulty of the exercises. Based on the objective difficulty of the exercises and the cognitive state generated by the student ID, the subjective difficulty of the exercises is modeled. A contrastive learning strategy is used to optimize the generation of subjective difficulty. The objective difficulty and subjective difficulty are then concatenated to obtain the comprehensive difficulty of the exercises. Specifically, S3 is as follows: First, integrate the exercises obtained in S2 with perceptual representations. After MLP projection, it serves as the potential objective difficulty of the exercises. The objective difficulty level of the exercises is generated by combining them with the exercise ID. By piecing together the data, we can obtain the objective difficulty level of the exercises. The calculation process is as follows: ; ; ; in, Let be the one-hot vector of the exercise. , For a trainable weight matrix, and This is a bias term.
[0037] Based on the objective difficulty of the exercises and cognitive states generated by student ID Subjective difficulty of modeling exercises The calculation process is as follows: ; ; ; ; ; in, For the student's one-hot vector, the gate value This indicates the student's level of acceptance of objective difficulty; offset value. This indicates the direction of fluctuation in the perceived difficulty level. To assess the overall difficulty of the exercises, we need to consider both the objective difficulty level and the overall difficulty level of the exercises. Subjective difficulty of exercises It was pieced together.
[0038] Further employing a contrastive learning strategy to optimize the generation of subjective difficulty in exercises, the specific calculation process is as follows: We introduce group cognitive consensus and stratified loss as monitoring signals for subjective difficulty, specifically for target exercises. The set of students who answered correctly is defined as the positive consensus group. The set of students who answered incorrectly is defined as the negative consensus group. .
[0039] To eliminate noise caused by individual errors or guesswork, a positive consensus group is computed. and negative consensus groups The centroid in the subjective difficulty semantic space is used to represent the average cognitive state of the two groups. The calculation process of the centroid is as follows: ; in, The model represents students In the exercises The personalized subjective difficulty vector generated above; and The number of students in the positive consensus group and the negative consensus group, respectively; and These represent the centroids of the positive and negative consensus groups in the subjective difficulty semantic space, respectively.
[0040] Based on the aforementioned centroid, cognitive consensus loss is adopted. Minimize the perceptual variance within the group: .
[0041] This invention employs layered loss Aiming to expand the positive consensus group and negative consensus groups To establish clear capability boundaries by addressing the perceived gap between them: ; in, The pre-set difficulty level differences. This represents the L2 norm, used to quantify the overall magnitude of the subjective difficulty vector.
[0042] S4. Train the cognitive diagnostic model, specifically as follows: A three-layer neural network is used to simulate the interaction between students and exercises, and the overall difficulty of the exercises obtained from S3 is analyzed. Students' cognitive status The enhanced continuous Q matrix obtained from S1 And the discrimination of exercises generated by exercise IDs The three-layer neural network is input for training.
[0043] The main loss function is the binary cross-entropy loss function. And combined with the contrastive learning loss function of S3 and the policy loss function for the enhancement of the continuous Q matrix. The overall loss function is obtained. .
[0044] Then based on the overall loss function The parameters of the neural network are optimized, and a well-trained cognitive diagnostic model is obtained through multiple iterations.
[0045] In this step, the exercise discrimination... From the one-hot vector of the exercise Obtain, and Similarly, the training process for cognitive diagnosis is as follows: ; in, For trainable Q-matrix A single line represents the association vector between a specific exercise and a knowledge concept.
[0046] Then, The three-layer neural network is trained using the following calculation process: ; ; ; in, The students output by the model For exercises The predicted score, obtained after training, is the student's score. Cognitive state vector .
[0047] The main training function of the cognitive diagnostic model of this invention is the binary cross-entropy loss function. Its expression is: ; in, For students For exercises The actual score.
[0048] To overcome the shortcomings of mislabeling and omission in traditional manual Q-matrices, S14 achieves enhanced continuity. Matrix. However, relying solely on semantic features cannot fully correct cognitive biases in the original annotation. Therefore, this invention designs a method for enhancing continuous matrix. Matrix-corrected policy loss function Its expression is: ; in, For trainable augmented continuous matrix.
[0049] By combining subjective difficulty comparison learning strategies, the overall loss function of the final cognitive diagnostic model is obtained. The expression is as follows: ; in, The weighting of the loss in the subjective difficulty comparison of the exercises. The weights of the strategy are adjusted for the Q matrix.
[0050] S5. During the testing phase, the test data is input into the trained cognitive diagnostic model to obtain the probability of students answering the questions correctly, and to predict their performance and acquire their cognitive status.
[0051] Example 1 This invention was evaluated on four cognitive diagnostic datasets, comparing its performance with state-of-the-art cognitive diagnostic methods. Tables 1 and 2 show the comparison results of this invention with state-of-the-art cognitive diagnostic methods on the Math1, ASSIST, MOOPer, and DS2023 datasets, respectively, with the best results indicated in bold.
[0052] Table 1. Comparative experimental results of different cognitive diagnostic methods on the Math1 and ASSIST datasets.
[0053] Table 2. Comparative experimental results of different cognitive diagnostic methods on the MOOPer and DS2023 datasets.
[0054] In Table 1, the hyphen "-" indicates that experiments could not be conducted because the dataset did not provide the exercise text data required for model training. On the Math1 and ASSIST datasets, which lack exercise text content, this invention performed exceptionally well. Except for a slightly lower AUC metric than the CCD model on the ASSIST dataset, it outperformed other baseline models in ACC, RMSE, and AUC. Specifically, on the Math1 dataset, this invention showed the most significant improvement in ACC, achieving a 1.54% improvement compared to the classic NCDM model and a 2.15% improvement compared to the KaNCD model, which also utilizes knowledge structure information. On the ASSIST dataset, this invention also maintained its performance advantage.
[0055] This demonstrates that even without textual semantic assistance, this invention can effectively utilize the structure-aware information provided by the enhanced exercise-knowledge concept graph. By mining the positional relationships of exercises within the knowledge architecture and the dependency logic between knowledge concepts, the model successfully constructs a high-quality objective difficulty representation of exercises, thus achieving more accurate diagnosis than models such as KaNCD even in purely structured data scenarios.
[0056] As shown in Table 2, the advantages of this invention are further amplified in the DS2023 and MOOPer datasets, which provide rich exercise texts, fully validating the effectiveness of introducing semantically aware information. On the DS2023 dataset, despite the high baseline performance of each model (NCDM's AUC reached 0.8379), this invention still achieved a significant breakthrough, with an AUC of 0.8618, a 2.39% improvement over NCDM, and outperformed the CNCD-F model, which also utilizes text features. On the MOOPer dataset, the performance improvement is particularly significant, with this invention achieving a 5.41% improvement in AUC compared to the NCDM model and a 2.43% improvement compared to the CNCD-F model.
[0057] In summary, this invention demonstrates state-of-the-art cognitive diagnostic performance on four different datasets, covering scenarios with and without provided exercise text, and student-exercise interaction records of small, medium, and large scales. The performance improvements across these datasets fully validate the significant contribution of this invention to enhancing cognitive diagnostic effectiveness.
[0058] Ablation Experiment Analysis: To verify the necessity and effectiveness of each module, as shown in Tables 3 and 4, this invention conducted ablation experiments on four datasets. By comparing the complete model with specific variants, the impact of semantic-aware representation of exercises, structure-aware representation of exercises, Q-matrix enhancement strategy, and subjective difficulty contrast learning strategy on model performance was explored. Due to the lack of exercise text data in the Math1 and ASSIST datasets, ablation experiments involving the w / -TEXT module could not be conducted on these two datasets. The best results are shown in bold. The NCD model was used as a baseline, and the contributions of enhancement modules to improving cognitive diagnostic performance were evaluated by progressively integrating them. The definitions of each variant are as follows: w / -EKG: indicates a variant that adds structure-aware representation of exercises to the baseline model. w / -TEXT: indicates a variant that adds semantic-aware representation of exercises to the baseline model. w / -CASSF: indicates a variant that incorporates structure-aware representation features, semantic-aware representation of exercises, and a fusion of these two representations into the baseline model. w / oQ: indicates a variant that removes the Q-matrix correction module. w / o-CL: This indicates that the variant of the comparative learning module that targets the subjective difficulty of exercises has been removed.
[0059] Table 3 Ablation experimental results of different models on Math1 and ASSIST datasets
[0060] Table 4. Ablation experimental results of different models on the DS2023 and MOOPer datasets.
[0061] Tables 3 and 4 show that, compared to the baseline model NCD, the variants introducing structure-aware representations of exercises (w / -EKG) and the variants introducing semantic-aware representations of exercises (w / -TEXT) significantly improve performance on all datasets. Particularly on the MOOPer dataset, the introduction of semantic information improves AUC by 4.31% and 4.63% with w / o-CL and w / oQ, respectively.
[0062] This indicates that the knowledge graph structural features and textual semantic features of the exercises contain rich information about the perceived difficulty of the exercises, effectively compensating for the shortcomings of relying solely on student responses. The variant that integrates both structural and semantic information (w / -CASSF) generally outperforms the single-feature variant in all metrics. This demonstrates that the structural and semantic information of the exercises is highly complementary, and combining them through the fusion mechanism designed in this invention can more comprehensively and three-dimensionally characterize the integrated features of the exercises. This invention achieves optimal results on all metrics and is significantly superior to the variant that removes the Q-matrix correction module (w / oQ) and the variant that removes the contrastive learning module (w / o-CL).
[0063] Sensitivity analysis. To systematically study the impact of hyperparameters, this invention conducted three types of experiments, focusing on the problem structure-aware feature dimension. Comparative learning hyperparameters and Q-matrix correction hyperparameter Impact on model performance, such as Figures 5-10 As shown.
[0064] Experimental results show that: (1) Graph embedding dimension The experimental results are as follows Figure 5 and Figure 6 As shown, for datasets with fewer interactions (i.e., Math1, ASSIST, and DS2023), the best performance is respectively at... =50、 =100 and The value is reached when =50, which is suitable for datasets with large interaction volumes (i.e., MOOPer). Optimal performance is achieved when the value is 200.
[0065] (2) Q-matrix correction coefficient hyperparameter The experimental results are as follows Figure 7 and Figure 8 As shown, for datasets with fewer interactions (i.e., Math1, ASSIST, and DS2023), the best performance is in [missing data]. =0.01 is reached, which is suitable for datasets with large interaction volumes (i.e., MOOPer). Optimal performance is achieved when the value is 0.001.
[0066] (3) Subjective difficulty comparison learning coefficient hyperparameter The experimental results are as follows Figure 9 and Figure 10 As shown, for datasets with fewer interactions (i.e., Math1, ASSIST, and DS2023), the best performance is in [missing data]. The value is reached when =0.1, which is suitable for datasets with large interaction volumes (i.e., MOOPer). Optimal performance is achieved when the value is 0.01.
[0067] Interpretability analysis of diagnostic results. To qualitatively evaluate effectiveness, this invention calculated the DOA (Directory of Assessment) of state-of-the-art models and this invention on the MOOPer and DS2023 datasets, such as… Figure 11As shown, the DOA metric measures the consistency of diagnostic results. The results demonstrate that this invention achieved the highest DOA values on both datasets, proving its superior interpretability. Although the traditional DINA model establishes a very high benchmark due to its binary assumptions and high consistency with the DOA metric logic, this invention still surpasses it. This indicates that by introducing exercise structure awareness and semantic awareness, this invention effectively constrains the latent spatial distribution of the neural network, successfully achieving diagnostic interpretability that is highly consistent with the actual response performance while maintaining high prediction accuracy.
[0068] In a horizontal comparison of neural network models, the DOA ranking shows a trend of this invention > CCD > CNCD-F > NCDM. The superiority of this invention over CNCD-F confirms that, compared to relying solely on shallow textual information, mining deep semantic logic through thought chain templates and combining it with graph structure constraints can effectively eliminate diagnostic ambiguity. Furthermore, by utilizing objective difficulty anchors and subjective difficulty offset mechanisms, the model can more reasonably explain the differences in responses among students with similar abilities. This invention significantly outperforms KaNCD and NCDM, further demonstrating that introducing multi-source auxiliary information (textual semantics and graph structure) can effectively supplement the diagnostic context and help the model find more reasonable explanation paths in complex response patterns.
[0069] Visual analysis of diagnostic cases. Exercise 21 from the Math1 dataset was selected as a case study for analysis. For example... Figure 12 As shown, traditional models often face the "interpretability paradox," namely, a logical misalignment between students' ability level (proficiency) and the difficulty of the exercises. For example, a high-achieving student (number 1110) has a proficiency (0.6770) in knowledge concept K6 that is lower than the objective difficulty of that knowledge concept (0.6943). Although the model predicts a high probability of correct answers, logically, an ability lower than the difficulty should lead to failure, creating a contradiction. This invention identifies the student's high ability and dynamically reduces the subjective difficulty to 0.5648, thus successfully restoring the reasonable logical relationship of "proficiency > difficulty" and resolving the aforementioned paradox.
[0070] Conversely, for the low-achieving student (number 143), the model increased the subjective difficulty to 0.8411, thus providing a reasonable explanation for their lower predicted score. Figure 13 This mechanism was further visualized using radar charts. Compared to the static objective difficulty (gray dashed line), the subjective difficulty profile of high-level students (cyan solid line) generally contracts inward, indicating that high proficiency effectively reduces students' perceived difficulty of the exercises; conversely, the subjective difficulty profile of low-level students (blue solid line) shows a trend of expansion outward. This significant morphological difference of "contraction for high-achieving students and expansion for struggling students" confirms that this invention can achieve adaptive personalized difficulty modeling based on individual cognitive states.
[0071] Therefore, this invention adopts the above-mentioned neurocognitive diagnostic method based on enhanced perception of exercise difficulty. By integrating knowledge graph structure and semantic information of exercise text, and combining Bloom cognitive hierarchy modeling and contrastive learning strategies, it accurately constructs the comprehensive difficulty of exercises, significantly improves the prediction accuracy and interpretability of cognitive diagnosis, outperforms existing advanced models on multiple datasets, is adaptable to different scales of educational scenarios, and provides more reliable technical support for smart education.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A neurocognitive diagnostic method based on enhanced perception of exercise difficulty, characterized in that, Includes the following steps: S1. Obtain the semantic-aware representation of the exercise, the structural-aware representation of the exercise, the Bloom vector of the exercise, and the enhanced continuous Q matrix, respectively; S2. Using the Bloom vector obtained in S1 as a priori, the semantic perception representation and the structural perception representation of the exercises are adaptively fused to obtain the final fused perception representation of the exercises. S3. The potential objective difficulty of the exercises is obtained by combining the sensory representation projection of the exercises with the basic objective difficulty of the exercises. The subjective difficulty of the exercises is modeled based on the objective difficulty of the exercises and the cognitive state generated by the student ID. A contrastive learning strategy is used to optimize the generation of subjective difficulty. The objective difficulty and subjective difficulty are then combined to obtain the comprehensive difficulty of the exercises. S4. A three-layer neural network is used to simulate the interaction process between students and exercises. The comprehensive difficulty of the exercises obtained in S3, the student's cognitive state, the enhanced continuous Q matrix obtained in S1, and the exercise discrimination generated by the exercise ID are input into the three-layer neural network for training to obtain a trained cognitive diagnostic model. S5. Input the test data into the trained cognitive diagnostic model to obtain the probability of students answering the questions correctly, and then predict their performance and acquire their cognitive status.
2. The neurocognitive diagnostic method based on enhanced perception of exercise difficulty as described in claim 1, characterized in that, In S1, the number of students in the dataset is defined as... The number of exercises is The number of knowledge concepts is The semantic-aware representation, structure-aware representation, Bloom vector, and enhanced continuous Q-matrix of exercises are obtained from the dataset, including the following steps: S11. Use BGE to embed the exercise text to obtain the semantic-aware representation of a specific exercise. ; S12. Based on a large language model, design relation proposer and relation discriminator templates. Mine semantic associations from exercise texts, knowledge concepts, and Q-matrices to construct an enhanced exercise-knowledge concept graph. Then, use TransE to perform embedding learning on the graph to obtain a structure-aware representation of a specific exercise. ; S13. Bloom cognitive hierarchy of exercises labeled using mind chain technology, transforming the output of the large model into multi-hot vectors. It employs a recursive hierarchical weighting strategy to model the dependencies between cognitive levels, obtaining Bloom vectors for specific exercises. ; S14. Based on the Q matrix, an enhanced continuous Q matrix is obtained using a contrastive learning fine-tuning strategy of BGE. .
3. The neurocognitive diagnostic method based on enhanced perception of exercise difficulty as described in claim 2, characterized in that, In S13, six basic cognitive vectors are defined. ,in A hierarchical weighting strategy is adopted, and the calculation process is as follows: ; ; in, , Representing dimension, This represents the specific dimension of each defined basic cognitive vector, while These are learnable scalar parameters used to control from lower-order dimensions. To the current dimension That is, the first Information inheritance rate of each cognitive vector; , Formula Each of them The vector formed represents the attention coefficient of each basic cognitive vector; This is a learnable weight matrix.
4. The neurocognitive diagnostic method based on enhanced perception of exercise difficulty as described in claim 3, characterized in that, S2 specifically refers to: Bloom vector obtained using S1 As a priori, the semantic perception representation of the exercises is... and the perceptual representation of exercise structure Projecting onto a unified metric space specifically involves: ; in, To unify the feature dimensions of the metric space, , , The projection parameter matrix is learnable; Then, attention is performed based on metrics. and Two-way alignment, specifically: ; ; ; in, This indicates two fully connected layers; The gating coefficients are used to perform bidirectional feature correction. and These are the corrected semantic-aware vector and the structural-aware vector, respectively. Complete the semantic perception representation of the exercises and the perceptual representation of exercise structure The adaptive fusion yields the final problem fusion perceptual representation. Specifically: ; ; in, The symbol [.;.] represents a concatenation operation along the feature dimension, mapping the hidden layer features to two scalars, corresponding to the text-aware weights. and structure-aware weights and satisfy ; The weight matrix is a learnable matrix; For bias terms, This is a layer normalization operation.
5. The neurocognitive diagnostic method based on enhanced perception of exercise difficulty as described in claim 4, characterized in that, S3 specifically refers to: The exercises obtained in S2 are integrated with perceptual representations. After MLP projection, it serves as the potential objective difficulty of the exercises. The objective difficulty level of the exercises is generated by combining them with the exercise ID. By piecing together the data, we can obtain the objective difficulty level of the exercises. The calculation process is as follows: ; ; ; in, Let be the one-hot vector of the exercise. , For a trainable weight matrix, and For bias terms; Based on the objective difficulty of the exercises and cognitive states generated by student ID Subjective difficulty of modeling exercises The calculation process is as follows: ; ; ; ; ; in, For the student's one-hot vector, the gate value This indicates the student's level of acceptance of objective difficulty; offset value. This indicates the direction of fluctuation in the perceived difficulty level. To assess the overall difficulty of the exercises, we need to consider both the objective difficulty level and the overall difficulty level of the exercises. Subjective difficulty of exercises It was pieced together.
6. The neurocognitive diagnostic method based on enhanced perception of exercise difficulty as described in claim 5, characterized in that, In S3, a contrastive learning strategy is used to optimize the generation of subjective difficulty in exercises. The specific calculation process is as follows: Introducing group cognitive consensus and stratified loss as monitoring signals for subjective difficulty, for target exercises... The set of students who answered correctly is defined as the positive consensus group. The set of students who answered incorrectly is defined as the negative consensus group. ; Computing a positive consensus group and negative consensus groups The centroid in the subjective difficulty semantic space is used to represent the average cognitive state of the two groups. The calculation process of the centroid is as follows: ; in, The model represents students In the exercises The personalized subjective difficulty vector generated above; and The number of students in the positive consensus group and the negative consensus group, respectively; and These represent the centroids of the positive and negative consensus groups in the subjective difficulty semantic space, respectively. Based on the aforementioned centroid, cognitive consensus loss is adopted. Minimize the perceptual variance within the group: ; Layered loss Expand the positive consensus group and negative consensus groups To establish clear capability boundaries by addressing the perceived gap between them: ; in, The pre-set difficulty level differences. This represents the L2 norm, used to quantify the overall magnitude of the subjective difficulty vector.
7. The neurocognitive diagnostic method based on enhanced perception of exercise difficulty as described in claim 6, characterized in that, S4 specifically refers to: A three-layer neural network is used to simulate the interaction between students and exercises, and the overall difficulty of the exercises obtained from S3 is analyzed. Students' cognitive status The enhanced continuous Q matrix obtained from S1 And the discrimination of exercises generated by exercise IDs The three-layer neural network is input for training; The main loss function is the binary cross-entropy loss function. And combined with the contrastive learning loss function of S3 and the policy loss function for the enhancement of the continuous Q matrix. The overall loss function is obtained. ; Based on the overall loss function The parameters of the neural network are optimized, and a well-trained cognitive diagnostic model is obtained through multiple iterations.
8. The neurocognitive diagnostic method based on enhanced perception of exercise difficulty as described in claim 7, characterized in that, In S4, the discrimination of exercises From the one-hot vector of the exercise The specific training process for acquiring cognitive diagnostics is as follows: ; in, For trainable Q-matrix One row represents the association vector between a specific exercise and a knowledge concept; Then, The three-layer neural network is trained using the following calculation process: ; ; ; in, The students output by the model For exercises The predicted score, obtained after training, is the student's score. Cognitive state vector .
9. The neurocognitive diagnostic method based on enhanced perception of exercise difficulty as described in claim 8, characterized in that, In S4, the binary cross-entropy loss function The expression is: ; in, For students For exercises The true score; For enhanced continuity Matrix-corrected policy loss function The expression is: ; in, For trainable augmented continuous matrix; By combining subjective difficulty comparison learning strategies, the overall loss function of the final cognitive diagnostic model is obtained. The expression is as follows: ; in, The weighting of the loss in the subjective difficulty comparison of the exercises. The weights of the strategy are adjusted for the Q matrix.