A Likelihood-Based Q-Matrix Estimation Method for Cognitive Diagnosis

By introducing AIC and BIC as model selection indicators and using an iterative optimization algorithm to estimate the Q matrix, the problems of time-consuming and mis-specified Q matrix construction by experts are solved, efficient and accurate Q matrix estimation is achieved, and the accuracy and applicability of the cognitive diagnosis model are improved.

CN120596861BActive Publication Date: 2025-10-14JINAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511106894.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-10-14
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

In the existing technology, the construction of Q matrix relies on expert experience, which has the limitations of strong subjectivity, high time consumption, high cost and missetting, affecting the accuracy and efficiency of cognitive diagnosis models.

Method used

By using AIC and BIC as model selection indicators and utilizing an iterative optimization algorithm to estimate the unknown part of the Q matrix, the dependence on experts is reduced and the accuracy and efficiency of Q matrix estimation are improved.

Benefits of technology

It improves the accuracy and efficiency of Q-matrix estimation, reduces labor costs, is applicable to the complex G-DINA model framework, and enhances the accuracy and applicability of cognitive diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596861B_ABST
    Figure CN120596861B_ABST
Patent Text Reader

Abstract

The application discloses a likelihood-based cognitive diagnosis Q matrix estimation method, belongs to the field of cognitive diagnosis evaluation, and comprises the following steps: obtaining response data of students to questions and a partially defined Q matrix, fitting a G-DINA model to obtain a posterior distribution of a student attribute vector, and calculating a marginal likelihood based on the posterior distribution. Furthermore, the fitting degrees of different q vectors are evaluated through AIC or BIC values, the optimal q vector is selected to gradually fill unknown parts of the Q matrix, and the process is continued until convergence. The method reduces the dependence on experts, improves the accuracy and efficiency of Q matrix estimation, and is suitable for large-scale education evaluation scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of cognitive diagnostic assessment, and particularly relates to a likelihood-based cognitive diagnostic Q-matrix estimation method. BACKGROUND

[0002] Cognitively Diagnostic Assessment (CDA) is an evaluation method for measuring students' mastery of knowledge points in a specific field. Its core lies in analyzing students' knowledge structure and ability level through a cognitive diagnostic model (CDM), thereby providing a basis for personalized teaching and learning. As an important component of CDMs, Q-matrix defines the relationship between items and attributes, and is a key tool for precise diagnosis.

[0003] In the prior art, Q-matrix is usually constructed manually by domain experts based on experience and knowledge. However, this method has obvious limitations. First, the subjectivity of expert judgment can lead to inaccurate Q-matrix. Second, as the number of items increases, the process of manually constructing and verifying Q-matrix by experts becomes extremely time-consuming and costly. In addition, the Q-matrix constructed by experts may have misjudgments, which can affect the fitting effect of the model and data. SUMMARY

[0004] Obtaining response data of students to items and a partially defined Q-matrix;

[0005] Taking AIC and BIC as model selection indicators, estimating the unknown part of the Q-matrix through an iterative optimization algorithm;

[0006] Outputting a final estimated Q-matrix based on the unknown part of the Q-matrix.

[0007] Preferably, the calculation formula of AIC and BIC is:

[0008] ;

[0009] ;

[0010] wherein, AIC is Akaike Information Criterion, BIC is Bayesian Information Criterion, L is the marginal likelihood of the model, P is the number of model parameters, and N is the sample size.

[0011] Preferably, the process of estimating the unknown part of the Q-matrix through an iterative optimization algorithm comprises:

[0012] Initializing the unknown part of the Q-matrix as an empty matrix;

[0013] using the fitted G-DINA model, obtaining the posterior distribution of the student attribute vector according to the student response data to the questions and the partially defined Q matrix;

[0014] based on the posterior distribution of the student attribute vector, calculating the marginal likelihood of each question;

[0015] for each question in the unknown part of the Q matrix, calculating the AIC or BIC value of all possible q vectors based on the marginal likelihood of the question, and selecting the q vector with the minimum AIC or BIC value as the optimal q vector of the question;

[0016] updating the unknown part of the Q matrix, and re-fitting the G-DINA model using the updated Q matrix until the unknown part of the Q matrix no longer changes;

[0017] outputting the final estimated Q matrix.

[0018] Preferably, the process of calculating the marginal likelihood of the model comprises:

[0019] calculating the conditional likelihood of each question according to the student response data to the questions and the posterior distribution of the student attribute vector;

[0020] weighting and averaging the conditional likelihood to obtain the marginal likelihood of each question.

[0021] Preferably, in the iterative optimization algorithm, the posterior distribution of the student attribute vector is dynamically updated according to the latest estimated value in each iteration.

[0022] Preferably, the termination condition of the iterative optimization algorithm is that the unknown part of the estimated Q matrix no longer changes, specifically: comparing the unknown part of the Q matrix of the current iteration and the last iteration, when they are consistent, stopping the iteration.

[0023] Preferably, the student response data to the questions includes the student's answering situation and the corresponding question information, which is used to fit the G-DINA model and calculate the marginal likelihood of the model.

[0024] Preferably, the partially defined Q matrix is pre-defined by domain experts and contains the known attribute relationships of part of the questions, which is used to guide the estimation of the unknown part of the Q matrix.

[0025] In another aspect, the present application also provides an electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.

[0026] In another aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method.

[0027] Compared with the prior art, the present application has the following advantages and technical effects:

[0028] The present application aims at the shortcomings of the traditional Q matrix construction method, and efficiently estimates the Q matrix through a data-driven method based on model fitting indicators (AIC and BIC), which is especially suitable for complex G-DINA model frameworks.

[0029] In terms of technical effects, the present application introduces AIC and BIC as Q matrix selection criteria under the G-DINA model framework, effectively avoids overfitting problems, and improves the accuracy of Q matrix estimation. Compared with the Q matrix constructed by experts, the data-driven method of the present application can punish the model complexity through AIC / BIC indicators, and then select the optimal Q vector combination, thereby improving the fitting degree of the model. Simulation experiments show that the Q matrix constructed based on the method of the present application performs well in complex cognitive diagnosis tests. When using AIC as the selection criterion, the recovery rate reaches 94.34%; when using BIC, the recovery rate can reach 99.99%. These recovery rates are much higher than the traditional expert-constructed Q matrix, and can remain stable under multiple experimental conditions.

[0030] The present application can effectively support the estimation of Q matrix under the G-DINA model framework. The G-DINA model is a widely used general framework, which can be compatible and extended to DINA, DINO and A-CDM and other cognitive diagnosis models. The application of the method of the present application in this framework makes the method applicable to more complex models and can be generalized to multiple application scenarios. By introducing AIC and BIC criteria, the calculation process of Q matrix estimation is simplified, and the computational complexity brought by the Bayesian method and Markov chain Monte Carlo method is avoided.

[0031] In terms of economic effects, the present application reduces the dependence on domain experts and significantly reduces the labor cost in the Q matrix construction process. Traditionally, Q matrix needs to be constructed by experts based on subjective judgment. Especially when the number of questions and attributes is large, the input cost of experts is extremely high. Using the method of the present application, only a small number of experts are needed to participate in defining part of the Q matrix, and the rest can be automatically estimated by data-driven methods, which greatly reduces the construction cost. Since experts do not need to construct and verify the relationship between questions and attributes one by one, the present application can significantly speed up the Q matrix construction process.

[0032] In terms of social effects, the efficient Q matrix estimation method of the application reduces the use threshold of complex cognitive diagnosis models, enabling more education and psychological evaluation institutions to widely adopt cognitive diagnosis models to provide more accurate personalized feedback, thereby improving the effectiveness of teaching and psychological diagnosis. BRIEF DESCRIPTION OF DRAWINGS

[0033] The accompanying drawings, which form a part of this application, are intended to provide further understanding of the application and are incorporated herein in their entirety. The schematic embodiments of the application and their descriptions are used to explain the application and do not constitute an improper limitation on the application. In the drawings:

[0034] Figure 1 Algorithmic diagram of the likelihood-based Q matrix estimation method of the embodiments of the application. DETAILED DESCRIPTION

[0035] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0036] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0037] Embodiment one

[0038] As Figure 1 shown, the likelihood-based cognitive diagnosis Q matrix estimation method provided in the present embodiment includes:

[0039] Part of the structure of the Q matrix is defined:

[0040] The Q matrix is divided into two parts: the defined part ( ) and the undefined part ( ). Predefined by domain experts, containing the attribute relationship of part of the questions; The undefined part needs to be estimated by a data-driven method.

[0041] Q matrix estimation method based on model fitting index:

[0042] Input data: student response data (X) to questions, defined Q matrix ( ).

[0043] Objective: Estimate the q vector of the undefined part .

[0044] Method: Utilize AIC and BIC as model selection indicators, estimate through iterative optimization algorithm .

[0045] Iterative optimization algorithm:

[0046] Initialization: Set to empty matrix.

[0047] Step 1: Use to fit G-DINA model and obtain posterior distribution of student attribute vector.

[0048] Step 2: For each item in , calculate AIC or BIC value for all possible q vectors, select the minimum value corresponding to the optimal solution.

[0049] Step 3: Update and use the new to refit the model until no longer changes.

[0050] Calculation of model fitting indicators:

[0051] Definition of AIC and BIC:

[0052] ;

[0053] ;

[0054] Where L is the marginal likelihood of the model, P is the number of model parameters, and N is the sample size.

[0055] Calculation of marginal likelihood: Calculate the marginal likelihood of each item through student response data and posterior distribution of attribute vector.

[0056] Dynamic updating mechanism:

[0057] In each iteration, update the posterior distribution of student attribute vector according to the latest estimation value, ensuring that the model can dynamically adjust according to the data.

[0058] Role of technical solutions:

[0059] Improve accuracy: Data-driven approach reduces reliance on expert experience, reducing human error.

[0060] Improve efficiency: Automated estimation process reduces time and cost of manually constructing and verifying Q matrix.

[0061] Enhance applicability: Support multiple CDM frameworks, including G-DINA model, suitable for different evaluation scenarios.

[0062] Example Two

[0063] A likelihood-based cognitive diagnostic Q-matrix estimation method is provided in this embodiment, which includes:

[0064] This embodiment is based on the G-DINA model framework, using AIC and BIC as model selection indicators, and estimating the unknown part of the Q matrix through iterative optimization algorithm. G-DINA model is a general cognitive diagnostic model, which can flexibly describe the complex relationship between items and attributes. AIC and BIC punish the complexity of the model, ensuring that the selected q vector achieves the best balance between model fitting and complexity. The main purpose of the simulation study is to verify the performance of the proposed likelihood-based Q matrix estimation method under different conditions, and to compare it with the existing G-DINA model discriminant index (GDI) method. Through simulation study, the applicability and accuracy of the method in different educational assessment scenarios can be evaluated, providing scientific basis for practical application.

[0065] Five factors were manipulated in the simulation study to comprehensively evaluate the performance of the likelihood-based Q matrix estimation method:

[0066] Model generation: Four cognitive diagnostic models were used to generate data, including DINA model, DINO model, A-CDM and saturated G-DINA model.

[0067] Number of attributes: Set to 3 or 5 attributes.

[0068] First part test length: According to the number of attributes, Set to 7, 14 (when K=3) or 25, 50 (when K=5).

[0069] Sample size: Set to 500, 1000 and 2000.

[0070] Item quality: Based on the lowest ( ) and highest ( ) success probability, divided into high item quality and low item quality.

[0071] Experimental conditions:

[0072] Number of attributes: K=3 or K=5.

[0073] First part test length: Or When K=3); Or When K=5).

[0074] Sample size: N=500, 1000, 2000.

[0075] Item quality: high item quality ( , ) and low item quality ( , ).

[0076] Generating model: DINA model, DINO model, A-CDM, and saturated G-DINA model.

[0077] Evaluation index:

[0078] The performance of Q matrix estimation methods is evaluated using the element recovery rate. The ERR formula is:

[0079] ;

[0080] where is the element of the estimated Q matrix, is the element of the true Q matrix.

[0081] The average recovery rate of Q matrix estimation methods based on likelihood methods performs well under different conditions. When using AIC, the average recovery rate ranges from 72.22% to 94.34%; when using BIC, the average recovery rate ranges from 72.00% to 99.99%; and when using GDI, the average recovery rate ranges from 59.08% to 100.00%.

[0082] The impact of sample size and item quality: When the sample size is larger or the item quality is higher, the recovery rate is higher. For example, when the generating model is the G-DINA model, the sample size is 500, and the item quality is low, the average recovery rate using AIC and BIC is 74.21% and 72.48%, respectively; while under the condition of high item quality, all recovery rates are higher than 91.59%.

[0083] The impact of the length of the first part of the test: When the length of the first part of the test is longer, the recovery rate is slightly higher. For example, when the item quality is high, , the recovery rate is slightly higher than .

[0084] The impact of the number of attributes: The number of attributes (K=3 or K=5) has no significant impact on the recovery rate, as long as contains all possible q vectors, enough information can be obtained from the data to estimate .

[0085] ​​Impact of the model of generation: Under A-CDM, the likelihood-based approach performs slightly worse, while it performs well under the other three models. This suggests that A-CDM might not be suitable for Q-matrix estimation using likelihood-based or GDI approaches.

[0086] Comparison with existing methods

[0087] Comparison with GDI approach: Under certain conditions, the recovery rate of the likelihood-based approach is higher than the best recovery rate of the GDI approach, especially when the number of attributes is large and the item quality is low. This suggests that the likelihood-based approach might be superior to the GDI approach under certain conditions.

[0088] Comparison of AIC and BIC: Under most conditions, AIC performs better than BIC when the item quality is low, while BIC performs better than AIC when the item quality is high. This suggests that the appropriate model selection criterion can be chosen according to the item quality.

[0089] Example Three

[0090] A likelihood-based cognitive diagnosis Q-matrix estimation method is provided in this example, which includes:

[0091] As Figure 1 shown, in theory, the correct q vector should provide the maximum likelihood value. However, since the likelihood value does not decrease when more attributes are specified in general CDM, the likelihood needs to be penalized according to the number of parameters for each item. Therefore, the model selection indicators AIC and BIC are used as criteria to select the correct q vector. These indicators penalize the number of parameters. The response of item j is conditionally independent of other item responses under the local independence assumption. Therefore, it is appropriate to estimate the q vector for each item in Part 2 separately. For each item in Part 2, the algorithm exhaustively compares the AIC and BIC values of all possible q vectors. The q vector with the smallest AIC or BIC value will be selected, because the correct q vector should provide the best fit and take into account the complexity of the model.

[0092]

[0093] or:

[0094]

[0095] Given , the AIC and BIC calculations are as follows:

[0096] ;

[0097] ;

[0098] where is the number of parameters for item j based on the G-DINA model, N is the number of examinees, is the marginal likelihood of the second part response data given the vector of attributes.

[0099] To obtain the information criterion, the marginal likelihood needs to be computed. This can be done by taking the product of the conditional likelihoods, as follows:

[0100] ;

[0101] where is the conditional distribution of the second part response data given the vector of attributes.

[0102] Given the vector , the marginal likelihood can be computed as because the probability of success in a simplified latent group is homogeneous within the group. is the probability that an examinee in a simplified latent group correctly answers item j in the second part, weighted by the prior distribution of can be expressed as:

[0103] ;

[0104] where, is defined as the probability of success for item j in the second part for a simplified latent group ; can be computed as:

[0105] ;

[0106] where, is the total number of expected examinees in a simplified latent group , is the number of expected examinees that correctly answer item j in a simplified latent group .

[0107] In this case, the prior distribution of latent groups , which in fact is the posterior probability of examinee i being in a simplified latent group , can be obtained by fitting a CDM to . This is done by first estimating, i.e. . This can then be used as the prior probability in the estimation of the Q matrix for the second part. Due to measurement error in the estimation, it cannot be considered as the true prior for the estimation of the Q matrix for the second part, which will lead to an erroneous estimation.

[0108] The posterior distribution of the examinee's attribute profile should be consistent between Part 1 and Part 2, because the whole test is administered to a group of examinees. An accurate estimate of the attribute profile posterior distribution is an important key to connect the two parts of the test. Therefore, it should be estimated based on the Empirical Bayes Expectation-Maximization (EM) method. Unlike the standard Bayesian method, the Empirical Bayes method uses the posterior distribution estimated in the previous iteration to update the prior distribution. Therefore, the proposed method employs the Empirical Bayes EM method to iteratively obtain and update the prior distribution by the newly estimated q vectors. When the estimated The algorithm stops when there is no change anymore.

[0109] As shown in FIG. 1, the designed algorithm is used to iteratively estimate the unspecified part of the Q matrix, the procedure of which is as follows: Figure 1 Step 1: Let

[0110] be the zero matrix of Part 2. Step 2: Fit the G-DINA model to

[0111] by using , and obtain .

[0112] Step 3: Calculate the AIC and BIC values of all q vectors to obtain . For select:

[0113]

[0114] or:

[0115] .

[0116] Step 4: Compare and . If is the same as , the algorithm stops; otherwise, perform Step 5.

[0117] Step 5: Fit the G-DINA model to X by using and , and obtain .

[0118] Step 6: Let = .

[0119] Step 7: Go back to Step 3 until the algorithm stops.

[0120] ​In another aspect, the present embodiment also provides an electronic device, comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.

[0121] In another aspect, the present embodiment also provides a computer readable storage medium, which stores a computer program, wherein the computer program is executable by a processor to implement the method.

[0122] The above merely provides the preferred embodiments of the present application, but the protection scope of the present application is not limited thereto, and any changes or replacements within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0123] The above merely provides the preferred embodiments of the present application, but the protection scope of the present application is not limited thereto, and any changes or replacements within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A likelihood-based cognitive diagnosis Q matrix estimation method, characterized in that: include: Obtain student response data to the questions and partially define the Q matrix; Using AIC and BIC as model selection indicators, the unknown part of the Q matrix is ​​estimated through an iterative optimization algorithm; Outputting a final estimated Q matrix based on the unknown portion of the Q matrix; The calculation formulas of AIC and BIC are: ; ; in, is the Akaike information criterion, is the Bayesian information criterion, L is the marginal likelihood of the model, P is the number of model parameters, and N is the sample size; The process of estimating the unknown part of the Q matrix by the iterative optimization algorithm includes: Initialize the unknown part of the Q matrix to an empty matrix; Use the fitted G-DINA model to obtain the posterior distribution of the student attribute vector based on the student response data and the partially defined Q matrix; Calculate the marginal likelihood of each question based on the posterior distribution of the student attribute vector; For each question in the unknown part of the Q matrix, calculate the AIC or BIC values ​​of all possible q vectors based on the marginal likelihood of each question, and select the q vector with the smallest AIC or BIC value as the optimal q vector for the question; Update the unknown part of the Q matrix and refit the G-DINA model using the updated Q matrix until the unknown part of the Q matrix no longer changes; Output the final estimated Q matrix; The calculation process of the marginal likelihood of the model includes: Calculate the conditional likelihood of each question based on the student's response data to the question and the posterior distribution of the student's attribute vector; The conditional likelihood is weighted averaged to obtain the marginal likelihood of each question.

2. The method according to claim 1, characterized in that In the iterative optimization algorithm, in each iteration, the posterior distribution of the student attribute vector is dynamically updated according to the latest estimated value.

3. The method according to claim 1, characterized in that The termination condition of the iterative optimization algorithm is that the unknown part of the estimated Q matrix no longer changes. Specifically, the unknown part of the Q matrix of the current iteration is compared with that of the previous iteration. When the two are consistent, the iteration is stopped.

4. The method according to claim 1, wherein The student's response data to the questions includes the student's answer status and the corresponding question information, which is used to fit the G-DINA model and calculate the marginal likelihood of the model.

5. The method according to claim 1, wherein The partially defined Q matrix is ​​pre-defined by domain experts, contains known attribute relationships of some questions, and is used to guide the estimation of the unknown part of the Q matrix.

6. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein: When the processor executes the computing program, the method according to any one of claims 1 to 5 is implemented.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Diagnostic report generation method, system and equipment

    CN117574876A

  • Cross-time-point cognitive diagnosis method considering student influence factors

    CN119830237A