A method for correcting cognitive diagnostic biases with covariates for multiple groups

By introducing a classification error probability matrix and a multi-group logistic regression model, and optimizing parameters, the flexibility and accuracy issues of existing cognitive diagnostic models in multi-group data analysis are resolved. This enables effective modeling of multi-group data and accurate characterization of covariates, thereby improving the applicability and accuracy of the model.

CN120495036BActive Publication Date: 2025-10-28JINAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510991117.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-28
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Existing cognitive diagnostic models suffer from insufficient structural flexibility, complex operation, difficulty in incorporating examinees' covariates, and estimation bias when processing multi-group data, thus affecting the accuracy and applicability of the models.

Method used

A cognitive diagnostic bias correction method with covariates for multiple groups is adopted. By introducing a classification error probability matrix and a multi-group logistic regression model, the model parameters are optimized, the classification error is corrected, and the flexibility and accuracy of the model are improved.

Benefits of technology

It improves the accuracy and stability of the model in multi-group data analysis, enabling it to better handle multi-group data, reveal the impact of covariates on cognitive attributes, and support personalized teaching and clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495036B_ABST
    Figure CN120495036B_ABST
Patent Text Reader

Abstract

This invention relates to the field of educational assessment technology, and particularly to a method for correcting cognitive diagnostic biases with covariates for multiple groups. The method includes: acquiring examinees' answer data and the relationship between question attributes; inputting the answer data and the relationship between question attributes into an attribute mastery probability interpretation model to obtain the examinees' attribute mastery probabilities for each cognitive attribute; the attribute mastery probability interpretation model is obtained by training multiple cognitive diagnostic models using a training set; during model training, the latent knowledge states of examinees are assigned based on the posterior distributions output by the multiple cognitive diagnostic models, and a classification error probability matrix is ​​calculated. Multiple latent logistic regression models are used to evaluate the influence of external covariate data on cognitive attribute mastery states, and the model parameters of the multiple cognitive diagnostic models are optimized using an objective function with the classification error probability matrix as the target weight. This invention solves several key problems in multi-group data processing and covariate modeling in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of educational assessment, psychometrics, and clinical diagnostics, and in particular to a method for correcting cognitive diagnostic biases with covariates for multiple groups. Background Technology

[0002] With the continuous development of educational measurement technology, cognitive diagnostic models (CDMs) are increasingly widely used in instructional assessment and personalized learning. They provide test-takers with more targeted and informative feedback, thus supporting instructional interventions and learning planning. These models are not only applied to various aspects of education, such as the assessment of proportional reasoning, spatial skills, and digital literacy, but are also gradually demonstrating their potential in other fields such as clinical diagnosis, becoming an important tool for analyzing individual cognitive states. Compared with traditional measurement methods, they offer higher diagnostic accuracy and interpretability, and have become an important tool in educational assessment research.

[0003] The G-DINA model, as a general cognitive diagnostic model, extends and expands upon the traditional DINA model. By relaxing the restrictions on attribute interaction effects, the G-DINA model can be equivalent to other cognitive diagnostic models based on alternative link functions under specific parameter settings. Therefore, the G-DINA model possesses good model adaptability and wide applicability, making it an ideal tool for analyzing candidates' mastery of multiple cognitive attributes.

[0004] Building upon this foundation, the MG-GDINA model was developed to meet the needs of analyzing cognitive differences among different groups of test takers. As a direct extension of the G-DINA model, this model can simultaneously consider the similarities and differences in cognitive attribute mastery among multiple groups (such as different schools or regions), thus providing a basis for educational decision-making and resource allocation. The introduction of the MG-GDINA model provides cognitive diagnostic models with greater adaptability and application value in scenarios involving the analysis of cognitive status across multiple groups.

[0005] In the field of educational measurement, using latent regression models to study the relationship between examinees' scale scores and covariates (such as gender, age, and family background) has always been a research focus. However, incorporating examinees' covariates into cognitive diagnostic models to analyze the relationship between these covariates and diagnostic results (such as attribute mastery levels or mastery patterns) has not yet been fully researched and applied. Existing techniques (such as the one-step method) can accurately assess the impact of examinees' covariates on diagnostic results, but these methods lack flexibility in model structure. In practice, every modification to any part of the model requires re-estimation of the entire process, which is cumbersome, inefficient, and limits its practicality and scalability.

[0006] In addition, to improve the flexibility of the model, some solutions adopt the "stepwise method" or the "three-step method". The three-step method mainly includes the following steps: CDM estimation: First, the cognitive diagnostic model is fitted to the test data of the candidates; latent category assignment: The latent categories of the candidates are assigned according to the fitting results; latent logistic regression: Finally, the influence of the candidates' covariates on their mastery of latent attributes is estimated by using a logistic regression model.

[0007] Although the three-step method enhances the flexibility of the model to some extent, it may still produce biased estimates because the classification error is not fully considered, resulting in an inaccurate assessment of the candidate's mastery of attributes.

[0008] More importantly, existing cognitive diagnostic models typically do not consider modeling issues in multi-group scenarios. Especially in practical educational assessment applications, it is often necessary to compare candidates from different genders and backgrounds. They lack the ability to process multi-group data, thus having certain limitations when comparing differences and similarities between multiple groups, affecting the accuracy of the analysis and the breadth of application.

[0009] In summary, existing technical methods still have the following significant drawbacks:

[0010] The model structure lacks flexibility, is complex to operate, has biases in estimating covariates of included candidates, and lacks support for modeling and comparing multi-group data across multiple systems.

[0011] Insufficient processing capacity for multiple datasets: Existing cognitive diagnostic models generally fail to adequately consider the heterogeneity among different groups (such as school type, region, and educational intervention conditions), making it difficult to effectively distinguish and compare the mastery status of candidates' potential cognitive attributes in multi-group scenarios. This limitation significantly affects the applicability and explanatory power of the models in various scenarios such as educational assessment, psychological evaluation, and clinical diagnosis.

[0012] The covariate modeling mechanism is imperfect: Although the traditional latent regression model has been widely used in item response theory, there is still a lack of effective methods for incorporating covariates into the estimation and modeling of latent attributes within the framework of cognitive diagnostic models. The current "one-step method" is rigid and difficult to extend flexibly. While the "three-step method" has a certain degree of flexibility, it is prone to introducing estimation bias in the process of latent category assignment and covariate regression modeling, which reduces the interpretability and accuracy of the model.

[0013] Estimation bias affects the determination of latent attributes: The traditional three-step method directly performs covariate regression analysis without ignoring classification error, which may cause a systematic bias in the mastery of latent attributes, affecting the scientificity and effectiveness of subsequent educational interventions or individual cognitive diagnosis.

[0014] Therefore, there is an urgent need for a new cognitive diagnostic analysis method to improve model flexibility, estimation accuracy, and group comparison capabilities, so as to meet the practical application needs in complex educational assessment scenarios. Summary of the Invention

[0015] To address the problems existing in the prior art, the present invention aims to provide a cognitive diagnostic bias correction method with covariates for multiple groups. Based on the MG-GDINA model, this method extends the existing method by combining a classification error correction mechanism with a covariate modeling strategy. This enables effective modeling and comparison of multi-group data, systematic correction of classification errors, and accurate characterization of the relationship between covariates and latent cognitive attributes. This comprehensively improves the accuracy and stability of latent attribute estimation, and solves the core technical bottlenecks of existing CDM models in multi-group data analysis and covariate processing.

[0016] To achieve the above objectives, the present invention provides the following solution:

[0017] A method for correcting cognitive diagnostic biases with covariates for multiple groups includes:

[0018] The candidate's answer data and the relationship between the question attributes are obtained. The answer data and the relationship between the question attributes are input into the attribute mastery probability interpretation model to obtain the candidate's attribute mastery probability on each cognitive attribute. The attribute mastery probability interpretation model is obtained by training multiple cognitive diagnostic models using a training set.

[0019] During model training, the potential knowledge states of examinees are assigned based on the posterior distributions output by the multiple sets of cognitive diagnostic models. At the same time, the classification error probability matrix is ​​calculated, and multiple sets of latent logistic regression models are used to evaluate the influence of external covariate data on the cognitive attribute mastery state. The model parameters of the multiple sets of cognitive diagnostic models are optimized using the objective function with the classification error probability matrix as the target weight.

[0020] Optionally, the multiple sets of cognitive diagnostic models include:

[0021]

[0022] in, Let represent the vector of a candidate's mastery of the attribute required for the j-th question in the g-th group. Let be the number of attributes required for problem j. This indicates that attribute k is a necessary attribute to answer question j. This represents the parameter values ​​of group g for question j. This represents the probability that a test taker answers a question correctly without knowing any of the attributes required for question j. This indicates that the main effect of attribute k on answering question j has been understood. This indicates that the attributes have been mastered. and The resulting second-order interaction effect This represents the comprehensive interactive effect of answering a question correctly after mastering all the attributes required for that question. This represents the level of understanding of the k-th attribute by the l-th candidate in the g-th group. This represents the level of understanding of the k'th attribute by the l-th candidate in the g-th group. This indicates the number of attributes that need to be mastered for question j.

[0023] Optionally, calculating the classification error probability matrix includes:

[0024]

[0025] in, The classification error probability matrix is... Indicates possible attribute category values, To achieve a state of true attribute mastery. To be assigned to the attribute mastery state. In order to observe the answer data of candidate i Then, the posterior probability of the attribute mastery state. For indicator functions, i.e., if Then take 1, otherwise take 0. Let g be the answer data of candidate i in the g-th group. Let i be the estimated attribute mastery level of candidate i. This reflects the candidate's true level of mastery of attributes. This represents the total number of candidates in group g.

[0026] Optionally, based on the classification error probability matrix, the proportion of candidates assigned to the attribute mastery state under the condition of true attribute mastery state includes:

[0027] The percentage of candidates in group g who correctly categorized the attributes they did not master:

[0028]

[0029] The percentage of candidates in group g who did not master the attribute under the misclassification:

[0030]

[0031] in, The classification error probability matrix is... Indicates possible attribute category values, To achieve a state of true attribute mastery. In order to observe the answer data of candidate i Then, the posterior probability of the attribute mastery state. For indicator functions, i.e., if Then take 1, otherwise take 0. Let g be the answer data of candidate i in the g-th group. Let i be the estimated attribute mastery level of candidate i. This reflects the candidate's true level of mastery of attributes. This represents the total number of candidates in group g.

[0032] Optionally, using the classification error probability matrix as the target weights includes:

[0033] The sample hierarchy correction weights for each candidate are obtained by calculating the column corresponding to the estimated attribute classification of the candidate in the classification error probability matrix:

[0034]

[0035] in, The classification error probability matrix is... For indicator functions, i.e., if If the result is 1, then take 1; otherwise, take 0.

[0036] Optionally, using the classification error probability matrix as the target weights includes:

[0037] Using the classification error probability matrix, the correction weights for the posterior distribution hierarchy are obtained:

[0038]

[0039] in, To determine the proportion of attribute k at the sample level, In order to observe the answer data of candidate i Then, the posterior probability of the attribute mastery state.

[0040] Optionally, assessing the strength of the influence of the external covariate data on the mastery status of cognitive attributes includes:

[0041]

[0042] in, To affect the intensity, Let g be the covariate matrix of the g-th group. Let be the regression coefficient of the g-th group. is the intercept term for the g-th group.

[0043] Optionally, optimizing the model parameters of the multiple cognitive diagnostic models using an objective function includes:

[0044] The model parameters of the multiple cognitive diagnostic models are optimized by maximizing the weighted log-likelihood function.

[0045]

[0046] in, For the sample-level correction weights and the posterior-level correction weights, Given candidate i as a covariate, the probability that candidate i possesses attribute k. Let i be the covariate of candidate i in the g-th group. Let be the likelihood function.

[0047] The beneficial effects of this invention are as follows:

[0048] This invention employs a three-step bias correction method, introducing a classification error probability matrix (CEP) and a multi-group logistic regression model, which reduces estimation bias and improves the model's accuracy and robustness. Compared to a single-group model, by introducing the MG-GDINA model and multi-group extensions of logistic regression, it can simultaneously process data from multiple groups, providing technical support for comparative studies of multiple datasets.

[0049] This invention significantly reduces computational complexity through a phased modeling strategy, minimizing the need to re-estimate the overall model and thus saving time on model tuning and re-estimation. By incorporating covariates, the model can more effectively reveal the impact of covariates on latent attributes, helping to understand the differences in attribute mastery among different covariates (such as gender, age, and family background).

[0050] This invention can accurately measure cognitive differences among different student groups, helping teachers to provide individualized instruction and improve the quality of education, thereby promoting educational equity and personalized learning. At the same time, in clinical diagnosis, this method can accurately assess the cognitive status of different patient groups, providing effective support for fields such as neuropsychological assessment and cognitive impairment screening, and promoting the progress of precision medicine. Attached Figure Description

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 This invention provides a method for correcting cognitive diagnostic biases with covariates for multiple groups. Detailed Implementation

[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0054] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0055] like Figure 1 As shown in the figure, this embodiment discloses a cognitive diagnostic bias correction method with covariates for multiple groups, including: obtaining the relationship between the examinee's answer data and the question attributes; inputting the answer data and the relationship between the question attributes into an attribute mastery probability interpretation model to obtain the examinee's attribute mastery probability on each cognitive attribute; the attribute mastery probability interpretation model is obtained by training multiple cognitive diagnostic models using a training set; during the model training process, the examinee's potential knowledge state is assigned based on the posterior distribution output by the multiple cognitive diagnostic models, and the classification error probability matrix is ​​calculated. Multiple latent logistic regression models are used to evaluate the influence strength of external covariate data on the cognitive attribute mastery state, and the model parameters of the multiple cognitive diagnostic models are optimized using an objective function with the classification error probability matrix as the target weight.

[0056] Specifically, this embodiment discloses a method for correcting cognitive diagnostic biases with covariates for multiple groups, including:

[0057] Step 1: Fitting a Multi-Group Cognitive Diagnostic Model (MG-GDINA): In the initial stage, based on the response data of different candidate groups, a multi-group cognitive diagnostic model (MG-GDINA) is constructed to estimate the mastery of cognitive attributes in each group of candidates. The model input is a standard format response matrix, where rows represent candidates, columns represent tasks / questions, and elements are 0 or 1, representing incorrect and correct answers, respectively. Simultaneously, candidates are divided into multiple pre-defined groups (e.g., group A, group B, etc.) based on their basic information (e.g., gender, age, educational background). During implementation, a Q-matrix is ​​first constructed to clarify the mapping relationship between tasks and cognitive attributes. The Q-matrix elements are binary, with 1 indicating that the task requires testing the corresponding attribute, and 0 indicating no. This Q-matrix forms the basic structure for subsequent model fitting. After data processing, the MG-GDINA model fitting process is executed. This model, as an extension of the G-DINA model, has the ability to handle multi-group data and can generate independent attribute mastery estimation results for different groups. By maximizing the likelihood function, the model outputs the attribute mastery probability distribution for each group, providing a foundation for subsequent steps.

[0058] Step 2: Calculate the latent attribute state category and classification error probability matrix (CEP): After fitting the MG-GDINA model, the latent knowledge states of the examinees are assigned based on the posterior distribution of the model output. Specifically, the expected posterior (EAP) or maximum a posteriori (MAP) strategy is used, setting a mastery threshold (e.g., 0.5) for each cognitive attribute to determine whether an individual has reached a mastery state. Simultaneously, the system calculates the classification error probability matrix (CEP). CEP reflects the probability of an examinee being misclassified in each latent cognitive state, and its calculation is based on the consistency between the examinee's actual response and the model's predicted response. For example, if a cognitive attribute is judged as "mastered," but the examinee's actual answer pattern does not conform to the response characteristics of mastery for that attribute, the corresponding CEP value will be increased. The CEP results will be used for further optimization of model accuracy, serving as a correction factor input to subsequent parameter adjustment steps, thereby improving the stability and reliability of state assignment.

[0059] Step 3: Introducing latent attribute correction and optimization modeling with covariates: In this stage, the CEP values ​​obtained in Step 2 are combined with external covariate data (such as gender, family background, socioeconomic status, etc.) and incorporated into the latent logistic regression modeling process to systematically optimize the estimation of cognitive attribute states.

[0060] Specifically, multiple latent logistic regression models are employed, with CEP as the weight, to estimate the impact of each covariate on cognitive attribute mastery. The weighting is designed to mitigate the interference of individuals with high classification errors on the model parameters, thereby improving the robustness of the overall estimation. This process obtains the optimal parameter solution by maximizing the weighted log-likelihood function.

[0061] Introducing covariates helps reveal the systematic impact of different background factors on knowledge mastery levels. For example, the model can output the parameters of each covariate to help analysts determine whether candidates' cognitive mastery of various attributes is affected by the covariates.

[0062] Ultimately, the model generates estimates of cognitive mastery across multiple groups, including covariate control and error correction, which can serve as a decision-making basis for subsequent personalized teaching recommendations, resource allocation optimization, and educational intervention design.

[0063] Furthermore, without considering covariates, a multi-group cognitive diagnostic model (MG-GDINA) was fitted to the response data based on group parameters:

[0064] A multi-set cognitive diagnostic model (MG-GDINA) was constructed to estimate the mastery status of candidates from different backgrounds across the cognitive attribute dimensions. The candidates' response data and the relationship between question attributes (Q matrix) provided the basic input for the model. In the response matrix, rows represent candidates and columns represent questions; the Q matrix defines the correspondence between questions and cognitive attributes.

[0065] In educational applications, test takers can be divided into different groups based on their class, school type, region, or educational intervention conditions. By jointly modeling the response data of each group, the MG-GDINA model can reveal the differences in mastery of various knowledge points (such as mathematical calculation, scientific reasoning, and language comprehension) among test takers from different educational backgrounds. This analysis helps education policymakers identify the weaknesses of specific groups in specific cognitive abilities, providing a basis for subsequent teaching.

[0066] For example, before fitting the MG-GDINA model, two key types of data need to be prepared: candidate data and the Q matrix. Candidate data is typically represented as a two-dimensional matrix, where rows represent candidates and columns represent their responses to each item (0 for incorrect, 1 for correct). This matrix has dimensions N×J, where N is the total number of candidates and J is the total number of items. Additionally, the group factor is a vector of length N indicating the group each candidate belongs to (e.g., numerical codes 1 and 2, representing group 1 and group 2). The Q matrix describes the relationship between each item and its latent attributes, with rows representing items and columns representing attributes. If an item requires an attribute, a 1 is entered in the corresponding position; otherwise, a 0 is entered. The Q matrix has dimensions J×K, where J is the number of items and K is the number of attributes. Preparing these two types of data provides the necessary input for fitting the MG-GDINA model.

[0067] The response function formula for the MG-GDINA model is as follows:

[0068]

[0069] set up Let represent the mastery vector of the attribute required for the j-th question by the examinees in the g-th group, where This indicates that attribute k is mastered. This indicates that the knowledge is not yet acquired. The number of attributes required to acquire knowledge for question j is denoted as... ,in This indicates that attribute k is a necessary attribute to answer question j. (Parameter vector) This represents the parameter of group g for question j, with the following specific meanings: This represents the probability that a candidate answers a question correctly without knowing any of the attributes required for question j. This indicates that the attributes have been mastered. The main effect of answering question j correctly; This indicates that the attributes have been mastered. and The resulting second-order interaction effect; This represents the comprehensive interactive effect of answering a question correctly after mastering all the attributes required for that question.

[0070] This model comprehensively covers single attributes, the interactions between multiple attributes, and their differences between groups. It is a natural extension of the G-DINA model in multi-group scenarios.

[0071] Furthermore, based on the classification error probability matrix, the proportion of candidates assigned to the attribute mastery state under the condition of true attribute mastery state includes: the proportion of candidates who did not master the attribute under correct classification in group g and the proportion of candidates who did not master the attribute under incorrect classification in group g.

[0072] Furthermore, using the classification error probability matrix as the target weight includes: calculating the sample level correction weight of the examinee by using the column in the classification error probability matrix that corresponds to the estimated attribute classification of the examinee; and using the classification error probability matrix to obtain the correction weight of the posterior distribution level.

[0073] Specifically, based on the results of the first step, the degree of attribute mastery or mastery pattern of individuals in each group is classified, and the classification error probability is calculated:

[0074] After model fitting, the posterior probability is used to classify each student's knowledge mastery status (whether they have mastered a certain cognitive attribute). This process employs the expected posterior (EAP) or maximum a posteriori (MAP) principle, and sets a threshold (e.g., 0.5) to determine whether a student has mastered an attribute. To improve the reliability of state assignment, the classification error probability (CEP) for each student is further calculated to assess the likelihood that a student is misclassified as having mastered or not mastered the attribute.

[0075] In educational settings, this metric can be used to quantify the reliability of a model's assessment of a student's true cognitive level. Subsequently, the CEP value is converted into a correction weight to correct errors in subsequent model parameter estimations, thereby improving the robustness of classification decisions.

[0076] The posterior distribution of the examinees is used to assign them to latent categories. In latent class analysis (LCA), examinee assignment can be proportional, pattern-based, or mean-based. In this invention, the vector of proportional assignment is equal to the estimated posterior distribution of the examinees, used to reflect the fuzziness and uncertainty of the examinees' mastery status, thereby improving the accuracy of subsequent model estimation or inference. Pattern and mean assignments correspond to the maximum a posteriori (MAP) and expected a posteriori (EAP) methods, respectively. The latter is used to calculate the marginalized attribute level probabilities.

[0077] The second model is as follows:

[0078] The attribute-level classification error probability matrix is ​​used to calculate the classification error probability of the latent class. This matrix generates a 2×2 contingency table, represented as follows. ,in This represents the possible attribute category value, which is either 0 or 1. The formula is as follows:

[0079]

[0080] This formula can be interpreted as representing a state where the true attributes are known. Under these conditions, they are assigned to the attribute mastery state. The percentage of test takers. For example:

[0081]

[0082]

[0083] The former can be interpreted as the proportion of candidates in group g who are correctly classified as not having mastered attribute k, while the latter is the proportion of candidates in group g who are incorrectly classified as not having mastered attribute k.

[0084] The sample hierarchical correction weight for candidate i is calculated using the column corresponding to the candidate's estimated attribute classification in the attribute level classification error probability matrix, as shown in the following formula:

[0085]

[0086] The formula for calculating the correction weights of the posterior distribution hierarchy is:

[0087]

[0088] in, This indicates the proportion of attribute k that is mastered at the sample level.

[0089] The calculation result SL is a 2×2 matrix with K members in each of the g groups, and PDL is a 2×2 matrix with K members in each of the g groups, which is used as the correction weight in the third step.

[0090] Furthermore, multiple sets of logistic regression functions were used to incorporate covariates as explanatory variables:

[0091]

[0092] in, Given the covariate matrix, the corrected objective function is:

[0093]

[0094] in, Represents a parameter vector. equal and Optimizing the objective function yields the parameters in... The corrected estimate in.

[0095] The final output is a C×K matrix of g groups, where C represents the number of covariates and K represents the number of measured attributes.

[0096] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for correcting cognitive diagnostic biases with covariates for multiple groups, characterized in that, include: Obtain the candidate's answer data and the relationship between the question attributes, input the answer data and the relationship between the question attributes into the attribute mastery probability interpretation model, and obtain the candidate's attribute mastery probability on each cognitive attribute; The attribute mastery probability interpretation model is obtained by training multiple sets of cognitive diagnostic models using a training set; During model training, the potential knowledge state of the examinee is allocated based on the posterior distribution output by the multiple sets of cognitive diagnostic models. At the same time, the classification error probability matrix is ​​calculated, and multiple sets of latent logistic regression models are used to evaluate the influence of external covariate data on the cognitive attribute mastery state. The model parameters of the multiple sets of cognitive diagnostic models are optimized using the objective function with the classification error probability matrix as the target weight. Calculating the classification error probability matrix includes: in, The classification error probability matrix is... Indicates possible attribute category values, To achieve a state of true attribute mastery. To be assigned to the attribute mastery state. In order to observe the answer data of candidate i Then, the posterior probability of the attribute mastery state. For indicator functions, i.e., if Then take 1, otherwise take 0. Let g be the answer data of candidate i in the g-th group. Let i be the estimated attribute mastery level of candidate i. This reflects the candidate's true level of mastery of attributes. This represents the total number of candidates in group g. Optimizing the model parameters of the multiple cognitive diagnostic models using an objective function includes: The model parameters of the multiple cognitive diagnostic models are optimized by maximizing the weighted log-likelihood function. in, For the sample-level correction weights and the posterior-level correction weights, Given candidate i as a covariate, the probability that candidate i possesses attribute k. Let i be the covariate of candidate i in the g-th group. Let be the likelihood function.

2. The cognitive diagnostic bias correction method with covariates for multiple groups according to claim 1, characterized in that, The multiple cognitive diagnostic models include: in, Let represent the vector of a candidate's mastery of the attribute required for the j-th question in the g-th group. Let be the number of attributes required for problem j. This indicates that attribute k is a necessary attribute to answer question j. This represents the parameter values ​​of group g for question j. This represents the probability that a test taker answers a question correctly without knowing any of the attributes required for question j. This indicates that the main effect of attribute k on answering question j has been understood. This indicates that the attributes have been mastered. and The resulting second-order interaction effect This represents the comprehensive interactive effect of answering a question correctly after mastering all the attributes required for that question. This represents the level of understanding of the k-th attribute by the l-th candidate in the g-th group. This represents the level of understanding of the k'th attribute by the l-th candidate in the g-th group. This indicates the number of attributes that need to be mastered for question j.

3. The cognitive diagnostic bias correction method with covariates for multiple groups according to claim 1, characterized in that, Based on the classification error probability matrix, the proportion of candidates assigned to the attribute mastery state under the condition of true attribute mastery state includes: The percentage of candidates in group g who correctly categorized the attributes they did not master: The percentage of candidates in group g who did not master the attribute under the misclassification: in, The classification error probability matrix is... Indicates possible attribute category values, To achieve a state of true attribute mastery. In order to observe the answer data of candidate i Then, the posterior probability of the attribute mastery state. For indicator functions, i.e., if Then take 1, otherwise take 0. Let g be the answer data of candidate i in the g-th group. Let i be the estimated attribute mastery level of candidate i. This reflects the candidate's true level of mastery of attributes. This represents the total number of candidates in group g.

4. The cognitive diagnostic bias correction method with covariates for multiple groups according to claim 3, characterized in that, The target weights, using the classification error probability matrix as the basis, include: The sample hierarchy correction weights for each candidate are obtained by calculating the column corresponding to the estimated attribute classification of the candidate in the classification error probability matrix: in, The classification error probability matrix is... For indicator functions, i.e., if If the result is 1, then take 1; otherwise, take 0.

5. The cognitive diagnostic bias correction method with covariates for multiple groups according to claim 3, characterized in that, The target weights, using the classification error probability matrix as the basis, include: Using the classification error probability matrix, the correction weights for the posterior distribution hierarchy are obtained: in, To determine the proportion of attribute k at the sample level, In order to observe the answer data of candidate i Then, the posterior probability of the attribute mastery state.

6. The cognitive diagnostic bias correction method with covariates for multiple groups according to claim 1, characterized in that, The assessment of the strength of the influence of the external covariate data on the mastery status of cognitive attributes includes: in, To affect the intensity, Let g be the covariate matrix of the g-th group. Let be the regression coefficient of the g-th group. is the intercept term for the g-th group.

Citation Information

Patent Citations

  • Cross-time-point cognitive diagnosis method considering student influence factors

    CN119830237A

  • Multi-group time sequence student cognitive diagnosis method based on joint probability model

    CN120235739A