Multi-group-oriented cognitive diagnosis deviation correction method with covariants

By introducing the MG-GDINA model and multi-group logistic regression model, combined with classification error correction, the flexibility and accuracy of the multi-group cognitive diagnostic model are solved, effective processing of multi-group data and covariate impact assessment are realized, and the accuracy and applicability of education and clinical diagnosis are improved.

CN120495036AActive Publication Date: 2025-08-15JINAN UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510991117.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-08-15
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

The existing cognitive diagnostic models lack flexibility and accuracy in multi-group scenarios, making it difficult to effectively deal with covariates, resulting in insufficient applicability and explanatory power of the model in educational assessment and clinical diagnosis.

Method used

The MG-GDINA model is used to combine the classification error correction mechanism and multiple sets of logistic regression models. By obtaining the candidate's answer data and the question attribute relationship, the classification error probability matrix is calculated, and the model parameters are optimized with it as the target weight to evaluate the impact of covariates on cognitive attributes.

Benefits of technology

It improves the flexibility and accuracy of the model, can effectively process multi-group data, reduces computational complexity, accurately measure cognitive differences among different student groups, and supports personalized learning and clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495036A_ABST
    Figure CN120495036A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of education evaluation, in particular to a multi-group-oriented cognitive diagnosis deviation correction method with covariants, which comprises the following steps: acquiring answer data and question attribute relations of examinees, inputting the answer data and the question attribute relations into an attribute mastering probability interpretation model, and obtaining a question mastering probability interpretation model; obtaining the attribute mastering probability of the examinee on each cognitive attribute; the attribute mastering probability interpretation model is obtained by training multiple groups of cognitive diagnosis models by using a training set; in a model training process, distribution of potential knowledge states is carried out on examinees based on posterior distribution output by multiple groups of cognitive diagnosis models, a classification error probability matrix is calculated, multiple groups of potential logic regression models are adopted, and the influence intensity of external covariable data on cognitive attribute mastering states is evaluated. And the classification error probability matrix is used as a target weight, and model parameters of multiple groups of cognitive diagnosis models are optimized by using a target function. According to the method, a plurality of key problems in the aspects of multi-group data processing and covariant modeling in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of educational assessment, psychological measurement and clinical diagnosis, and in particular to a cognitive diagnosis bias correction method with covariates for multiple groups. Background Art

[0002] With the continuous development of educational measurement technology, cognitive diagnostic models (CDMs) are increasingly being used in teaching assessment and personalized learning. They can provide students with more targeted and informative feedback, thereby supporting teaching interventions and learning planning. These models are not only used in various aspects of education, such as the assessment of proportional reasoning, spatial skills, and digital literacy, but are also gradually demonstrating their potential in other fields such as clinical diagnosis, becoming a valuable tool for analyzing individual cognitive status. Compared with traditional measurement methods, they offer higher diagnostic accuracy and interpretability, making them a key tool in educational assessment research.

[0003] As a general cognitive diagnostic model, the G-DINA model generalizes and expands upon the traditional DINA model. By relaxing restrictions on attribute interactions, the G-DINA model can, under specific parameter settings, be equivalent to other cognitive diagnostic models based on alternative link functions. Therefore, the G-DINA model possesses excellent adaptability and broad applicability, making it an ideal tool for analyzing examinees' mastery of multiple cognitive attributes.

[0004] On this basis, the MG-GDINA model was developed to address the need for analyzing cognitive differences between different groups of examinees. As a direct extension of the G-DINA model, this model can simultaneously consider the similarities and differences in cognitive attributes across multiple groups (e.g., different schools, regions, etc.), thus providing a basis for educational decision-making and resource allocation. The introduction of the MG-GDINA model provides greater adaptability and application value for cognitive diagnostic models in multi-group cognitive analysis scenarios.

[0005] In the field of educational measurement, using latent regression models to study the relationship between test-taker scale scores and covariates (such as gender, age, and family background) has long been a research focus. However, incorporating test-taker covariates into cognitive diagnostic models to analyze the relationship between these covariates and diagnostic outcomes (such as attribute mastery or mastery patterns) has not been fully studied and applied. While existing technical solutions (such as the one-step method) can accurately assess the impact of test-taker covariates on diagnostic outcomes, these solutions lack flexibility in terms of model structure. In practice, every modification to any part of the model requires re-estimation of the entire process, which is cumbersome and inefficient, limiting its practicality and scalability.

[0006] In addition, to improve the flexibility of the model, some schemes adopted the "step-by-step method" or "three-step method", in which the steps of the three-step method mainly include: CDM estimation: first fitting the cognitive diagnostic model to the test data of the examinees; latent category allocation: allocating the latent categories of the examinees according to the fitting results; latent logistic regression: finally, estimating the influence of the examinees' covariates on their latent attribute mastery status through the logistic regression model.

[0007] Although the three-step method enhances the flexibility of the model to a certain extent, the classification error is not fully considered, which may still lead to biased estimates and inaccurate assessment of the model's mastery of examinees' attributes.

[0008] More importantly, existing cognitive diagnostic models usually do not consider modeling issues in multi-group scenarios. Especially in actual educational assessment applications, it is often necessary to compare candidates from different genders and backgrounds. They lack the ability to process multi-group data. Therefore, there are certain limitations when comparing the differences and similarities between multiple groups, which affects the accuracy of the analysis and the breadth of application.

[0009] In summary, the existing technical methods still have the following significant defects:

[0010] The model structure lacks flexibility, is complex to operate, has biased estimates of covariates for included candidates, and lacks multi-system support for modeling and comparing multi-group data:

[0011] Inadequate multi-group data processing capabilities: Existing cognitive diagnostic models generally fail to fully account for heterogeneity across different groups (e.g., school type, region, and educational intervention conditions), making it difficult to effectively differentiate and compare the mastery of students' latent cognitive attributes across multiple groups. This limitation significantly impacts the models' applicability and explanatory power in diverse scenarios, including educational assessment, psychological evaluation, and clinical diagnosis.

[0012] Imperfect covariate modeling mechanism: Although traditional latent regression models have been widely used in item response theory, there is still a lack of effective methods to incorporate covariates into the estimation and modeling of latent attributes within the framework of cognitive diagnostic models. The current "one-step method" has a rigid structure and is difficult to expand flexibly. Although the "three-step method" has certain flexibility, it is easy to introduce estimation bias in the process of latent category allocation and covariate regression modeling, reducing the interpretability and accuracy of the model.

[0013] Estimation bias affects the determination of latent attributes: The traditional three-step method directly conducts covariate regression analysis while ignoring classification errors, which may cause systematic deviations in the mastery of latent attributes and affect the scientificity and effectiveness of subsequent educational interventions or individual cognitive diagnosis.

[0014] Therefore, a new cognitive diagnostic analysis method is urgently needed to improve model flexibility, estimation accuracy and group comparison capabilities to meet practical application needs in complex educational assessment scenarios. Summary of the Invention

[0015] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a cognitive diagnostic bias correction method with covariates for multiple groups. It is expanded on the basis of the MG-GDINA model, combining the correction mechanism of classification errors with the covariate modeling strategy to achieve effective modeling and comparison of multi-group data, systematic correction of classification errors, and accurate characterization of the relationship between covariates and potential cognitive attributes, thereby comprehensively improving the accuracy and stability of potential attribute estimation and solving the core technical bottlenecks of existing CDM models in multi-group data analysis and covariate processing.

[0016] To achieve the above object, the present invention provides the following solutions:

[0017] A cognitive diagnostic bias correction method with covariates for multiple groups, including:

[0018] Obtaining the relationship between the examinee's answer data and question attributes, inputting the answer data and the question attribute relationship into an attribute mastery probability interpretation model to obtain the examinee's attribute mastery probability for each cognitive attribute; the attribute mastery probability interpretation model is obtained by training multiple groups of cognitive diagnosis models using a training set;

[0019] During the model training process, the examinees are assigned potential knowledge states based on the posterior distributions output by the multiple groups of cognitive diagnostic models, and the classification error probability matrix is calculated at the same time. A multiple group of latent logistic regression models is used to evaluate the influence of external covariate data on the cognitive attribute mastery state. The classification error probability matrix is used as the target weight, and the model parameters of the multiple groups of cognitive diagnostic models are optimized using the objective function.

[0020] Optionally, the multiple groups of cognitive diagnostic models include:

[0021]

[0022] in, represents the mastery vector of the attributes required for the jth question by the examinees in the gth group, is the number of attributes required to master question j, Indicates that attribute k is a necessary attribute to answer question j, represents the item parameters of group g for question j, represents the probability that the examinee will answer the question without mastering any of the attributes required for question j, represents the main effect of mastering attribute k on answering question j. Indicates mastery of attributes and The second-order interaction effect caused by It represents the comprehensive interactive effect of answering the question after mastering all the attributes required for the question. Indicates the mastery of the k-th attribute by the l-th examinee in the g-th group, Indicates the mastery of the k'th attribute by the lth examinee in the gth group, Indicates the number of attributes required to master question j.

[0023] Optionally, calculating the classification error probability matrix includes:

[0024]

[0025] in, is the classification error probability matrix, represents the possible attribute classification values, To grasp the state in the true attributes, To be assigned to the attribute master state, is the observed answer data of candidate i After that, the posterior probability of the attribute mastering the state, is the indicator function, that is, if If yes, it takes 1, otherwise it takes 0. is the answer data of candidate i in group g, is the estimated attribute mastery status of candidate i, For the candidates i real attribute master status, is the total number of candidates in group g.

[0026] Optionally, obtaining the proportion of examinees assigned to the attribute mastery state under the condition of the true attribute mastery state according to the classification error probability matrix includes:

[0027] The proportion of candidates who correctly classified the attributes they did not master in group g:

[0028]

[0029] The proportion of candidates who did not master the attribute in the incorrect classification in group g:

[0030]

[0031] in, is the classification error probability matrix, represents the possible attribute classification values, To grasp the state in the true attributes, is the observed answer data of candidate i After that, the posterior probability of the attribute mastering the state, is the indicator function, that is, if If yes, it takes 1, otherwise it takes 0. is the answer data of candidate i in group g, is the estimated attribute mastery status of candidate i, For the candidates i real attribute master status, is the total number of candidates in group g.

[0032] Optionally, taking the classification error probability matrix as the target weight includes:

[0033] The columns corresponding to the estimated attribute classifications of the examinees in the classification error probability matrix are used to perform calculations to obtain the sample-level correction weights of the examinees:

[0034]

[0035] in, is the classification error probability matrix, is the indicator function, that is, if If yes, it takes 1; otherwise, it takes 0.

[0036] Optionally, taking the classification error probability matrix as the target weight includes:

[0037] Using the classification error probability matrix, the correction weights of the posterior distribution level are obtained:

[0038]

[0039] in, In order to grasp the proportion of attribute k at the sample level, is the observed answer data of candidate i After that, the attribute grasps the posterior probability of the state.

[0040] Optionally, evaluating the influence strength of the external covariate data on the cognitive attribute mastery state includes:

[0041]

[0042] in, To influence the intensity, is the covariate matrix of the g-th group, is the regression coefficient of the g-th group, is the intercept term of the gth population.

[0043] Optionally, optimizing the model parameters of the multiple groups of cognitive diagnosis models using an objective function includes:

[0044] The model parameters of the multiple groups of cognitive diagnosis models are optimized by maximizing the weighted log-likelihood function:

[0045]

[0046] in, are the sample-level correction weights and the posterior distribution-level correction weights, is the probability that a candidate has mastered attribute k given the covariate of candidate i, is the covariate of candidate i in group g, is the likelihood function.

[0047] The beneficial effects of the present invention are:

[0048] This paper uses a three-step bias correction approach, along with the classification error probability matrix (CEP) and a multi-group logistic regression model, to reduce estimation bias and improve model accuracy and robustness. Compared to single-group models, the introduction of the MG-GDINA model and the multi-group extension of multi-group logistic regression enable simultaneous processing of data from multiple groups, providing technical support for comparative studies of multi-group data.

[0049] This invention significantly reduces computational complexity through a phased modeling strategy, reducing the need to re-estimate the entire model, thereby saving time on model adjustments and re-estimation. By incorporating covariates, the model can more effectively reveal the impact of covariates on underlying attributes, helping to understand the differences in attribute capture based on different covariates (such as gender, age, and family background).

[0050] This invention can accurately measure the cognitive differences among different groups of students, helping teachers to teach students in accordance with their aptitude and improve the quality of education, thereby promoting the development of educational equity and personalized learning. At the same time, in clinical diagnosis, this method can accurately assess the cognitive status of different patient groups, providing effective support for fields such as neuropsychological assessment and cognitive impairment screening, and promoting the advancement of precision medicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0052] Figure 1 This is a cognitive diagnostic bias correction method with covariates for multiple groups according to an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0054] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0055] like Figure 1 As shown, this embodiment discloses a cognitive diagnosis bias correction method with covariates for multiple groups, including: obtaining the relationship between the examinee's answer data and question attributes, inputting the answer data and question attribute relationship into an attribute mastery probability explanation model, and obtaining the examinee's attribute mastery probability on each cognitive attribute; the attribute mastery probability explanation model is obtained by training multiple groups of cognitive diagnosis models using a training set; during the model training process, the examinee's potential knowledge state is allocated based on the posterior distribution output by the multiple groups of cognitive diagnosis models, and the classification error probability matrix is calculated at the same time. A multiple group latent logistic regression model is used to evaluate the influence of external covariate data on the cognitive attribute mastery state, and the classification error probability matrix is used as the target weight, and the model parameters of the multiple groups of cognitive diagnosis models are optimized using the objective function.

[0056] Specifically, this embodiment discloses a method for correcting cognitive diagnostic bias with covariates for multiple groups, including:

[0057] Step 1: Fitting the Multi-Group Cognitive Diagnostic Model (MG-GDINA): Initially, a multi-group cognitive diagnostic model (MG-GDINA) is constructed based on the response data from different groups of examinees. This model estimates each group's mastery of cognitive attribute dimensions. The model input is a standard response matrix, with rows representing examinees and columns representing tasks / items, with elements 0 or 1 representing incorrect and correct responses, respectively. Furthermore, examinees are divided into predefined groups (e.g., Group A, Group B, etc.) based on their basic information (e.g., gender, age, and educational background). During implementation, a Q matrix must first be constructed to clearly define the mapping between tasks and cognitive attributes. Elements in the Q matrix are binary, with 1 indicating that the task tested the corresponding attribute and 0 indicating that it did not. This Q matrix serves as the foundation for subsequent model fitting. After data organization is complete, the MG-GDINA model fitting process is executed. As an extension of the G-DINA model, this model can handle multi-group data and generate independent attribute mastery estimates for different groups. By maximizing the likelihood function, the model outputs a probability distribution of attribute mastery for each group, providing a foundation for subsequent steps.

[0058] Step 2: Calculate Latent Attribute State Categories and the Classification Error Probability (CEP) Matrix: After fitting the MG-GDINA model, candidates are assigned to latent knowledge states based on the posterior distribution of the model output. This method employs either the expected a posteriori (EAP) or maximum a posteriori (MAP) strategy, setting a mastery threshold (e.g., 0.5) for each cognitive attribute to determine whether an individual has reached mastery. While assigning states, the system simultaneously calculates the classification error probability (CEP) matrix. The CEP reflects the probability of a candidate being misclassified for each latent cognitive state and is calculated based on the consistency between the candidate's actual response and the model's predicted response. For example, if a cognitive attribute is judged as "mastered," the corresponding CEP value will increase if the candidate's actual response pattern does not conform to the response characteristics of mastery for that attribute. The CEP results are used to further optimize model accuracy and serve as correction factors in subsequent parameter adjustment steps, thereby improving the stability and reliability of state assignments.

[0059] Step 3: Latent attribute correction and optimization modeling with the introduction of covariates: In this stage, the CEP values obtained in the second step are incorporated into the latent logistic regression modeling process together with external covariate data (such as gender, family background, socioeconomic status, etc.) to systematically optimize the cognitive attribute state estimation.

[0060] Specifically, a multi-group latent logistic regression model was used, using CEP as weights, to estimate the impact of each covariate on cognitive attribute mastery. The weights were set to mitigate the influence of individuals with high classification error on the model parameters, thereby improving the robustness of the overall estimate. This process maximized the weighted log-likelihood function to obtain the optimal parameter solution.

[0061] The introduction of covariates helps reveal the systematic impact of different background factors on knowledge mastery. For example, the model can output the parameters of each covariate to help analysts determine whether the covariate influences the cognitive mastery of various attributes of the examinee.

[0062] Ultimately, the model forms an estimation result of cognitive mastery status for multiple groups, including covariate control and error correction, which can serve as the decision basis for subsequent personalized teaching recommendations, resource allocation optimization and educational intervention design.

[0063] Furthermore, without considering the covariates, the multi-group cognitive diagnostic model (MG-GDINA) was fitted to the response data according to the group parameters:

[0064] A multi-group cognitive diagnostic model (MG-GDINA) was constructed to estimate the mastery of cognitive attributes across test-taker backgrounds. The model's basic inputs are test-taker response data and the relationship between item attributes (the Q matrix). The Q matrix represents the rows of the response matrix and the columns of the question. The Q matrix defines the correspondence between items and cognitive attributes.

[0065] In educational applications, test-takers can be divided into different groups based on their class, school type, region, or educational intervention. By jointly modeling the response data from each group, the MG-GDINA model can reveal differences in the mastery of various knowledge points (such as mathematical calculation, scientific reasoning, and language comprehension) among test-takers from different educational backgrounds. This analysis helps educational decision-makers identify specific cognitive abilities that are weak in specific groups, providing a foundation for subsequent instruction.

[0066] For example, before fitting the MG-GDINA model, two key data types must be prepared: examinee data and a Q matrix. Examinee data are typically represented as a two-dimensional matrix, with rows representing examinees and columns representing their responses to each item (0 for an error, 1 for a correct answer). This matrix has dimensions N × J, where N is the total number of examinees and J is the total number of tasks. Furthermore, the group factor is a vector of length N, indicating the group to which each examinee belongs (for example, the numerical codes 1 and 2 represent Groups 1 and 2). The Q matrix describes the relationship between each item and the underlying attribute, with rows representing items and columns representing attributes. If an item requires a certain attribute, the corresponding position is filled with 1, and otherwise, with 0. The Q matrix has dimensions J × K, where J is the number of items and K is the number of attributes. By preparing these two types of data, we can provide the necessary input for fitting the MG-GDINA model.

[0067] The response function formula of the MG-GDINA model is as follows:

[0068]

[0069] set up represents the mastery vector of the attributes required for the jth question by the examinees in the gth group, where Indicates mastering attribute k, Indicates that the problem has not been mastered. The number of attributes required to master question j is recorded as ,in Indicates that attribute k is a necessary attribute to answer question j. Parameter vector Represents the item parameters of group g for question j. The specific meanings are as follows: represents the probability that the examinee will answer the question without mastering any of the attributes required for question j; Indicates mastery of attributes The main effect of answering the correct question j; Indicates mastery of attributes and The second-order interaction effects caused by It represents the comprehensive interactive effect on answering the question correctly after mastering all the attributes required for the question.

[0070] This model form completely covers the interaction between single attributes and multiple attributes and their differences between groups. It is a natural extension of the G-DINA model in multi-group scenarios.

[0071] Furthermore, according to the classification error probability matrix, the proportion of examinees assigned to the attribute mastery state under the condition of the true attribute mastery state is obtained, including: the proportion of examinees who did not master the attribute under the correct classification in the g-th group and the proportion of examinees who did not master the attribute under the incorrect classification in the g-th group.

[0072] Furthermore, using the classification error probability matrix as the target weight includes: using the columns in the classification error probability matrix corresponding to the estimated attribute classification of the examinee to perform calculations to obtain the examinee's sample level correction weight; using the classification error probability matrix to obtain the correction weight of the posterior distribution level.

[0073] Specifically, the attribute mastery degree or mastery pattern of each group of individuals is classified according to the results of the first step, and the classification error probability is calculated:

[0074] After model fitting, the posterior probability is used to classify each student's knowledge mastery status (whether or not a particular cognitive attribute is mastered). This process employs the expected a posteriori (EAP) or maximum a posteriori (MAP) principles, with a threshold (e.g., 0.5) set to determine whether a particular attribute is mastered. To improve the reliability of the status assignment, the classification error probability (CEP) is further calculated for each student to assess the likelihood that a student will be misclassified as having mastered or not mastered the attribute.

[0075] In educational settings, this metric can be used to quantify the confidence of the model in determining a student's true cognitive level. Subsequently, the CEP value is converted into a correction weight to correct errors in subsequent model parameter estimates, improving the robustness of classification decisions.

[0076] The posterior distribution of candidates is used to assign them to latent classes. In latent class analysis (LCA), candidate assignment can be proportional, mode, or mean. In this invention, the vector of proportional assignments is equal to the estimated candidate posterior distribution, which is used to reflect the ambiguity and uncertainty of the candidate's mastery state, thereby improving the accuracy of subsequent model estimation or inference. Mode and mean assignments correspond to the maximum a posteriori (MAP) and expected a posteriori (EAP) methods, respectively. The latter is used to calculate marginalized attribute-level probabilities.

[0077] The second step model is as follows:

[0078] The attribute level classification error probability matrix is used to calculate the classification error probability of the latent class. This matrix generates a 2×2 contingency table represented as ,in Represents the possible attribute classification value, which is 0 or 1. The formula is as follows:

[0079]

[0080] This formula can be interpreted as Under the conditions, it is assigned to the attribute mastery state For example:

[0081]

[0082]

[0083] The former can be interpreted as the proportion of candidates in group g who are correctly classified as not having mastered attribute k, while the latter is the proportion of candidates in group g who are incorrectly classified as not having mastered attribute k.

[0084] The sample-level correction weight for candidate i is calculated using the column in the attribute-level classification error probability matrix that corresponds to the candidate's estimated attribute classification, using the following formula:

[0085]

[0086] The correction weight calculation formula of the posterior distribution level is:

[0087]

[0088] in, Indicates the proportion of attribute k at the sample level.

[0089] The calculation result SL is a 2×2 matrix of K people in each of the g groups, and PDL is a 2×2 matrix of the number of people included in the current group of K people in the g groups, which is used as the correction weight in the third step.

[0090] Furthermore, a multi-group logistic regression function was used to include covariates as explanatory variables:

[0091]

[0092] in, is a known covariate matrix, and the corrected objective function is:

[0093]

[0094] in, represents the parameter vector, equal and Optimizing the objective function can obtain the parameters The corrected estimate in .

[0095] The final output result is g groups of C×K matrices, where C represents the number of covariates and K represents the number of measured attributes.

[0096] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A cognitive diagnostic bias correction method with covariates for multiple groups, characterized by: include: Obtaining the relationship between the examinee's answer data and the question attribute, inputting the answer data and the question attribute relationship into the attribute mastery probability interpretation model, and obtaining the examinee's attribute mastery probability for each cognitive attribute; The attribute mastery probability explanation model is obtained by training multiple groups of cognitive diagnosis models using a training set; During the model training process, the examinees are assigned potential knowledge states based on the posterior distributions output by the multiple groups of cognitive diagnostic models, and the classification error probability matrix is calculated at the same time. A multiple group of latent logistic regression models is used to evaluate the influence of external covariate data on the cognitive attribute mastery state. The classification error probability matrix is used as the target weight, and the model parameters of the multiple groups of cognitive diagnostic models are optimized using the objective function.

2. The method for correcting cognitive diagnostic bias with covariates for multiple groups according to claim 1, characterized in that: The multiple groups of cognitive diagnostic models include: in, represents the mastery vector of the attributes required for the jth question by the examinees in the gth group, is the number of attributes required to master question j, Indicates that attribute k is a necessary attribute to answer question j, represents the item parameters of group g for question j, represents the probability that the examinee will answer the question without mastering any of the attributes required for question j, represents the main effect of mastering attribute k on answering question j. Indicates mastery of attributes and The second-order interaction effect caused by It represents the comprehensive interactive effect of answering the question after mastering all the attributes required for the question. Indicates the mastery of the k-th attribute by the l-th examinee in the g-th group, Indicates the mastery of the k'th attribute by the lth examinee in the gth group, Indicates the number of attributes required to master question j.

3. The method for correcting cognitive diagnostic bias with covariates for multiple groups according to claim 1, characterized in that: Calculating the classification error probability matrix includes: in, is the classification error probability matrix, represents the possible attribute classification values, To grasp the state in the true attributes, To be assigned to the attribute master state, is the observed answer data of candidate i After that, the posterior probability of the attribute mastering the state, is the indicator function, that is, if If yes, it takes 1, otherwise it takes 0. is the answer data of candidate i in group g, is the estimated attribute mastery status of candidate i, For the candidates i real attribute master status, is the total number of candidates in group g.

4. The method for correcting cognitive diagnostic bias with covariates for multiple groups according to claim 1, characterized in that: According to the classification error probability matrix, obtaining the proportion of examinees assigned to the attribute mastery state under the condition of the true attribute mastery state includes: The proportion of candidates who correctly classified the attributes they did not master in group g: The proportion of candidates who did not master the attribute in the incorrect classification in group g: in, is the classification error probability matrix, represents the possible attribute classification values, To grasp the state in the true attributes, is the observed answer data of candidate i After that, the posterior probability of the attribute mastering the state, is the indicator function, that is, if If yes, it takes 1, otherwise it takes 0. is the answer data of candidate i in group g, is the estimated attribute mastery status of candidate i, For the candidates i real attribute master status, is the total number of candidates in group g.

5. The method for correcting cognitive diagnostic bias with covariates for multiple groups according to claim 4, characterized in that: Taking the classification error probability matrix as the target weight includes: The columns corresponding to the estimated attribute classifications of the examinees in the classification error probability matrix are used to perform calculations to obtain the sample-level correction weights of the examinees: in, is the classification error probability matrix, is the indicator function, that is, if If yes, it takes 1; otherwise, it takes 0.

6. The method for correcting cognitive diagnostic bias with covariates for multiple groups according to claim 4, characterized in that: Taking the classification error probability matrix as the target weight includes: Using the classification error probability matrix, the correction weights of the posterior distribution level are obtained: in, In order to grasp the proportion of attribute k at the sample level, is the observed answer data of candidate i After that, the attribute grasps the posterior probability of the state.

7. The method for correcting cognitive diagnostic bias with covariates for multiple groups according to claim 1, characterized in that: Evaluating the impact of the external covariate data on the cognitive attribute mastery status includes: in, To influence the intensity, is the covariate matrix of the g-th group, is the regression coefficient of the g-th group, is the intercept term of the gth population.

8. The method for correcting cognitive diagnostic bias with covariates for multiple groups according to claim 1, characterized in that: Optimizing the model parameters of the multiple groups of cognitive diagnosis models using the objective function includes: The model parameters of the multiple groups of cognitive diagnosis models are optimized by maximizing the weighted log-likelihood function: in, are the sample-level correction weights and the posterior distribution-level correction weights, is the probability that a candidate has mastered attribute k given the covariate of candidate i, is the covariate of candidate i in group g, is the likelihood function.

Citation Information

Patent Citations

  • English listening and reading ability grade evaluation system based on cognitive diagnosis technology

    CN118968830A

  • Cross-time-point cognitive diagnosis method considering student influence factors

    CN119830237A

  • Multi-group time sequence student cognitive diagnosis method based on joint probability model

    CN120235739A

  • Method for estimating examinee attribute parameters in cognitive diagnosis models

    US20060040247A1

  • Assessment and management system for rehabilitative conditions and related methods

    US20190096513A1