A method for evaluating the discrimination of test questions based on the GRM model

Through the GRM model-based test question division evaluation method, the K-Means clustering algorithm and ICC graph are used to solve the problem of time-consuming and laborious and poor usability of test question division evaluation in the existing technology, and realizes automated and accurate test question division evaluation, which is suitable for learning systems.

CN115829377BActive Publication Date: 2025-07-29GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211434999.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2025-07-29
Estimated Expiration
2042-11-16

AI Technical Summary

Technical Problem

The existing method of differentiation evaluation of test questions is time-consuming, labor-intensive, subjective and poor use. Especially for non-professional users, it is difficult to effectively evaluate the differentiation of test questions.

Method used

The test question distinction evaluation method based on the GRM model is used to obtain data from the learning system database, use the K-Means clustering algorithm and the GRM model to generate an ICC chart, and combine the user answer data to evaluate the test question distinction, including filtering data, determining K value, calculating difficulty coefficients and user types.

Benefits of technology

It improves the ease of use of test question distinction evaluation, can automatically classify categories according to the intrinsic relationship of the data, reduces professional knowledge dependence, provides intuitive ability and score relationship analysis, and improves the accuracy and efficiency of test question distinction evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115829377B_ABST
    Figure CN115829377B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for evaluating the item discrimination based on the GRM model, which includes the following steps: S1: Obtain the information of question ID, user ID, the number of assessment points, and the number of passed assessment points from the database of the learning system; S2: Screen out the users who have not participated in answering questions; S3: If the user participates in answering questions, the number of times the user answers questions is incremented by 1; if it is determined that the number of passed assessment points of the user is not equal to the number of assessment points, return to S2 until all questions are judged; S4: If the number of passed assessment points of the user is equal to the number of assessment points, the number of correctly answered questions is incremented by 1; S5: Use the screened data for the K-Means clustering algorithm; S6: Determine the value of K; S7: Determine the difficulty coefficient of the GRM model and generate an ICC graph based on GRM; at the same time, judge the user type; S8: Evaluate and obtain the item discrimination. The present invention can make the similarity between samples of the same category high and the similarity between samples of different categories low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method for evaluating the discrimination of test questions based on the GRM model. Background Art

[0002] In today's society, education is constantly developing. The "Action Plan for Education Informatization 2.0" points out that personalized learning goals of "teaching students in accordance with their aptitude" should be formulated for learners, which includes providing appropriate test question resources according to the personalized characteristics and needs of learners. Test question resources are a necessary condition for accurately evaluating the knowledge state of learners and a key factor for learners to achieve personalized learning. The discrimination of test questions is an important parameter of test question resources. Reasonably providing test question resources with appropriate difficulty for learners can, on the one hand, support learners to achieve timely and accurate evaluation of their knowledge state. On the other hand, it can also stimulate the learning potential of learners and help them cross the nearest development zone of their own cognition.

[0003] Existing discrimination prediction methods are mainly divided into two categories: manual calibration and data-driven. Manual calibration means that teachers or experts label the discrimination of test questions according to their experience, subjective opinions, etc. At the same time, in recent years, with the development of learning platforms, marking systems and question bank systems, a boom in automatic prediction methods for test question difficulty driven by data has been evoked. In recent years, the main automatic prediction methods include classical measurement theory methods based on statistics and prediction methods based on natural language processing.

[0004] The existing manual calibration prediction technology marks the difficulty of test questions subjectively by teachers according to their own experience, intentions, etc., which has obvious disadvantages such as time-consuming and laborious, and the difficulty marking is highly subjective and not accurate enough.

[0005] The classical measurement theory method based on statistics predicts factors such as the discrimination of test questions and the ability of candidates from a statistical perspective, but its mathematical model is simple and cannot describe complex logical relationships. It often requires professional knowledge and largely relies on manually marked data.

[0006] The research objects of prediction methods based on natural language processing are mostly limited to multiple-choice questions, and other question types are not considered enough, resulting in insufficient practicality. At the same time, the current text encoding methods mainly use traditional methods such as Word2vec and ELMo, which are difficult to represent the depth of test question texts and integrate more context semantic information.

[0007] Therefore, if users do not have strong professional capabilities related to data, they cannot classify and evaluate test questions ideally, and its usability is poor. Summary of the Invention

[0008] The present invention overcomes the deficiencies of the prior art and provides a method for evaluating the discrimination of test questions based on the GRM model. By adopting a new method, it enables users without strong professional capabilities related to data to classify test questions more ideally, so as to solve the technical problem of poor usability of the method for evaluating test questions.

[0009] The object of the present invention is achieved through the following technical solutions:

[0010] A method for evaluating the discrimination of test questions based on the GRM model includes the following steps:

[0011] S1: Obtain information on question IDs, user IDs, the number of assessment points, and the number of passed assessment points from the database of the learning system;

[0012] S2: Sequentially determine whether each question has been answered by a user, and screen out users who have not participated in answering;

[0013] S3: If a user participates in answering, then determine whether the number of passed assessment points of the user is equal to the number of assessment points, and at the same time, the number of times the user answers the question +1; if it is determined that the number of passed assessment points of the user is not equal to the number of assessment points, then delete the answering data, return to step S2, and continue to judge until all questions have been judged;

[0014] S4: If the number of passed assessment points of the user is equal to the number of assessment points, then the number of correctly answered questions +1;

[0015] S5: Use the screened data as a new data set for the K-Means clustering algorithm;

[0016] S6: Determine the value of K through the silhouette coefficient;

[0017] S7: Determine the difficulty coefficient of the GRM model through the value of K, introduce the ability judgment parameter of the user to generate an ICC graph based on GRM; at the same time, judge the user type according to which question is answered correctly for the first time;

[0018] S8: Evaluate and obtain the discrimination of the test questions.

[0019] Preferably, in step S4, when the user answers a question correctly for the first time, record which time it is, and count the total number of times of answering the question and the answering accuracy rate.

[0020] Preferably, in step S5, the screened data includes the ID of the participating user, the question ID, the number of correctly answered questions, the total number of answered questions, the accuracy rate, and which question is answered correctly for the first time.

[0021] Preferably, in step S6, determine to divide into K clusters through the silhouette coefficient S, and the calculation formula is as follows:

[0022]

[0023] Where a represents the average distance of a sample point from all other points in the same cluster, that is, the similarity of the sample point to other points in the same cluster; b represents the average distance of the sample point from all points in the next nearest cluster, that is, the similarity of the sample point to other points in the next nearest cluster, and S ∈ (-1, 1).

[0024] The closer S is to 1, the better the clustering effect; the closer it is to -1, the worse the clustering effect. In order to make the difference within each cluster small and the difference between clusters large for each cluster, after importing the data, select the silhouette coefficient that is closer to 1, and then determine the value of K, that is, determine how many categories this question can be divided into.

[0025] Preferably, in step S7, each question has a discrimination parameter, and there are more than 1 item difficulty parameters. The item difficulty parameters have a size order, and the higher the level, the greater the difficulty. The probability that user j gets t points or more on question i is:

[0026]

[0027] Where A represents the ability level, D is a constant,, a i represents the item discrimination of the i-th question, A j represents the ability of user j, B i represents the difficulty of question i.

[0028] More preferably, the judgment of the ability level should be based on the data of the number of times when the user answers the question correctly for the first time and the correct rate of the student's answer. The earlier the number of times the user answers the question correctly and the higher the correct rate, the stronger the ability of the user;

[0029] The calculation of the probability score of user j getting t points in question i is:

[0030] P(X ji = t) = P(X ji ≥ t) - P(X ji ≥ t + 1)

[0031] Where P(X ji ≥ t) represents the probability that user j scores greater than or equal to t points in question i, and P(X ji ≥ t + 1) represents the probability that user j scores greater than or equal to t + 1 points in question i.

[0032] Subtracting the two can obtain the probability that user j gets t points in question i.

[0033] Preferably, in the method for evaluating the discrimination of test questions based on the GRM model, in step S7, if there are n difficulty coefficients, there are n + 1 types of scores for the user, that is, K-Means clustering yields n + 1 categories.

[0034] Preferably, in step S7, the value of the number of user types is n + 1, which is equal to the number of categories obtained by K-Means clustering.

[0035] The beneficial effects of the present invention are as follows:

[0036] 1. The present invention performs clustering through the K-Means clustering algorithm, improving the usability of the method for evaluating the discrimination of test questions. The present invention does not require the user to have strong professional capabilities related to data. It can also relatively ideally divide the data into K categories. Even without knowing any sample labels in advance, it can divide the samples into several categories according to the internal relationship between the data, making the similarity between samples in the same category high and the similarity between samples in different categories low. Therefore, even if the user does not understand relevant professional knowledge, they can classify the learners participating in the same question.

[0037] 2. The present invention conducts data simulation based on the item response model GRM, and can intuitively see the relationship between the ability of learners and the score levels from the final ICC curve. The ability of learners depends on factors such as the number of times they answer a question correctly for the first time and the answer correct rate. When the number of times a learner answers a question correctly for the first time is earlier and the answer correct rate is higher, then the learning ability is stronger. If, as the ability of the learner increases, it can be seen that the probability of getting a low score decreases, while the probability of getting a high score increases continuously, making this question have good discrimination. The discrimination of a test question can be evaluated based on this. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The present invention is further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the following drawings.

[0039] Figure 1 is a schematic flowchart of the method for evaluating the discrimination of test questions based on the GRM model in an embodiment of the present invention;

[0040] Figure 2 is the ICC curve graph based on the GRM model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The following further describes in detail the method for evaluating the discrimination of test questions based on the GRM model with specific embodiments. These embodiments are only for comparison and explanation purposes, and the present invention is not limited to these embodiments.

[0042] Embodiment

[0043] As Figure 1 shown, the method for evaluating the item discrimination based on the GRM model provided by the present invention includes the following steps:

[0044] S1: Obtain the information of the question ID, user ID, number of assessment points, and number of passed assessment points from the database of the learning system;

[0045] S2: Judging in turn whether each question has been answered by a user, and screening out the users who have not participated in the answer;

[0046] S3: If the user participates in the answer, then judge whether the number of passed assessment points of the user is equal to the number of assessment points, and at the same time, the number of times the user answers the question +1; if it is judged that the number of passed assessment points of the user is not equal to the number of assessment points, then delete the answer data, return to step S2, and continue to judge until all questions are judged;

[0047] S4: If the number of passed assessment points of the user is equal to the number of assessment points, then the number of correct questions +1;

[0048] S5: Use the screened data as a new data set for the K-Means clustering algorithm;

[0049] S6: Determine the value of K through the silhouette coefficient;

[0050] S7: Determine the difficulty coefficient of the GRM model through the value of K, introduce the ability judgment parameter of the user to generate the ICC graph based on GRM; at the same time, judge the user type according to which question is answered correctly for the first time;

[0051] S8: Evaluate and obtain the item discrimination.

[0052] Preferably, in step S4, when the user answers the question correctly for the first time, record which time it is, and count the total number of times of answering the question and the answering accuracy rate.

[0053] Preferably, in step S5, the screened data includes the ID of the participating user, the question ID, the number of correct questions, the total number of answered questions, the accuracy rate, and which question is answered correctly for the first time.

[0054] Preferably, in step S6, determine to divide into K clusters through the silhouette coefficient S, and the calculation formula is as follows:

[0055]

[0056] Among them, a represents the average distance between a sample point and all other points in the same cluster, that is, the similarity between the sample point and other points in the same cluster; b represents the average distance between the sample point and all points in the next nearest cluster, that is, the similarity between the sample point and other points in the next nearest cluster, and S ∈ (-1, 1).

[0057] After determining the division into K clusters through the silhouette coefficient, the data is clustered, and the main process is as follows:

[0058] 2 Randomly select K centers, denoted as

[0059] ② Define the loss function:

[0060]

[0061] Among them, x i represents the i-th sample, C i is the cluster to which x i belongs, represents the center point corresponding to the cluster, M is the total number of samples, and J is the defined loss function; J(C, t)

[0062] ③ Let z = 0, 1, 2...... be the number of iteration steps, and repeat the following process until J converges:

[0063] a: For each sample x i Assign it to the nearest center

[0064]

[0065] b: For each class center K, recalculate the center of this class

[0066]

[0067] First, fix the center point, adjust the category to which each sample belongs to reduce J, then fix the category of each sample, adjust the center point to continue to reduce J, and the two processes alternate in a loop. J monotonically decreases until the minimum value, and the center point and the category of the sample division converge simultaneously.

[0068] Preferably, in step S7, each question has a discrimination parameter, and there are more than 1 item difficulty parameters. The item difficulty parameters have a size order, and the higher the level, the greater the difficulty. The probability that user j gets t and above t points on question i is:

[0069]

[0070] Among them, A represents the ability level, D is a constant, generally taking the value of 1.702, a i represents the item discrimination of the i-th question, A jDenote the ability of user j as B i Denote the difficulty of test question i

[0071] More preferably, the judgment of the ability level is based on two pieces of data: the order in which the user answers the question correctly for the first time and the correct rate of the student's answer. The earlier the number of times the user answers the question correctly and the higher the correct rate, the stronger the ability of the user

[0072] The calculation of the probability score probability of t points for user j in test question i is as follows

[0073] P(X ji =t)=P(X ji ≥t)-P(X ji ≥t + 1)

[0074] where P(X ji ≥t) represents the probability that user j scores greater than or equal to t points in question i, and P(X ji ≥t + 1) represents the probability that user j scores greater than or equal to t + 1 points in question i

[0075] Preferably, in S7, if there are n difficulty coefficients, then there are n + 1 types of scores for the user, that is, K-Means clustering obtains n + 1 categories

[0076] Preferably, in step S7, the value of the number of user types is n + 1, which is equal to the category value obtained by K-Means clustering

[0077] If there are two difficulty coefficients, then there are three types of scores for the user, which can be 0 points, 1 point, and 2 points. These three types here actually correspond to the three categories in the clustering. For example Figure 2 After clustering, it can be divided into 3 categories. Simulate the ICC curve of GRM, where the horizontal axis represents the ability level of the learner, the vertical axis represents the probability, curve ③ represents a score of 2 points, curve ② represents a score of 1 point, and curve ① represents a score of 0 points

[0078] From Figure 2 it can be obtained that when the ability of the test taker (user) is stronger, the probability of getting two points will increase. When the ability of the test taker is at a medium level, the probability of getting 1 point is also the largest. When the ability of the test taker is weaker, the probability of getting 0 points is also the largest. Therefore, through this model, the discrimination of the test questions can be evaluated. The scoring probabilities of test takers with different abilities are different. When the ability of the test taker is stronger, then they will belong to a higher level of score, that is, the higher the score. Based on this, students with different learning abilities can be distinguished, and thus the discrimination of a test question can be evaluated

[0079] The setting of the difficulty coefficient should be based on the number of clusters divided by the K-Means clustering algorithm. For example, if there are 5 clusters, then it can be divided into 5 categories, that is, four difficulty coefficients, and the scoring situations can be 0, 1, 2, 3, 4 from low to high.

[0080] Finally, by comparing whether there is a difference in the scores obtained by students with different ability values when answering this question. For example, if a student with strong ability gets a relatively high score when answering this question, and a student with weak ability gets a relatively low score when answering this question, it can indicate that the discrimination of this question is good. On the contrary, if both students with strong ability and students with weak ability get relatively high or low scores when answering this question, it indicates that the discrimination of this question is poor.

[0081] The method for evaluating the discrimination of test questions provided in the above embodiments of the present invention focuses on clustering through the K-Means clustering algorithm, and does not require the user to have strong professional ability related to data. It can also ideally divide the data into K categories. Even when not knowing any sample labels in advance, it can divide the samples into several categories according to the internal relationship between the data, so that the similarity between samples in the same category is high, and the similarity between samples in different categories is low. Therefore, even if the user does not understand the relevant professional knowledge, they can classify the learners participating in the same question, improving its usability.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A method for evaluating the discrimination of test questions based on the GRM model, characterized in that, It includes the following steps: S1: Obtain the information of question ID, user ID, number of assessment points, and number of passed assessment points from the database of the learning system; S2: Sequentially determine whether each question has been answered by users, and screen out users who have not participated in answering; S3: If a user participates in answering, determine whether the number of passed assessment points of the user is equal to the number of assessment points, and at the same time, the number of answering times of this user is incremented by 1; if it is determined that the number of passed assessment points of the user is not equal to the number of assessment points, then delete the answering data, return to step S2, and continue to judge until all questions have been judged; S4: If the number of passed assessment points of the user is equal to the number of assessment points, then the number of correctly answered questions is incremented by 1; S5: Use the filtered data as a new data set for the K-Means clustering algorithm; S6: Determine the value of K through the silhouette coefficient; S7: Determine the difficulty coefficient of the GRM model through the value of K, introduce the ability judgment parameter of the user to generate an ICC graph based on the GRM; at the same time, judge the user type according to which time the user answers correctly for the first time; S8: Evaluate and obtain the discrimination of the test questions; In the above step S7, each question has a discrimination parameter, and there are more than 1 item difficulty parameters. The item difficulty parameters have a size order, and the higher the level, the greater the difficulty. The probability that user j gets t and above t points on question i is: Where A represents the magnitude of ability, D is a constant, and a i represents the item discrimination of the i-th question, and A j represents the ability of user j, and B i represents the difficulty of question i; The judgment of the ability level should be based on the data of which time the user answers correctly for the first time and the correct rate of the student's answering. The earlier the number of times the user answers correctly and the higher the correct rate, the stronger the ability of the user; The calculation of the probability score of user j getting t points in test question i is: P(X ji = t) = P(X ji ≥ t) - P(X ji ≥ t + 1) Where P(X ji ≥t) represents the probability that user j scores greater than or equal to t points in question i, P(X ji ≥t+1) represents the probability that user j scores greater than or equal to t+1 points in question i.

2. The method for evaluating the discrimination of test questions based on the GRM model according to claim 1, wherein In the above step S4, when the user answers a question correctly for the first time, record which time it is, and count the total number of answering times and the answering correct rate.

3. The test question discrimination evaluation method based on the GRM model according to claim 1 is characterized in that: In the above step S5, the filtered data includes the ID of the participating user, question ID, number of correctly answered questions, total number of answered questions, correct rate, and which question the user answers correctly for the first time.

4. The item discrimination evaluation method based on the GRM model according to claim 1, characterized in that In the above step S6, determine to divide into K clusters through the silhouette coefficient S, and the calculation formula is as follows: Where a represents the average distance between the sample point and all other points in the same cluster, that is, the similarity between the sample point and other points in the same cluster; b represents the average distance between the sample point and all points in the next nearest cluster, that is, the similarity between the sample point and other points in the next nearest cluster, and S ∈ (-1, 1).

5. The method for evaluating the item discrimination based on the GRM model according to claim 1, characterized in that In the above step S7, if there are n difficulty coefficients, then there are n + 1 types of scores for the user, that is, the K-Means clustering obtains n + 1 categories.

6. The method for evaluating the item discrimination based on the GRM model according to claim 1, wherein, In the above step S7, the value of the number of user types is n + 1, which is equal to the category value obtained by the K-Means clustering.

Citation Information

Patent Citations

  • Prediction method, system, terminal and server of test score

    CN106682768A

  • Learning ability evaluation method and system based on cognitive diagnosis

    CN114491050A