Difficulty level calculation device, difficulty level calculation method, and test creation system

The difficulty level calculation device uses AI to analyze question data and historical answer rates to determine the difficulty of new exam questions, addressing the challenge of standardized difficulty level calculation in exams.

JP7833229B1Active Publication Date: 2026-03-19FORESIGHT CO LTD
View PDF 24 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing systems struggle to accurately calculate the difficulty level of exam questions, especially when new questions are created without historical data, and fail to standardize difficulty levels across multiple practice exams.

Method used

A difficulty level calculation device that analyzes existing problem data to determine index values for factors influencing difficulty, using generative artificial intelligence to identify units and answer formats, and calculates the difficulty level of new questions based on these factors and historical answer rates.

Benefits of technology

Enables accurate calculation of difficulty levels for new questions, ensuring consistency across exams and facilitating the creation of tests with desired difficulty levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007833229000001_ABST
    Figure 0007833229000001_ABST
Patent Text Reader

Abstract

Calculate the difficulty level of the exam questions. [Solution] The difficulty level calculation device includes: a storage unit that stores existing problem data including the question texts of multiple existing problems and existing problem answer data relating to the answers of multiple test takers to the existing problems; a first processing unit that calculates index values ​​for multiple factors influencing the difficulty level of each of the existing problems based on the content and / or format of the question text of the existing problem, calculates the correct answer rate based on the existing problem answer data, and determines weights indicating the degree of influence of each of the multiple index values ​​on the correct answer rate; and a second processing unit that calculates index values ​​for multiple factors of a target problem based on target problem data including the question text of the target problem whose difficulty level is to be evaluated, and calculates the difficulty level of the target problem based on the index values ​​and the weights.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a technique for calculating the difficulty level of test questions.

Background Art

[0002] Various tests such as qualification tests are composed of a plurality of questions with various difficulty levels. For mock tests, it is desirable to set the difficulty level according to this test. Also, in order to evaluate the growth of learners in mock tests conducted multiple times, it is desirable to align the difficulty levels.

[0003] Patent Document 1 discloses an exam optimization system that can suppress the variation in the difficulty level of questions to be presented and reflect the test results in question selection. The exam optimization system of Patent Document 1 stores the statistical information of each question registered as a candidate for test questions, the questions based on the test implementation history, and each test result. Then, the exam optimization system acquires a question pool including each question that is a candidate for question presentation, and according to the designation of the test operator, uses either one of the question selection methods including a method based on a predetermined average method and a method based on a predetermined function that can output each question to be a candidate as a probability value, selects a question from the question pool, and generates a test question. Further, the exam optimization system calculates the individual statistical information about the test results of the generated test questions and reflects it in the stored statistical information.

[0004] Patent Document 2 discloses a test question analysis system that evaluates the passing influence degree, which is the degree of correctness required to pass the test for a question. The test question analysis system of Patent Document 2 discriminates the prerequisite knowledge required to solve the target question to be analyzed, calculates the question sentence length, which is the length of the sentence of the target question, specifies the answer format for the target question, and evaluates the passing influence degree based on the prerequisite knowledge, the question sentence length, and the answer format.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

[0006] The question optimization system described in Patent Document 1 selects questions of appropriate difficulty from a question pool containing statistical information such as success rates based on the history of exam implementation, conducts the exam, calculates individual statistical information based on the exam results, and reflects it in the stored statistical information. Therefore, when a new question is created, it is not possible to calculate the difficulty level because there is no statistical information for that question.

[0007] The pass / fail influence calculated by the examination question analysis system of Patent Document 2 represents the degree to which a question must be answered correctly in order to pass the exam. Therefore, while the pass / fail influence is useful for learners to determine which questions to focus on and to what extent, it is unsuitable for use by those who create or provide examination questions, such as setting the difficulty level of the exam or each question, or standardizing the difficulty level of multiple practice exams. One of the purposes included in this disclosure is to provide a technique for calculating the difficulty level of exam questions. [Means for solving the problem]

[0008] A difficulty level calculation device according to one embodiment included in this disclosure includes: a storage unit that stores existing problem data including the question texts of a plurality of existing problems, and existing problem answer data relating to the answers of a plurality of test takers to the existing problems; a first processing unit that calculates index values ​​for a plurality of factors influencing the difficulty level of each of the existing problems based on the content and / or format of the question text of the existing problem, calculates the correct answer rate based on the existing problem answer data, and determines weights indicating the degree of influence of each of the plurality of index values ​​on the correct answer rate; and a second processing unit that calculates index values ​​for a plurality of the factors of a target problem based on target problem data including the question text of the target problem whose difficulty level is to be evaluated, and calculates the difficulty level of the target problem based on the index values ​​and the weights. [Effects of the Invention]

[0009] According to one aspect of this disclosure, it is possible to calculate the difficulty level of a problem. [Brief explanation of the drawing]

[0010] [Figure 1] This is a block diagram of the difficulty level calculation device. [Figure 2] This is a block diagram of the learning processing unit. [Figure 3] This is a block diagram of the estimation processing unit. [Figure 4] This diagram shows the hardware configuration of the computer that makes up the difficulty level calculation device. [Figure 5] This figure shows an example of a question frequency score determination table. [Figure 6] This figure shows an example of a prerequisite knowledge score determination table. [Figure 7] This figure shows an example of a multiple-choice question. [Figure 8] This figure shows an example of a combinatorial problem. [Figure 9] This figure shows an example of a multiple-choice question. [Figure 10] This figure shows an example of a calculation problem. [Figure 11]It is a diagram showing an example of a quantity problem. [Figure 12] It is a diagram showing an example of a descriptive problem. [Figure 13] It is a diagram showing an example of an answer format score determination table. [Figure 14] It is a diagram showing an example of a number of confusing limbs score determination table. [Figure 15] It is a diagram showing an example of a character count score determination table. [Figure 16] It is a diagram showing an example of learning data. [Figure 17] It is a block diagram showing the configuration of a test creation system using a difficulty calculation device.

Embodiments for Carrying Out the Invention

[0011] Embodiments of the present invention will be described with reference to the drawings. In the present embodiment, a difficulty calculation device for calculating the difficulty of the problems of a given test is disclosed.

[0012] FIG. 1 is a block diagram of the difficulty calculation device according to the present embodiment. The difficulty calculation device 10 has a data storage unit 20, a learning processing unit 30, and an estimation processing unit 40. The difficulty calculation device 10 receives an input of target problem data 51 of a target problem for which the difficulty is to be calculated, calculates the difficulty of the target problem, and outputs it as difficulty 52.

[0013] The data storage unit 20 stores existing problem data, which includes the question texts of multiple existing problems, and existing problem answer data, which includes the answers of multiple test takers to the existing problems. Existing problems are problems that have been administered in the past and for which answers have been obtained from test takers. For example, these include problems that have been asked in the past N (e.g., the past 10) main examinations, and / or problems that have been asked in mock examinations conducted in the past. For main examination problems, the examination institution has created correct answers and explanations, and the test takers' answers have been collected and stored. For mock examination problems, the examination institution has prepared correct answers and explanations, and the test takers' answers have been collected and stored. The existing problem data includes the question texts of the existing problems. The existing problem data may also include explanations for the existing problems. The existing problem answer data includes the answers collected from test takers to the existing problems. Existing problems have correct answers, and based on these correct answers, it is possible to calculate the correct answer rate for each existing problem, which is the percentage of correct answers out of all possible answers. The data stored in the data storage unit 20 is used for processing by the learning processing unit 30.

[0014] The learning processing unit 30 determines the weights that indicate the degree to which each of the index values ​​for multiple factors influencing the difficulty of a problem has an impact on the correct answer rate, based on the existing problem data 22 and existing problem answer data 21 stored in the data storage unit 20. Specifically, for each existing problem, the learning processing unit 30 calculates index values ​​for multiple factors influencing the difficulty of the problem based on the content and format of the problem statement. Each factor relates to the difficulty of the problem from its own perspective, and the difficulty of the problem is determined by the combined influence of these factors.

[0015] Furthermore, the learning processing unit 30 calculates the correct answer rate for each of the existing problems based on the existing problem answer data. The learning processing unit 30 then determines weights that indicate the degree of influence of each of the multiple indicator values ​​on the correct answer rate. The weights determined in this way by the learning processing unit 30 are used in the processing by the estimation processing unit 40.

[0016] The estimation processing unit 40 calculates index values ​​for multiple factors of the target problem based on the target problem data 51, which includes the problem statement of the target problem to be evaluated for difficulty. Based on the calculated index values ​​and the weights determined by the learning processing unit 30, the estimation processing unit 40 calculates the difficulty level 52 of the target problem. The estimation processing unit 40 corresponds to the second processing unit in the claim. The target problem is a problem whose difficulty level and correct answer rate are unknown. For example, a problem created for a mock exam to be conducted may be the target problem.

[0017] As described above, the difficulty level calculation device 10 calculates the difficulty level 52 of the target problem based on the content and format of the problem and the correct answer rate of existing problems for which there is a track record of answers. Therefore, even if the target problem is one for which there is no track record of answers, the difficulty level of the target problem can be accurately identified.

[0018] Figure 2 is a block diagram of the learning processing unit. Figure 3 is a block diagram of the estimation processing unit. Figure 4 is a diagram showing the hardware configuration of the computer that makes up the difficulty level calculation device.

[0019] Referring to Figure 2, the learning processing unit 30 includes a question frequency evaluation unit 31, a prerequisite knowledge evaluation unit 32, an answer format evaluation unit 33, a confusion option number evaluation unit 34, a question text length evaluation unit 35, a correct answer rate calculation unit 36, and a weight determination unit 37. The learning processing unit 30 receives existing problem answer data 21, existing problem data 22, learning content data 23, and a score determination table group 24 as input. The learning processing unit 30 outputs learning data 25 and weight information 26.

[0020] The question frequency evaluation unit 31 calculates an index value for the frequency of questions for each existing question, based on the existing question data 22 and the learning content data 23, for each unit, which is a coherent unit of information that learners should learn and acquire. The learning content data 23 is data that records the learning content that learners should acquire for this exam. The learning content is equivalent to a learning text, and the information that learners should acquire is recorded in units.

[0021] For example, the question frequency evaluation unit 31 calculates how many times a given unit of knowledge that learners should acquire in order to solve the existing problem has appeared in the past N (e.g., the past 10) exams. In doing so, the question frequency evaluation unit 31 may extract words that appear in the question text of the existing problem and identify the unit from the extracted words. Alternatively, the question frequency evaluation unit 31 may use an unillustrated generative artificial intelligence model to identify the unit from the question text. The generative artificial intelligence model could be, for example, a large-scale language model or a small-scale language model. For example, the generative artificial intelligence model could be input with the question text and learning content, and output which unit of the learning content the question text corresponds to. Furthermore, the question frequency evaluation unit 31 calculates an index value by referring to the question frequency score judgment table included in the score judgment table group 24.

[0022] Figure 5 shows an example of a question frequency score determination table. The question frequency score determination table 241 defines a score A corresponding to the number of times each unit has appeared in the past 10 main examinations. The question frequency evaluation unit 31 refers to the question frequency score determination table 241 and obtains a score A corresponding to the number of times an existing unit has appeared.

[0023] In the example in Figure 5, for instance, if a certain unit is presented 10 times, the frequency evaluation unit 31 obtains a score of 10 as A. Similarly, if a certain unit is presented 9 times, the frequency evaluation unit 31 obtains a score of 9 as A. If a certain unit is presented 2 times, the frequency evaluation unit 31 obtains a score of 2 as A. If a certain unit is presented 1 time or 0 times, the frequency evaluation unit 31 obtains a score of 1 as A.

[0024] Here, as an example, a lower score of A indicates a higher difficulty level. Topics that are frequently tested tend to have higher correct answer rates because learners have studied them well, while topics that are rarely tested tend to have lower correct answer rates because learners have studied them very little. The question frequency evaluation unit 31 then stores the score A of the factors related to question frequency as an index value in the learning data 25.

[0025] The prerequisite knowledge evaluation unit 32 calculates an index value for each of the existing problems, based on the existing problem data 22 and the learning content data 23, regarding the breadth of the prerequisite knowledge required to solve the problem.

[0026] For example, the prerequisite knowledge evaluation unit 32 identifies the number of units of knowledge that a learner must learn and acquire in order to solve the existing problem in the learning content. In this case, the prerequisite knowledge evaluation unit 32 may identify the units from the words that appear in the problem statement of the existing problem. Alternatively, the question frequency evaluation unit 31 may use an unillustrated generative artificial intelligence model to identify the units or the number of units from the problem statement. The generative artificial intelligence model may be, for example, a large-scale language model or a small-scale language model. For example, the generative artificial intelligence model may be input with the problem statement and learning content and output how many units of the learning content the problem statement corresponds to. Furthermore, the prerequisite knowledge evaluation unit 32 refers to the prerequisite knowledge score judgment table included in the score judgment table group 24 and extracts an index value corresponding to the number of units.

[0027] Figure 6 shows an example of a prerequisite knowledge score determination table. The prerequisite knowledge score determination table 242 defines a score B corresponding to the number of prerequisite knowledge units as the breadth of the prerequisite knowledge required to answer the question.

[0028] For example, if the number of prerequisite knowledge units for a given problem is 1, the prerequisite knowledge evaluation unit 32 obtains a score of 10 as a B score. If the number of prerequisite knowledge units for a given problem is 2, the prerequisite knowledge evaluation unit 32 obtains a score of 8 as a B score. If the number of prerequisite knowledge units for a given problem is 3, the prerequisite knowledge evaluation unit 32 obtains a score of 6 as a B score. If the number of prerequisite knowledge units for a given problem is 4, the prerequisite knowledge evaluation unit 32 obtains a score of 4 as a B score. If the number of prerequisite knowledge units for a given problem is 5, the prerequisite knowledge evaluation unit 32 obtains a score of 2 as a B score.

[0029] Here, as an example, a smaller score B indicates a higher difficulty level. In the prerequisite knowledge score determination table 242, a smaller number of units corresponds to a larger score B. This reflects the tendency for questions to become easier as the scope of prerequisite knowledge narrows. When a test-taker solves a question from a single unit, they only need to recall that single unit and apply it to each option. On the other hand, if, for example, each option is drawn from multiple different units, the test-taker needs to recall a wide range of knowledge to solve the problem, making it more difficult. The score B calculated by the prerequisite knowledge evaluation unit 32 is stored as an index value for factors related to prerequisite knowledge in the learning data 25.

[0030] The answer format evaluation unit 33 calculates an index value for the answer format of each existing problem based on the existing problem data 22. Typical answer formats include multiple-choice questions, matching questions, selection questions, calculation problems, counting problems, and written response questions. Even for questions concerning knowledge in the same field, the ease of arriving at the correct answer may vary depending on the answer format. The format of the question is also determined to some extent by the answer format. Therefore, the answer format evaluation unit 33 analyzes the question to identify the answer format.

[0031] For example, the answer format evaluation unit 33 can determine which answer format a question belongs to by determining whether the question statement meets predetermined conditions for each answer format. Alternatively, the answer format evaluation unit 33 can identify the answer format from the question statement by using a generative artificial intelligence model. Generative artificial intelligence models include, for example, large-scale language models and small-scale language models. For example, the generative artificial intelligence model may be input with the question statement and templates for each answer format, and output which answer format the question statement corresponds to.

[0032] Below are examples of exam questions for each answer format. Figure 7 shows an example of a multiple-choice question. Multiple-choice questions are questions that present a set number of sentences (options) as choices, and the test-taker must select one correct answer (correct option) from among these options. The number of options (options) varies depending on the test, for example, four options, five options, etc. Figure 7 shows an example of a multiple-choice question that asks the test-taker to determine which of the five sentences (options) is correct. The correct option is the correct answer. Another example is a multiple-choice question that asks the test-taker to identify one incorrect option from among several options. In this case, the incorrect option is the correct answer. A characteristic of multiple-choice questions is that the question text contains multiple options, and these options themselves are the answer choices.

[0033] In addition, multiple-choice questions in legal exams and other similar tests include knowledge-based questions, case-based questions, and precedent-based questions. Case-based questions ask about specific cases. A characteristic of case-based questions is that the question text includes words that indicate the specific entities involved in the case (case subject terminology). Typical case subject terminology that appear in legal exams include A, B, C, D, Company A, Company B, Company C, and Company D. Precedent-based questions ask about understanding of specific precedents. A characteristic of precedent-based questions is that the question text includes words or combinations of words specific to those precedents. Knowledge-based questions ask about knowledge of specific matters. Questions that do not fall under either precedent-based or case-based questions can be presumed to be knowledge-based questions.

[0034] Figure 8 shows an example of a combinatorial problem. A combination problem presents multiple options consisting of several sentences, and the question asks the student to select the combination of sentences that satisfies the conditions given in the problem statement. This includes cases with zero or one sentence. Figure 8 shows an example of a combination problem where there are three items, A, B, and C, and four options (1, 2, 3, and 4), and the student must select the incorrect combination. A key characteristic of combination problems is that the problem statement contains multiple sentences, and each of the options consists of a combination of zero or more sentences.

[0035] Figure 9 shows an example of a multiple-choice question. Multiple-choice questions are those in which the question contains several blanks, and a separate list of options containing words that can fill in those blanks is provided. The questioner is asked to select the correct word or phrase from the options to fill in each blank. Figure 9 shows an example where the question contains four blanks labeled A, B, C, and D, and the questioner is asked to select the correct word or phrase from 20 options, numbered 1 to 20. Multiple-choice questions are characterized by having a question containing multiple blanks and a separate section listing the possible words or phrases that can fill in the blanks as options.

[0036] Figure 10 shows an example of a calculation problem. Calculation problems are problems in which you perform numerical calculations according to the conditions given in the problem statement to arrive at the correct answer. Figure 10 shows an example problem in which four amounts, 1 to 4, are given as options, and you are asked to calculate the amount required in the problem statement and select the option that matches the calculated amount. Calculation problems are characterized by the fact that each of the answer options contains a number, and the problem statement also contains the numbers used in the calculation.

[0037] Figure 11 shows an example of a counting problem. A counting problem is a type of problem where the problem statement contains multiple options, and the number of options is given as a set of choices. The question asks the reader to select the number of options that satisfy the conditions given in the problem statement. The number here includes zero. Figure 10 shows an example of a counting problem where the problem statement has three options, A, B, and C, and four choices, 1 (1), 2 (2), 3 (3), and 4 (none). The question asks the reader to select how many of the three options are incorrect. A characteristic of counting problems is that each of the multiple choices is a combination of symbols that correspond to multiple options. The symbols here can be anything that identifies each option, such as numbers, letters of the alphabet, katakana, hiragana, kanji, or shapes.

[0038] Figure 12 shows an example of a written response question. Descriptive questions require students to write about specific matters based on the question text. Figure 12 shows an example of a descriptive question that requires students to write about the matters required in the question text in a sentence of a specified length. A characteristic of descriptive questions is that they contain specific words or phrases that require written responses. "Write a response" in the question text shown in Figure 12 is an example of a specific word or phrase that requires written responses. Other examples include "Write a response," "Discuss the matter," and "Explain."

[0039] The answer format evaluation unit 33 identifies the answer format of an existing problem and calculates an index value by referring to the answer format score determination table 243 included in the score determination table group 24.

[0040] Figure 13 shows an example of an answer format score determination table. The answer format score determination table 243 defines the score C corresponding to the answer format. When the answer format evaluation unit 33 identifies the answer format, it refers to the answer format score determination table 243 to obtain the score C corresponding to the answer format.

[0041] In the example of the answer format score determination table 243 in Figure 13, a score of 10 is associated with multiple-choice knowledge questions and combination questions. Multiple-choice knowledge questions and combination questions are in a format that can be solved with simple memorization or basic calculations. Also, a score of 8 is associated with multiple-choice case questions. Multiple-choice case questions are in a format that requires the application of knowledge. A score of 6 is associated with multiple-choice case law questions and selection questions. Multiple-choice case law questions and selection questions are in a format that requires reading comprehension. A score of 4 is associated with calculation problems and counting problems. Calculation problems and counting problems are in a format that requires multiple pieces of knowledge or accurate knowledge. A score of 2 is associated with descriptive questions. Descriptive questions are in a format that requires advanced interpretation and application skills. Thus, as an example here, the lower the score C, the more difficult the question.

[0042] The score C calculated by the answer format evaluation unit 33 is stored as an index value of the factors related to the answer format in the learning data 25.

[0043] The Confusing Choices Evaluation Unit 34 calculates an index value for each of the existing problems based on the existing problem data 22, relating to the number of choices in a multiple-choice question where it is not easy to determine whether the choice is correct or incorrect (confusing choices). The Confusing Choices Evaluation Unit 34 calculates the index value by referring to the confusing choices score determination table included in the score determination table group 24.

[0044] Figure 14 shows an example of a confusing answer score determination table. The confusing answer score determination table 244 defines a score D corresponding to the number of confusing answer choices in a multiple-choice question. Furthermore, the confusing answer score determination table 244 also defines a score D for questions with answer formats other than multiple-choice questions.

[0045] The confusion count evaluation unit 34 identifies the number of options whose correctness is not easy to determine, based on the existing problem data 22, when the existing problem is a multiple-choice question, and obtains a score D corresponding to the confusion count by referring to the confusion count score determination table 244. For each option, the determination of whether it is easy or incorrect can be made, for example, by judging it as easy if it can be determined whether it is correct or incorrect from the content of the learning content data 23, and judging it as not easy if it cannot be determined whether it is correct or incorrect from the content of the learning content data 23.

[0046] In the confusion score determination table 244, for a four-choice multiple-choice question, a score of 10 is assigned to one confusing option, a score of 7.5 is assigned to two options, a score of 5 is assigned to three options, and a score of 2.5 is assigned to four options. Similarly, for a five-choice multiple-choice question, a score of 10 is assigned to one confusing option, a score of 8 is assigned to two options, a score of 6 is assigned to three options, a score of 4 is assigned to four options, and a score of 2 is assigned to five options.

[0047] Furthermore, if an existing question is not a multiple-choice question, the score D is set to 1. In other words, combination questions, selection questions, calculation questions, counting questions, and written response questions are all assigned a score of 1 as D.

[0048] In multiple-choice questions, the probability of arriving at the correct answer through elimination tends to change depending on the number of options that are easy to determine as correct or incorrect, and the number of options that are not. In the Confusing Option Score Judgment Table 244, a smaller score D is assigned to options that are difficult to determine as correct or incorrect, reflecting the tendency that the more options that are difficult to determine as correct or incorrect, the less likely the elimination method is to be used, making the question more difficult.

[0049] Thus, as an example here, a smaller score D indicates a higher difficulty level. The score D calculated by the confusion count evaluation unit 34 is stored as an index value of the factors related to the confusion count in the training data 25.

[0050] The problem sentence length evaluation unit 35 calculates an index value for the length of the sentence in each of the existing problems based on the existing problem data 22. The problem sentence length evaluation unit 35 calculates the index value by referring to the character count score determination table included in the score determination table group 24.

[0051] Figure 15 shows an example of a character count score determination table. The character count score determination table 245 defines a score E corresponding to a range of character counts. The range of character counts is defined by what percentage of all questions the character count of one question falls into. The question text length evaluation unit 35 calculates the character count of the question text of an existing question based on the existing question data 22, identifies its relative position among all questions stored in the existing question data 22, and obtains a score E corresponding to the range of character counts by referring to the character count score determination table 245.

[0052] In the character count score determination table 245 in Figure 15, a score of 10 is assigned to the character count if it falls in the range of 80 percent to 100 percent, i.e., the top 20 percent with the most characters. A score of 8 is assigned to the character count if it falls in the range of 60 percent to less than 80 percent. A score of 6 is assigned to the character count if it falls in the range of 40 percent to less than 60 percent. A score of 4 is assigned to the character count if it falls in the range of 20 percent to less than 40 percent. A score of 2 is assigned to the character count if it falls in the range of 0 percent to less than 20 percent, i.e., the bottom 20 percent with the fewest characters.

[0053] For example, problems with a large number of characters tend to be more difficult because they require more consideration of the issues, while problems with fewer characters tend to be easier because they require fewer considerations. In the character count score determination table 245, a larger score E is assigned to problems with a larger number of characters, reflecting the tendency for problems to become more difficult as the number of issues to be considered increases with the number of characters.

[0054] Thus, as an example, a smaller score E indicates a higher difficulty level. The index value calculated by the problem text length evaluation unit 35 is stored as the index value E of the factor related to problem text length in the training data 25.

[0055] The correct answer rate calculation unit 36 ​​calculates the correct answer rate for each of the existing problems based on the existing problem answer data 21. The correct answer rate calculated by the correct answer rate calculation unit 36 ​​is stored as the correct answer rate R in the training data 25.

[0056] The learning data 25 is generated in the manner described above. Figure 16 shows an example of the learning data. The learning data 25 stores the intrinsic difficulty index values ​​A, B, C, D, and E and the correct answer rate R for each existing problem. The intrinsic difficulty index values ​​include the score A calculated by the question frequency evaluation unit 31, the score B calculated by the prerequisite knowledge evaluation unit 32, the score C calculated by the answer format evaluation unit 33, the score D calculated by the number of confusing options evaluation unit 34, and the score E calculated by the question text length evaluation unit 35. The correct answer rate R is the value calculated by the correct answer rate calculation unit 36. In the example shown in Figure 16, for problem 001, the index value A is 5, the index value B is 6, the index value C is 10, the index value D is 7.5, the index value E is 6, and the correct answer rate R is 32 percent. The same index values ​​and correct answer rates are stored for existing problems from problem 002 onwards.

[0057] The weight determination unit 37 determines weights a, b, c, d, and e, which represent the degree of influence of each index value A, B, C, D, and E on the accuracy rate R for multiple existing problems stored in the training data 25. Specifically, the weight determination unit 37 determines the weights a, b, c, d, and e such that the difference between the sum of the values ​​obtained by weighting the multiple index values ​​A, B, C, D, and E in the existing problems by the weights a, b, c, d, and e and the accuracy rate R is minimized. In other words, the weight determination unit 37 determines the weights a, b, c, d, and e such that the difference between the intrinsic difficulty level DIF, expressed by the formula A×a+B×b+C×c+D×d+E×e=DIF, and the accuracy rate R is minimized. The method for determining the weights is not particularly limited, but for example, it can be implemented using linear regression, least squares method, etc. The weights a, b, c, d, and e determined by the weight determination unit 37 are output as weight information 26. These weights a, b, c, d, and e are used by the estimation processing unit 40 to calculate the difficulty level of the target problem.

[0058] Referring to Figure 3, the estimation processing unit 40 includes a question frequency evaluation unit 41, a prerequisite knowledge evaluation unit 42, an answer format evaluation unit 43, a confusion option number evaluation unit 44, a question text length evaluation unit 45, and a difficulty calculation unit 46. The estimation processing unit 40 receives target question data 51, learning content data 23, a score determination table group 24, and weight information 26 as input. The target question data 51 includes the question text of the target question for which the difficulty level is to be calculated. The target question data 51 may also include an explanation of the target question. The estimation processing unit 40 calculates and outputs the difficulty level 52 of the target question.

[0059] The question frequency evaluation unit 41 calculates an index value for the frequency of questions for each unit for the target questions, based on the target question data 51 and the learning content data 23. The question frequency evaluation unit 41 calculates the index value in the same way as the question frequency evaluation unit 31 of the learning processing unit 30. That is, the question frequency evaluation unit 41 calculates how many times a unit of material that learners should learn and acquire in order to solve the target question has been asked in the past 10 actual exams. Then, the question frequency evaluation unit 41 refers to the question frequency score determination table 241 included in the score determination table group 24 and obtains a score A corresponding to the number of times the unit of the target question has been asked. The index value for question frequency calculated in this way is output to the difficulty calculation unit 46.

[0060] The prerequisite knowledge evaluation unit 42 calculates an index value for the breadth of prerequisite knowledge required to solve a target problem, based on the target problem data 51 and the learning content data 23. The prerequisite knowledge evaluation unit 42 calculates the index value in the same way as the prerequisite knowledge evaluation unit 32 of the learning processing unit 30. That is, the prerequisite knowledge evaluation unit 42 identifies the number of units of knowledge that the learner must learn and acquire in order to solve the target problem in the learning content. Then, the prerequisite knowledge evaluation unit 42 obtains a score B corresponding to the number of units by referring to the prerequisite knowledge score determination table 242 included in the score determination table group 24. The index value for the breadth of prerequisite knowledge calculated in this way is output to the difficulty calculation unit 46.

[0061] The answer format evaluation unit 43 calculates an index value for the format of the answer to a given problem based on the target problem data 51. The answer format evaluation unit 43 calculates the index value in the same way as the answer format evaluation unit 33 of the learning processing unit 30. That is, the answer format evaluation unit 43 analyzes the problem statement of the target problem to identify the answer format, and obtains a score C corresponding to the answer format by referring to the answer format score determination table 243 included in the score determination table group 24. The index value for the answer format calculated in this way is output to the difficulty calculation unit 46.

[0062] The Confusing Choice Evaluation Unit 44 calculates an index value for the number of choices in a multiple-choice question that are difficult to determine as correct or incorrect, based on the target question data 51. The Confusing Choice Evaluation Unit 44 calculates the index value in the same way as the Confusing Choice Evaluation Unit 34 of the Learning Processing Unit 30. That is, the Confusing Choice Evaluation Unit 44 identifies the number of choices in a multiple-choice question that are difficult to determine as correct or incorrect, and obtains a score D corresponding to the number of confusing choices by referring to the Confusing Choice Score Judgment Table 244 included in the Score Judgment Table Group 24. The index value for the number of confusing choices calculated in this way is output to the Difficulty Calculation Unit 46.

[0063] The problem sentence length evaluation unit 45 calculates an index value for the length of the sentence in a problem for a target problem based on the target problem data 51. The problem sentence length evaluation unit 45 calculates the index value in the same way as the problem sentence length evaluation unit 35 of the learning processing unit 30. That is, the problem sentence length evaluation unit 45 calculates the number of characters in the problem sentence of the target problem, identifies its relative position among all problems, and obtains a score E corresponding to the character count range by referring to the character count score determination table 245 included in the score determination table group 24. The index value for the problem sentence length calculated in this way is output to the difficulty calculation unit 46.

[0064] The difficulty level calculation unit 46 calculates the difficulty level 52 of the target question based on the score A calculated by the question frequency evaluation unit 41, the score B calculated by the prerequisite knowledge evaluation unit 42, the score C calculated by the answer format evaluation unit 43, the score D calculated by the number of confusing options evaluation unit 44, and the score E calculated by the question text length evaluation unit 45, along with the weights a, b, c, d, and e included in the weight information 26. Specifically, the difficulty level calculation unit 46 calculates the difficulty level of the target question by summing the values ​​obtained by weighting multiple indicator values ​​A, B, C, D, and E of the target question by the weights a, b, c, d, and e. That is, the difficulty level calculation unit 46 calculates the difficulty level 52 of the target question using the formula A×a+B×b+C×c+D×d+E×e. The difficulty level 52 calculated in this way is output from the estimation processing unit 40.

[0065] In this way, the estimation processing unit 40 calculates the difficulty level 52 of the target problem based on the content and format of the problem and the correct answer rate of existing problems for which there is a track record of solving. Therefore, even if the target problem is one for which there is no track record of solving, the difficulty level of the target problem can be accurately identified.

[0066] Figure 17 is a block diagram showing the configuration of a test creation system using a difficulty calculation device. The test creation system 70 includes a difficulty level calculation device 10, a control device 71, and a question generation device 72. The test creation system 70 operates based on the actions of a user 90 who intends to create test questions.

[0067] The control device 71 receives a request from the user 90 to generate test questions that specify difficulty requirements. This generation request includes requirements regarding the number, format, and difficulty level of questions to be included in the test. For example, the requirements may specify the number of questions to be included in the test, the number of questions for each answer format, and the number of questions for each difficulty range. Upon receiving the generation request, the control device 71 instructs the question generation device 72 to generate the test questions.

[0068] The problem generation device 72 generates the number and format of problems instructed by the user 90. The test problems generated by the problem generation device 72 are sent to the difficulty level calculation device 10 via the control device 71. The difficulty level calculation device 10 calculates the difficulty level of each problem generated by the problem generation device 72. The difficulty level calculated by the difficulty level calculation device 10 is sent to the control device 71.

[0069] The control device 71 determines whether the difficulty level calculated by the difficulty level calculation device 10 meets the difficulty level requirements instructed by the user 90. If the difficulty level of all questions meets the specified difficulty level requirements, the control device 71 presents the test questions to the user 90. On the other hand, if the difficulty level calculated by the difficulty level calculation device 10 does not meet the difficulty level requirements instructed by the user 90, the control device 71 instructs the question generation device 72 to regenerate the questions. At this time, all questions may be regenerated, or some questions may be regenerated. In this way, the control device 71 repeats the generation of questions by the question generation device 72 and the calculation of difficulty levels by the difficulty level calculation device 10 until questions that meet the difficulty level requirements are obtained.

[0070] In this way, the test creation system 70 generates questions of a difficulty level that meet the specified conditions, making it possible to efficiently create test questions with the desired difficulty level.

[0071] Furthermore, the difficulty calculation device 10 of this embodiment can also be realized by having a computer execute a software program that defines the processing procedures for each part. Figure 4 is a block diagram showing the hardware configuration of the computer that constitutes the difficulty calculation device.

[0072] The computer 60 comprises a processor 61, main memory 62, auxiliary storage 63, communication device 64, input device 65, and output device 66, which are interconnected via an internal bus 67.

[0073] The processor 61 is an arithmetic processing unit such as a CPU (Central Processing Unit), and functions as a part of the difficulty calculation device 10 according to this embodiment by executing a software program stored in the main memory 62 or auxiliary memory 63.

[0074] The main memory 62 is volatile memory such as RAM (Random Access Memory) and temporarily stores software programs executed by the processor 61 and the data necessary for their processing.

[0075] The auxiliary storage device 63 is a non-volatile memory such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or ROM (Read Only Memory), and stores software programs and various data semi-permanently. In the difficulty calculation device 10, the input data, the data used, and the generated data are recorded in the main memory device 62 and / or the auxiliary storage device 63.

[0076] The communication device 64 is a device that enables communication with external devices and services via a communication network (not shown). In this embodiment, for example, when the difficulty calculation device 10 uses a generative artificial intelligence model as an external service, it uses the generative artificial intelligence model through communication via the communication device 64.

[0077] The input device 65 is an input device such as a keyboard, mouse, and touch panel operated by the user 90. In this embodiment, when the user 90 directly operates the difficulty calculation device 10, information from the user 90 is input to the input device 65. The output device 66 is a display device that displays video, images, and text, such as a liquid crystal display. When the user 90 directly operates the difficulty calculation device 10, information is displayed to the user 90 on the output device 66. If the user 90 does not directly operate the difficulty calculation device 10, the input device 65 and output device 66 do not need to be provided.

[0078] Furthermore, the difficulty calculation device 10 may be configured not only as a single computer 60, but also as a distributed system in which multiple computers work in cooperation. In addition, part or all of the difficulty calculation device 10 may be built on the cloud.

[0079] In this embodiment, we have shown examples of using factors that have meaning that humans can understand as factors influencing the difficulty level, such as the frequency of questions per unit, the scope of prerequisite knowledge (number of units), the answer format, the number of ambiguous options, and the length of the question text (number of characters). However, the present invention is not limited to this. For example, multiple factors such as the frequency of questions, the length of the question text, and the scope of prerequisite knowledge may be correlated with each other. In that case, the contribution of a particular factor may be overestimated or underestimated in weight learning. As another example, the learning processing unit 30 may apply principal component analysis to the factors. Instead of the frequency of questions, the length of the question text, and the scope of prerequisite knowledge, the learning processing unit 30 may use principal component parameters whose correlations have been resolved by principal component analysis as factors, and the learning processing unit 30 and the estimation processing unit 40 may use these principal component parameters instead of the above-mentioned factors. In that case, the principal component parameters will be a mixture of multiple original factors, so their meaning will be difficult for humans to understand, and it will also be difficult to directly extract values ​​from the question text. Therefore, we can use the principal component transformation matrix to calculate the principal component parameters from the original factors extracted from the problem statement, etc.

[0080] The embodiments described above are illustrative for explaining the present invention and are not intended to limit the scope of the invention to those embodiments only. Those skilled in the art can implement the present invention in various other forms without departing from the spirit of the invention. Furthermore, the embodiments described above include the following items. However, the items included in these embodiments are not limited to those listed below.

[0081] (Item 1) The system includes: a storage unit that stores existing problem data, which includes the question texts of multiple existing problems; existing problem answer data, which includes the answers of multiple test takers to the existing problems; a first processing unit that calculates index values ​​for multiple factors influencing the difficulty of each of the existing problems based on the content and / or format of the question text, calculates the correct answer rate based on the existing problem answer data, and determines weights indicating the degree of influence of each of the multiple index values ​​on the correct answer rate; and a second processing unit that calculates index values ​​for multiple factors of a target problem based on target problem data, which includes the question text of the target problem whose difficulty level is to be evaluated, and calculates the difficulty level of the target problem based on the index values ​​and the weights. As a result, the difficulty level of the target problem is calculated based on the content and / or format of the problem and the correct answer rate of existing problems for which there is a record of answers, so even if the target problem is one for which there is no record of answers, the difficulty level of the target problem can be accurately identified.

[0082] (Item 2) In the difficulty level calculation device described above, the factors include the frequency of questions on each unit, which is a unit of material to be learned. Units that are frequently tested tend to have a high correct answer rate because learners have studied them well, while units that are rarely tested tend to have a low correct answer rate because learners have studied them very little. By using the frequency of questions on each unit as one of the factors, it becomes possible to accurately calculate the difficulty level while taking the above trends into account.

[0083] (Item 3) In the difficulty level calculation device described above, the factors include the breadth of prerequisite knowledge required to solve the problem. When solving a problem, if the question is taken from a single unit, one only needs to recall that unit and apply it to each option. On the other hand, if, for example, each option is taken from multiple different units, one needs to recall a wide range of knowledge to solve the problem, which increases the difficulty level. Therefore, by including the breadth of prerequisite knowledge required to solve the problem as one of the factors, it becomes possible to accurately calculate the difficulty level while taking the above tendencies into account.

[0084] (Item 4) In the difficulty level calculation device described above, the first processing unit and the second processing unit define the number of units, which are units of learning that include the knowledge necessary to arrive at the correct answer, as the breadth of the prerequisite knowledge required to solve the problem.

[0085] (Item 5) In the difficulty level calculation device described above, the factors include the format of the answer to the question. When an exam includes questions with multiple answer formats, the difficulty level of the question tends to change depending on the answer format. By making this answer format one of the factors, it becomes possible to accurately calculate the difficulty level while taking the above-mentioned trend into account.

[0086] (Item 6) In the difficulty level calculation device described above, the first processing unit determines whether the existing problem is a multiple-choice question in which one correct answer is selected from multiple options, or, if it is a multiple-choice question, whether each option of the existing problem is a knowledge question in which knowledge about a specific matter is required to be answered, a case question in which an answer is required to be given to a specific case, a case law question in which an answer is required to be given based on an understanding of precedents, a combination question in which multiple options are included in which a combination of options that meet predetermined conditions is required to be selected, or the problem statement has multiple blanks and multiple phrases including the words that correspond to the blanks are presented separately from the problem statement, and The system determines whether a question is a multiple-choice question requiring the selection of a word or phrase to fill in the blank from a set of words, a calculation problem requiring the answer to be the result of a numerical calculation based on the problem statement, a counting problem requiring the selection of the number of statements that meet a predetermined condition from a set of statements, or a descriptive question requiring the description of a predetermined matter based on the problem statement. The values ​​associated with each of the following types of questions—knowledge questions in the multiple-choice format, case studies in the multiple-choice format, precedents in the multiple-choice format, combination problems, multiple-choice questions, calculation problems, counting problems, and descriptive questions—are used as indicator values ​​for the format of the answer to the question.

[0087] (Item 7) In the difficulty level calculation device described above, the factors include the number of options in a multiple-choice question that are not easy to determine as correct or incorrect. In multiple-choice questions, the probability of obtaining the correct answer by elimination tends to change depending on the number of options that are easy to determine as correct or incorrect and those that are not. Therefore, by including the number of options that are not easy to determine as correct or incorrect as one of the factors, it becomes possible to accurately calculate the difficulty level that takes the above tendency into account. (Item 8)

[0088] In the difficulty calculation device described above, the length of the problem statement is included as one of the factors. For example, problems with a large number of characters tend to be more difficult because there are more points to consider, while problems with a small number of characters tend to be easier because there are fewer points to consider. Therefore, by including the length of the problem statement as one of the factors, it becomes possible to accurately calculate the difficulty level while taking the above-mentioned tendencies into account.

[0089] (Item 9) In the difficulty level calculation device described above, the first processing unit determines the weights such that the difference between the sum of the values ​​obtained by weighting multiple indicator values ​​in the existing problem by the weights and the correct answer rate is minimized, and the second processing unit calculates the difficulty level of the target problem as the sum of the values ​​obtained by weighting multiple indicator values ​​in the target problem by the weights.

[0090] (Item 10) A method for a computer to store existing problem data, which includes the problem statements of multiple existing problems, and existing problem answer data, which includes the answers of multiple test takers to the existing problems; calculate index values ​​for each of the existing problems, based on the content and / or format of the problem statement, for multiple factors that influence the difficulty of the problem; calculate the correct answer rate based on the existing problem answer data; determine weights indicating the degree of influence of each of the multiple index values ​​on the correct answer rate; calculate index values ​​for multiple factors of a target problem, based on target problem data, which includes the problem statement of the target problem whose difficulty is to be evaluated; and calculate the difficulty of the target problem based on the index values ​​and the weights.

[0091] (Item 11) A test creation system comprising: a difficulty level calculation device as described above; a question generation device that generates test questions; and a control device that receives a request to generate test questions specifying difficulty level requirements, causes the question generation device to generate test questions, causes the difficulty level calculation device to calculate the difficulty level of the generated test questions, determines whether the calculated difficulty level satisfies the specified requirements, causes the question generation device to regenerate test questions if the calculated difficulty level does not satisfy the requirements, and outputs the test questions if the calculated difficulty level satisfies the requirements. [Explanation of Symbols]

[0092] 10...Difficulty Calculation Device 20...Data storage unit 21… Existing problem answer data 22…Existing problem data 23…Learning content data 30…Learning Processing Unit 31…Frequency Evaluation Section 32…Prerequisite Knowledge Evaluation Department 33…Answer Format Evaluation Department 34… Department for evaluating the number of vagrant limbs 35… Question text length evaluation section 36... Correct answer rate calculation unit 37…Decision Section 40…Estimation Processing Unit 41…Frequency Evaluation Section 42…Prerequisite Knowledge Evaluation Department 43…Answer Format Evaluation Department 44… Department for evaluating the number of vagrant limbs 45… Question text length evaluation section 46...Difficulty Calculation Section 60… Computer 61… Processor 62…Main memory 63…Auxiliary storage device 64...Communication device 65...Input device 66…Output device 67...Internal bus 70… Exam creation system 71...Control device 72…Problem generation device

Claims

1. A storage unit that stores existing problem data including the question texts of multiple existing problems, existing problem answer data relating to the answers of multiple test takers to the said existing problems, and score determination information that defines index values ​​for multiple factors that affect the difficulty of the problem in relation to the content and / or format of the question text, A first processing unit calculates index values ​​for multiple factors based on the content and / or format of the question text of each of the aforementioned existing problems by referring to the score judgment information, calculates the correct answer rate based on the answer data of the aforementioned existing problems, and determines the weights that indicate the degree of influence of each of the multiple index values ​​on the correct answer rate. A second processing unit calculates the difficulty level of a target problem based on target problem data including the problem statement of the target problem, by referring to the score determination information, and by referring to the score determination information, and calculates the difficulty level of the target problem based on the index values ​​and the weights. A difficulty level calculation device having the following features.

2. The aforementioned factors include the frequency of questions related to each unit, which is a coherent unit of material to be learned. The difficulty level calculation device according to claim 1.

3. The aforementioned factors include the breadth of prerequisite knowledge required to solve the problem. The difficulty level calculation device according to claim 1.

4. The first and second processing units define the number of units, which are coherent units of learning that include the knowledge necessary to arrive at the correct answer, as the breadth of the prerequisite knowledge required to solve the problem. The difficulty level calculation device according to claim 3.

5. The aforementioned factors include the format of the answer to the problem. The difficulty level calculation device according to claim 1.

6. The first processing unit is, Whether the aforementioned existing question is a multiple-choice question in which one correct answer is selected from multiple options, or if it is a multiple-choice question, whether each option in the existing question is a knowledge question in which knowledge about a specific matter is required, a case question in which answers are required to respond to a specific case, or a case law question in which answers are required to respond based on an understanding of precedents, Is it a combination problem that includes multiple options and asks the user to select the combination of options that meet the given conditions? Is it a multiple-choice question where the question statement contains multiple blanks, and a separate set of phrases containing words that fit into the blanks is presented, and the questioner is asked to select the word that fits into the blank from among the set of phrases? Is it a calculation problem that requires the answer to be the result of a numerical calculation based on the problem statement? Is this a counting problem that asks the user to select the number of statements that meet a given condition, including multiple statements? Determine whether it is a written question that requires you to describe specific matters based on the problem statement. The values ​​associated with each of the following types of questions—knowledge questions in the multiple-choice format, case studies in the multiple-choice format, precedents in the multiple-choice format, combination questions, selection questions, calculation questions, counting questions, and descriptive questions—are used as indicator values ​​for the format of the answers to the aforementioned questions. The difficulty level calculation device according to claim 5.

7. The aforementioned factors include the number of options in a multiple-choice question where it is not easy to determine whether they are correct or incorrect. The difficulty level calculation device according to claim 1.

8. The aforementioned factors include the length of the sentence in question. The difficulty level calculation device according to claim 1.

9. The first processing unit determines the weights such that the difference between the sum of the values ​​obtained by weighting the multiple index values ​​in the existing problem by the weights and the correct answer rate is minimized. The second processing unit calculates the difficulty level of the target problem by summing the values ​​obtained by weighting multiple index values ​​of the target problem according to the weights. The difficulty level calculation device according to claim 1.

10. Computers The system stores existing problem data, which includes the question texts of multiple existing problems; existing problem answer data, which includes the answers of multiple test takers to the said existing problems; and score determination information, which defines index values ​​for multiple factors that influence the difficulty of a problem in relation to the content and / or format of the question text. For each of the aforementioned existing problems, by referring to the score determination information, index values ​​for multiple factors based on the content and / or format of the problem statement of the aforementioned existing problem are calculated, the correct answer rate is calculated based on the answer data of the aforementioned existing problem, and weights indicating the degree of influence of each of the multiple index values ​​on the correct answer rate are determined. Based on the target problem data, which includes the problem statement of the target problem to be evaluated for difficulty, index values ​​for multiple factors of the target problem are calculated by referring to the score determination information, and the difficulty of the target problem is calculated based on the index values ​​and the weights. A method for calculating the difficulty of carrying out a task.

11. The difficulty level calculation device according to claim 1, A question generation device that generates test questions, A control device that receives a request to generate test questions with specified difficulty requirements, causes the question generation device to generate test questions, causes the difficulty level of the generated test questions to be calculated by the difficulty level calculation device, determines whether the calculated difficulty level satisfies the specified requirements, causes the question generation device to regenerate the test questions if the calculated difficulty level does not satisfy the requirements, and outputs the test questions if the calculated difficulty level satisfies the requirements. A test creation system that has the following features.

Citation Information

Patent Citations

  • Teaching support system

    JP1993281899A

  • Method and system for performing simulated examination while utilizing communication network

    JP2002072857A

  • Learning program and learning system

    JP2005134484A

  • Capability estimation system, method, program and recording medium

    JP2008242637A

  • Learning system

    JP2009086203A