Sentence grading device, sentence grading method, and sentence grading program
The essay scoring device addresses the issue of score variations by applying multiple score prompts to different evaluation regions within the essay scoring device, resulting in a standardized and reliable scoring process.
Patent Information
- Application Number
- JP2023212901
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2043-12-18
AI Technical Summary
Existing essay scoring methods struggle to eliminate variations in scores due to different focuses emphasized by scorers, leading to inconsistent evaluations.
An essay scoring device that inputs essays, applies multiple score prompts to various evaluation regions, calculates evaluation region scores, and computes an overall score to standardize scoring and reduce bias.
The device effectively eliminates variations in scores by standardizing the evaluation process across different scorers, enhancing the reliability of overall scores.
Smart Images

Figure 2025096910000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an essay scoring device, an essay scoring method, and an essay scoring program for scoring essay answers created by examinees such as mock tests for entrance examinations.
Background Art
[0002] An ability diagnosis system is known that obtains levels such as "knowledge and skills", "thinking ability, judgment ability, and expression ability", and "creativity" in the form of scores for each thinking code number and presents them to examinees (Patent Document 1, abstract).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] It is desired to score not only problems with a single solution but also essays such as compositions and short essays by focusing on levels such as "knowledge and skills", "thinking ability, judgment ability, and expression ability", and "creativity". On the other hand, when a person scores an essay, it is impossible to completely eliminate potential biases such as different focuses emphasized by the scorers, and the scores may vary depending on the scorers.
[0005] In view of the above circumstances, an object of the present disclosure is to eliminate variations in the focuses emphasized by scorers and eliminate variations in scores by scorers caused by variations in focuses when scoring essays.
Means for Solving the Problems
[0006] An essay scoring device according to one aspect of the present disclosure is an essay input unit that inputs an essay to be scored, A score prompt for individually calculating the scores of one or more evaluation regions among a plurality of evaluation regions, wherein the one or more evaluation regions for which a plurality of score prompts calculate scores are different for each score prompt, a score prompt execution unit that executes each of the plurality of score prompts on the text, an evaluation region score calculation unit that calculates an evaluation region score, which is the score of each evaluation region, based on the plurality of scores of the plurality of evaluation regions calculated by executing each of the plurality of score prompts on the text, and an overall score calculation unit that calculates an overall score of the text based on the evaluation region scores of each of the plurality of evaluation regions. It is provided with.
Effect of the Invention
[0007] According to the present disclosure, when scoring a text, it is possible to eliminate the variation in the points of interest that scorers consider important and the variation in scores by scorers caused by the variation in points of interest.
[0008] Note that the effects described here are not necessarily limited, and any of the effects described in the present disclosure may be applicable.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Mode for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0011] 1. Configuration of the Essay Scoring Device
[0012] FIG. 1 shows the configuration of an essay scoring device according to an embodiment of the present disclosure.
[0013] An essay 10 to be scored, which is created by a test taker (essay writer) such as in a mock test for an entrance examination, is input into the essay scoring device 1. The essay scoring device 1 scores the input essay 10 to be scored. The essay scoring device 1 outputs the scoring result to the electronic device 20. The electronic device 20 displays the received scoring result on a built-in or external display of the electronic device 20, or prints the scoring result on a printing medium (typically paper).
[0014] The essay 10 to be scored is, for example, a relatively long essay such as a composition or a short essay. The essay 10 to be scored is typically text data obtained by performing OCR (optical character recognition) on scanned data of an answer handwritten by a test taker (essay writer). Alternatively, the essay 10 to be scored may be text data created using the word processor function of a computer.
[0015] The electronic device 20 may be any one or more of a terminal device (personal computer, tablet computer, smartphone, etc.) of a test taker (essay writer) taking a mock test for an entrance examination, a terminal device (personal computer) of a test institution such as a cram school where the test taker (essay writer) belongs, or a school. Further, the electronic device 20 may be a server device on the side of the mock test organizer, a printer for printing the scoring result, a computer connected to the printer, or the like. An electronic device serving as an interface may be interposed between the essay scoring device 1 and the electronic device 20 (not shown).
[0016] The essay scoring device 1 operates as an essay input unit 110, a score prompt execution unit 120, an evaluation area score calculation unit 130, an overall score calculation unit 140, a pass index calculation unit 150, an extended evaluation area determination unit 160, a comment prompt execution unit 170, and a GUI generation unit 180 by loading and executing an essay scoring program recorded in the ROM into the RAM by a CPU constituting the computer.
[0017] 2. Thinking Code
[0018] Figure 2 shows the thinking code.
[0019] The text scoring device 1 scores the text using the thinking code 200. The thinking code 200 includes a plurality (nine in this example) of evaluation regions A1, A2, A3, B1, B2, B3, C1, C2, and C3. The text scoring device 1 calculates the overall score of the text based on the scores of the plurality of evaluation regions A1, A2, A3, B1, B2, B3, C1, C2, and C3.
[0020] The plurality of evaluation regions A1, A2, A3, B1, B2, B3, C1, C2, and C3 are in a 3x3 matrix form constituted by a first axis 211 which is the horizontal axis and a second axis 212 which intersects (orthogonally in this example) the first axis 211. The first axis 211 represents the level of thinking (three levels in this example). The second axis 212 represents the level of knowledge (three levels in this example).
[0021] Taking the evaluation region A1 as the origin, the first axis 211 indicates that the thinking level increases in three levels as it moves away from the evaluation region A1 (in the order of moving to the evaluation regions B1 and C1). Specifically, "A" represents the perspective of "Do you have knowledge and can you understand?", "B" represents the perspective of "Can you apply it and think and express logically?", and "C" represents the perspective of "Can you think and express critically and creatively?".
[0022] Taking the evaluation region A1 as the origin, the second axis 212 indicates that the knowledge level increases in three levels as it moves away from the evaluation region A1 (in the order of moving to the evaluation regions A2 and A3). Specifically, "1" represents the scale of "simple knowledge level problems", "2" represents the scale of "complex knowledge level problems", and "3" represents the scale of "transformative (higher-order) knowledge level problems".
[0023] Evaluation area A1 represents the ability of "having knowledge of or understanding simple knowledge-level questions", specifically regarding general scoring criteria. Evaluation area A2 represents the ability of "having knowledge of or understanding complex knowledge-level questions", specifically regarding the clarity of claims. Evaluation area A3 represents the ability of "having knowledge of or understanding transformative (higher-order) knowledge-level questions", specifically regarding the interestingness of ideas.
[0024] Evaluation area B1 represents "being able to apply in simple knowledge-level questions and think and express logically", specifically regarding specific examples, data, and episodes. Evaluation area B2 represents "being able to apply in complex knowledge-level questions and think and express logically", specifically regarding the connection between claims and reasons, counterarguments to claims, and the reasons for persuading against those counterarguments. Evaluation area B3 represents "being able to apply in transformative (higher-order) knowledge-level questions and think and express logically", specifically regarding multi-faceted thinking and the degree of utilization of the five perspectives (the perspective of a bird, an insect, a fish, a bat, and a heart).
[0025] Evaluation area C1 represents "being able to think and express critically and creatively in simple knowledge-level questions", specifically regarding specifying the conditions of how important it is for whom. Evaluation area C2 represents "being able to think and express critically and creatively in complex knowledge-level questions", specifically regarding the historical background and story that support the reasons. Evaluation area C3 represents "being able to think and express critically and creatively in transformative (higher-order) knowledge-level questions", specifically regarding original ideas.
[0026] 3. Score Prompt
[0027] A description is given of a plurality (five in this example) of score prompts 121-125 executed by the score prompt execution unit 120. The score prompts 121-125 are each a machine learning model. That is, when the text 10 to be scored is input to each of the score prompts 121-125, scores for the text 10 are independently output from each of the score prompts 121-125.
[0028] The score prompts 121-125 include the Toulmin model prompt 121, the Goodman model prompt 122, the fifth model prompt 123, the Bloom taxonomy model prompt 124, and the general scoring criteria prompt 125.
[0029] The score prompts 121-125 each individually calculate scores for one or more of the evaluation areas A1, A2, A3, B1, B2, B3, C1, C2, C3. The one or more evaluation areas for which the score prompts 121-125 are to calculate scores differ for each score prompt. For example, as will be described later, the Toulmin model prompt 121 calculates scores for the evaluation areas A2, A3, B2, B3, C1 respectively. On the other hand, the Goodman model prompt 122 calculates scores for the evaluation areas A2, B1, B2, B3, C1, C3 respectively. The evaluation areas A2, A3, B2, B3, C1 that are the calculation targets of the Toulmin model prompt 121 are different from the evaluation areas A2, B1, B2, B3, C1, C3 that are the calculation targets of the Goodman model prompt 122.
[0030] The Toulmin model prompt 121 is a prompt based on the Toulmin model. The Toulmin model is a stable model that is used in schools for exploration and essays. It is suitable for analyzing whether items such as claims, grounds, warrants, rebuttals, backing, and qualifiers are written and raising points for improvement. Therefore, the Toulmin model prompt 121 is suitable for calculating the scores of evaluation areas A2, A3, B2, B3, and C1 respectively. The Toulmin model prompt 121 is, for example, a prompt like "Please analyze the following passage from the six perspectives of the Toulmin model (claim, ground, warrant, qualifier, rebuttal, backing)."
[0031] The Goodman model prompt 122 is a prompt based on the Goodman model. The Goodman model is also said to be a sense that can be used when considering the mathematics of national universities with difficult admissions from the perspective of rational thinking in the integration of liberal arts and science. It is a conceptual lens for generating ideas of art thinking and design thinking. Therefore, the Goodman model prompt 122 is suitable for calculating the scores of evaluation areas A2, B1, B2, B3, C1, and C3 respectively. The Goodman model prompt 122 is, for example, a prompt like "Please analyze the following passage from the following five perspectives. 1) Is the claim clearly written? 2) Is the reason for supporting the claim written? 3) Are specific examples for supporting the claim written? 4) Does it show a way of thinking that is contrary to the claim and prove that it is less persuasive compared to the claim? 5) Is the content of the claim original?"
[0032] The five - eye model prompt 123 is a prompt based on "observing with the five eyes", which is advocated by the "Observe" in the first "O" of the OODA thinking management model. It can check compound - eye thinking. Therefore, the five - eye model prompt 123 is suitable for calculating the scores of the evaluation areas C1 and C2 respectively. The five - eye model prompt 123 is, for example, a prompt like "Please analyze the following text with the following five eyes: 1) the eyes of a bird, 2) the eyes of an insect, 3) the eyes of a bat, 4) the eyes of a fish, 5) the eyes of the heart. The definitions of each eye are: 1) the eyes of a bird: the eyes that can see new solutions by grasping things from an overlooking and overall perspective like the eyes of a bird flying in the sky; 2) the eyes of an insect: the eyes when clearly defining what to do and concentrating on the things in front; 3) the eyes of a fish: the perspective of reading the trend, so to speak, considering time; 4) the eyes of a bat: the eyes of looking at things upside - down; 5) the eyes of the heart: the eyes of seeing the essence of things."
[0033] The Bloom's Taxonomy model prompt 124 is a prompt based on the Bloom's Taxonomy model. The Bloom's Taxonomy model classifies the process by which learners acquire knowledge and skills into six categories (knowledge, comprehension, application, analysis, synthesis, evaluation), and hierarchically captures the process of knowledge acquisition and thinking. Therefore, the Bloom's Taxonomy model prompt 124 is suitable for calculating the scores of the evaluation areas A2, B2, B3, C2, and C3 respectively. The Bloom's Taxonomy model prompt 124 is, for example, a prompt like "Please analyze the following text with the six cognitive areas of Bloom's Taxonomy: 1) knowledge, 2) comprehension, 3) application, 4) analysis, 5) evaluation, 6) creation."
[0034] The general scoring criterion prompt 125 is a prompt based on general scoring criteria. The general scoring criteria are whether the precautions for writing a composition are followed and whether the writing rules such as correct spelling and no omission of characters are appropriately utilized. Therefore, the general scoring criterion prompt 125 is suitable for calculating the score in evaluation area A1. The general scoring criterion prompt 125 is, for example, a prompt such as "Please analyze the following passage from the following four perspectives: 1) Whether the meaning is clear, 2) Whether there are spelling mistakes or omissions, 3) Whether the style and terms are appropriate, 4) Whether it is written from the perspective of the reader."
[0035] 4. Operation Flow of the Passage Scoring Device
[0036] Figure 3 shows the operation flow of the passage scoring device. Figure 4 is a diagram for explaining the score calculation method.
[0037] The passage input unit 110 inputs the passage to be scored (step S1).
[0038] The score prompt execution unit 120 executes a plurality (five in this example) of score prompts 121 - 125 for the passage respectively (step S2). The Toulmin model prompt 121 calculates "100, 20, 60, 100, 100" as the scores in evaluation areas A2, A3, B2, B3, and C1 respectively. The Goodman model prompt 122 calculates "100, 100, 60, 100, 20, 80" as the scores in evaluation areas A2, B1, B2, B3, C1, and C3 respectively. The fifth model prompt 123 calculates "30, 70" as the scores in evaluation areas C1 and C2 respectively. The Bloom's taxonomy model prompt 124 calculates "100, 100, 100, 100, 100" as the scores in evaluation areas A2, B2, B3, C2, and C3 respectively. The general scoring criterion prompt 125 calculates "70" as the score in evaluation area A1.
[0039] That is, for each evaluation area A1 - C3, the score prompt execution unit 120 individually calculates N scores (where N is an integer and N ≥ 1) using N score prompts 121 - 125. The value of N for at least some of the evaluation areas is different. For example, for evaluation area A1, the score prompt execution unit 120 calculates N = 1 score (70) using N = 1 score prompt (general scoring criteria prompt 125). For evaluation area C2, the score prompt execution unit 120 calculates N = 2 scores (70, 100) using N = 2 score prompts (the fifth model prompt 123, Bloom's taxonomy model prompt 124).
[0040] Based on the multiple scores of the multiple evaluation areas A1, A2, A3, B1, B2, B3, C1, C2, C3 calculated by executing the multiple score prompts 121 - 125 on the text respectively, the evaluation area score calculation unit 130 calculates the score of each evaluation area (evaluation area score) (step S3). Specifically, the evaluation area score calculation unit 130 calculates (the sum of the N scores of a certain evaluation area) / N as the evaluation area score of this evaluation area.
[0041] For example, the evaluation area score calculation unit 130 calculates 70 / 1 = 70 as the evaluation area score a1 of evaluation area A1 (N = 1). The evaluation area score calculation unit 130 calculates (100 + 100 + 100) / 3 = 100 as the evaluation area score a2 of evaluation area A2 (N = 3). The evaluation area score calculation unit 130 calculates 20 / 1 = 20 as the evaluation area score a3 of evaluation area A3 (N = 1).
[0042] The evaluation area score calculation unit 130 calculates 100 / 1 = 100 as the evaluation area score b1 of evaluation area B1 (N = 1). The evaluation area score calculation unit 130 calculates (60 + 60 + 100) / 3 = 73 as the evaluation area score b2 of evaluation area B2 (N = 3). The evaluation area score calculation unit 130 calculates (100 + 100 + 100) / 3 = 100 as the evaluation area score b3 of evaluation area B3 (N = 3).
[0043] The evaluation area score calculation unit 130 calculates (100 + 20 + 30) / 3 = 50 as the evaluation area score c1 of the evaluation area C1 (N = 3). The evaluation area score calculation unit 130 calculates (70 + 100) / 2 = 85 as the evaluation area score c2 of the evaluation area C2 (N = 2). The evaluation area score calculation unit 130 calculates (80 + 100) / 2 = 90 as the evaluation area score c3 of the evaluation area C3 (N = 2). In this example, the full score of the evaluation area score is 100.
[0044] Here, for at least one of the first axis 211 and the second axis 212, the total value of N of each of the plurality of evaluation areas A1, A2, A3, B1, B2, B3, C1, C2, C3 at the lowest level is preferably smaller than the total value of N of each of the plurality of evaluation areas A1, A2, A3, B1, B2, B3, C1, C2, C3 at the highest level.
[0045] In this example, for the first axis 211 (horizontal axis), the total value of N of the evaluation areas A1, B1, C1 at the lowest level is 1 + 1 + 3 = 5. The total value of N of the evaluation areas A3, B3, C3 at the highest level is 1 + 3 + 2 = 6. Comparing the total values of N at the lowest level and the highest level, 5 < 6. Compared with the evaluation areas A1, B1, C1 at the lowest level, for the evaluation areas A3, B3, C3 at the highest level, the evaluation area scores can be calculated based on a larger number of score prompts 121 - 125, and the reliability is high. On the other hand, the evaluation areas A1, B1, C1 at the lowest level have lower complexity and lower scoring difficulty compared to the evaluation areas A3, B3, C3 at the highest level. Therefore, even if the evaluation area scores are calculated based on a small number of score prompts 121 - 125, the possibility of a decrease in reliability is relatively low.
[0046] In this example, for the second axis 212 (vertical axis), the total value of N in the evaluation regions A1, A2, and A3 at the lowest level is 1 + 3 + 1 = 5. The total value of N in the evaluation regions C1, C2, and C3 at the highest level is 3 + 2 + 2 = 7. Comparing the total values of N at the lowest and highest levels, 5 < 7. Compared with the evaluation regions A1, A2, and A3 at the lowest level, for the evaluation regions C1, C2, and C3 at the highest level, the level can calculate the evaluation region score based on a larger number of score prompts 121 - 125, and the reliability is high. On the other hand, the evaluation regions A1, A2, and A3 at the lowest level have lower complexity and lower difficulty in scoring compared to the evaluation regions C1, C2, and C3 at the highest level. Therefore, even if the evaluation region score is calculated based on a small number of score prompts 121 - 125, the possibility of a decrease in reliability is relatively low.
[0047] The overall score calculation unit 140 calculates the overall score of the text based on the evaluation region scores a1, a2, a3, b1, b2, b3, c1, c2, and c3 of the plurality of evaluation regions A1, A2, A3, B1, B2, B3, C1, C2, and C3 (step S4). In this example, the full score of the overall score is 100, which is equal to the full score of each evaluation region score. Specifically, the overall score calculation unit 140 calculates (the sum of the evaluation region scores of all evaluation regions) / (the total number of evaluation regions) as the overall score S of the text. In this example, the overall score calculation unit 140 calculates 76 as the overall score S of the text, where (70 + 100 + 20 + 100 + 73 + 100 + 50 + 85 + 90) / 9 = 688 / 9 = 76. In other words, the evaluation region score is calculated as the average value of the scores calculated by N score prompts respectively, and further, the overall score is calculated as the average value of the 9 evaluation region scores. So to speak, the overall score is calculated by the "average of averages".
[0048] If the overall score is simply calculated as the average value of all values without depending on the evaluation area scores, then (70 + 100 + 100 + 100 + 20 + 100 + 60 + 60 + 100 + 100 + 100 + 100 + 100 + 20 + 30 + 70 + 100 + 80 + 100) / 19 = 1510 / 19 = 79.4 ≈ 79. For example, in the case of 79.5, if rounding off the decimal part to 80, the final evaluation will change from B to A. It is necessary to consider the discrepancy with the previous B evaluation of 76. Therefore, in this embodiment, by calculating the evaluation area scores and obtaining the average thereof as the overall score, the error will be reduced.
[0049] When the average value of the evaluation area scores of a plurality of evaluation areas A1 - C3 of a plurality of passers of a certain test is X, the average value of the evaluation area scores of a plurality of evaluation areas A1 - C3 of a plurality of failers of the test is Y (X > Y), and the evaluation area scores of a plurality of evaluation areas A1 - C3 of the writer is Z, the pass index calculation unit 150 calculates the pass index of the writer for the test related to the test based on (Y - Z) / (X - Y) (step S5). That ratio becomes the index for pass / fail of a certain test. The essay scoring device 1 may similarly calculate for examinees of tests in other schools and compare the pass / fail indexes to discriminate the school pass characteristics of examinees or recommend schools for the examinees.
[0050] When the average value of the overall scores of a plurality of passers of a certain test is x, the average value of the evaluation area scores of the overall scores of a plurality of failers of the test is y (x > y), and the overall score of the examinee (writer) is z, the pass index calculation unit 150 may calculate the pass index of the writer for the test related to the test based on (y - z) / (x - y). The pass rate standard may be scored by stipulating the school - specific scores based on the essay thinking codes of existing passers.
[0051] The elongation evaluation area determination unit 160 determines which of the plurality of evaluation areas A1 - C3 with an evaluation area score lower than the threshold should be the subject of elongation effort based on the evaluation area scores of the evaluation areas A1 - C3 around the plurality of evaluation areas A1 - C3 with an evaluation area score lower than the threshold (step S6). For example, when there are a plurality of evaluation areas A1 - C3 with a low evaluation area score, the elongation evaluation area determination unit 160 may determine that any one of the evaluation areas A1 - C3 with a relatively high evaluation area score among the surrounding evaluation areas A1 - C3 should be the subject of elongation effort, assuming that the score is likely to increase (the growth rate is high).
[0052] The comment prompt execution unit 170 executes the comment prompt 171 on the text 10 to be scored. The comment prompt 171 generates a comment on the text 10 to be scored (step S7). The comment is a comment regarding ethics and futurity. The comment prompt execution unit 170 is, for example, a prompt such as "Please analyze the following text from the two perspectives of ethics and futurity." Comments regarding ethics and futurity are not the basis for score calculation, but may be useful when providing feedback on the comments. Although schools each have an educational philosophy and a spirit of founding a school, since they are common regarding ethics and futurity, it is considered preferable not to deviate from ethics and futurity.
[0053] Figure 5 shows an example of the GUI.
[0054] The GUI generation unit 180 generates the GUI 300 and outputs it to the electronic device 20 (step S8). The electronic device 20 can display the GUI 300 on a built-in or external display via a web application or the like.
[0055] The GUI 300 displays the evaluation area score group 310, the comprehensive evaluation 320, the passing criterion 330, the elongation evaluation area 340, and the comment 350. The layout of each element 310 - 350 is not limited to the example in the figure.
[0056] The evaluation area score group 310 displays evaluation area scores and the like for each of a plurality of 3x3 matrix-shaped evaluation areas A1-C3. For example, within the quadrilateral representing evaluation area A2, the evaluation area score "100", the content of evaluation area A2 "clarity of claim", and the class "A" when the evaluation area score is classified as A-E (highest - lowest) are displayed. The same applies to other evaluation areas. Symbols A1-C3 indicating the evaluation areas may be displayed within each quadrilateral. The arrangement of the evaluation area score group 310 (the arrangement of the plurality of 3x3 matrix-shaped evaluation areas A1-C3) is the same as the arrangement of the thinking code 200. As a modification, the thinking code 200 itself may be used as part of the GUI. That is, within the quadrilaterals of evaluation areas A1-C3 of the thinking code 200, the evaluation area score, the content of the evaluation area "clarity of claim", and classes A-E may be described.
[0057] The comprehensive evaluation 320 displays at least the overall score "76". The comprehensive evaluation 320 may further display the class "Evaluation B" when the overall score is classified as Evaluation A - Evaluation E (highest - lowest). Thereby, the comprehensive writing ability of the examinee can be evaluated, and the feedback of the comment 350 can be automated in a web application or the like.
[0058] 5. Conclusion
[0059] When a person grades an article such as a composition or a short paper, there is a possibility that the scores will vary among graders because potential biases such as different focuses emphasized by the graders cannot be completely eliminated. For example, there may be a bias where a certain grader intentionally or potentially emphasizes evaluation area A3 "Does the person have knowledge of and can understand problems at a transformative (higher-order) knowledge level?", and another grader intentionally or potentially emphasizes evaluation area C3 "Can the person think critically and express themselves or think creatively and express themselves in problems at a transformative (higher-order) knowledge level?". If there is a bias in the emphasized evaluation areas, as a result, the scores given by the graders may vary. That is, when grading the same article, it is impossible to completely eliminate the possibility that the overall scores for the same article will be different between a grader who emphasizes evaluation area A3 and a grader who emphasizes evaluation area C3.
[0060] In contrast, according to the present embodiment, for each evaluation region A1 - C3, N scores (where N is an integer greater than or equal to 1) are individually calculated using N score prompts 121 - 125. Then, based on the N scores calculated for a certain evaluation region, the evaluation region score of that evaluation region is calculated. As a result, since the evaluation region scores of each evaluation region are calculated from one or more viewpoints, the reliability of the evaluation region scores of each evaluation region is high. Also, evaluation region scores are calculated for all the evaluation regions A1 - C3, and an overall score is calculated based on all (9) of the evaluation region scores. This can prevent the reliability of the overall score from decreasing due to the result that the emphasis is biased towards some evaluation regions and the overall score varies.
[0061] As described above, according to the present embodiment, when grading an article, it is possible to eliminate the variation in viewpoints that scorers value, and eliminate the variation in scores (overall scores) by scorers caused by the variation in viewpoints. As a result, the reliability of the overall score of the article is high.
[0062] 6. Application of the Present Embodiment
[0063] Nearest - developed region to a model article for passing: The article grading device 1 may re - grade what the examinee has rewritten based on the feedback of the comment 350, and further propose a model article that takes into account the intention of the examinee on the web and in the app.
[0064] Direction of strategic learning for passing: The article grading device 1 may, at the same time as re - grading, timely determine whether the passing rate for the target school has increased. Thereby, it is possible to form an increase in motivation through realistic reflection.
[0065] Visualizing the potential of writing thinking and expression ability after enrollment: The writing scoring device 1 can determine the degree of effort to improve thinking and expression ability after enrollment by comparing the ideal score of the writing thinking code (100 for all 9 regions) with the score of the examinee's writing thinking code (average of 9 regions). Ideal score ÷ examinee's score. For example, 100 ÷ 100 = 1, 100 ÷ 50 = 2, so examinees with a score of 50 can know that they need to work twice as hard. Moreover, it can provide feedback on the order of which code regions to work on (starting from the areas of strength).
[0066] Separate sales competency (skill) evaluation system = Automatically distribute questions to train writing thinking and expression ability for career advancement: According to the order of the areas to be trained, questions for training each area can be provided, and the writing evaluation device 1 can be applied to score these questions.
[0067] Apply to the short essay writing in college entrance examinations: The writing evaluation device 1 can be applied to score the short essays in college entrance examinations.
[0068] Development into a talent discovery system (evaluating talent from essays): The writing evaluation device 1 can be applied to score short essays for working adults. It can also be used in a writing scoring system for examinees, high school students, college students, and teachers.
[0069] Although each embodiment and each modification of the present technology have been described above, the present technology is not limited only to the above-described embodiments, and it goes without saying that various changes can be made without departing from the gist of the present technology.
Explanation of symbols
[0070] 1 Writing scoring device 110 Writing input section 120 Score prompt execution section 121 Toulmin model prompt 122 Goodman model prompt 123 Fifth model prompt 124 Bloom's taxonomy model prompt 125 General scoring criteria prompt 130 Evaluation area score calculation unit 140 Overall score calculation unit 150 Pass criterion calculation unit 160 Extended evaluation area determination unit 170 Comment prompt execution unit 171 Comment prompt 180 GUI generation unit 20 Electronic device
Claims
1. A sentence input unit for inputting a sentence to be scored, A score prompt for individually calculating scores for one or more of a plurality of evaluation areas, wherein the one or more evaluation areas for which a plurality of score prompts calculate scores are different for each score prompt, and a score prompt execution unit for respectively executing the plurality of score prompts on the sentence, An evaluation area score calculation unit for calculating an evaluation area score, which is the score of each evaluation area, based on the plurality of scores of the plurality of evaluation areas calculated by respectively executing the plurality of score prompts on the sentence, An overall score calculation unit for calculating an overall score of the sentence based on the evaluation area scores of each of the plurality of evaluation areas A sentence scoring device comprising the above.
2. The sentence scoring device according to Claim 1, The score prompt execution unit individually calculates N scores (N is an integer greater than or equal to 1) for each evaluation area using N score prompts, and the values of N for at least some of the evaluation areas are different, The evaluation area score calculation unit calculates (the sum of the N scores of a certain evaluation area) / N as the evaluation area score of this evaluation area, The full score of each evaluation area score is equal to the full score of the overall score, The overall score calculation unit calculates (the sum of the evaluation area scores of all evaluation areas) / (the total number of evaluation areas) as the overall score of the sentence A sentence scoring device.
3. The sentence scoring device according to Claim 2, The plurality of evaluation areas are in a matrix form constituted by a first axis representing the level of thinking and a second axis intersecting the first axis and representing the level of knowledge A sentence scoring device.
4. The sentence scoring device according to Claim 3, For at least one of the first axis and the second axis, the total value of N for each of the plurality of evaluation areas at the highest level is greater than the total value of N for each of the plurality of evaluation areas at the lowest level A sentence scoring device.
5. The sentence scoring device according to any one of Claims 1 to 4, Let the average value of the evaluation area scores of the plurality of evaluation areas of a plurality of passers of a certain test be X, Let the average value of the evaluation area scores of the plurality of evaluation areas of a plurality of failers of the test be Y (X > Y), When the evaluation area scores of the plurality of evaluation areas of the sentence creator are Z, A passing index calculation unit that calculates a passing index of the author of the text for a test related to the test based on (Y - Z) / (X - Y). A text scoring device further comprising the above.
6. The text scoring device according to any one of Claims 1 to 4, An extended evaluation area determination unit that determines which of the plurality of evaluation areas with an evaluation area score lower than the threshold should be the subject of extension effort based on the evaluation area scores of the evaluation areas around the plurality of evaluation areas with an evaluation area score lower than the threshold. A text scoring device further comprising the above.
7. The text scoring device according to any one of Claims 1 to 4, A comment prompt execution unit that executes a comment prompt for generating a comment on the text with respect to the text, A text scoring device further comprising the above.
8. The text scoring device according to Claim 3 or 4, A GUI generation unit that generates a GUI for displaying an evaluation area score for each of the plurality of matrix-shaped evaluation areas. A text scoring device further comprising the above.
9. By a computer, A step of inputting a text to be scored, A score prompt for individually calculating the scores of one or more of the plurality of evaluation areas, wherein the one or more evaluation areas for which the plurality of score prompts calculate scores are different for each score prompt, and a step of executing each of the plurality of score prompts with respect to the text, A step of calculating an evaluation area score, which is the score of each evaluation area, based on the plurality of scores of the plurality of evaluation areas calculated by executing each of the plurality of score prompts with respect to the text, A step of calculating an overall score of the text based on the evaluation area scores of each of the plurality of evaluation areas A text scoring method for executing the above.
10. In the computer of the text scoring device, A step of inputting a text to be scored, A score prompt for individually calculating the scores of one or more of the plurality of evaluation areas, wherein the one or more evaluation areas for which the plurality of score prompts calculate scores are different for each score prompt, and a step of executing each of the plurality of score prompts with respect to the text, A step of calculating an evaluation area score, which is the score of each evaluation area, based on the plurality of scores of the plurality of evaluation areas calculated by executing each of the plurality of score prompts with respect to the text, A step of calculating an overall score of the text based on the evaluation area scores of each of the plurality of evaluation areas A text scoring program for causing execution thereof.
Citation Information
Patent Citations
Answering scoring method and device, computer equipment and storage medium
CN113723774A
Description transition display device and program
JP2018022083A
Created sentence evaluation device
JP2022032319A
System and method for automatically scoring answer of sentence to question
JP2023086037A
Ability Assessment System
JP3227380U