Scoring system and scoring method

The scoring system automates the grading of handwritten answers by integrating handwriting recognition and similarity evaluation, addressing the inefficiencies and errors in traditional grading methods, and providing immediate feedback to improve the grading process.

JP2026036553APending Publication Date: 2026-03-05NAT UNIV CORP TOKYO UNIV OF AGRI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024139228
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing grading systems for essay-style questions are time-consuming, prone to errors, and lack immediate feedback, especially in large-scale examinations, and current multiple-choice tests do not adequately assess deep understanding or written responses.

Method used

A scoring system and method that utilizes handwriting recognition, similarity evaluation, and scoring criteria to automatically grade handwritten answers, providing immediate feedback through a user interface for examinees and graders, with options for correction and confirmation.

Benefits of technology

Reduces grading time and effort while increasing scoring reliability, enabling immediate feedback and reducing grading errors, thus enhancing examinee engagement and commitment to the grading process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026036553000001_ABST
    Figure 2026036553000001_ABST
Patent Text Reader

Abstract

To provide a scoring system or the like capable of reducing the labor and time required for scoring and at the same time increasing the reliability of the scoring. [Solution] The scoring system includes a handwriting recognition unit that recognizes handwritten answers; a similarity evaluation unit that evaluates the similarity between the answer and at least one example correct answer; a scoring unit that determines the answer as correct, incorrect, or rejected based on the evaluation result of the similarity evaluation unit and given scoring criteria; and a providing unit that provides the test taker with a test taker user interface that displays the scoring results of the scoring unit and accepts reports of incorrect scores, and provides the grader with a grader user interface that displays the scoring results and accepts corrections of incorrect scores and confirmation of rejection scores.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a scoring system and a scoring method. [Background technology]

[0002] Academic tests, whether they are in-class quizzes, midterm or final exams, nationwide mock tests, or entrance exams, are essential for measuring ability and learning outcomes. However, grading requires a great deal of time and effort. Furthermore, human grading is prone to errors and variability (hereafter, these are collectively referred to as "grading errors," and machine-generated grading errors are referred to as "mis-grading" to distinguish them). Furthermore, the longer the delay in returning graded results, the less opportunity and motivation test takers have to review.

[0003] As a means to solve these problems, multiple-choice testing has been adopted for large-scale simultaneous examinations, and CBT / WBT (Computer Based Testing / Web Based Testing) is beginning to become popular for qualification exams and small- to medium-sized exams. However, currently, only questions that test takers can mark or type into the keyboard are allowed, and written questions that test their thinking ability and deep understanding are excluded. There are also concerns that multiple-choice questions have a side effect on children's problem-solving behavior (encouraging them to choose without thinking).

[0004] The fundamental solution to these problems is to provide essay-style questions that test examinees' thinking ability and deep understanding, to implement automatic scoring and scoring support for essay-style questions, to reduce the time and effort required for scoring, to increase the reliability of scoring, and to provide feedback on the scoring results while they are still fresh in examinees' minds.In addition, examinees should be able to check the scoring results and immediately request corrections or ask questions by entering their ID and password.

[0005] Automatic scoring and scoring support for essay-style questions would also be useful for adding essay-style questions to CBT / WBT. Tablets, electronic paper, or devices that integrate display and handwritten input and allow direct instruction, direct operation, and direct entry (hereafter referred to as tablets, etc.) are suitable input methods for this purpose. However, since many tests are still conducted using writing implements such as paper and pencil, automatic scoring and scoring support for paper tests are also necessary.

[0006] The inventor attempted to automatically score 120,000 answer sheets in a trial test introducing a written answer format into the Common University Entrance Examination (Non-Patent Documents 1 and 2). In this test, a certain number of answer sheets for each question had to be manually graded in advance, and an automatic scoring method had to be machine-learned based on the pre-graded answer sheets and the scoring. This is not suitable for small-scale tests. [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] Haruki Oka, Hung Tuan Nguyen, Cuong Tuan Nguyen, Masaki Nakagawa and Tsunenori Ishioka: Fully Automated Short Answer Scoring of the Trial Tests for Newly Conducted Common Entrance Examinations for Japanese University, to appear Proc. 23rd International Conference on Artificial Intelligence in Education (AIED 2022), Durham University, UK (2022, 7). [Non-patent document 2] Hung Tuan Nguyen, Cuong Tuan Nguyen, Haruki Oka, Tsunenori Ishioka and Masaki Nakagawa: Handwriting Recognition and Automatic Scoring for Descriptive Answers in Japanese Language Tests, Proc. 18th International Conference on Frontiers in Handwriting Recognition (ICFHR 2022), Hyderabad, India, pp.274-284 (2022.12). Summary of the Invention [Problem to be solved by the invention]

[0008] An object of the present invention is to provide a scoring system etc. that can reduce the time and effort required for scoring and at the same time increase the reliability of the scoring. [Means for solving the problem]

[0009] (1) The present invention relates to a scoring system including a handwriting recognition unit that recognizes a handwritten answer; a similarity evaluation unit that evaluates the similarity between the answer and at least one example of a correct answer; a scoring unit that determines the answer as correct, incorrect, or rejected based on the evaluation result of the similarity evaluation unit and given scoring criteria; and a providing unit that provides a test taker with a user interface for displaying the scoring results of the scoring unit and accepting reports of incorrect scoring, and a grader user interface for displaying the scoring results and accepting corrections of incorrect scoring and confirmation of a rejected score, to the grader.

[0010] The present invention also relates to a scoring method including a handwriting recognition step of recognizing a handwritten answer; a similarity evaluation step of evaluating the similarity between the answer and at least one example correct answer; a scoring step of determining the answer as correct, incorrect, or rejected based on the evaluation result of the similarity evaluation step and given scoring criteria; and a providing step of providing an examinee with a user interface for examinees that displays the scoring results of the scoring step and accepts reports of incorrect scoring, and providing a grader with a user interface for graders that displays the scoring results and accepts corrections of incorrect scoring and confirmation of rejection.

[0011] (2) In the scoring system and scoring method according to the present invention, the scoring unit (in the scoring step) determines whether the degree of similarity is greater than or equal to a threshold value α c If the similarity is equal to or greater than the threshold value β, the answer is determined to be correct. c (β c <α c ), the answer is determined to be an incorrect answer, and the similarity is less than a threshold α c is less than the threshold β c If the answer does not match any of the correct answer examples, the similarity is determined to be a threshold value α w If the similarity is equal to or greater than the threshold value β, the answer is determined to be correct. w (β w <α w ), the answer is determined to be an incorrect answer, and the similarity is less than a threshold α w is less than the threshold β w If so, the answer may be rejected.

[0012] (3) In the scoring system and scoring method according to the present invention, the handwriting recognition unit (in the handwriting recognition step) may recognize the answer using a recognition method specified in the scoring criteria from among a plurality of recognition methods.

[0013] (4) In the scoring system and scoring method according to the present invention, the similarity assessment unit (in the similarity assessment step) may assess the similarity based on the confidence level of character recognition and the edit distance of character strings.

[0014] (5) In the scoring system and scoring method according to the present invention, the similarity assessment unit (in the similarity assessment step) may assess the similarity using a Siamese network.

[0015] (6) In addition, in the scoring system and scoring method of the present invention, the scoring unit (in the scoring step) may, when it is specified that the similarity is not to be used in the scoring criteria, determine the answer to be correct if the answer matches any of the correct answer examples, and may determine the answer to be incorrect if the answer does not match any of the correct answer examples. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 2 is a diagram schematically showing the flow of automatic scoring and scoring support in the scoring system according to the present embodiment. [Figure 2] FIG. 1 is a diagram showing the configuration of a scoring system according to an embodiment of the present invention. [Figure 3] FIG. 10 is a diagram showing an example of a question and a handwritten answer. [Figure 4] FIG. 4 shows an example of a correct answer to the question shown in FIG. 3 and an example of a description of the grading criteria. [Figure 5] A diagram showing the configuration of a Siamese network. [Figure 6] FIG. 10 is a diagram showing the generation of a provisional symbol-positional relationship sequence. [Figure 7] A diagram showing Transformer training. [Figure 8] 10 is a flowchart showing the flow of a process for determining whether an answer is correct, incorrect, or rejected based on similarity. [Figure 9] 10 is a flowchart showing the flow of a process for determining whether an answer is correct, incorrect, or rejected based on similarity. [Figure 10] 10 is a flowchart showing the flow of a correct answer, incorrect answer, and rejection determination process. [Figure 11] FIG. 10 is a diagram showing an example of the UI of a question creation tool. [Figure 12] 12A and 12B are diagrams showing examples of control statements output in accordance with the example of the automatic marking control flag shown in FIG. 11; [Figure 13] FIG. 10 is a diagram showing example question statements and examples of control flags automatically generated from each question statement. [Figure 14] FIG. 10 is a diagram showing an example of a UI for examinees. [Figure 15] FIG. 10 is a diagram showing an example of a UI for a grader. [Figure 16] FIG. 10 is a diagram showing an example of a UI for a grader. [Figure 17] FIG. 10 is a diagram showing an example of a threshold adjustment UI for manually adjusting a similarity threshold. [Figure 18] The bar graphs show the correct score rate, rejection rate, declaration rate, and risk score rate when answers are divided into five groups based on the accuracy rate and automatic scoring is performed using the basic method and the confidence-based similarity method with editing operations. DETAILED DESCRIPTION OF THE INVENTION

[0017] Hereinafter, an embodiment of the present invention will be described. Note that the embodiment described below does not unduly limit the content of the present invention described in the claims. Furthermore, not all of the configurations described in the embodiment are necessarily essential constituent elements of the present invention.

[0018] 1. Configuration Figure 1 is a schematic diagram illustrating the flow of automatic scoring and scoring support in the scoring system of this embodiment. As shown in Figure 1, the scoring system functions as automatic scoring by returning the scoring results to the examinee before the grader confirms and corrects them. The grader confirms and corrects the scoring results before returning them to the examinee. The grader switches between automatic scoring and scoring support with a switch. The addition of a process in which the grader confirms and corrects the examinee's responses, in addition to the process of examinee answers, graders' answers, and examinee confirmation, changes the paradigm of exam scoring and increases examinee commitment to grading. Because all answers are stored electronically, there is no risk of tampering. The scoring support system prevents grading errors by automatically scoring the answers before the grader actually grades them. The advantages of automatic scoring are that it provides immediate feedback to examinees and reduces the grading effort.

[0019] 2 is a diagram showing the configuration of the scoring system according to this embodiment. Scoring system 1 includes processing unit 100, storage unit 110, and communication unit 120. Scoring system 1 may be configured as a single server, or may be configured as multiple servers that distribute and execute the processing of each unit of processing unit 100.

[0020] The storage unit 110 stores programs and various data for causing the computer to function as each unit of the processing unit 100, and also functions as a work area for the processing unit 100, and this function can be realized by a hard disk, RAM, etc.

[0021] The communication unit 120 performs various controls to communicate with terminals (terminals of test takers, graders, and question creators) via a network (Internet), and its functions can be realized by hardware such as various processors or communication ASICs, or programs.

[0022] The processing unit 100 performs processes such as character recognition, similarity evaluation, scoring, and sending and receiving information to and from a terminal via the communication unit 120. The functions of the processing unit 100 can be realized by hardware such as various processors (CPU, DSP, etc.) and ASICs (gate arrays, etc.), or by programs. The processing unit 100 includes a handwriting recognition unit 101, a similarity evaluation unit 102, a scoring unit 103, and a providing unit 104.

[0023] The handwriting recognition unit 101 recognizes answers handwritten on the examinee's device (e.g., tablet) (converting them into a group of candidate code strings). Handwritten answer recognition utilizes a deep neural network model trained using learning patterns. The handwriting recognition unit 101 has online (electronic ink) and offline (image) recognition functions for handwritten Japanese, English, and mathematical expressions. Online recognition is resistant to continuations and irregularities, while offline recognition is resistant to incorrect stroke order and duplicate writing. For both online and offline recognition, the top k candidates (1st to kth) with the highest scores (confidence, likelihood) and their scores are output as recognition results. For offline recognition, the electronic ink stroke strings are connected, weighted, and converted into an image, which is then subjected to offline recognition. The handwriting recognition unit 101 recognizes handwritten answers using one of several recognition methods (e.g., online recognition methods for Japanese, English, and mathematical expressions, offline recognition methods, a combined online and offline recognition method, etc.) specified by the scoring criteria (rubric) described below. For example, online recognition is useful for questions that require writing kanji characters in the correct stroke order. Offline recognition is also useful for questions that require the correctness of words regardless of stroke order, or questions that often require incorrect stroke order, additional strokes, or double writing (e.g., arithmetic problems in the lower grades of elementary school). It is also possible to use both online and offline recognition methods (recognizing characters using a score obtained by weightedly adding the scores from online and offline recognition). It is also possible to use multiple recognition methods for both online and offline recognition, and to use a majority vote method (ensemble method). The handwriting recognition unit 101 can also use context processing such as trigrams for recognition. The handwriting recognition unit 101 also has a function for determining stroke order, stops, hooks, and pressure points to assess handwriting practice. In this embodiment, an example of answering questions using a tablet or the like will be described, but the present invention can also be applied to a paper-based exam in which test takers answer questions distributed to them, and the answers collected from the test takers are scanned (imaged) and sent to the scoring system 1.In this case, the questions and handwritten answers are separated from the scanned image of the answer sheet, and the handwritten answer images are recognized offline. The examinee then enters their ID and password on a PC or tablet to check the scoring results and report any incorrect scores.

[0024] The similarity evaluation unit 102 evaluates the similarity between the handwritten answer and at least one correct answer example. The similarity evaluation unit 102 may evaluate the similarity based on the recognition result by the handwriting recognition unit 101 and the edit distance of the character string, or may evaluate the similarity between the answer pattern and the correct answer pattern using a Siamese network. The similarity evaluation unit 102 may also evaluate the similarity between the features of the answer pattern extracted in the pattern recognition stage and the features of the correct answer pattern, or the similarity between the code string of the answer and the code string of the correct answer example. The similarity may also be calculated by weighted addition of similarities evaluated by multiple methods. Semantic analysis may also be used to recognize equivalence (the same value despite different expressions), synonymy (the same meaning despite different expressions), and implication (different meanings included), or to evaluate the content of short sentences.

[0025] The scoring unit 103 judges the answer as either correct, incorrect, or rejected based on the recognition result of the handwriting recognition unit 101 and the evaluation result of the similarity evaluation unit 102, and by referring to the correct answer examples and the scoring criteria. The descriptive questions include not only words and formulas but also questions with diagram answers, such as connecting related items with lines or adding lines to complete a geometric figure. Diagram answers are recognized using a method suited to each of them, without using the handwriting recognition unit 101. When the use of similarity is explicitly or implicitly specified in the scoring criteria, the scoring unit 103 uses similarity to judge the answer as either correct, incorrect, or rejected. In this case, if the answer matches one of the correct answer examples, the scoring unit 103 judges the answer as either correct, incorrect, or rejected based on the similarity if the similarity is greater than or equal to a threshold α c If the similarity is greater than or equal to the threshold value β c (β c <α c ), the answer is determined to be incorrect, and the similarity is less than the threshold α c is less than the threshold β cIf the answer does not match any of the correct answers, the similarity is determined to be equal to or greater than the threshold α w If the similarity is greater than or equal to the threshold value β w (β w <α w ), the answer is determined to be incorrect, and the similarity is less than the threshold α w is less than the threshold β w If the similarity is equal to or greater than the threshold value α, the scoring unit 103 determines the answer to be rejected. Alternatively, the scoring unit 103 may determine the answer to be correct if the similarity is equal to or greater than the threshold value β (β<α), determine the answer to be incorrect if the similarity is less than the threshold value β (β<α), and reject the answer if the similarity is less than the threshold value α and equal to or greater than the threshold value β. Furthermore, if it is explicitly or implicitly specified that the similarity is not to be used in the scoring criteria, the scoring unit 103 does not use the similarity and determines the answer to be correct if the answer matches any of the correct answer examples, and determines the answer to be incorrect if the answer does not match any of the correct answer examples.

[0026] The providing unit 104 provides an examinee UI (user interface) to the examinee (by displaying it on the display unit of the examinee's terminal), and provides a grader UI to the grader (by displaying it on the display unit of the grader's terminal). The providing unit 104 acquires questions, sample correct answers, and scoring criteria entered by a question creator (such as a tool developer, publisher, or teacher) using a question creation tool, displays the acquired questions on the examinee UI (by distributing them to the examinee), transmits the acquired sample correct answers to the similarity evaluation unit 102 and the scoring unit 103, and transmits the acquired scoring criteria to the handwriting recognition unit 101 and the scoring unit 103. The providing unit 104 also acquires handwritten answers entered on the examinee UI, and transmits the acquired handwritten answers to the handwriting recognition unit 101 and the similarity evaluation unit 102. The providing unit 104 also displays the scoring results of the scoring unit 103 on the examinee UI and the grader UI. Furthermore, the providing unit 104 acquires the declaration of an incorrect marking received in the examinee UI, displays the acquired declaration content in the grader UI, acquires the correction of the incorrect marking and confirmation of the reject marking (confirming it as a correct answer or confirming it as an incorrect answer) received in the grader UI, and reflects the acquired correction content and confirmation content in the marking results displayed in the examinee UI. When the marking results of the marking unit 103 are displayed in the grader UI, the correct answer group, the incorrect answer group, and the rejected answer group are displayed in a distinguishable manner, and further the rejected answer group and the incorrect answer group (and the correct answer group if there are multiple correct answer examples or partially correct answers) are displayed in descending or ascending order of similarity (or displayed in clustering), thereby facilitating marking (correction, confirmation) for each group.

[0027] 2. Method of this embodiment FIG. 3 shows an example of a question and a handwritten answer, and FIG. 4 shows an example of a correct answer for the question in FIG. 3 and an example of a description of the grading criteria. In this embodiment, the correct answer and the grading criteria are described in XML. This description data is called QARA (Question-Answer-Rubric-Annotation), and the description language is called QARA-ML. FIG. 4(a) is an example of a QARA for the question in FIG. 3(a) (Japanese language for first grade elementary school), and FIG. 4(b) is an example of a QARA for the question in FIG. 3(b) (Japanese language for fourth grade elementary school).

[0028] In QARA, the question is identified by the "id" attribute of the "question" element, and the type of question is specified by the "type" attribute. For correct answers, the "answer" element is used, the "eaid" attribute is used to identify the correct answer, and the "value" attribute is used to describe the correct answer. The "answer" element can also describe the score for correct answers, partially correct answers, and their scores. For scoring criteria, the "rubric" element is used, specifying the recognition method with the "recognizer" attribute, and the "score" attribute to determine whether to perform a perfect match (perfect matching) or synonyms and implications based on semantic analysis (semantic). The "rubric" element can specify whether to use contextual processing in recognition, whether to use similarity (rejection), as described below, the similarity calculation method and threshold, the allowable string edit distance, and the keywords to match. If no settings are specified, the implicit (default) settings are used. In addition, in the "annotation" element, the "x" and "y" attributes specify the X and Y coordinates of the top left corner of the answer input field, and the "width" and "height" attributes specify the width and height of the answer input field.

[0029] In the examples of Figures 3(a) and 4(a), the "recognizer" attribute of the "rubric" element is "online japanese", which specifies that the online Japanese recognition method should be used as the recognition method. In addition, the "strokeOrder" and "shape" attributes are set to "true" as additional attributes to evaluate the stroke order and character shape, so contextual processing will not be used. In the examples of Figures 3(b) and 4(b), the "recognizer" attribute of the "rubric" element is "online & offline japanese", which specifies that the recognition method should use a combination of online and offline Japanese recognition methods. In addition, the "stringDirection" attribute of the "annotation" element is "ttb (top to bottom)" and the "lineDirection" attribute is "rtl (right to left)", which specifies that the direction of characters entered in the answer input field is from top to bottom and the direction of lines is from right to left.

[0030] In the scoring method of this embodiment, the similarity between the answer and the correct answer example is evaluated, and if the top k candidates of the answer recognition match any of the correct answer examples, the similarity is equal to or exceeds the correct answer threshold (α c ), it is judged as the correct answer, while if the similarity is less than the correct answer threshold, it is rejected. This prevents incorrect answers from being judged as correct (hereinafter referred to as false positive (FP)). In addition, if the top k candidates for the recognized answer do not match any of the correct answers, the similarity is judged as being less than the correct answer threshold (α w ) or more, it is judged to be the correct answer. This prevents a correct answer from being erroneously judged as incorrect (hereinafter referred to as a false negative (FN)). In addition, if the top k candidates for the recognized answer do not match any of the correct answers, the similarity is judged to be equal to or greater than the incorrect answer threshold (β w ), it is judged as an incorrect answer, and the similarity is less than the correct answer threshold (α w ) and is less than the incorrect answer threshold (β w ) or more, the result is rejected. This prevents both false positive and false negative mis-scoring.

[0031] Note that false positive mis-scoring (FP) has a high risk of remaining as an incorrect score without being reported by the test taker, so the ratio of the number of false positive mis-scoring (FP) to the total number of answers N is called the risky scoring rate. On the other hand, false negative mis-scoring (FN) is likely to be reported by the test taker, so the ratio of the number of false negative mis-scoring (FN) to the total number of answers N is called the reporting rate. Also, because the grader will likely mark false negative mis-scoring (FN) and rejections, the ratio of the total number of false negative mis-scoring (FN) and rejections to the total number of answers N is called the human scoring rate. In reality, false positive mis-scoring (FP) will be reported and false negative mis-scoring (FN) will not be reported, so these are estimated values. In addition, correctly marking a correct answer as correct is called a true positive (TP), and correctly marking an incorrect answer as incorrect is called a true negative (TN). The ratio of the total number of true positives (TP) and true negatives (TN) to the total number of answers N is called the correct mark rate.

[0032] The following explains the method for evaluating similarity. When an answer pattern (AP) that is a handwritten answer is given, the answer pattern AP and the number of correct answers (Expected Answers: EA) are compared using the following formula: j ,1≦j≦l) with each other. j ) is calculated as the similarity. j ) may be used as the similarity.

[0033]

number

[0034] The similarity (one2one similarity) between AP and EA depends on whether the answer pattern AP is (1) a normal character string pattern, (2) a character string that is not recognized but treated as a handwritten pattern or a line diagram answer pattern (general handwritten pattern) that is not a character string, or (3) a structural pattern such as a mathematical formula or a chemical formula. j ) can be calculated in different ways.

[0035] If the answer pattern AP is (1) a character string pattern, the top k candidates (RC i ,1≦i≦k) and the score of each candidate (RS i ,0≦RS i ≦1) and candidate RC i and correct answer EA j Using the edit distance between AP and EA, the similarity is calculated using one of the following formulas: j ) is found.

[0036]

number

[0037] Here, EditDistance is a function that takes two arguments S1 and S2 and calculates the number of edit operations (insertion, deletion, substitution) required to match S1 with S2 (the edit distance). Also, length is a function that calculates the length of the argument (number of characters). In the above formula, the correct answer is EA j Candidate RC that matches (has an edit distance of 0) i Score of RS i The maximum value of the similarity (hereinafter referred to as the similarity using the confidence factor without editing operation) is calculated. i and correct answer EA j The weight of the candidate RC is smaller as the ratio of the edit distance to the number of characters increases. i Score of RS i The similarity is calculated by multiplying the result by 1 and the maximum value is calculated as the similarity (hereinafter referred to as the confidence-based similarity with edit operation). When answering a sentence or clause, it is possible to calculate the edit distance between the top answer pattern candidate and the correct answer example, or to use the bit string title of the keyword, without using multiple candidates.

[0038] If the answer pattern AP is (2) a general handwritten pattern, the distance between the answer pattern AP and the correct answer pattern (such as similarity or difference) can be measured. j A simple process such as dynamic programming (DP-matching) can be used to match the feature sequences.

[0039]

number

[0040] In Japanese language answers consisting of one or several characters, mis-scoring can occur if a non-existent character is recognized as a valid character and the recognized character matches the correct answer. To prevent this, a method can be used to evaluate similarity using a deep distance learning Siamese network without using the recognition results. Here, we explain the case of evaluating the similarity of Japanese language answers consisting of one or several characters, but it can also be used to evaluate the similarity of English, mathematical expressions, and diagram answers. As shown in Figure 5, a Siamese network consists of subnetworks with identical structures, share the same weights, and are trained simultaneously. One of two patterns is input to one subnetwork, and the other pattern is input to the other subnetwork, and the similarity (distance) between the two is output. When training a Siamese network, there are two types of input data (pairs of two answer patterns): positive pairs and negative pairs. Positive pairs are patterns with the same label, and use two correct answer patterns with the same correct answer example. On the other hand, negative pairs are patterns with different labels, and pairs of correct and incorrect answer patterns of the same correct answer example are used. However, since there are many questions with few incorrect answers, a correct answer example is paired with a correct answer pattern of another correct answer example with a different label. In order to train the Siamese network to output similarity, it is trained so that the distance D becomes small when the input data is a positive pair, and so that the distance D becomes large when the input data is a negative pair. Answer pattern AP and correct answer example EA j When calculating the similarity of , first select N correct patterns and put them into a set U EAj Then, the answer pattern AP and each correct answer pattern are input to the Siamese network, N similarities (siamese distances) are calculated, and the average of the N similarities is found. Note that the maximum value of the N similarities may also be found.

[0041]

number

[0042] This method does not require retraining the Siamese network for new correct answers; it only requires preparing a few correct answer patterns.

[0043] If the answer pattern AP is (3) a structural pattern, symbols and their positional relationships can be extracted by handwriting recognition using a deep neural network, and similarity can be calculated by conditional decoding.

[0044]

number

[0045] In this case, as shown in Figure 6, features are extracted from the time-series input pattern of the answer (a coordinate sequence in the case of online handwriting, or a sequence of extracted symbols in the case of offline handwriting), and a BLSTM (Bidirectional Long Short-term Memory) network is used to extract the positional relationships between symbols, outputting them along with the time-series classification probability. The LaTeX (sequence of LaTeX symbols) of the correct answer is input into a time-series transformer (Sequence-to-sequence Transformer), which generates a provisional symbol-positional relationship sequence in which the symbols and their positional relationships are arranged according to the stroke order. For example, "x Sup 2 NoRel + Right a" indicates a provisional symbol-positional relationship sequence in which an "x" is written, a "2" is written to its right (Sup), a "+" is written regardless of position (NoRel), and an "a" is written to its right (Right). Since there are multiple possible stroke orders for this expression, all of them are generated. Next, the similarity between the time-series classification probability and the provisional symbol-positional relationship sequence is calculated. To do this, Using the CTC (Connectionist Temporal Classifier) ​​alignment method (CTC forward backward algorithm), we align the time-series classification probability with the provisional symbol-positional relationship sequence, and output the probability of P(provisional symbol-positional relationship sequence | time-series classification probability) as the similarity. To generate the provisional symbol-positional relationship sequence, we need to train a time-series Transformer. As shown in Figure 7, we generate provisional symbol-positional relationship sequences with multiple different stroke orders using the random scanning method (Cuong Tuan Nguyen, Thanh-Nghia Truong, Hung Tuan Nguyen, Masaki Nakagawa: Global Context for Improving Recognition of Online Handwritten Mathematical Expressions, Proc. ICDAR 2021, Lausanne, Switzerland, pp. 617-631, 2021.) using the SRT (Symbol Relation Tree) label data of the training pattern.In addition, you can also use the BERT (Bidirectional Encoder Representations from Transformers) model, which calculates the similarity between the answer and the correct answer example.

[0046] Next, the determination of correct answer, incorrect answer, or rejection (automatic scoring) based on similarity will be described with reference to the flowcharts of Figs. 8 and 9. Fig. 8 shows an example of the determination process using the recognition result. The similarity evaluation unit 102 calculates the above-mentioned similarity using the confidence factor without editing operation or the similarity using the confidence factor with editing operation, and calculates the answer pattern AP and the correct answer example EA j The similarity between the two is calculated (step S10).

[0047] Next, the scoring unit 103 scores the top k candidates RC that are the recognition results of the answer pattern AP. i The correct answer is EA j It is determined whether the top k candidate RCs match any of the above (step S11). i The correct answer is EA j If the similarity matches any of the above (Y in step S11), the scoring unit 103 checks whether the similarity is greater than or equal to the correct answer threshold α c It is determined whether or not the correct answer threshold α c If the similarity is equal to or greater than the correct answer threshold α c If the similarity is less than the incorrect answer threshold β c (β c <α c ) (step S13), and c If it is less than the threshold value β (Y in step S13), it is determined to be an incorrect answer. c If it is equal to or greater than this (N in step S13), it is rejected.

[0048] Top k candidate RCs i The correct answer is EA j If the similarity does not match any of the above (N in step S11), the scoring unit 103 checks whether the similarity is greater than the correct answer threshold α wIt is determined whether or not the correct answer threshold α w If the similarity is equal to or greater than the correct answer threshold α w If the similarity is less than the incorrect answer threshold β w (β w <α w ) (step S15), and w If it is less than the threshold value β (Y in step S15), it is determined to be an incorrect answer. w If it is equal to or greater than this (N in step S15), it is rejected. c , α w , β c and β w is a value between 0 and 1, and α c and α w , β c and β w are different values.

[0049] 9 shows an example of a determination process that does not use the recognition result (using only the similarity). The similarity evaluation unit 102 uses the above-mentioned DP-matching or Siamese network to determine the answer pattern AP and the correct answer example EA. j The similarity between the two is calculated (step S20).

[0050] Next, the scoring unit 103 determines whether the similarity is equal to or greater than the correct answer threshold α (step S21), and if it is equal to or greater than the correct answer threshold α (Y in step S21), it determines the answer as correct. If the similarity is less than the correct answer threshold α (N in step S21), the scoring unit 103 determines whether the similarity is less than the incorrect answer threshold β (β<α) (step S22), and if it is less than the incorrect answer threshold β (Y in step S22), it determines the answer as incorrect, and if it is equal to or greater than the incorrect answer threshold β (N in step S22), it rejects the answer. Note that the correct answer threshold α and the incorrect answer threshold β can be made equal to prevent the answer from being rejected.

[0051] The similarity threshold is set from the training answer patterns, but there is also a method of training the deep neural network by tagging the answer patterns and correct answer patterns with tags indicating whether they should be rejected or not. It is also possible to train the threshold sequentially while the system is in operation.

[0052] Next, the details of the correct answer, incorrect answer, and rejection judgment (automatic scoring) will be explained using the flowchart in Figure 10. First, the scoring unit 103 refers to the scoring criteria (the "rubric" element of QARA) to determine whether or not a handwriting learning judgment has been specified (step S30). If a handwriting learning judgment has been specified (Y in step S30), the handwriting recognition unit 101 judges the correctness of the stroke order, stops, hooks, and pressure points for the answer pattern AP (step S31), and outputs the judgment result to the scoring unit 103.

[0053] If the handwriting learning assessment is not specified (N in step S30), the scoring unit 103 refers to the scoring criteria and determines whether or not the requirement for correct stroke order and correct character shape is specified (step S32). If the requirement for correct stroke order and correct character shape is specified (Y in step S32), the handwriting recognition unit 101 recognizes the answer pattern AP using an online recognition method that does not use context processing (step S33). Next, the scoring unit 103 refers to the scoring criteria and determines whether or not to use similarity (step S34). If the requirement for similarity is specified (Y in step S34), the process proceeds to the similarity-based assessment process (FIGS. 8 and 9). If the requirement for similarity is not specified (N in step S34), the scoring unit 103 determines whether or not the top k candidates, which are the recognition results of the answer pattern AP, match any of the correct answer examples (correct answer examples described in the "answer" element of the QARA) (step S35). If they match (Y in step S35), the answer is determined to be correct, and if they do not match (N in step S35), the answer is determined to be incorrect. When similarity is used, there are cases where a correct answer is judged to be rejected without changing the risk mark rate due to the characteristics of the question and answer, and in such cases, judgment without using similarity is effective.

[0054] If the correct stroke order or the correct character shape is not required (N in step S32), the scoring unit 103 refers to the scoring criteria and determines whether or not to use contextual processing (step S36). If the contextual processing is not to be used (Y in step S36), the scoring unit 103 determines whether or not to use similarity (step S37). If the similarity is not to be used (N in step S37), the scoring unit 103 determines whether or not the recognition method specified in the scoring criteria (the "recognizer" attribute) is a single recognition method (one of the online recognition method, the offline recognition method, or the combined online and offline recognition method) (step S38). If the specified recognition method is a single recognition method (Y in step S38), the handwriting recognition unit 101 recognizes the answer pattern AP using the specified recognition method (step S39). The scoring unit 103 determines whether the top k candidates, which are the recognition results, match any of the correct answer examples (step S40). If they match (Y in step S40), it is determined to be the correct answer, and if they do not match (N in step S40), it is determined to be an incorrect answer. If multiple recognition methods are used (N in step S38), the handwriting recognition unit 101 recognizes the answer pattern AP using each of the specified multiple recognition methods (step S41). The scoring unit 103 determines whether the top k candidates, which are the recognition results of any of the specified multiple recognition methods, match any of the correct answer examples (step S42). If they match (Y in step S42), it is determined to be the correct answer, and if they do not match (N in step S42), it is determined to be an incorrect answer. If similarity is used (Y in step S37), the handwriting recognition unit 101 recognizes the answer pattern AP using the recognition method specified in the scoring criteria (step S43), and proceeds to the similarity judgment process (FIGS. 8 and 9). If the specified recognition method uses multiple recognition methods, and the results of the similarity judgment using the recognition results of each recognition method match, the result is adopted, and if they do not match, the result is rejected.

[0055] If context processing is used (N in step S36), the answer pattern AP is recognized using context processing according to the recognition method / authentication method specified in the scoring criteria (step S44). The scoring unit 103 determines, according to the scoring criteria, whether the character string of the recognition result exactly matches the character string of the correct answer example, whether the edit distance between the character string of the recognition result and the character string of the correct answer example is within an allowable range, whether the character string of the recognition result contains a keyword, and performs a semantic analysis of the character string of the recognition result using natural language processing to determine whether the character string of the recognition result and the character string of the correct answer example are equivalent, synonymous, or implicational.

[0056] Next, the user interface (UI) will be described. FIG. 11 is a diagram showing an example of the UI of a question creation tool. The question creator enters a question statement in the question statement text box TB and enters multiple correct answer examples in the text box TB for each correct answer. For correct answer examples, it is possible to set whether or not partial points are awarded and the percentage of partial points relative to the score for a completely correct answer. Furthermore, by setting attention detection to ON, it is possible to detect specific characters and use them for instruction. The question creator also sets the scoring criteria (automatic scoring control flag). In this example, the scoring criteria can be set to include whether or not to judge calligraphy practice, whether or not to require correct stroke order, whether or not to use context processing, whether or not to use rejection (whether or not to use similarity), and the type of character recognition engine (Japanese / English). FIG. 12 shows an example of control statements (inkML) output for the example of the automatic scoring control flag shown in FIG. 11.

[0057] In addition, the grade (and the year of admission to identify the corresponding curriculum guidelines) and subject can be set, and correct answer examples for questions and automatic scoring control flags can be automatically generated according to the set grade and subject. For example, in science, the katakana spelling of plant names is considered the correct answer, but for kanji, the question creator and grader will have to make the decision. Also, for subjects other than Japanese, the calligraphy learning judgment and the requirement for correct stroke order are not set to ON (present). Furthermore, when the question creator presses the control flag generation button FB, an automatic scoring control flag can be automatically generated from the entered question sentence using natural language processing. Figure 13 shows example question sentences and examples of control flags automatically generated from each question sentence.

[0058] Figure 14 shows an example of a test-taker UI. The test-taker UI shown in Figure 14 displays the results of automatic grading and accepts test-taker reports of incorrect scores and pending grading (rejected scores). In this example, the answer to the equation marked with a "△" (△) indicates that the answer "6 + 2 = 8" is pending grading, while the answer marked with a "○" (○) indicates that the answer "8" is correct. Test-taker can select (tap) the box for the pending answer "6 + 2 = 8," enter a comment indicating that it is the correct answer (in this case, "I think it's correct"), and press the "Send feedback" button to declare that the answer "6 + 2 = 8" is correct. This function ensures transparency in grading, encourages review immediately after the test, and is expected to increase test-taker initiative. It also increases test-taker commitment and promotes communication with graders to resolve problems. In addition, for each answer to a question, a record of the marks, reports of incorrect marks and marks pending, and corrections by the marker will be saved and made available to both the examinee and the marker.

[0059] 15 and 16 are diagrams showing examples of the grader UI. The grader UIs shown in FIGS. 15 and 16 are screens that display the grading results and accept the grader's correction of incorrect marks and confirmation of pending grading (rejected marks). The grader UI shown in FIG. 15 is used in grading support before the marks are returned to the examinee, and the grader UI shown in FIG. 16 is used in automatic grading after the marks are returned to the examinee (after the examinee reports). The grader UI displays each examinee's answer for each question, categorized into a pending grading group marked with a "△" (a group of answers determined to be rejected), a correct answer group marked with a "○" (a group of answers determined to be correct), and an incorrect answer group marked with an "×" (a group of answers determined to be incorrect). Here, automatic grading is performed using similarity, and answers in the pending grading group and the incorrect answer group are displayed in descending order of similarity. By selecting an answer (by hovering the mouse or pen over the answer, or by clicking or tapping on the answer), the grader can display that answer along with the question on the left side of the grader UI. In the example shown in Figure 15, the first answer in the group waiting to be graded is selected. Then, a "○" mark and an "×" mark are displayed in the box of that answer, and if the grader selects the "○" mark, the answer is marked as correct, and if the grader selects the "×" mark, the answer is marked as incorrect. The same applies when correcting incorrect marks for answers in the correct answer group or incorrect answer group.

[0060] In the example shown in Figure 16, the third answer in the pending grading group has been commented on by a test-taker, so the answer is displayed in a bold frame. The question and answer are displayed on the left side of the grader UI, and the correct answer, the recognition result, and the test-taker's comment ("I think it's correct.") are displayed at the top of the grader UI. The grader can enter a response comment (here, "It was correct.") and confirm the answer as correct by selecting the "○" mark directly below, or confirm the answer as incorrect by selecting the "×" mark. These confirmations of the score and response comments are displayed in the test-taker UI. If multiple test-takers have commented on incorrect or pending grading, the comments are displayed in order, and the next comment is displayed once that processing is complete. Alternatively, multiple comments can be displayed, and the grader can select an answer and confirm the score of the answer. This comment and response function can also be used to provide individual guidance to individual test-takers regarding their incomprehension or misunderstanding.

[0061] The confirmation of the score by the scorer can be used to sequentially learn the threshold for determining whether a given answer is correct, incorrect, or rejected based on similarity in order to reduce future incorrect scores. For example, if an incorrect answer is confirmed (corrected) as correct, there may be a correct answer in the top ranking, but the similarity is lower than the incorrect answer threshold β c If the answer is judged to be incorrect because it is less than β c The learning is done in the direction of lowering the threshold β w If the answer is judged to be incorrect because it is less than β w Also, if a correct answer is confirmed (corrected) as an incorrect answer, the correct answer is higher in the recognition ranking and the similarity is lower than the correct answer threshold α c If the answer is judged to be correct because it is equal to or greater than α c The learning is done in the direction of increasing the similarity threshold α w If the answer is judged to be correct because it is equal to or greater than α w Also, if the rejection is confirmed as the correct answer, the correct answer is higher in the recognition ranking, but the similarity is lower than the correct answer threshold α c Less than the incorrect answer threshold β c If the result is rejected because it is greater than or equal to α cThe learning is done in the direction of lowering the threshold α w Less than the incorrect answer threshold β w If the result is rejected because it is greater than or equal to α w Also, if the rejection is confirmed as an incorrect answer, the correct answer is higher in the recognition, but the similarity is lower than the correct answer threshold α c Less than the incorrect answer threshold β c If the result is rejected because it is greater than or equal to β c The learning is done in the direction of increasing the similarity threshold α w Less than the incorrect answer threshold β w If the result is rejected because it is greater than or equal to β w During learning, the learning rate is multiplied by the threshold change range so that the learning gradually converges.

[0062] Although the above-mentioned automatic learning of thresholds is expected to improve the accuracy of automatic scoring overall, it is also possible to provide a function to manually adjust the strictness of scoring (similarity threshold) for individual tests or questions. By adjusting the similarity threshold, it is possible to adjust the range of how far an answer deviates from the model answer before it is judged to be correct, or whether it is an incorrect answer or rejected, and it is possible to select from simple to strict scoring standards. Figure 17 shows an example of a threshold adjustment UI for manually adjusting the similarity threshold. The threshold adjustment UI shown in Figure 17 displays the question statement, sample correct answers, and the similarity thresholds (α w , β w , α c , β c ), and a display field DC for displaying whether the sample answer is correct, incorrect, or rejected. When each threshold is adjusted using the slider SL, the judgment process shown in Figure 8 is performed on the sample answer based on the adjusted threshold, and the sample answer is displayed in the display field DC as either correct, incorrect, or rejected according to the judgment result. This function is used by question creators and graders when creating and grading questions.

[0063] 3. Evaluation Experiment An experiment was conducted to evaluate the scoring method of this embodiment. In this experiment, answer data for Japanese, English, and arithmetic provided by 50 fifth- and sixth-grade elementary school students was used. Although the rubric (QARA) allows detailed specification of scoring criteria for each question, in order to verify the feasibility of the method, a consistent method was used for all questions, rather than individual settings for each question. Because the answers for Japanese, English, and arithmetic provided by fifth- and sixth-grade elementary school students consisted of words and formulas, we attempted to determine exact matches and similarity.

[0064] In this experiment, regardless of stroke order or character shape (N in step S32 in Figure 10), without using context processing (Y in step S36), with or without using similarity (Y in step S37), a perfect match with the correct answer example or a judgment using similarity was made, and online recognition, offline recognition, and a combined online and offline recognition method (hereinafter referred to as combined recognition) were compared as a single recognition method. Combined recognition is a simple method in which the answer is judged to be correct if the top candidate from either recognition method matches one of the correct answer examples.

[0065] Table 1 shows the performance (FN rate, FP rate) of automatic scoring for answers in three subjects when exact matches are obtained without using similarity, and when online, offline, and combined recognition are used as the single recognition method (steps S39 and S40). Here, the FN rate is the rate at which correct answers are mistakenly marked as incorrect (the ratio of false negatives (FN) to the sum of false negatives (FN) and true positives (TP)), and is called the false negative rate. The FP rate is the rate at which incorrect answers are mistakenly marked as correct (the ratio of false positives (FP) to the sum of false positives (FP) and true negatives (TN), and is called the false positive rate. The FN rate and FP rate are indicators of the performance of the scoring system, and are useful even when the ratio of correct to incorrect answers changes.

[0066] [Table 1]

[0067] Online recognition shows a lower FN rate than offline recognition, and offline recognition shows a lower FP rate than online recognition. The purpose of this embodiment is to make these available individually for each question, but even simple combined recognition shows an even lower FN rate. On the other hand, combined recognition does not improve the FP rate much, and it tends to score questions more easily.

[0068] The FP rate, an important indicator, is quite low for mathematical formulas (arithmetic) but high for Japanese (national language) and English. In actual elementary school exams, the error rate for each question is about 20% overall. Multiplying this by 5%, the error rate accounts for about 5% of all answers, which is considered a sufficient level for use, but there is room for improvement. Because the Japanese and English recognition methods are trained on correctly written characters and phrases, even if the use of explicit context is stopped, it is possible that the system is implicitly learning context from the training pattern. As a result, even minor handwritten errors are correctly recognized. Furthermore, handwritten kanji with missing or extra strokes are often recognized as correct kanji. Therefore, while simple combined recognition is effective in reducing the FN rate, it is counterproductive in reducing the FP rate.

[0069] Table 2 shows the performance of automatic scoring (FN rate, FP rate, manual scoring rate, risky scoring rate, correct scoring rate, rejection rate, and declaration rate) for answers in three subjects when using similarity (with rejection) and combined recognition as the single recognition method (step S43), and when performing the judgment process of FIG. 8 using the confidence-based similarity without editing operations and the confidence-based similarity with editing operations. Here, the rejection rate is the ratio of the number of rejected answers to the total number of answers N. Note that the basic method used as the comparison standard is when the same combined recognition is used as the single recognition method and an exact match is obtained (no rejection). Regarding the formulas, results when using confidence-based similarity without editing operations are not shown because a sufficiently low FP rate is already achieved without rejection.

[0070] [Table 2]

[0071] When rejection is adopted using confidence-based similarity without editing operations, the FP rate is halved for Japanese and English, and therefore the risky marking rate is reduced. On the other hand, the FN rate increases, and the human marking rate and reporting rate increase. Furthermore, when confidence-based similarity with editing operations is used, the FN rate decreases for all subjects, with a particularly large reduction for English and mathematical formulas, leading to a significant reduction in the reporting rate. On the other hand, the FP rate remains unchanged for Japanese but increases for English and mathematics, and the risky marking rate increases for English and mathematics accordingly. However, it is believed that future improvements in each area will bring this within an acceptable range.

[0072] These experiments show that it is effective to use different methods depending on the subject, such as whether or not to use rejection (using similarity), and to use different similarity methods. In this experiment, we did not use a rubric to specify the method for each question, but since the effectiveness of the method differs depending on the subject, we found that it is necessary to specify the method for each question using a rubric.

[0073] Table 3 shows the performance of automatic scoring (FN rate, FP rate, human scoring rate, risky scoring rate, correct scoring rate, rejection rate, and declaration rate) when the judgment process shown in Figure 9 is performed on single-character answers in Japanese using similarity scores from the Siamese network.

[0074] [Table 3]

[0075] On average for each grade, the risky marking rate is below 1.8%, the reported marking rate is below 0.9%, the false marking rate is below 12.4%, and the correct marking rate is above 86%.

[0076] Table 4 shows the performance of automatic scoring (FN rate, FP rate, human scoring rate, risky scoring rate, correct scoring rate, rejection rate, and declaration rate) for each setting of the correct answer threshold α and incorrect answer threshold β when performing the judgment process shown in Figure 9 using similarity based on the Siamese network for single-character answers in Japanese.

[0077] [Table 4]

[0078] When the correct answer threshold α and incorrect answer threshold β are both set to 0.5 (no rejection), the correct marking rate is high and the human marking rate is low, but the risky marking rate is 2.94% and the claim rate is 1.52%. It can be seen that lowering the incorrect answer threshold β reduces the claim rate, but increases the human marking rate and rejection rate, resulting in a lower correct marking rate. It can be seen that increasing the correct answer threshold α reduces the risky marking rate, but increases the human marking rate and rejection rate, resulting in a lower correct marking rate. Therefore, when setting the thresholds, it is necessary to take into account each indicator of automatic marking. By adopting rejection and setting the thresholds appropriately (β = 0.3, α = 0.8), it can be seen that the risky marking rate is 1.57%, the claim rate is 0.69%, the human marking rate is 11.26%, and the correct marking rate is 87.17%.

[0079] Figure 18 shows the correct score rate, rejection rate, declaration rate, and risky score rate for each group of 6th grade English answers divided into five groups based on the accuracy rate (GD1: 80% or higher accuracy rate, GD2: 60% to less than 80% accuracy rate, GD3: 40% to less than 60% accuracy rate, GD4: 20% to less than 40% accuracy rate). The bar graphs show the correct score rate, rejection rate, declaration rate, and risky score rate for each group when automatic scoring is performed using the basic method (no rejection) and the confidence-based similarity with editing operations (with rejection) shown in Table 2. Note that there were no questions with a correct score rate of GD4. For easy questions with a correct score of 80% or higher (GD1), using confidence-based similarity with editing operations results in some correct scores being rejected. Since the scoring of questions may not always be in line with the scoring criteria, it may be possible to first try applying automatic scoring that does not use rejection, regardless of the scoring criteria, and return the results as is for the answer group for questions with a very large number of correct answers, while applying automatic scoring that uses similarity for the answer group for other questions.

[0080] As seen in the evaluation experiments, the scoring method of this embodiment reduces the time and effort required to grade answers for essay-style questions that test examinees' thinking ability and deep understanding, while also increasing the reliability of the scoring and providing feedback on the scoring results while the results are still fresh in the examinee's memory. However, performance, such as the risky scoring rate, the reporting rate, and the manual scoring rate, varies depending on whether the question is Japanese (Japanese language), English, or arithmetic, and whether the answer is a word, phrase, or sentence, and whether it requires one or several Japanese kanji characters. If these vary too much depending on the subject or question, the work and impressions of examinees and graders will become inconsistent. To avoid this, the scoring criteria specify whether to use rejection and what similarity to use for each subject and question. This allows examinees and graders to be provided with a fairly consistent level of automatic scoring and scoring support performance. This is extremely important from the perspective of user experience.

[0081] The present invention is not limited to the above-described embodiments, and various modifications are possible. The present invention includes configurations that are substantially the same as those described in the embodiments (for example, configurations with the same functions, methods, and results, or configurations with the same purpose and effects). The present invention also includes configurations in which non-essential parts of the configurations described in the embodiments are replaced. The present invention also includes configurations that achieve the same effects as the configurations described in the embodiments, or that can achieve the same purpose. The present invention also includes configurations in which publicly known technology is added to the configurations described in the embodiments. [Explanation of symbols]

[0082] 1...Scoring system, 100...Processing unit, 101...Handwriting recognition unit, 102...Similarity evaluation unit, 103...Scoring unit, 104...Providing unit, 110...Storage unit, 120...Communication unit

Claims

1. a handwriting recognition unit that recognizes handwritten answers; a similarity evaluation unit that evaluates the similarity between the answer and at least one correct answer example; a scoring unit that judges the answer as correct, incorrect, or rejected based on the evaluation result of the similarity evaluation unit and given scoring criteria; A scoring system including: a providing unit that provides a test taker with a test taker user interface that displays the scoring results of the scoring unit and accepts reports of incorrect scoring, and a grader user interface that displays the scoring results and accepts corrections of incorrect scoring and confirmation of rejection of scoring, to the grader.

2. In claim 1, The scoring unit If the answer matches any of the correct answer examples, the similarity is equal to or exceeds a threshold value α c If the similarity is equal to or greater than the threshold value β, the answer is determined to be correct. c (β c <α c ), the answer is determined to be an incorrect answer, and the similarity is less than a threshold α c is less than the threshold β c If the answer is more than 1, the answer is rejected. If the answer does not match any of the correct answers, the similarity is equal to or less than a threshold α w If the similarity is equal to or greater than the threshold value β, the answer is determined to be correct. w (β w <α w ), the answer is determined to be an incorrect answer, and the similarity is less than a threshold α w is less than the threshold β w If so, the scoring system rejects the answer.

3. In claim 1, The handwriting recognition unit A scoring system that recognizes the answer using a recognition method designated by the scoring criteria from among a plurality of types of recognition methods.

4. In claim 1, The similarity evaluation unit A scoring system that evaluates the similarity based on character recognition confidence and string edit distance.

5. In claim 1, The similarity evaluation unit A scoring system that uses a Siamese network to evaluate the similarity.

6. In claim 1, The scoring unit When the scoring criteria specify that the similarity should not be used, the scoring system judges the answer to be correct if the answer matches any of the correct answer examples, and judges the answer to be incorrect if the answer does not match any of the correct answer examples.

7. a handwriting recognition step for recognizing a handwritten answer; a similarity evaluation step of evaluating a similarity between the answer and at least one correct answer example; a scoring step of determining the answer as correct, incorrect, or rejected based on the evaluation result of the similarity evaluation step and a given scoring criterion; A scoring method including a providing step of providing a test taker with a test taker user interface that displays the scoring results of the scoring step and accepts reports of incorrect scoring, and providing a grader with a grader user interface that displays the scoring results and accepts corrections of incorrect scoring and confirmation of rejection of scoring.