Scoring support device, program

The AI-powered scoring support device enhances the accuracy and efficiency of scoring handwritten answers by recognizing answer frames, correcting errors, and managing processing load based on question type, addressing misrecognition and load issues in existing systems.

JP7859724B1Active Publication Date: 2026-05-15MICROSIMULATION CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
MICROSIMULATION CO LTD
Filing Date
2026-01-14
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing AI-based scoring systems struggle with accurately recognizing handwritten answers in unique handwriting styles or characters that resemble learned characters, leading to misrecognition and increased processing load due to insufficient consideration of question context and excessive natural language processing.

Method used

An AI-powered scoring support device that includes an acquisition unit for converting handwritten answers into digital data, a conversion unit that recognizes answer frames and sets scoring criteria, and a scoring unit that processes information based on question type, using a conversion engine to correct errors and switch scoring modes to manage processing load.

Benefits of technology

Improves recognition accuracy by preventing misrecognition of characters and reducing processing load through context-aware scoring, ensuring accurate and efficient evaluation of handwritten answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007859724000001_ABST
    Figure 0007859724000001_ABST
Patent Text Reader

Abstract

This suppresses the increase in processing load and the decrease in the accuracy of handwritten answer recognition and scoring. [Solution] The conversion unit 32 of the scoring support server 30 recognizes the answer frame, which has a set question type (answer format) and scoring criteria, from the answer sheet image data in which the answer corresponding to the question is expressed, and converts the portion expressed in the recognized answer frame into digital data in an answer format that matches the question type. If the evaluation value of the conversion is low, it refers to the question text, etc., and replaces it with the correct characters, etc. The scoring unit 33 switches between a simple automatic scoring mode, an AI automatic scoring mode requiring natural language processing, and a manual scoring mode depending on the type of question, and processes the information represented by the digital data according to the scoring criteria in the selected mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a scoring support technology for supporting the scoring process for answers to problem statements.

Background Art

[0002] In recent years, attempts have been made to automatically score descriptive answers to problem statements using AI (Artificial Intelligence) technology. For example, Patent Document 1 discloses a technology for improving the accuracy of correct answer determination by an AI performing an approximation determination on the input result of an essay answer question through natural language classification. This technology changes the combination of the premise (subject and reason part) and the conclusion part of the sentence, and not only for correct answers but also for incorrect answers, makes the AI remember the answer patterns, thereby improving the determination accuracy through machine learning with a small number of patterns and shortening the time until the scoring is completed.

[0003] Patent Document 2 discloses a technology in which an AI (artificial intelligence) creates a scoring criterion according to the solution process from a free description question to the correct answer based on free description question information indicating a free description question related to a test subject, correct answer information indicating the correct answer, and scoring information indicating the score for the free description question information, and scores the answer of the answerer from the answer information indicating the answer of the answerer based on the scoring criterion.

[0004] However, even with the technologies disclosed in Patent Documents 1 and 2, handwritten answers expressed in unique handwriting styles that fall outside the AI's learning range, or handwritten answers expressed in characters that closely resemble learned characters but have different meanings, may not be correctly recognized (scored). For example, the number "7" may be misrecognized as "17" (a combination of "1" and "7") depending on the handwriting. Also, the number "3" may be misrecognized as the katakana character "ヨ" or the hiragana character "ろ", the number "1" may be misrecognized as the uppercase letter "I" or the katakana character "エ", and the handwritten English word "modern" may be misrecognized as "modem". Many of these misrecognitions stem from the fact that the context (background information) of the question type (mathematical formula / descriptive / multiple choice, etc.) is not sufficiently considered in character recognition. Furthermore, some question types do not necessarily require natural language processing. If such question types are not distinguished from those that require natural language processing, the processing load becomes excessive, and the processing waiting time increases. [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] Patent No. 7387101 [Patent Document 2] Patent No. 7598513 [Overview of the project] [Problems that the invention aims to solve]

[0006] The primary objective of the present invention is to provide a technology that can improve the recognition accuracy of the portion to be scored and suppress the increase in processing load during scoring. Other objectives of the present invention will become clear from the description herein. [Means for solving the problem]

[0007] One aspect of the present invention includes an acquisition unit that acquires answer image data in which the answer corresponding to the problem statement is expressed, and a conversion unit that recognizes an answer frame in which the problem type and scoring criteria are set from the answer image data and converts the portion expressed in the recognized answer frame into digital data in an answer format that conforms to the problem type. One of several scoring modes with different processing loads during scoring. Switch according to the type of problem. along Switched Grading The system includes a scoring unit that processes the information represented by the digital data in accordance with the scoring criteria, The conversion unit is equipped with a conversion engine that has learned the correct word forms, usages, and grammar, as well as the easily erroneous word forms, usages, and grammar. When the question type is a short-answer or long-answer descriptive answer format, the conversion engine replaces any descriptive portion of the answer box where the conversion evaluation value is lower than a predetermined value with words used in the question text or model answer for that question type. It is a scoring support device. The "question type" is information used to distinguish whether a question requires a multiple-choice answer (e.g., a symbol) or a written answer.

[0008] Another aspect of the present invention includes a scoring support device accessible to a grader terminal having a display screen, wherein the grader terminal is a communication information terminal that converts an image of an answer sheet on which answers corresponding to the question text are expressed in handwriting into answer image data, and the scoring support device includes an acquisition unit that acquires the answer image data from the grader terminal, an acquisition unit that recognizes an answer frame on the answer image data on which the question type and scoring criteria are set, converts the portion expressed in the recognized answer frame into digital data in an answer format suitable for the question type, and displays the answer image data on the display screen, in which the digital data is expressed in the answer frame. The system switches between several scoring modes, each with a different processing load during scoring, according to the type of question. The system includes a scoring unit that processes the information represented by the digital data in accordance with the scoring criteria in the switched scoring mode. The conversion unit is equipped with a conversion engine that has learned the correct word forms, usages, and grammar, as well as the easily erroneous word forms, usages, and grammar. When the question type is a short-answer or long-answer descriptive answer format, the conversion engine replaces the descriptive portion of the answer box where the conversion evaluation value is lower than a predetermined value with words used in the question text or model answer for that question type. This is an answer sheet grading system.

[0009] Another aspect of the present invention is a program that causes a computer to operate as the scoring support device described above. [Effects of the Invention]

[0010] According to the present invention, the conversion unit recognizes the answer box with the question type set and converts it into digital data in an answer format that matches the recognized question type. For example, in the case of a question type that requires the entry of a number less than "10", the number "7" will not be mistakenly recognized as "17". Also, the number "3" will not be recognized as the hiragana character "ro" or the katakana character "yo". Furthermore, the scoring mode is switched according to the question type and scoring is performed according to the set scoring criteria, so it is possible to suppress an increase in processing load during the scoring process while preventing a decrease in scoring accuracy. [Brief explanation of the drawing]

[0011] [Figure 1] Schematic diagram of the AI-powered automatic scoring system. [Figure 2] A sequence diagram showing the overall procedure of tasks and processes performed between the grader's terminal and the grading support server. [Figure 3] A diagram illustrating the recognition process procedure. [Figure 4] This diagram illustrates the recognition results and replacement examples for descriptive answers where the recognition accuracy was below the threshold. (a) shows the question text, (b) shows the original handwritten answer written by the student in the answer field, (c) shows the candidate words that lowered the evaluation score, and (d) shows the answer data after replacement with the standard candidate words. [Figure 5] A diagram illustrating the procedure for setting the recognition frame. [Figure 6] An example diagram of answer image data with a recognition frame set. [Figure 7] A diagram illustrating the procedure for setting scoring criteria. [Figure 8] An example diagram of a detailed settings template. [Figure 9] An example diagram of a detailed answer image data. [Figure 10] A diagram illustrating the procedures for scoring, tabulation, and analysis. [Figure 11] An example diagram of the preview screen for an answer sheet displaying model answer data. [Figure 12] An example diagram of model answer data used during the evaluation content setting process. [Figure 13] An example of the scoring screen for "Question 4(2)" in Figure 12. [Figure 14] Exemplary diagram of detailed scoring results.

Embodiments for Carrying Out the Invention

[0012] Hereinafter, an example of an embodiment when the present invention is applied to an AI answer scoring system that can be used by scorers in a learning school to which a plurality of students belong will be described. [Overall Configuration] FIG. 1 is a diagram showing an example of the overall configuration of the AI answer scoring system 1. The AI answer scoring system 1 can be configured, for example, as SaaS (Software as a Service) specialized for answer scoring. The AI answer scoring system 1 includes, as a component, a scoring support server 30 to which a plurality of scorer terminals 10 can be respectively accessed via a network 20. The network 20 is a digital communication network, and in the present embodiment, it will be described as the Internet, but this is not limiting.

[0013] [Scorer Terminal] The scorer terminal 10 is, for example, a PC (Personal Computer) having a communication function. The PC is equipped with a data input / output interface for a keyboard, a mouse, a display, a printer, a scanner, a camera, a fixed storage such as an SSD (Solid State Drive), and a portable memory such as a USB (Universal Serial Bus) flash memory. The scanner is used when scanning the answer sheets described later and uploading the scanned data to the scoring support server 30. The same applies to the camera.

[0014] Note that the PC for uploading the scanned data and the PC for setting input for scoring may be separated. Instead of the PC for setting input, a tablet terminal or a PDA (Personal Digital Assistant) can also be used.

[0015] In this embodiment, the person who fills in the answers on the answer sheets to be graded is assumed to be a student attending a cram school. Furthermore, the person who operates the grader terminal 10 is assumed to be a grader who has been granted access rights by the grader support server 30, i.e., an instructor (teacher, lecturer) who presents the question text to multiple students, collects the papers, and grades them based on the model answers.

[0016] In the following explanation, scanned data of the question text may be referred to as "question text image data," scanned data of the answer sheet as "answer sheet image data," and scanned data of the model answer as "model answer image data." Additionally, digital data converted from "question text image data" may be referred to as "question text data," digital data converted from the portion of the answer sheet image data represented in the answer box may be referred to as "answer data," and digital data converted from "model answer image data" may be referred to as "model answer data."

[0017] <Scoring support server> The scoring support server 30 is a computer with server functionality that serves as an example of a scoring support device, and is connected to a DB unit (DB stands for "database," the same applies hereinafter) 40 built on a large-capacity SSD. The scoring support server 30 has the functions of an acquisition unit 31, a conversion unit 32, a scoring unit 33, and a control unit 34. These functions 31-34 and the DB unit 40 are realized by a processor with memory executing a scoring program.

[0018] The processor is typically one or more GPUs (Graphics Processing Units). GPUs are parallel-processing processors specialized for simultaneously processing large amounts of data, such as natural language processing. Inside the processor, there is DRAM (Dynamic Random Access Memory) and VRAM (Video Random Access Memory) available as working memory for processing. The scoring program is transferred from a program storage medium accessible by the processor to the VRAM and executed there.

[0019] The acquisition unit 31 is part of the front-end function that provides various UI (user interface) operating environments to the grader terminal 10. The front-end function displays, for example, an authentication screen, an upload request screen, a download request screen, a scoring progress screen, and an overlay screen for text, symbols, shapes, markings, comments, etc., on the display of the grader terminal 10. It also acquires various setting data and scoring data from manual scoring from the grader terminal 10 through such display screens and stores them in the DB unit 40. Alternatively, it may acquire various data stored in the DB unit 40 as appropriate and display them on the display of the grader terminal 10.

[0020] The conversion unit 32 recognizes the characters and other elements displayed in multiple frames shown in the problem statement data and answer image data (including model answer image data) acquired by the acquisition unit 31, and converts them into digital data. In this embodiment, AI-OCR (Artificial Intelligence-Optical Character Recognition) is used to realize this function. The AI-OCR is equipped with a conversion engine that has learned the shape features, usage, or grammar of correct characters and phrases, as well as the shape features, usage, or grammar of characters and phrases that are prone to errors. The conversion engine has learned the above correct characters and easily mistaken characters using, for example, a "few-shot" method, and its conversion accuracy is significantly higher than that of OCR using a "zero-shot" method.

[0021] The conversion engine can continuously recognize long texts of 200 characters or more. Furthermore, the conversion engine not only recognizes text by matching it against predefined patterns, but also infers pattern shape, words, word meanings, sentence meanings, and context. This makes it easy to identify portions of the converted characters, words, and sentences where the evaluation value representing the likelihood of conversion falls below a pre-set threshold. The conversion unit 32 can automatically replace parts that lower the evaluation score with characters, phrases, or sentences from the inference results, or prompt the grader to confirm or make the replacement through the grader terminal 10.

[0022] The scoring unit 33 is equipped with a natural language processing engine, for example, one that uses a large language model. The scoring unit 33 selectively switches to one of several scoring modes depending on the type of question, and in the switched mode, processes the information represented by the digital data according to the scoring criteria. In this embodiment, the multiple scoring modes are a simple automatic scoring mode that does not require natural language processing, an AI automatic scoring mode that requires natural language processing, and a manual scoring mode that waits for scoring input from a scorer operating the scorer terminal 10, but the system is not intended to be limited to these modes.

[0023] These scoring modes are automatically switched by the natural language processing engine according to the type of question it recognizes. Alternatively, the scorer can manually set (switch) the scoring mode for each type of question recognized by the conversion unit 32. Furthermore, in the AI ​​automatic scoring mode, if the scoring unit 33 determines that the information is ambiguous and requires judgment by the scorer, it can prompt the scorer terminal 10 to switch to manual scoring mode for that answer box.

[0024] In AI automatic scoring mode, in addition to automatically scoring the correctness of context and meaning, it can also analyze the calculation process to determine its correctness, analyze correctness trends, and automatically generate practice problems for review based on the analysis results. In this way, the scoring mode can be switched according to the type of problem, which prevents the natural language processing engine from becoming overloaded.

[0025] The control unit 33 controls communication with the grader terminal 10 and comprehensively controls the operation timing of the acquisition unit 31, conversion unit 32, and scoring unit 33. It also enables data reading / writing and updating processing in the DB unit 40.

[0026] The DB unit 40 includes an examination-related information DB 41, an answer history information DB 42, and a respondent information DB 43. The examination-related information DB 41 stores the examination questions and model answers acquired by the acquisition unit 31 from the grader terminal 10 and recognized by the conversion unit 32. The examination-related information DB 41 also contains, The system also stores information related to exams, such as learning assessment criteria set by the Ministry of Education, Culture, Sports, Science and Technology, school-specific learning assessment criteria, scoring standards, characteristics and trends of various qualification exams, and examples of characteristics of entrance exam questions for each school. This information can be useful for generating practice problems and for supplementary lessons for students based on scoring results.

[0027] The answer history information DB42 stores, for each student, the content of the exams they have taken (question text, model answers, etc.) and their scoring results from past exams. This information can be useful for understanding the degree of improvement in a student's academic ability and for comparing their scoring with that of other students on the same questions.

[0028] The respondent information DB43 stores various information about the respondent, such as the school they currently attend (elementary school, junior high school, high school, or are a ronin / preparing for university entrance exams), their desired school, the number of times they have taken the exam, whether they are a student of a cram school, and if so, the length of time they have been enrolled. This information is obtained when the cram school registers its users and can be useful in setting scoring criteria.

[0029] [Example of operation mode] Next, we will describe an example of the operation of the AI ​​answer scoring system 1 configured as described above. Figure 2 is a sequence diagram of the overall procedure of work and processing performed between the student, the grader, the grader terminal 10, and the scoring support server 30.

[0030] <Question> In Figure 2, the grader distributes to each student one or more question statements created to evaluate the students' learning assessment criteria, and an answer sheet for expressing their answers to those questions (Question, S101). The layout of the distributed answer sheets is at the student's discretion. However, the layout of the answer sheets distributed to all students will be the same. In this example, students are given answer sheets that display a space for their name, student number, question number for each question, answer spaces corresponding one-to-one with the question number (question type), a total score space, and a comment space. In other words, the answer sheet displays multiple answer spaces, including answer spaces for different question types. One of the questions is assumed to be a question requiring a written answer.

[0031] The grader writes the model answer (including "selection by symbol," "selection by character," "word," and "written statement") on the distributed answer sheet and scans it with the grader terminal 10 to create model answer image data (S102). The model answer should be written in a word processor, but it may also be handwritten. Note that the model answer image data may be created before distribution to the students.

[0032] Students write their name in the name box and their student number in the student number box on the answer sheet, then read the question, think about it, and write their answer by hand in the answer box corresponding to the question number. Typically, the time available to write the answer in the answer box is shorter than the time limit. Therefore, regardless of the type of question, the characters and other writing expressed in the answer box almost inevitably become messy cursive images, and students' writing habits (including differences in stroke order) tend to be easily revealed.

[0033] <Acquisition of answer sheet image data, etc.> The grader collects answer sheets from all students and scans each one on the grader terminal 10 to convert them into answer image data (S103). Then, all student answer image data is put into a single folder along with the question text image data and model answer image data, and a folder ID (ID is an abbreviation for identification information; the same applies hereinafter) is associated with it. After that, it is uploaded to the grading support server 30 (S104). The grading support server 30 acquires the uploaded answer image data, etc., and starts the recognition process (S105).

[0034] <Recognition processing and accuracy verification> The recognition process (S105) is performed, for example, by the conversion unit 32 in the procedure shown in Figure 3. That is, the conversion unit 32 converts the problem statement image data into problem statement data and saves it (S201). The save location is the examination-related information DB 41, but it may also be the work memory of the processor. The same applies to saving the answer data and model answer data hereafter.

[0035] Next, the conversion unit 32 recognizes the layout of the various frames displayed in the answer image data. It also recognizes and saves the answer frames and the corresponding question types based on the question text data (S202). The correspondence between the question type and the answer box is established by associating the hierarchically displayed question numbers. Alternatively, the question type may be inferred from the analysis results of the portion displayed in the answer box. Furthermore, each answer box is associated with, for example, a two-dimensional coordinate on the answer sheet, so that during grading, only that box can be cut out and placed alongside those of other students.

[0036] The conversion unit 32 then recognizes the model answer image data for each answer box and saves the recognition result as model answer data (S203). This example demonstrates how the type of question can be inferred from the character type displayed in the answer box (model answer data) using a shape analysis process. Specifically, if the model answer data is the number "3," the system determines that the question type requires a single digit to be displayed in the answer box. Then, when recognizing the portion of the answer box in subsequent student answer image data, the system prompts the conversion engine to avoid outputting results other than single digits (for example, the hiragana "ro," the katakana "yo," or the number "17"). The same applies to characters, symbols, phrases, clauses, and written text other than single digits. Furthermore, the conversion unit 32 may recognize and save not only the answer box and question type, but also the parts displayed in the name box, ID box, and other boxes.

[0037] The conversion unit 32 then recognizes each student's answer data for each answer box (S204), and at that time, it checks for factors that affect the evaluation value (likelihood) of recognition accuracy, regardless of the type of question (S205). These factors are called the "first factors." The first factors include, for example, the original image having too little contrast, blurring during paper scanning, partial loss (fading) of characters, and remaining overwriting or corrections of characters, and are technical elements that induce misrecognition regardless of the type of question or the student's handwriting style.

[0038] The conversion unit 32 also checks for factors that are specific to the type of problem and affect the evaluation value (likelihood) of recognition accuracy (S206). These factors are called "second factors" to distinguish them from the first factors. The second factors include, for example, characters with special shapes (the presence or absence of shape features such as closed ring vertical lines, intersections, and branches), and are factors in which the cause of misrecognition strongly depends on the type of problem and the student's handwriting habits.

[0039] After confirming the first and second elements that induce such misconversions, the conversion unit 32 checks the recognition accuracy (evaluation value) for each answer box (S207). The evaluation value is output by the conversion engine for each character type, phrase, segment, or sentence, so it is sufficient to check that. If the evaluation value is above the threshold for all parts (S208:Y), the conversion unit 32 considers the converted part of the answer box to be the student's answer data and saves that answer data associated with the student's name (S214).

[0040] If there is a part in S208 with accuracy below the threshold (S208:N), the conversion unit 32 determines whether the first element is lowering the evaluation value (S209). If it is the first element (S209:Y), the conversion unit 32 prompts the grader to switch to manual grading mode (S210) and proceeds to the process in S214. In other words, the conversion unit 32 associates a message with the answer frame or answer data indicating that manual grading should be set rather than automatic grading based on the results of the conversion, and completes the confirmation of the recognition accuracy for that answer data.

[0041] In S209, the second factor is lowering the evaluation score, and if the question type does not require a written answer (S209:N, S211:N), for example, if a question type requires a single digit but the student recognizes a different digit, it is highly likely that the student has made a mistake, and the process immediately moves to S214.

[0042] On the other hand, in S209, if the second element is lowering the evaluation score, and the question type is one that requires a written answer (S209:N, S211:Y), the conversion unit 32 identifies the word or phrase that is lowering the evaluation score (S212). In automatic recognition of written answers in Japanese using natural language processing, one of the factors that lowers the evaluation score is an error in morphological analysis, but in reality, the use of synonyms (expressing automobile as car, vehicle, or vehicle) and errors in the use of okurigana (such as arawasu, hyōsu, hyōwasu) can also significantly lower the evaluation score.

[0043] The conversion unit 32 replaces such low-evaluation words or phrases with correct words or phrases (S213). For convenience, the candidates considered to be correct words or phrases are called "criteria candidates." While it is possible to select criterion candidates using natural language processing (meaning and contextual analysis) with a large dictionary, applying natural language processing to all written answers from multiple students would result in an excessive processing load. Therefore, in this example, the processing load is reduced by selecting criterion candidates from words or phrases present in the question text data or model answer data.

[0044] After the substitution, the recognition accuracy of the student's answer box is checked again (S207). If the evaluation value does not improve after the substitution, the criteria candidate may be changed and the processes in S212 and S213 may be performed, or a message indicating that manual scoring should be set rather than automatic scoring based on the results of the conversion by the conversion unit 32 may be linked to the answer box or answer data, and the confirmation of the recognition accuracy of that answer data may be completed.

[0045] The substitution process will be explained in detail with reference to Figure 4. Figure 4 is an explanatory diagram showing the recognition results and substitution examples for descriptive answers where the recognition accuracy was below the threshold, where (a) is the question text data, (b) is the original handwritten answer written by the student in the answer field, (c) is the word or phrase that lowered the evaluation score, and (d) is the answer data after substitution with the standard candidate.

[0046] The problem statement data in Figure 4(a) asks students to "explain, in approximately 120-160 characters, the heavy snowfall in Fukui Prefecture and the countermeasures being taken, referring to data on the number of cars owned by residents." As shown in Figure 4(b), the original handwritten answers of the students have low contrast with the background and are written sloppily. As a result, as shown in Figure 4(c), the evaluation values ​​for the parts of the recognition results "▲Fukui Prefecture" w1 and "total number of cars owned" w2 were lowered. This is because "▲Fukui Prefecture" does not exist in Japan, and "total number of cars owned" is not a normal usage of the Japanese language. However, it is strongly inferred that "▲Fukui Prefecture" is a misrecognition of "Fukui Prefecture" as it appears in the problem statement, and "total number of cars owned" is a rephrasing of "number of cars owned" as it appears in the problem statement. Therefore, as shown in Figure 4(d), the conversion unit 32 replaces "▲Fukui Prefecture" w1 with "Fukui Prefecture" c1 and "Number of Owned Units" w2 with "Number of Owned Units" c2, and then checks the recognition accuracy (evaluation value) of S205 again.

[0047] The conversion unit 32 highlights the replaced parts c1 and c2. The highlighting can be, for example, by underlining or highlighting. The conversion unit 32 also displays the digital data together with the original answer data. This allows for verification, for example during scoring, whether the recognition accuracy or replacement by the conversion unit 32 was appropriate. Note that Figure 4 shows an example where a candidate criterion exists in the problem statement, but the candidate criterion may also be identified from the model answer data.

[0048] The conversion unit 32 repeats the processing from S205 to S209 for the recognition results of the answer sheet image data of other students (S210:N). Then, when the recognition accuracy check for all students is completed (S210:Y), the conversion unit 32 finishes the recognition process.

[0049] <Recognition frame setting process> After completing the recognition process (S105), the scoring unit 33 performs a recognition frame setting process for each answer sheet image data (S106). The recognition frame is primarily a frame that allows the scoring unit 33 to recognize which part of the answer sheet image data corresponds to which frame. Since the model answer data has the same layout as the answer sheet image data, in this example, the recognition frame setting process is performed on the model answer image data, and the result is then applied to the student's answer sheet image data.

[0050] Figure 5 is a diagram illustrating the procedure for setting the recognition frame. Here, the manual setting procedure performed by the grader through the operation of the grader terminal 10 is explained, but the scoring unit 33 can also perform this automatically by referring to the problem statement data and model answer data.

[0051] The scoring unit 33 displays an image file of one student's answer sheet, such as an image file of a model answer, on the display screen (S301). It then sets the name, ID, answer, and score fields (S302), and sets the student's name (or the instructor's name in the case of model answer data) in the name field, the student's attendance number (or the instructor's number in the case of model answer data) in the ID field, the hierarchical question number in the answer field, and the score in the score field. If there are other fields, it sets the content corresponding to those fields and finishes the process (S303).

[0052] Figure 6 is an example of answer sheet image data with recognition frames set. In the example shown, the student's name is set in the name frame 310 and the student number is set in the ID frame 311. In addition, the question numbers (1)321, (2)322, (3)323... (36)336, (37)337, and (38)338 are set sequentially in the answer frames 320. The set information is saved and referenced when grading the student's answer data.

[0053] <Scoring content setting process> Returning to Figure 2, after completing the recognition frame setting process (S106), the scoring unit 33, in cooperation with the scorer terminal 10, executes the scoring content setting process in the procedure illustrated in Figure 7 (S107). Specifically, the scoring unit 33 sets the scoring content in detail for each question type corresponding to the answer frame within the recognition frame of the model answer image data (which may also be answer sheet image data) (S401). It also sets the point allocation and the content of the correct answer for each answer frame (S402) and saves it. The content of the correct answer is the model answer data for that answer frame.

[0054] The details of the scoring content set in S401 may include, for example, the question type, the scoring mode and scoring criteria that are appropriate for that question type. In this example, the details of the scoring content can be selected for each question number from a pre-created detailed setting template. Figure 8 is an example of a detailed settings template. This detailed settings template can be displayed as a speech bubble window on the display screen of the grader terminal 10, while the answer sheet image data with the answer boxes set is displayed. In the example shown, the question type can be set to one of the following for each question type: "Selection by Symbol," "Short Answer," "Long Answer," "Mathematical Formula," "Kanji Writing," or "Manual Grading."

[0055] "Symbol selection" questions are a type of question where the correct answer is chosen from a set of options. For this type of question, scoring is performed by determining whether the selected symbol matches the symbol of the model answer. In other words, complex processing such as contextual analysis is unnecessary, and scoring can be done without natural language processing. Therefore, "simplified automatic scoring" is performed for "symbol selection" questions to suppress an increase in processing load.

[0056] "Short-answer questions" are short-answer questions requiring the writing of words or phrases, while "long-answer questions" are long-answer questions requiring the writing of complete sentences. Both require natural language processing such as text analysis. Therefore, "AI automatic scoring," which requires natural language processing, is performed for both "short-answer questions" and "long-answer questions." In addition, "AI automatic scoring" can display the scoring results on the screen of the scorer's terminal 10 as needed to prompt confirmation. This is because if there are ambiguous parts in the answer, the scorer may need to make a judgment.

[0057] "Mathematical formulas" are problems that involve mathematical formulas and calculations, and do not require complex processing such as contextual analysis. Therefore, the "simplified automatic scoring" described above is performed.

[0058] "Kanji writing" is a type of problem that requires the ability to judge the precise shape of kanji characters, making accurate scoring difficult even with AI. Therefore, "manual scoring" is performed, where a human grader scores the answers. "Manual scoring" may also be applied to problems that do not belong to pre-defined categories (such as "symbol selection" or "kanji writing").

[0059] In the example in Figure 8, "Long written response" is set for question XX. In other words, for question XX, the scoring unit 33 will perform "AI automatic scoring". The scoring content can be set manually by the scorer, but the scoring content can also be set automatically by the generating AI by loading model answer data into the generating AI.

[0060] The detailed settings template also defines the scoring criteria. The scoring criteria can be set by the grader according to the teaching objectives. However, as shown in the lower part of Figure 8, the scoring unit 33 can also automatically generate the scoring criteria using the "AI scoring criteria generation function". The example shown in Figure 8 shows that, using the "AI scoring criteria generation function", if a student describes specific actions related to the given theme that are in line with that theme, 4 points will be added to the score as a bonus element. It also shows that, regarding logical thinking ability, if the logical development that forms the basis of the answer is clear and consistent, 3 points will be added to the score. The content of each bonus element can be deleted or added as appropriate through the grader terminal 10. Figure 9 is an example of detailed setting answer sheet image data. In the example shown, 2 to 4 points are allocated to each answer box.

[0061] For question types where "Manual Scoring" is set in Figure 8, the system switches to manual scoring mode. In manual scoring mode, if necessary, the system can extract only the answer frames for that question type from all answer image data and display them in a list along with the answer frames of the model answers. In this case, the portion expressed in the answer frame may be converted (partially replaced) by the conversion unit 32, or the conversion by the conversion unit 32 may be avoided.

[0062] Furthermore, even when the AI ​​automatic scoring mode is set, the scoring unit 32 may prompt the scorer to switch to manual scoring mode when scoring answer boxes for question types in which the information converted by the conversion unit 32 is judged to be ambiguous information requiring human judgment. This prevents a decrease in scoring accuracy.

[0063] <Scoring, tabulation, and analysis processing> Returning to Figure 2, after completing the scoring content setting process (S107), the scoring unit 33 performs the scoring, aggregation, and analysis process (S108). Specifically, this process is performed according to the procedure shown in Figure 10. That is, the scoring unit 33 acquires model answer data (S501) and prepares to score each student's answer data. Then, the scoring unit 33 acquires answer sheet image data with scoring content set in the recognized answer frame (S502). Then, using the model answer data as a base, it sequentially scores the answer frame portion of the answer sheet image data according to the scoring criteria and saves it (S503). This process is repeated for all students' answer sheet image data (S504:N). When scoring of all students' answer sheet image data is completed (S504:Y), the scoring unit 33 aggregates the scoring results for all students for each question (S505).

[0064] The scoring unit 33 then performs an analysis of the answer trends of all students based on the aggregated results (S506). From the results of the answer trend analysis, it uses natural language processing to identify each student's tendency to make mistakes, areas to review, and areas for improvement, and presents them to the scorer terminal 10 (S507). After presenting the results of the scoring, aggregation, and analysis processing to the scorer terminal 10 (S507), the scoring unit 33 associates the scoring results with the student ID fields read from the answer fields, and under the control of the control unit 34, stores each student's answers and scoring results in the answer history information DB 42. By storing the answer history in this way, scorers can check the history of each student's answers and scoring results.

[0065] The control unit 34 provides after-sales support, for example, through a chatbot function (S109). The chatbot function receives questions as text data and provides answers in text format.

[0066] Returning to Figure 2, the grader terminal 10 presents the grading results to each student. It also performs overall evaluations of each student, creates instructional plans, and generates new problem statements (S110). In addition, it provides individual instruction as needed (S111). Students check the grading results for their submitted answers and receive individual instruction as needed (S112).

[0067] [Examples of other operational methods] The scoring support server 30 can perform partial replacement processing of recognition results by the conversion unit 32 and scoring processing by the scoring unit 33 even for handwritten answer sheet image data where answers in foreign languages ​​are mixed with answers in Japanese. The following describes an example of scoring processing for an answer sheet for the English subject, which is one of the subjects taught at cram schools.

[0068] Figure 11 is an example of a preview screen of an answer sheet displaying model answer data. This preview screen shows the results of the recognition processing (S105) and recognition frame setting processing (S106) for characters entered by the instructor (grader) on an answer sheet 50 with the same layout as the student's answer sheet. This preview screen can be displayed on the display of the grader terminal 10 that is accessing the system. In the illustrated example, the answer sheet 50 contains a mixture of answer frames corresponding to multiple question types, each represented as digital data, including hierarchically listed question numbers, correct single-digit numbers, correct single-character katakana, correct English words, correct Japanese phrases, and correct short and long English sentences separated by periods.

[0069] The model answer data is presented in printed form in the answer field. Therefore, it is converted into answer data almost exactly as it was originally written. For example, the correct English short sentences 511, 513, and 514 are displayed along with the converted digital data in the conversion unit 32. This allows the grader to confirm that the converted model answer data is correct.

[0070] Figure 12 is an example of model answer data during the evaluation content setting process. Here, the point allocations 521, 522, 523, and 524 for each question type ("Question 1" to "Question 4") are set all at once, and then the detailed content for the answer box 530 of "Question 2" is set through the perspective setting window 531.

[0071] The perspective setting window 531 allows you to set scoring elements based on the learning evaluation perspectives defined by the Ministry of Education, Culture, Sports, Science and Technology. In the example shown, "Knowledge and Skills" 532 is set as the perspective. However, the content of the perspective setting window 531 can also be changed to reflect the unique scoring criteria of the school taking the exam.

[0072] Figure 13 is an example of the scoring screen for "Question 4(2)" in Figure 12. Each is a handwritten English text with a distinctive writing style. By extracting and displaying the original texts of multiple students' answers from the answer sheet image data, it is possible to understand the different answer tendencies among students. In the figure, the "Scoring Complete" tag 611 becomes active when scoring of the answer sheet image data is complete. When the "Go to Answer Image" tag 612 is operated, the answer data of the answer box converted by the conversion unit 32 or the content of the original answer text expressed in that answer box is displayed. When the "Display Answer List" tag 613 is operated, the answer data or original text of all students' answer boxes is displayed in a list. In the example shown, the state in which the "Go to Answer Image" tag 612 has been operated is shown.

[0073] At the top of Figure 13, the question number and model answer data 614 are displayed. Below that, the same question number as model answer data 614 and the original student answers (handwritten) 615, 616, and 617 are displayed. The graded results 6151, 6161, and 6171 are shown for the original student answers. Graders can grade while checking model answer data 614 and each student's original answer 615, 616, and 617, thus reducing variations in grading.

[0074] Figure 14 is an example of detailed scoring results. The scoring screen from Figure 13 is used as a background image 710 with reduced brightness, and the detailed scoring results for student number 6 are superimposed in the detailed scoring results window 712 on top of it. This detailed scoring results window 712 shows that manual verification is required, the correct English sentence, the student's answer data, the total score, and the score results for each scoring criterion. To move to the detailed scoring results window of the previous student, the user operates the back button 714, and to move to the detailed scoring results window of the next student, the user operates the forward button 716.

[0075] In this embodiment, the conversion unit 32 of the scoring support server 30 recognizes the answer frame from the answer sheet image data, where the question type (answer format) and scoring criteria are set, and converts the portion expressed in the recognized answer frame into digital data in an answer format that matches the question type. In other words, since the answer frame portion is recognized after interpreting the question type (context), misconversion is prevented even if the characters are handwritten. For example, if the question type requires a numerical answer (e.g., "3"), non-numerical characters such as "з" (ze: Cyrillic letter) and "ろ" (hiragana) are excluded from the conversion candidates. Also, if the answer format is an English word or sentence, misconversion of "I" (alphabet) to "エ" (katakana) is prevented. Therefore, even if the answer sheet contains a mixture of different question types, the answer data can be converted correctly.

[0076] In this embodiment, the scoring unit 33 also switches between a simplified automatic scoring mode, an AI automatic scoring mode, and a manual scoring mode depending on the type of question, and processes the information represented by the digital data according to the scoring criteria in the selected mode. Therefore, for answer types of questions that can be scored without requiring natural language processing, such as "symbol selection," the processing load can be reduced by setting the simplified automatic scoring mode. Conversely, if there are ambiguous parts in the answer, setting the manual scoring mode makes it possible to score more accurately without excessively increasing the processing load.

[0077] [Differentiation] In this embodiment, a method has been described in which problem statement data, model answer data, and answer image data are uploaded from the grader terminal 10 to the grading support server 30 via the network 20. However, this data may also be acquired by the grading support server 30 via a portable recording medium such as a USB memory stick. Furthermore, the grading support server 30 may be equipped with input / output devices such as a display and a keyboard, and various settings and grading may be performed through these input / output devices.

[0078] In this embodiment, an example of grading answer sheets for an English subject as a foreign language subject has been described, but the same method can be applied to subjects in Chinese or other languages.

[0079] Furthermore, while this embodiment describes an example of application to an AI answer scoring system usable by graders at cram schools, it can be similarly applied to answer sheets for qualification exams, certification exams, entrance exams, and other examinations.

Claims

1. An acquisition unit that acquires answer image data in which the answer corresponding to the problem statement is expressed, A conversion unit recognizes an answer box from the answer image data in which the question type and scoring criteria have been set, and converts the portion represented in the recognized answer box into digital data in an answer format that conforms to the question type. The system includes a scoring unit that switches between multiple scoring modes, each with a different processing load during scoring, according to the type of problem, and processes the information represented by the digital data in the switched scoring mode according to the scoring criteria, The conversion unit is equipped with a conversion engine that has learned the correct word forms, usages, and grammar, as well as the easily erroneous word forms, usages, and grammar. The conversion engine, when the question type is a short-answer or long-answer descriptive answer format, replaces the descriptive portion of the answer box where the conversion evaluation value is lower than a predetermined value with words used in the question text or model answer of that question type, thus providing a scoring support device.

2. The scoring unit is equipped with a natural language processing engine, and the natural language processing engine switches between one of the following modes depending on the type of question: a simplified automatic scoring mode that does not require contextual analysis, an AI automatic scoring mode that requires natural language processing, or a manual scoring mode that waits for scoring input from the operator of the scoring terminal, as described in claim 1.

3. The conversion unit highlights the replaced portion of the digital data. The scoring support device according to claim 1.

4. The conversion unit displays the digital data together with the original answer data. The scoring support device according to claim 1.

5. The answer image data displays multiple answer boxes, including answer boxes for different question types, and the scoring unit enables switching between the simplified automatic scoring mode, the AI ​​automatic scoring mode, and the manual scoring mode for each answer box. The scoring support device according to claim 2.

6. The answer box can be supplementarily displayed with the scoring results for each scoring criterion and the reasons for those results. The scoring support device according to claim 5.

7. The scoring unit, in the AI ​​automatic scoring mode, will prompt the scorer to switch to the manual scoring mode when scoring answer boxes for question types that it has determined to have ambiguous information requiring human judgment. The scoring support device according to claim 2.

8. The scoring unit extracts only the answer frames for the relevant question type from all the answer sheet image data and displays them in a list along with the answer frames that show the correct answers. A scoring support device according to any one of claims 1 to 6.

9. If the aforementioned question type requires the respondent to provide a written response in a foreign language, and the manual scoring mode is set for that question type, The acquisition unit also acquires model answer data, which represents the model answer, along with the image data of the answers submitted by multiple respondents to the same problem statement. The conversion unit converts the model answer data into the answer frame and the digital data, and displays the answer frame in which the model answer is expressed, and a plurality of answer frames in which the answers of the respondents in the answer image data are expressed, side by side on a predetermined display screen. The scoring support device according to claim 2.

10. It has a scoring support device that is accessible to the scorer's terminal which has a display screen, The aforementioned grader terminal is a communication information terminal that converts an image of an answer sheet in which the answer corresponding to the question is written by hand into answer image data. The scoring support device is, An acquisition unit that acquires the answer sheet image data from the grader's terminal, A conversion unit recognizes an answer frame from the answer image data in which the question type and scoring criteria have been set, converts the portion represented in the recognized answer frame into digital data in an answer format that conforms to the question type, and displays the digital data representing the answer image data on the display screen. The system includes a scoring unit that switches between multiple scoring modes, each with a different processing load during scoring, according to the type of problem, and processes the information represented by the digital data in the switched scoring mode according to the scoring criteria, The conversion unit is equipped with a conversion engine that has learned the correct word forms, usages, and grammar, as well as the easily erroneous word forms, usages, and grammar. When the question type is a short-answer or long-answer descriptive answer format, the conversion engine replaces any descriptive portion of the answer box where the conversion evaluation value is lower than a predetermined value with words or phrases used in the question text or model answer for that question type. Answer sheet grading system.

11. A program that causes a computer to operate as a scoring support device as described in claim 1.