Intelligent correction method, device, equipment, medium and product
By combining structured processing of the answer area in pure card scenarios with a multimodal large language model, the problems of accuracy and depth feedback in intelligent grading in pure card scenarios are solved, achieving more efficient grading results and reducing the workload of teachers.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-24
AI Technical Summary
Existing intelligent grading technologies are not effective in pure card-based scenarios. OCR modules have large recognition errors, rule-based grading has low accuracy, and multi-mode solutions lack question stem information, resulting in poor performance and failing to meet teachers' grading requirements.
By acquiring the initial image and standard answer of the pure card answering area, the image and text recognition model is used for structured processing to extract triple features. The reference questions are retrieved from the question bank by combining a multi-path recall strategy, and the target grading results are generated by a multimodal large language model, taking into account the answering logic and semantics.
It improves the accuracy and in-depth feedback capabilities of intelligent grading, reduces the burden on teachers, and provides more targeted and interpretive grading results.
Smart Images

Figure CN121921804A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to an intelligent correction method, apparatus, device, medium, and product. Background Technology
[0002] With the rise of large-scale models, more and more industries are being empowered by AI, and the education sector is also actively exploring the practical applications of large-scale models, with exam marking being a particularly important area. In China's current basic education system, exams remain the primary way to assess students' knowledge acquisition. Under these circumstances, teachers bear a heavy workload in marking papers. AI can completely solve this problem through artificial intelligence technology, rather than requiring teachers to spend a significant amount of time and energy. Therefore, utilizing computer-assisted marking to reduce the workload of manual marking is of great significance to the teaching process.
[0003] Currently, intelligent grading mainly involves inputting images of students' answers into the system via photo or scan. The system then identifies and grades the students' answers and test information, and finally outputs the grading results for the current student.
[0004] However, the above method has poor grading results. Summary of the Invention
[0005] The main purpose of this application is to provide an intelligent correction method, device, equipment, medium and product that can improve the intelligent correction effect.
[0006] To achieve the above objectives, firstly, this application provides an intelligent grading method, comprising: Obtain the initial image and corresponding standard answer for the current pure card answering area; The initial image is processed in a structured manner to obtain the target image corresponding to the current pure card answering area; Based on the standard answer, obtain reference test questions that match the standard answer; Based on the target image, standard answer, and reference test questions, a target grading result matching the pure card answer content is generated through a pre-set multimodal large language model.
[0007] In one embodiment, the initial image is subjected to structured processing to obtain the target image corresponding to the current card-based answering area, including: The initial image is structured using a pre-defined image recognition model to obtain an empty granular image corresponding to each answer region in the initial image. Each empty granular image corresponding to each answer region contains the corresponding answer text. The empty granularity images corresponding to each answer region in the initial image are summarized to obtain all empty granularity images corresponding to the initial image. Use all empty granularity images corresponding to the initial image as the target image.
[0008] In one embodiment, obtaining reference test questions that match the standard answer based on the standard answer includes: Feature extraction is performed on the standard answer to obtain the triplet features of the standard answer; the triplet features include invariant components, variant components, and complementary components; The triplet features, as well as the subject, question type, and grade level information corresponding to the pure card test questions, are used as retrieval features; Based on retrieval features, a multi-path recall strategy is used to retrieve reference questions that match the standard answers from a pre-built question bank.
[0009] In one embodiment, the reference test questions include a first reference test question and a second reference test question; Based on retrieval features, a multi-path recall strategy is employed to retrieve reference questions matching the standard answers from a pre-built question bank, including: Based on retrieval features, a multi-path recall strategy is used to retrieve items from a pre-built question bank to obtain a recall question set. Different weights are assigned to different recall paths for weighted ranking. First and second reference questions are determined from the recalled question set. The first reference question is the question with the highest weighted score, and the second reference question is the question with the lowest weighted score. The weighted score is proportional to the matching degree of the standard answer.
[0010] In one embodiment, based on the target image, standard answer, and reference test questions, a target grading result matching the pure card answer content is generated using a preset multimodal large language model, including: Obtain the question stems of the first and second reference questions; The grading prompts, which include the target image, the standard answer, the stem of the first reference question, and the stem of the second reference question, are input into a preset multimodal large language model to obtain the first probability distribution corresponding to the first reference question and the second probability distribution corresponding to the second reference question. The probability distribution is used to characterize the grading result category for the reference question. Based on the first probability distribution and the second probability distribution, the target grading results that match the content of the pure card answers are obtained.
[0011] In one embodiment, based on a first probability distribution and a second probability distribution, obtaining a target grading result that matches the content of the pure card response includes: Obtain the probability of the grading result category from the first probability distribution and the second probability distribution; The probability difference between the grading result categories is obtained by subtracting the probability of the grading result category from the probability of the first probability distribution and the probability of the grading result category from the probability of the second probability distribution. Select the category of the grading result with the largest probability difference from the probability differences of the grading result categories as the target grading result.
[0012] In one embodiment, the method further includes: Initial grading results are obtained based on rule-based grading strategies; If the initial grading results meet the preset grading conditions, the step of obtaining reference test questions that match the standard answer will be triggered.
[0013] Secondly, embodiments of this application provide an intelligent correction device, comprising: The acquisition module is used to acquire the initial image and the corresponding standard answer for the current pure card answering area; The processing module is used to perform structured processing on the initial image to obtain the target image corresponding to the current pure card answering area; The retrieval module is used to retrieve reference test questions that match the standard answer. The grading module is used to generate target grading results that match the pure card answer content based on the target image, standard answer, and reference test questions through a preset multimodal large language model.
[0014] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the methods described above.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the methods described above.
[0016] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the methods described above.
[0017] This application provides an intelligent grading method, apparatus, device, medium, and product, comprising: first, acquiring an initial image corresponding to the current pure card answer area and the corresponding standard answer; then, performing structured processing on the initial image to obtain a target image corresponding to the current pure card answer area; and then, based on the standard answer, acquiring reference questions that match the standard answer. Based on the target image, standard answer, and reference questions, a target grading result matching the pure card answer content is generated through a preset multimodal large language model. This application retrieves reference questions matching the standard answer, providing richer context and scoring criteria for the grading process. It can not only determine the correctness of the answer but also understand the logical structure, knowledge application methods, and common error types of the answer, thus providing more in-depth feedback. Furthermore, relying on the comprehensive reasoning capabilities of the multimodal large language model, this application can simultaneously process text, images, and even implicit answering thought processes, generating targeted and highly interpretable grading results. This not only reduces the repetitive grading burden on teachers but also improves the effectiveness of intelligent grading. Attached Figure Description
[0018] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application. In the drawings: Figure 1 This is a schematic diagram of a pure card response image provided in an embodiment of this application; Figure 2 This is a flowchart illustrating an intelligent grading method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating an intelligent grading method provided in another embodiment of this application; Figure 4 This is a flowchart illustrating an intelligent grading method provided in another embodiment of this application; Figure 5 This is a schematic diagram of the structure of an intelligent correction device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the computer device provided in the embodiments of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0020] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein.
[0021] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0022] It should be understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0023] It should be understood that in this application, "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product or device.
[0024] It should be understood that in this application, "multiple" refers to two or more. "And / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, "and / or B" can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "Contains A, B, and C", "Contains A, B, and C" means that all three A, B, and C are contained; "Contains A, B, or C" means that one of A, B, and C is contained; "Contains A, B, and / or C" means that any one, two, or three of A, B, and C are contained.
[0025] It should be understood that in this application, "B corresponding to A", "B corresponding to A", "A corresponds to B", or "B corresponds to A" means that B is associated with A, and B can be determined based on A. Determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information. Matching A and B is defined as a similarity between A and B that is greater than or equal to a preset threshold.
[0026] Depending on the context, "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection."
[0027] The data involved in this application may be data authorized by the tester or fully authorized by all parties. The collection, dissemination, and use of the data shall comply with the relevant laws, regulations and standards of the relevant countries and regions. The implementation methods / executives of this application may be combined with each other.
[0028] The technical solutions of this application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0029] The present application will now be described in conjunction with the accompanying drawings and specific embodiments.
[0030] Grading is a crucial part of a teacher's daily work, primarily involving correcting student assignments and tests. From elementary to high school, teachers are burdened by a heavy workload of grading, resulting in less time for personalized tutoring. Therefore, reducing teachers' grading workload is a pressing need. With the development of artificial intelligence, various grading solutions have emerged in the education field. However, grading in purely card-based scenarios remains a challenge and a weakness of many solutions. The reason is simple: card-based scenarios provide too little information, leading to consistently poor grading effectiveness.
[0031] With the rise of large-scale models, more and more industries are being empowered by AI, and the education sector is also actively exploring the practical applications of large-scale models, with exam marking being a particularly important area. In China's current basic education system, exams remain the primary way to assess students' knowledge acquisition. Under these circumstances, teachers bear a heavy workload in marking papers. AI can completely solve this problem through artificial intelligence technology, rather than requiring teachers to spend a significant amount of time and energy. Therefore, utilizing computer-assisted marking to reduce the workload of manual marking is of great significance to the teaching process.
[0032] Based on the rapidly developing artificial intelligence technology, a systematic solution for grading student assignments and exams is entirely feasible: it involves taking photos or scanning images of students' answers. Figure 1 The system inputs the student's answers and test information, identifies and grades the student's answers and test information, and finally outputs the graded results for the current student.
[0033] For test grading scenarios, the main approaches are rule-based grading, model-based grading, and a combination of both, such as... Figure 2 As shown, the details are as follows: The rule-based grading scheme is as follows: After allowing a certain degree of error tolerance, the standard answer and the user's answer are hard-matched. Error tolerance is necessary because handwritten answers may be messy or have scribbling, which can lead to errors in OCR recognition of certain characters, such as recognizing 'n' as 'h' or 'v' as 'r'. The error tolerance process involves replacing the error-prone characters in the word with candidate characters (e.g., 'stard' -> 'stand'), forming a new word that is also considered a student's answer (both 'stard' and 'stand' are considered answers). If any answer matches the standard answer, it is considered correct.
[0034] Model-based grading typically employs a two-stage approach: OCR recognition of the answers, followed by grading based on the standard answer and the student's responses, with the model then outputting the grading results.
[0035] Of course, there are also solutions that use image input for direct grading (multi-model model). This is generally used for scenarios where the question stem contains images. By using a multi-model approach, more information about the question stem can be obtained, thereby improving the grading effect.
[0036] The disadvantages of existing technical solutions are as follows: For the two-stage grading scheme, the OCR module suffers from cascading errors: recognition of student answers in pure card scenarios exhibits significant errors due to the denser layout of the cards, such as in English fill-in-the-blank questions. Combined with the effectiveness of rule-based grading on subjective questions, whether the overall grading results can meet teachers' requirements is also a major risk.
[0037] The accuracy of rule-based grading needs improvement: Aside from errors introduced by OCR recognition, rule-based grading can only address a limited number of issues. Firstly, while it can grade objective questions, it struggles with more subjective questions such as fill-in-the-blank and sentence translation, where many answers may not be unique. In these cases, rule matching alone proves insufficient. Furthermore, various error-tolerance mechanisms can also lead to incorrect or incorrect judgments.
[0038] Model-based grading is ineffective because, lacking question stem information, the biggest advantage of the model over rules is its ability to compare the written answer with the standard answer based on the semantics of the response. However, semantic comparison is only applicable to humanities subjects; it is no longer suitable for science subjects. Furthermore, semantic equality between the standard and written answers does not necessarily guarantee a correct grade and can vary depending on different requirements. For example, stricter schools may require strict consistency between the standard and written answers, while more lenient schools may accept semantic equality as correct.
[0039] In summary, current multi-model solutions are all geared towards the question stem, essentially aiming to better extract information from the question stem and the answer. For example, for images in the question stem, multi-model solutions can better extract information from the images. However, for scenarios purely based on card-based questions without a question stem, the effectiveness of multi-model solutions will be significantly reduced.
[0040] To address the aforementioned issues, this application proposes an intelligent grading method.
[0041] Please see Figure 3 , Figure 3 This is a flowchart illustrating an intelligent grading method provided in an embodiment of this application. Figure 3 As shown, it includes the following steps: Step S301: Obtain the initial image and the corresponding standard answer for the current pure card answering area.
[0042] In the current pure answer sheet scenario, to perform intelligent grading, it is first necessary to obtain the initial image and its standard answer corresponding to the current answer sheet's answer area. The initial image comes from the student's completed answer sheet, which only contains the answer area and question number identifiers, but does not contain the question stem.
[0043] like Figure 4 As shown, in an actual exam scenario, the answer sheet is captured by a high-speed scanner or document camera to generate an initial image. This image contains all the answer areas and is also called a pure answer sheet image. Simultaneously, based on the subject, question type, grade level, and specific question number and blank index of the current exam question, the standard answer text corresponding to that answer area is matched and extracted from a pre-stored standard answer database. The standard answers are usually pre-entered into the system by the exam setter and are strictly bound to the question structure (such as question number and blank order) to ensure accurate retrieval during grading. At this point, the initial image of the area to be graded and its corresponding standard answer have been obtained simultaneously.
[0044] Step S302: Perform structured processing on the initial image to obtain the target image corresponding to the current pure card answering area.
[0045] The initial image is structured to obtain the target image corresponding to the current card answering area. First, the initial image is structured using a preset image recognition model to obtain an empty granular image corresponding to each answering area in the initial image. Each empty granular image corresponding to the answering area contains the corresponding answering text. Then, the empty granular images corresponding to each answering area in the initial image are summarized to obtain all empty granular images corresponding to the initial image. Finally, all empty granular images corresponding to the initial image are used as the target image.
[0046] Specifically, the initial image of the answer sheet to be processed (i.e., a scanned image of the entire answer sheet) and a pre-trained image recognition model (e.g., a large-scale OCR model based on deep learning) are first loaded. This model has the ability to analyze page layout, detect text, and recognize it. During processing, the model first performs a global layout analysis on the initial image to identify the position and boundaries of all answer areas. Each answer area usually corresponds to a specific blank under a question number, such as "Question 3, Blank 2". Based on the detected bounding boxes, the model accurately crops each answer area from the original image to form an independent empty-granular image (e.g., an image containing only the writing area of that blank). At the same time, the model performs text recognition on the handwritten content within this area, outputs the recognized text result, and associates and stores the text with the corresponding empty-granular image.
[0047] Subsequently, all detected answer areas are traversed, and their corresponding empty granularity images and recognized text are structurally summarized according to question number and empty position order to form a complete set of empty granularity images corresponding to the initial image. This set contains both the visual image information of each empty position and the preliminary recognized text information, together constituting the "target image" data required for subsequent grading.
[0048] It should be noted that if a region cannot be effectively segmented or recognized due to image blurring, writing exceeding boundaries, or severe smearing, that region will be marked as an "abnormal empty space." The image will be retained in the target image set, but the labeled recognition text will be empty or unreliable for further processing in the subsequent multi-modal grading stage. Through this process, the transformation from the overall answer sheet image to a structured, empty-granularity answer image set is achieved.
[0049] Step S303: Based on the standard answer, obtain reference test questions that match the standard answer.
[0050] To obtain reference test questions that match the standard answer, the standard answer must first be feature extracted to obtain its triplet features. These triplet features include invariant components, variant components, and supplementary components. Then, the triplet features, along with the subject, question type, and grade level information corresponding to the pure card test questions, are used as retrieval features. Based on these retrieval features, a multi-path recall strategy is employed to retrieve reference test questions that match the standard answer from a pre-built question bank.
[0051] When the reference questions include a first reference question and a second reference question, a multi-path recall strategy is used based on retrieval features to retrieve reference questions that match the standard answer from a pre-built question bank. This includes: retrieving a retrieved question set from the pre-built question bank based on retrieval features using a multi-path recall strategy; assigning different weights to different recall paths for weighted ranking; and determining the first reference question and the second reference question from the retrieved question set. The first reference question is the question with the highest weighted score, and the second reference question is the question with the lowest weighted score. The weighted score is proportional to the matching degree of the standard answer.
[0052] Specifically, this application extracts triplet features based on a fine-tuned generative large model and calls a generative large language model (based on Spark 2.6B) that has undergone two-stage fine-tuning training.
[0053] First, enter the standard answer to the current blank answer question (e.g., blankStdAnswer is “He spent two hours finishing his homework yesterday.”) and the question type information (e.g., topicCode is “Chinese to English”).
[0054] The prompt for the test item generation capability of the model in the first stage of training is as follows: You will play the role of a "Subject-Specific Test Question Generation Expert." Your core task is to reverse engineer the core test points based on the given standard answer {blankStdAnswer} and question type information {topicCode}, and generate test questions that match the knowledge points, question types, and difficulty levels. The specific execution requirements are as follows: Step 1: In-depth analysis of the standard answer (the basis of question setting) Extract core test points: Locate 2-3 key pieces of information (such as core concepts, key data, logical relationships, operational steps, etc.) from the standard answer and mark them as "must-know points" to ensure that the test questions do not deviate from the core of the standard answer. Organize information hierarchy: Differentiate between "basic memory points" (such as definitions, numerical values), "understanding and application points" (such as logical deduction, scenario association), and "expansion and extension points" (such as comparison with similar knowledge, integration with practical cases) to prepare for tiered difficulty levels.
[0055] Step Two: Question Design Rules (Covering Diverse Examination Scenarios) At least three common question types must be generated (you can choose subjective fill-in-the-blank questions, short answer questions, and problem-solving questions, matching the subject of the standard answer). Each question type must follow these rules: Question prompt: Focus on the logical chain, core features, key steps, and core test points of the standard answer; use concise language and avoid irrelevant information. Answer guidance: Ensure that candidates organize their answers around the key information in the standard answer, avoiding overly broad questions that lead to deviations.
[0056] Third step: Difficulty gradient control (matching different assessment objectives) The generated test questions need to cover 3 difficulty levels, each with a specific assessment objective: Basic Level (Testing Memory): Questions directly target the "basic memory points" in the standard answer. Candidates only need to recall the standard answer to answer (e.g., "The core ingredient of XX is ______").
[0057] Advanced Level (Assessing Comprehension): Questions will focus on "understanding application points," requiring candidates to understand the logic of the standard answer before responding (e.g., "Based on the principle of XX, explain why the raw material of XX cannot be replaced by XX").
[0058] Advanced Level (Assessing Application): Questions will be designed based on "Extension Points," requiring candidates to connect the standard answer with real-world scenarios and similar knowledge (e.g., "In a certain scenario, the product of XX does not match the standard answer; analyze the possible reasons").
[0059] IV. Output Requirements The output strictly follows the format of "1. Standard answer breakdown → 2. Questions by question type and difficulty (each question is marked with question type, difficulty, correct answer / key points)" to ensure that each question can be traced back to the core test points of the standard answer, without deviation or ambiguity.
[0060] The second phase trains the triplet extraction capability, and the prompt is as follows: You will play the role of a "Text Triple Component Extraction Specialist," responsible for accurately extracting "invariant components," "variant components," and "supplementary components" from the given text {blankStdAnswer}. Specific execution requirements are as follows: I. Clear Core Definition Immutable components: Core, fixed, and irreplaceable information in the text. Replacing them will change the core focus of the text (e.g., in the sentence "What city is the capital of Anhui?", if "Anhui" is replaced with "Jiangsu", the core focus will change completely).
[0061] Variant component: The text can be replaced by synonyms or near-synonyms, and the replacement does not change the core information the text points to (e.g., in "What city is the capital of Anhui Province?", "capital" can be replaced with "administrative center" or "provincial capital").
[0062] Supplementary components: These are used to supplement information such as the background, attributes, and scope of the "invariant components" or "variant components". Removing them will not affect the core semantics of the text (e.g., "province" and "central China city" in "What city is the capital of Anhui?").
[0063] II. Extraction and Execution Steps Step 1: Read through the text, locate the "unchanging components", confirm that they are the core fixed information of the text, mark them and explain the reasons.
[0064] Step 2: Based on the "invariant components", filter out the replaceable "variant components", list 1-2 reasonable replacement statements, and explain why the core pointer does not change after replacement.
[0065] Step 3: Discover the implicit or explicit "supplementary components" in the text, mark the objects they supplement (invariant / variant components), and explain the specific content of the supplement (background / attributes / scope, etc.).
[0066] III. Example Reference Given the text: What is the capital city of Anhui Province? Unchanging component: Anhui, because it is the region that the core text points to, and replacing it with other provinces would change the core of the problem.
[0067] Variant component: provincial capital, replaced with: administrative center / provincial capital, reason: the replacement still points to "the core administrative region of Anhui", without changing the semantics of the problem.
[0068] Additional components: Province (add attributes for "Anhui"), Central China cities (add geographical scope for "Anhui"), because removing them does not affect the core semantics of "inquiring about the core administrative region of Anhui".
[0069] IV. Output Requirements Please strictly adhere to the output format "[{"type":"Invariant","key":"","values":[""],"reason":""},{"type":"Variant","key":"","values":[""],"reason":""},{"type":"Supplementary","key":"","values":[""],"reason":""}]" to ensure that the definition and reason for each component match and that no key information is omitted.
[0070] The output example is as follows: [ { "type": "Invariant", key: "Anhui", "values": [ "" ], "reason": "The core region pointed to by the text; replacing it with another province will change the core of the problem." }, { "type": "Variant", "key": "provincial capital", "values": [ "Administrative Center" "provincial capital" ], "reason": "The replacement still refers to 'the core administrative region of Anhui,' without changing the semantics of the question." }, { "type": "Supplementary", "key": "Capital of Anhui Province", "values": [ "Cities in Central China" ], "reason": Removing this does not affect the core semantics of "inquiring about the core administrative region of Anhui" and serves as a supplement to the attribute of "the capital of Anhui Province". } ] The model extracts a specific prompt based on the triples from the second stage, performs in-depth analysis of the standard answer, and outputs triple information that strictly adheres to the specified JSON format: Invariant component: The extracted result is {"type": "Invariant", "key": "spent...finishing","values": [""], "reason": "'spend time doing sth' is a core fixed collocation, and the syntactic structure and core semantics change after replacement"}.
[0071] Variant component: The extracted result is {"type": "Variant", "key": "two hours", "values":["a couple of hours", "120 minutes"], "reason": "Time length description, which can be replaced synonymously without changing the core event semantics"}.
[0072] Supplementary component: The extracted result is {"type": "Supplementary", "key": "Tense and time adverbs", "values": ["Past tense", "yesterday"], "reason": "Supplement the time context of the action, removing it does not affect the core structural semantics of 'spending time doing something'"}.
[0073] At the same time, metadata information such as the subject (e.g., English), question type (e.g., fill-in-the-blank), and grade level (e.g., junior high school) of the test question is obtained.
[0074] Then, the aforementioned triplet information, subject, question type, grade level, and the vector representation of the standard answer (AnswerVector) are used together as retrieval features.
[0075] This application also includes a pre-built question bank (such as Milvus), where the fields and their corresponding meanings are shown in Table 1: Table 1. Fields and their meanings in the pre-built question bank
[0076] Execute a two-stage retrieval process, including recall and ranking, within a pre-built question bank (such as Milvus): Recall phase (multi-channel recall): Recall 1: Precise retrieval based on grade level, question type, subject, and complete standard answer text.
[0077] Recall 2: Similarity retrieval based on vectors of learning stages, question types, subjects, and standard answers (AnswerVector).
[0078] Recall 3: Search based on grade level, question type, subject, and variant components extracted from the standard answer (such as "two hours" and a list of synonyms).
[0079] All recalled test questions were weighted and ranked to determine the most relevant and least relevant reference test questions: Determine the first reference question (the most relevant question): During ranking, assign the highest weight (e.g., 0.5) to the result of recall 1, the second highest weight (e.g., 0.3) to recall 2, and the lowest weight (e.g., 0.2) to recall 3. After calculating the weighted total score for each question, select the question with the highest score as the first reference question. For example, if the question stem is "Last week she spent three hours reading," it will have the highest weighted score because its standard answer perfectly matches the structure "spend time doing" and the tense is consistent.
[0080] Determine the second reference question (the least relevant question): Use the opposite weighting strategy for ranking: assign the highest weight (e.g., 0.5) to recall 3, the second highest weight (e.g., 0.3) to recall 2, and the lowest weight (e.g., 0.2) to recall 1. After calculating the weighted total score, select the question with the highest score as the second reference question. This question is usually unrelated to the core test point of the current question; for example, it might be a question testing the fixed phrase "make up one's mind" in English-to-Chinese translation.
[0081] Regarding the two-stage fine-tuning strategy adopted in this case, the first stage trains the model's ability to generate test questions based on standard answers. The training data does not require manual annotation; the question stems themselves serve as the annotation results. The second stage trains the ability to extract triplet information. This is based on training data primarily derived from two sources: a. These questions have standard answers; b. Building upon a, triplet information is generated using a large-scale Spark model and then corrected by curriculum researchers. Its composition is as follows: [ { "instruction": "Prompt for the first step of the design", "input": "What is the capital city of Anhui Province?" "output":"[{\"type\":\"Invariant\",\"key\":\"Anhui\",\"values\":[\"\"],\"reason\":\"The core region of the text; replacing it with another province will change the core of the question\"},{\"type\":\"Variant\",\"key\":\"Provincial Capital\",\"values\":[\"Administrative Center\",\"Provincial Capital\"],\"reason\":\"After replacement, it still points to "Anhui's core administrative region" and does not change the semantics of the question\"},{\"type\":\"Supplementary\",\"key\":\"Anhui Provincial Capital\",\"values\":[\"Central China City\"],\"reason\":\"Removing it does not affect the core semantics of "Inquiring about Anhui's core administrative region" and is a supplementary attribute for "Anhui Provincial Capital"\"}]" }, { "instruction": "Prompt for the first step of the design", "input": "xx", "output": "xx" } ] Before using the training data for model training, the training data is first cleaned to create a dataset that is "highly pure, widely distributed, and structurally balanced." The cleaning rules are as follows: For English test questions, first remove non-alphabetic characters from the marked answers, then count the number of words (segmented by spaces), and filter out data with marked answers of 3 words or less; filter out essay questions.
[0082] For Chinese language test questions, we first delete non-Chinese characters in the standard answers, then use jieba to segment words and count the number of words, filtering out data with 3 or fewer words; at the same time, we filter out Chinese composition questions.
[0083] For science subjects, only problem-solving questions will be retained, while fill-in-the-blank and multiple-choice / true / false questions will be filtered out.
[0084] Based on the FastText tool, duplicate data in test questions is removed, including test questions with essentially the same test points and semantic meaning.
[0085] Regarding the amount of training data, a total of 6000 training data points are used, including 1000 for English, 1000 for Chinese, and 1000 each for Mathematics, Physics, Chemistry, and Biology. In addition, a small amount of first-stage data, such as 500 data points, is retained to prevent overfitting and preserve the model's ability to answer questions.
[0086] Step S303: Based on the target image, standard answer, and reference test questions, generate target grading results that match the pure card answer content through a preset multimodal large language model.
[0087] Based on the target image, standard answer, and reference questions, a pre-set multimodal large language model is used to generate target grading results that match the pure card answer content. First, the stems of the first reference questions and the second reference questions need to be obtained. Then, the grading prompts containing the target image, standard answer, the stems of the first reference questions, and the stems of the second reference questions are input into the pre-set multimodal large language model to obtain the first probability distribution corresponding to the first reference questions and the second probability distribution corresponding to the second reference questions. The probability distribution is used to represent the grading result category for the reference questions. Then, based on the first probability distribution and the second probability distribution, the target grading results that match the pure card answer content are obtained.
[0088] The step of obtaining a target grading result that matches the pure card answer content based on a first probability distribution and a second probability distribution includes: obtaining the probability of the grading result category of the first probability distribution and the second probability distribution; subtracting the probability of the grading result category of the first probability distribution from the probability of the grading result category of the second probability distribution to obtain the probability difference of the grading result category; and selecting the grading result category corresponding to the largest probability difference from the probability differences of the grading result categories as the target grading result.
[0089] Specifically, first, integrate all necessary information to construct a grading prompt message that conforms to a preset format. This prompt strictly follows the "Second Stage Multimodal Grading Prompt" designed in the handover document, as follows: # Require You are a senior expert in primary and secondary school teaching and research. Please grade the provided test questions based on your professional knowledge of the subject {subject} and in accordance with the requirements below.
[0090] # Correction Standards ① The grading results analyze the correctness of students' answers in each answer area, and then give the grading results based on the analysis results. The grading results include the following four types: - "Correct": The student's answer is completely correct. - "Error": The student's answer contained an error. - "Half right": The student's answer hit some of the key points of the answer key. - "Rejection": This applies to situations where the student's answer is blank or where it's impossible to determine the correctness of the answer. See the rejection instructions below for details. This requires special attention. ② When grading, base your work on the reference questions and focus on the core test points behind the current question. Minor errors unrelated to the core test points can be ignored, such as negligible punctuation errors and a few typos in long answers to humanities short answer questions; negligible capitalization issues and a few singular / plural errors in English-to-Chinese translation; and unsimplified fractions in science short answer questions.
[0091] ③ Note that when judging a student's answer as "correct," the content of the answer does not need to be completely identical to the correct answer. For example, in open-ended questions, a student's answer only needs to be relevant and reasonable; in humanities questions, a student's answer only needs to be equivalent to the correct answer; in science questions, if there are no specific requirements for the form of the answer, equivalent fractions, ratios, and decimals can be considered correct answers (e.g., 0.25, 1:4, 1 / 4 are equivalent); and in multiple-choice questions, as long as the student's selected options match the set of standard answers (no more or fewer selections), even if the order of the answers is different, it should be judged as correct. It is particularly important to note that for political short-answer questions, it is not necessary for the student's answer to be completely correct; as long as the student's answer is reasonable and hits some of the key points of the answer, it can be judged as "correct."
[0092] ④ First, analyze and correct the student's answer text in each answer area of the test questions to be corrected. If the result of the analysis is "incorrect" for the answer area, it is because the given student answer text may have recognition errors (such as easily confused English letters, unrecognized text outside the answer area, line breaks in the answer content, etc.), or the student's writing is not standardized or illegible. Therefore, you need to further combine the given test question image, locate the corresponding answer area in the image, re-identify the student's handwritten answer in that answer area in the image, and then re-analyze and correct the answer based on your recognition results to obtain the final correction result.
[0093] # Rejection Explanation ## Rejection Situation Before making corrections, you must check if any of the following conditions exist. If any of these conditions exist, the correction result must be "Rejected": ① Student answers are blank, empty strings, or contain only spaces. ② The answer area contains images of test questions awaiting grading that are blurry, incomplete, obscured, or distorted, making it impossible to accurately understand the questions. ③ The test questions to be graded are missing reference questions or have incomplete key information (such as the absence of standard answers). ## Special Notes on Rejection ① Rejection decisions must be made independently for each answer area. ② For blank answer areas: Regardless of whether there are answers in other answer areas for the same question, the grading result for the blank answer area must be "rejected". ③ For incomplete questions: All answer areas must be marked as "rejected". ④ Blank answer areas must not be marked as "incorrect". # Reference Example ## Sample Question Stem Presented in plain text format, as shown in the example below. The answer area in the question stem may be indicated by different forms of placeholders, such as "<blank 1-1> __", "__(1)__", etc.
[0094] Example: Write the inflections of the following words according to the requirements. 1. Appear →<blank 1-1> ______(antonym)\n2.modern→<blank 2-1> ______(antonym)\n3. sell→ <blank3-1>______(past tense)→<blank 3-2> ______(past participle). ## Student Answers The list format is used, where each element represents a response area as a dictionary, as shown in the example below. The keys are: "blankIdx" is the response area number, and "studentAnswer" is the student's response.
[0095] Example: [{"blankIdx": "1-1","studentAnswer": ""}, {"blankIdx": "2-1","studentAnswer": "ancient"}, {"blankIdx": "3-1","studentAnswer": "selled"}, {"blankIdx": "3-2","studentAnswer": "sold"}] ## Output Format The format is a list, where each element represents a response area as a dictionary, as shown in the example below. The keys are: "blankIdx" is the response area number, and "result" is the grading result.
[0096] Example: [{"blankIdx": "1-1", "result": "rejected"}, {"blankIdx": "2-1", "result": "correct"}, {"blankIdx": "3-1", "result": "incorrect"}, {"idx": "3-2", "result": "correct"}] # Questions to be graded ## Refer to the question stem {referenceTopicContent} ## Student Answers {studentAnswerList} # Task Please grade the students' answers to the above-mentioned questions one by one, based on the provided test question images and the text of the sample test questions, according to the grading standards and rejection instructions. Note that your output must strictly adhere to the output format defined in the reference example; do not output any irrelevant content.
[0097] In addition to the above prompts, the following content needs to be further specified: Target image: The blank image of the answer area at the current empty granularity, such as an image of a student's handwritten "He took two hours to finish his homework yesterday."
[0098] Standard answer: The standard answer to the current question (blankStdAnswer), for example, "He spent two hours finishing his homework yesterday."
[0099] Student's answer text: Text recognized by OCR from the target image (blankUserAnswer), for example, "He took two hours to finish his homework yesterday."
[0100] The first reference question stem: the complete stem of the most relevant question retrieved ({referenceTopicContent}), for example, the stem of a Chinese-to-English translation question: "She spent three hours reading last week." The second reference question stem: the complete stem of the least relevant question found (as noise reference), for example, the stem of an English-to-Chinese translation question: "They made up their minds to study hard." This structured prompt is then fed into a pre-defined multimodal large language model (based on the Spark 13B model) that has undergone two stages of fine-tuning training. This model is capable of understanding both textual and image information simultaneously.
[0101] After receiving the Prompt, the model will perform two parallel inference computations (or achieve equivalent processing in one forward propagation through a specific attention masking mechanism): Reasoning based on the most relevant questions: The model focuses on reasoning about the "first reference question stem" in the Prompt and its relevance to the current grading task, and generates a first probability distribution (P1) at the output layer for the preset grading result category tokens (i.e., "correct", "incorrect", "partially correct", "rejected"). For example: P1(correct)=0.05, P1(incorrect)=0.40, P1(partially correct)=0.53, P1(rejected)=0.02.
[0102] Reasoning based on the least relevant questions: The model focuses on the noise source of the "second reference question stem" in the Prompt for reasoning and generates a second probability distribution (P2). For example: P2(correct) = 0.20, P2(incorrect) = 0.10, P2(half correct) = 0.60, P2(rejected) = 0.10.
[0103] The two probability distributions mentioned above are calibrated at the model's output layer (after the Softmax layer) to remove biases introduced by noisy test questions. Calculate the probability difference: Subtract the first probability distribution (P1) from the second probability distribution (P2) element by element for each category of graded results.
[0104] Δ(Correct) = P1(Correct) - P2(Correct) = 0.05 - 0.20 = -0.15 Δ(error) = P1(error) - P2(error) = 0.40 - 0.10 = 0.30 Δ(half a pair) = P1(half a pair) - P2(half a pair) = 0.53 - 0.60 = -0.07 Δ(rejection) = P1(rejection) - P2(rejection) = 0.02 - 0.10 = -0.08 From the four calculated probability differences, select the one with the largest value. In this example, the largest probability difference is 0.30, corresponding to the grading result category of "error".
[0105] Therefore, the student's answer to this blank was determined to be "incorrect". This conclusion is consistent with the core test point (the requirement is to use the "spend time doing" structure, while the student used the "take time to do" structure), and the calibration process effectively suppressed the interference of the artificially high "half-correct" probability (0.60) caused by noisy test questions.
[0106] In one embodiment, the method further includes: obtaining an initial grading result based on a rule-based grading strategy; if the initial grading result meets a preset grading condition, then triggering the step of obtaining reference test questions that match the standard answer based on the standard answer.
[0107] Specifically, the "response text" (i.e., OCR recognition result) and "standard answer" obtained after structured processing are first subject to rule-based rapid grading.
[0108] The core of rule-based grading is text matching and error tolerance. The specific steps are as follows: Clean the student's answer text and the standard answer (e.g., remove leading and trailing spaces, unify capitalization, convert full-width and half-width characters), then determine if the cleaned answer text is completely identical to the standard answer. If they are identical, the initial grading result is directly marked as "correct"; if they are not completely identical, the error tolerance rule base is activated. For example: for English, a mapping of easily confused characters is established (e.g., "n"). "h", "v" The algorithm generates candidate answer strings for matching. If any candidate string matches the standard answer, the initial grading result is marked as "correct"; otherwise, it is marked as "incorrect".
[0109] For mathematics, both the answers and the responses are converted into standardized mathematical expressions (such as unifying the decimal and fraction forms) before comparison.
[0110] If the response text is empty, consists entirely of spaces, or contains unrecognizable garbled characters, the initial grading result will be marked as "rejected".
[0111] The subsequent process is determined based on the initial grading results and the pre-set subject-based allocation strategy. The pre-set grading conditions are: the initial grading result is "incorrect" and the "direct final judgment" rule is not met.
[0112] For humanities subjects (such as English and Chinese): The initial grading result is "Incorrect," and the standard answer for that blank contains more than 3 words or vocabulary words. For example, if the standard answer for an English translation sentence has 8 words, and the rule judges it as incorrect, then subsequent retrieval and multi-model grading will be triggered. If the initial grading result is "Correct" or "Rejected," then that result will be directly output as the final grading result. Alternatively, although it is judged as "Incorrect," if the standard answer contains ≤3 words (such as a single word fill-in-the-blank), the rule grading is considered reliable enough, and this "Incorrect" result will be directly used as the final result without triggering further processes.
[0113] For science-related (e.g., mathematics, physics) problem-solving questions: the initial grading result is "incorrect". Due to the complex structure and semantic diversity of answers to science-related problem-solving questions, the reliability of rule matching is low. Therefore, once a rule is judged to be incorrect, subsequent retrieval and multi-modal grading are triggered by default. If the initial grading result is "correct" or "rejected", the process terminates and outputs the result.
[0114] Once the above triggering conditions are met, the complete process described in the previous embodiment will be automatically executed: Based on the standard answer of the current test question, triplet feature extraction and multi-way retrieval will be initiated to obtain the first and second reference test questions. Then, the target image, standard answer, student answer text, and retrieved reference test question stems will be used to construct a Prompt input multimodal large language model. The probability distributions generated by the model based on relevant test questions and noisy test questions will be obtained, and after probability difference calibration, the final target grading result (such as "incorrect" or "partially correct") will be obtained.
[0115] This application provides an intelligent grading method, comprising: first, acquiring the initial image and corresponding standard answer corresponding to the current pure card answer area; then, performing structured processing on the initial image to obtain the target image corresponding to the current pure card answer area; and then, based on the standard answer, acquiring reference questions that match the standard answer. Based on the target image, standard answer, and reference questions, a target grading result matching the pure card answer content is generated through a preset multimodal large language model. This application retrieves reference questions matching the standard answer, providing richer context and scoring criteria for the grading process. It can not only determine the correctness of the answer but also understand the logical structure, knowledge application methods, and common error types of the answer, thus providing more in-depth feedback. Furthermore, relying on the comprehensive reasoning capabilities of the multimodal large language model, this application can simultaneously process text, images, and even implicit answering thought processes, generating targeted and highly interpretable grading results. This not only reduces the repetitive grading burden on teachers but also improves the effectiveness of intelligent grading.
[0116] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0117] The following are device embodiments of this application. For details not described in detail, please refer to the corresponding method embodiments described above.
[0118] Figure 5 This diagram illustrates the structure of an intelligent correction device according to an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. The intelligent correction device includes: The acquisition module 501 is used to acquire the initial image and the corresponding standard answer for the current pure card answering area; Processing module 502 is used to perform structured processing on the initial image to obtain the target image corresponding to the current pure card answering area; The retrieval module 503 is used to retrieve reference test questions that match the standard answers. The grading module 504 is used to generate target grading results that match the pure card answer content based on the target image, standard answer, and reference test questions through a preset multimodal large language model.
[0119] In one embodiment, the processing module 502 is further configured to perform structured processing on the initial image using a preset image recognition model to obtain an empty granular image corresponding to each answer region in the initial image, wherein the empty granular image corresponding to each answer region contains corresponding answer text. The empty granularity images corresponding to each answer region in the initial image are summarized to obtain all empty granularity images corresponding to the initial image. Use all empty granularity images corresponding to the initial image as the target image.
[0120] In one embodiment, the retrieval module 503 is further configured to extract features from the standard answer and obtain the triplet features of the standard answer; wherein, the triplet features include invariant components, variant components and supplementary components; The triplet features, as well as the subject, question type, and grade level information corresponding to the pure card test questions, are used as retrieval features; Based on retrieval features, a multi-path recall strategy is used to retrieve reference questions that match the standard answers from a pre-built question bank.
[0121] In one embodiment, the reference test questions include a first reference test question and a second reference test question; The retrieval module 503 is also used to retrieve a recall set of test questions from a pre-built test question bank based on retrieval features and employing a multi-path recall strategy. Different weights are assigned to different recall paths for weighted ranking. First and second reference questions are determined from the recalled question set. The first reference question is the question with the highest weighted score, and the second reference question is the question with the lowest weighted score. The weighted score is proportional to the matching degree of the standard answer.
[0122] In one embodiment, the grading module 504 is further configured to obtain the stem of the first reference test question and the stem of the second reference test question; The grading prompts, which include the target image, the standard answer, the stem of the first reference question, and the stem of the second reference question, are input into a preset multimodal large language model to obtain the first probability distribution corresponding to the first reference question and the second probability distribution corresponding to the second reference question. The probability distribution is used to characterize the grading result category for the reference question. Based on the first probability distribution and the second probability distribution, the target grading results that match the content of the pure card answers are obtained.
[0123] In one embodiment, the grading module 504 is further configured to obtain the probability of the grading result category of the first probability distribution and the second probability distribution; The probability difference between the grading result categories is obtained by subtracting the probability of the grading result category from the probability of the first probability distribution and the probability of the grading result category from the probability of the second probability distribution. Select the category of the grading result with the largest probability difference from the probability differences of the grading result categories as the target grading result.
[0124] In one embodiment, the apparatus further includes: a return execution module, which is used to obtain an initial grading result based on a rule-based grading strategy; If the initial grading results meet the preset grading conditions, the step of obtaining reference test questions that match the standard answer will be triggered.
[0125] This application provides an intelligent grading device, specifically used for: first acquiring the initial image and corresponding standard answer corresponding to the current pure card answering area; then performing structured processing on the initial image to obtain the target image corresponding to the current pure card answering area; and then, based on the standard answer, acquiring reference questions that match the standard answer. Based on the target image, standard answer, and reference questions, a target grading result matching the pure card answering content is generated through a preset multimodal large language model. This application retrieves reference questions matching the standard answer, providing richer context and scoring criteria for the grading process. It can not only determine the correctness of the answer but also understand the logical structure, knowledge application methods, and common error types of the answer, thus providing more in-depth feedback. Furthermore, relying on the comprehensive reasoning capabilities of the multimodal large language model, this application can simultaneously process text, images, and even implicit answering thought processes, generating targeted and highly interpretable grading results. This not only reduces the repetitive grading burden on teachers but also improves the effectiveness of intelligent grading.
[0126] This application Figure 6 A schematic diagram of a computer device is provided. (Example) Figure 6 As shown, the computer device 6 in this embodiment includes a processor 601, a memory 602, and a computer program 606 stored in the memory 602 and executable on the processor 601. When the processor 601 executes the computer program 606, it implements the steps described in the various intelligent correction method embodiments above, for example... Figure 3 Steps 301 to 304 are shown. Alternatively, when processor 601 executes computer program 606, it implements the functions of each module / unit in the above-described intelligent correction device embodiments, for example... Figure 5 The functions of modules / units 501 to 504 are shown.
[0127] This application also provides a readable storage medium storing a computer program, which, when executed by a processor, is used to implement the intelligent correction method provided in the various embodiments described above.
[0128] The readable storage medium can be a computer storage medium or a communication medium. A communication medium includes any medium that facilitates the transfer of computer programs from one location to another. A computer storage medium can be any available medium accessible to a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application-Specific Integrated Circuit (ASIC). Alternatively, the ASIC can be located in a user equipment. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0129] This application also provides a program product including executable instructions stored in a readable storage medium. At least one processor of the device can read the executable instructions from the readable storage medium, and the execution of the executable instructions by the at least one processor causes the device to implement the intelligent correction methods provided in the various embodiments described above.
[0130] In the embodiments of the above-described device, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0131] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An intelligent grading method, characterized in that, include: Obtain the initial image and corresponding standard answer for the current pure card answering area; The initial image is subjected to structured processing to obtain the target image corresponding to the current pure card answering area; Based on the standard answer, obtain reference test questions that match the standard answer; Based on the target image, the standard answer, and the reference test questions, a target grading result matching the pure card answer content is generated through a preset multimodal large language model.
2. The intelligent grading method according to claim 1, characterized in that, The step of performing structured processing on the initial image to obtain the target image corresponding to the current pure card answering region includes: The initial image is structured using a preset image recognition model to obtain an empty granular image corresponding to each answer region in the initial image, wherein the empty granular image corresponding to each answer region contains the corresponding answer text. The empty granularity images corresponding to each answer region in the initial image are summarized to obtain all empty granularity images corresponding to the initial image; The target image is defined as all empty granularity images corresponding to the initial image.
3. The intelligent grading method according to claim 1, characterized in that, The process of obtaining reference test questions that match the standard answer, based on the standard answer, includes: Feature extraction is performed on the standard answer to obtain the triplet features of the standard answer; wherein, the triplet features include invariant components, variant components, and complementary components; The triplet features, along with the subject, question type, and grade level information corresponding to the pure card test questions, are used as retrieval features; Based on the retrieval features, a multi-path recall strategy is used to retrieve reference questions that match the standard answers from a pre-built question bank.
4. The intelligent grading method according to claim 3, characterized in that, The reference test questions include a first reference test question and a second reference test question; The step of using a multi-path recall strategy to retrieve reference questions matching the standard answers from a pre-built question bank based on the retrieval features includes: Based on the aforementioned retrieval features, a multi-path recall strategy is employed to retrieve data from a pre-constructed question bank, thereby obtaining a recalled question set. Different weights are assigned to different recall paths for weighted sorting. A first reference question and a second reference question are determined from the recalled question set. The first reference question is the question with the highest weighted score, and the second reference question is the question with the lowest weighted score. The weighted score is proportional to the matching degree of the standard answer.
5. The intelligent grading method according to claim 4, characterized in that, The process of generating a target grading result that matches the pure card answer content based on the target image, the standard answer, and the reference test questions using a preset multimodal large language model includes: Obtain the stem of the first reference test question and the stem of the second reference test question; The grading prompt information, which includes the target image, the standard answer, the stem of the first reference test question, and the stem of the second reference test question, is input into the preset multimodal large language model to obtain the first probability distribution corresponding to the first reference test question and the second probability distribution corresponding to the second reference test question. The probability distribution is used to characterize the grading result category for the reference test question. Based on the first probability distribution and the second probability distribution, a target grading result matching the content of the pure card response is obtained.
6. The intelligent grading method according to claim 5, characterized in that, The step of obtaining the target grading result matching the pure card answer content based on the first probability distribution and the second probability distribution includes: Obtain the probability of the grading result category from the first probability distribution and the second probability distribution; The probability difference between the grading result categories is obtained by subtracting the probability of the grading result category from the probability of the first probability distribution and the probability of the grading result category from the probability of the second probability distribution. The category of the grading result with the largest probability difference among the grading result categories is selected as the target grading result.
7. The intelligent grading method according to any one of claims 1 to 6, characterized in that, The method further includes: Initial grading results are obtained based on rule-based grading strategies; If the initial grading result meets the preset grading conditions, then the step of obtaining reference test questions that match the standard answer based on the standard answer is triggered.
8. An intelligent correction device, characterized in that, include: The acquisition module is used to acquire the initial image and the corresponding standard answer for the current pure card answering area; The processing module is used to perform structured processing on the initial image to obtain the target image corresponding to the current pure card answering area; The retrieval module is used to retrieve reference test questions that match the standard answer. The grading module is used to generate a target grading result that matches the pure card answer content based on the target image, the standard answer, and the reference test questions, using a preset multimodal large language model.
9. A computer device, characterized in that, Includes a memory, and one or more processors communicatively connected to the memory; The memory stores instructions that can be executed by the one or more processors to cause the one or more processors to implement the intelligent correction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Includes a program or instructions that, when run on a computer, implement the intelligent correction method according to any one of claims 1 to 7.
11. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the intelligent correction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Examination paper automatic appraisal method and device, and electronic equipment
CN108960149A
Answer sheet identification method and device, storage medium and electronic equipment
CN114708598A
Correction method and device, equipment and medium
CN117636368A
Physical, chemical and biological experiment report scoring method and system based on multi-modal large model
CN119887470A
Reading method and device, electronic equipment and storage medium
CN120106024A