Handwritten blank-filling question correcting method, device and equipment for artificial intelligent paper marking and computer readable storage medium

By combining multimodal deep models and generative language models, the problem of automated scoring of math fill-in-the-blank questions was solved, achieving high-precision recognition and automated scoring of handwritten fill-in-the-blank questions, thus improving the efficiency of marking.

CN121617099APending Publication Date: 2026-03-06SHENZHEN SEA SKY LAND TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511752992.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

In online marking systems, the answers to math fill-in-the-blank questions are concise, but due to the large differences in students' writing habits and diverse forms of expression, the recognition accuracy is low, making it difficult to achieve automated scoring by the system. It still requires manual marking, which is inefficient.

Method used

This method combines multimodal deep models and generative language models. By acquiring target answer sheet images and matching answer regions, the method converts the images into original answer text using a visual encoder, cross-modal adapter, and language decoder. Then, it performs semantic understanding and normalization processing through generative language models to generate structured standard mathematical representations. Finally, it compares the results with preset reference answers in multiple dimensions and outputs the scoring results.

Benefits of technology

It achieves high-precision recognition and automated scoring of handwritten math answers, reducing problems of inaccurate recognition and inability to automate scoring, and improving grading efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121617099A_ABST
    Figure CN121617099A_ABST
Patent Text Reader

Abstract

The invention discloses a handwritten blank-filling question correcting method, device and equipment for artificial intelligent paper marking and a computer readable storage medium. The method comprises the steps that a target answer sheet image is collected; matching the target answer sheet image with a preset sheet surface template to obtain an answer area map corresponding to each gap filling question; converting the handwritten answer content in the answer area map into an original answer text through a multi-modal depth model; performing semantic understanding and standardization processing on the original answer text through a generative language model to generate a structured standard mathematical representation; and performing multi-dimensional comparison on the standard mathematical representation and a preset reference answer, and outputting a scoring result. Therefore, through combination of the multi-modal depth model and the finished language model, the handwritten mathematical answers are converted into structured standard representation, and automatic scoring is performed from multiple dimensions, so that the problems of inaccurate recognition and incapability of automatic scoring are effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image recognition and processing, and more specifically, to a method, apparatus, device, and computer-readable storage medium for grading handwritten fill-in-the-blank questions using artificial intelligence for marking. Background Technology

[0002] Online marking systems use Optical Character Recognition (OCR) to identify and process answer sheets. However, while the answers to math fill-in-the-blank questions are brief, the wide variety of writing habits and expressions among test-takers, including fractions, decimals, questions with and without units, and equivalent expressions, makes recognition accuracy low and hinders automated marking. This necessitates manual marking, resulting in low efficiency. Summary of the Invention

[0003] In view of the above problems, this application proposes a method, apparatus, equipment and computer-readable storage medium for grading handwritten fill-in-the-blank questions using artificial intelligence, which can solve the above problems.

[0004] In a first aspect, embodiments of this application provide a method for grading handwritten fill-in-the-blank questions using artificial intelligence. The method includes: acquiring a target answer sheet image; matching the target answer sheet image with a preset answer sheet template to obtain a corresponding answer area image for each fill-in-the-blank question; converting the handwritten answer content in the answer area image into original answer text using a multimodal deep model; the multimodal deep model integrates a visual encoder, a cross-modal adapter, and a language decoder; performing semantic understanding and standardization processing on the original answer text using a generative language model to generate a structured standard mathematical representation; the standard mathematical representation includes LaTeX format, standardized numerical values, unit identifiers, and semantic tags; and comparing the standard mathematical representation with a preset reference answer in multiple dimensions to output the scoring result.

[0005] Secondly, this application also provides a handwritten fill-in-the-blank question grading device for artificial intelligence-based marking. The device includes: a data acquisition module for acquiring a target answer sheet image; a first execution module for matching the target answer sheet image with a preset answer sheet template to obtain a corresponding answer area map for each fill-in-the-blank question; a second execution module for converting the handwritten answer content in the answer area map into original answer text using a multimodal deep model; the multimodal deep model integrates a visual encoder, a cross-modal adapter, and a language decoder; a third execution module for performing semantic understanding and normalization processing on the original answer text using a generative language model to generate a structured standard mathematical representation; the standard mathematical representation includes LaTeX format, standardized numerical values, unit identifiers, and semantic tags; and an output module for comparing the standard mathematical representation with a preset reference answer in multiple dimensions and outputting the scoring result.

[0006] Thirdly, this application embodiment also provides a handwritten fill-in-the-blank question grading device for artificial intelligence grading, including a processor, a memory, and one or more application programs; the one or more application programs are stored in the memory and configured to be executed by the processor to implement the above-described handwritten fill-in-the-blank question grading method for artificial intelligence grading.

[0007] Fourthly, this application embodiment also provides a computer-readable storage medium storing program code, wherein the above-mentioned handwritten fill-in-the-blank question grading method for artificial intelligence grading is executed when the program code is run by a processor.

[0008] The technical solution provided in this application includes the following method: acquiring a target answer sheet image; matching the target answer sheet image with a preset answer sheet template to obtain a corresponding answer area image for each fill-in-the-blank question; converting the handwritten answer content in the answer area image into original answer text using a multimodal deep model; the multimodal deep model integrates a visual encoder, a cross-modal adapter, and a language decoder; performing semantic understanding and normalization processing on the original answer text using a generative language model to generate a structured standard mathematical representation; the standard mathematical representation includes LaTeX format, normalized numerical values, unit identifiers, and semantic tags; comparing the standard mathematical representation with a preset reference answer in multiple dimensions, and outputting the scoring result. Thus, by combining a multimodal deep model with a generative language model, handwritten mathematical answers are transformed into a structured standard representation and automatically scored from multiple dimensions, effectively reducing problems of inaccurate recognition and the inability to automate scoring. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments and drawings obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0010] Figure 1 The illustration shows a flowchart of a method for grading handwritten fill-in-the-blank questions using artificial intelligence in an embodiment of this application.

[0011] Figure 2 The diagram shows a structural schematic of a handwritten fill-in-the-blank question grading device for artificial intelligence marking, provided in an embodiment of this application.

[0012] Figure 3 The diagram shows a structural schematic of a handwritten fill-in-the-blank question grading device for artificial intelligence marking, provided in an embodiment of this application.

[0013] Figure 4 This illustration shows a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0015] Online marking systems use Optical Character Recognition (OCR) to identify and process answer sheets. However, while the answers to math fill-in-the-blank questions are brief, the wide variety of writing habits and expressions among test-takers, including fractions, decimals, questions with and without units, and equivalent expressions, makes recognition accuracy low and hinders automated marking. This necessitates manual marking, resulting in low efficiency.

[0016] To address the aforementioned issues, this application provides a method, apparatus, device, and computer-readable storage medium for grading handwritten fill-in-the-blank questions using artificial intelligence. The method includes: acquiring a target answer sheet image; matching the target answer sheet image with a preset answer sheet template to obtain a corresponding answer area image for each fill-in-the-blank question; converting the handwritten answer content in the answer area image into original answer text using a multimodal deep model; the multimodal deep model integrates a visual encoder, a cross-modal adapter, and a language decoder; performing semantic understanding and standardization processing on the original answer text using a generative language model to generate a structured standard mathematical representation; the standard mathematical representation includes LaTeX format, standardized numerical values, unit identifiers, and semantic tags; and comparing the standard mathematical representation with a preset reference answer in multiple dimensions to output the scoring result.

[0017] Therefore, by combining multimodal deep models with formal language models, handwritten mathematical answers can be transformed into structured standard representations and automatically scored from multiple dimensions, effectively reducing problems of inaccurate recognition and inability to automate scoring.

[0018] Please see Figure 1 , Figure 1 This illustration shows a flowchart of a method for grading handwritten fill-in-the-blank questions using artificial intelligence in an examination paper, as provided in an embodiment of this application. Figure 1 As shown, the method may include steps 110 to 150.

[0019] In step 110, the target answer sheet image is acquired.

[0020] The standardized answer sheets completed by candidates are digitally imaged using high-speed scanning equipment or high-resolution image acquisition devices to obtain clear and complete target answer sheet images. These target answer sheet images include all answers to fill-in-the-blank questions and have sufficient resolution (usually no less than 300 dpi) to ensure that the details of handwritten characters and mathematical symbols are legible.

[0021] In some implementations, the image acquisition device may be a dedicated document scanner, a flatbed scanner, or an industrial camera equipped with a document capture mode.

[0022] In some implementations, the target questionnaire image is typically stored in grayscale or color (e.g., JPEG, PNG, or TIFF format) and preprocessed after acquisition, such as denoising, binarization, or contrast enhancement, to obtain the final target questionnaire image, thereby improving the recognition accuracy of the target questionnaire image in subsequent steps.

[0023] In some implementations, the target answer sheet image is a standardized answer sheet used in large-scale standardized tests (e.g., middle school entrance exams, college entrance exams, academic proficiency tests), with a fixed layout to facilitate template matching.

[0024] After acquiring the target answer sheet image, since the target answer sheet image contains the entire content of the answer sheet (including the candidate information area, the multiple choice question filling area, the answer area for multiple fill-in-the-blank questions, etc.), and this invention only needs to identify and score the handwritten answer part of the math fill-in-the-blank questions, it is necessary to first accurately locate the answer area corresponding to each fill-in-the-blank question.

[0025] In step 120, the target answer sheet image is matched with the preset answer sheet template to obtain the answer area image corresponding to each fill-in-the-blank question.

[0026] In some implementations, the preset answer sheet template can be a standard answer sheet template pre-stored by the system. This standard answer sheet template is a structured reference image constructed based on the fixed format of the answer sheet used in the exam, and it includes the coordinate position, size, and relative layout information of each fill-in-the-blank answer box.

[0027] In some implementations, the system spatially aligns the target answer sheet image with a preset answer sheet template using an image alignment algorithm based on feature point detection (e.g., SIFT, ORB) or template sliding matching, and then crops out the answer area image corresponding to each math fill-in-the-blank question.

[0028] In other words, the answer area map only contains the handwritten answer to a single fill-in-the-blank question, thus effectively eliminating interference from irrelevant areas and providing high-quality input for high-precision recognition in subsequent steps.

[0029] In step 130, the handwritten answers in the answer region map are converted into the original answer text using a multimodal deep model.

[0030] Firstly, the recognition of ordinary Chinese or English text in fill-in-the-blank questions mainly focuses on semantic structure and relatively linear feature types, emphasizing the correct order of character sequences and semantic coherence. In contrast, handwritten mathematical formulas have highly nested and hierarchical semantic structures, involving spatial layout and often featuring complex structures such as subscripts, superscripts, fractions, square roots, integrals, and matrices.

[0031] Secondly, the character sets corresponding to ordinary Chinese or English text are relatively large, potentially reaching tens of thousands of characters. Each character is a flat and independent two-dimensional symbol with complex stroke structures and a relatively fixed spatial layout. In contrast, the character sets corresponding to handwritten mathematical formulas often use special symbols such as Latin letters, Greek letters, operators, parentheses, and arrows.

[0032] Its structural definition is as follows: structural symbols are often defined using “^” (which represents a superscript), “_” (which represents a subscript), “\frac{}{}” (which represents a fraction), “\sqrt{}” (which represents a square root), etc. These symbols themselves do not represent specific recognition content, but only define the formula structure. The recognized formula text often has spatial relationships of up and down, left and right, and nesting, and can be rendered using relevant tools.

[0033] It is evident that handwritten mathematical formulas possess unique complexity in terms of structure, semantics, and spatial layout, making them unsuitable for general character recognition technologies. However, the multimodal deep model proposed in this application can simultaneously receive and process information from different modalities (e.g., images and text), thereby improving the accuracy of recognizing handwritten mathematical formulas, Chinese characters, and English characters.

[0034] Specifically, a multimodal deep model is an end-to-end neural network architecture that integrates visual understanding and language generation capabilities. Specifically, a multimodal deep model integrates a visual encoder, a cross-modal adapter, and a language decoder.

[0035] By working together with three core components—visual encoder, cross-modal adapter, and language decoder—a multimodal deep model can simultaneously understand the spatial structure and mathematical semantics of an image, thereby significantly improving the recognition accuracy of handwritten formulas, mixed units, and non-standard writing styles.

[0036] The original answer text is a preliminary recognition result generated by a multimodal deep model based on the answer region map. The original answer text is in the form of free text that mixes natural language and mathematical symbols.

[0037] While the original answer texts generally reflect the candidates' handwritten responses, they often exhibit inconsistencies in handwriting, grammatical or formatting errors, and / or ambiguities or minor errors in identification. For example, the same numerical value might be interpreted as "1 / 2," For example, "0.5" or "50%". Another example is "5cm" not being interpreted as a numerical value or unit. Yet another example is the handwritten "α" being misinterpreted as "a", or the cursive script being misread. It was identified as "v4".

[0038] In other words, the original answer text is a non-standardized recognition result output by a multimodal model, containing diverse writing styles and potential ambiguities. It needs to be semantically normalized before it can be used for equivalence judgment and intelligent scoring.

[0039] Furthermore, in some implementations, the step "converting the handwritten answer content in the answer region map into the original answer text using a multimodal depth model" may include the following steps:

[0040] (1) Global feature extraction is performed on the response region map by a visual encoder to generate a high-dimensional semantic feature sequence;

[0041] (2) Semantic alignment and dimension mapping of high-dimensional semantic feature sequences are performed through cross-modal adapters, and they are converted into visual context vector sequences that are compatible with the embedding space of the language decoder;

[0042] (3) The language decoder takes the visual context vector sequence as input and combines it with preset prompts to generate the original answer text.

[0043] In some implementations, the visual encoder employs a Vision Transformer-based architecture, which divides the response region map into multiple image patches and models the spatial dependencies between the individual image patches through a global self-attention mechanism; the Vision Transformer architecture is selected from at least one of SigLIP-ViT, CLIP-ViT, or MobileViT.

[0044] Specifically, the input response region image (e.g., a 224×224 pixel grayscale image) is uniformly divided into fixed-size image patches. Each image patch is converted into a vector representation after linear embedding. Subsequently, these vectors are added to learnable positional codes and input into a multi-layer Transformer encoder.

[0045] By incorporating a global receptive field through a global self-attention mechanism, this encoder can simultaneously model long-range dependencies between any two image patches, thereby effectively capturing two-dimensional spatial structure information in handwritten mathematical formulas. Examples include the correspondence between content above and below fraction lines, the relationship between subscripts and superscripts and their corresponding base characters, and the coverage area of ​​square roots.

[0046] Finally, the visual encoder outputs a high-dimensional semantic feature sequence consisting of [CLS] tags and multiple patch features. The dimensions of the high-dimensional semantic feature sequence are typically N×D (e.g., N=197, D=768). This allows for deep semantic understanding of the response region map through the visual encoder, and the high-dimensional semantic feature sequence serves as the input to the cross-modal adapter.

[0047] In some implementations, a cross-modal adapter is a lightweight neural network module used to map a high-dimensional semantic feature sequence output by a visual encoder to the semantic embedding space of a language decoder, achieving alignment between visual and textual modalities. The cross-modal adapter is selected from at least one of Q-Former, LoRA projection adapter, or feature map linear projection layer.

[0048] Specifically, the cross-modal adapter uses a high-dimensional semantic feature sequence as a query and generates a set of visual context vectors with the same embedding dimension as the language decoder through a lightweight Transformer layer or a low-rank projection structure (e.g., LoRA projection layer, Q-Former query mechanism, or fully connected mapping layer). This process does not rely on the answer region map input, but only performs semantic extraction and spatial transformation based on visual features, thereby providing the language decoder with an understandable conditional context to support its autoregressive generation of the original answer text.

[0049] In other words, the cross-modal adapter builds a bridge for mapping visual to textual information, achieving efficient alignment of "visual features to language context". The visual context vector sequence output by the cross-modal adapter is used as the input to the language decoder.

[0050] In some implementations, the language decoder employs a large-scale pre-trained language model; the large-scale pre-trained language model is selected from at least one of DeepSeek-MoE, Qwen3-0.5B, Qwen3-1.5B, or Qwen3-4B.

[0051] Specifically, the visual context vector sequence is injected into the Transformer layer of the language decoder as prefix embeddings or key-value pairs for cross-attention, enabling it to perceive visual semantic information in the image at each step of word prediction. Simultaneously, to guide the model to focus on the answer format for math fill-in-the-blank questions, the system also appends a pre-defined prompt to guide the language model in generating answers according to a specific format, content, and style. For example: "Please output the candidate's answer based on the handwritten content in the image, containing only mathematical expressions, numbers, or units, without explanation."

[0052] Under this constraint, the language decoder gradually generates a string of free text, i.e., the original answer text (such as "1 / 2", "0.5", "5cm" or "x=2"). Although this text can reflect the candidate's intention, it has not yet been standardized and may contain ambiguities, inconsistent formats, or slight recognition errors. Therefore, it needs to be semantically normalized by a generative language model and transformed into a structured standard mathematical representation.

[0053] Thus, we can see that, firstly, VisionTransformer is used as the visual encoder, leveraging its global self-attention mechanism to effectively capture complex two-dimensional spatial structures such as subscripts, superscripts, fractions, and square roots in formulas; secondly, a cross-modal adapter accurately maps visual features to the semantic space of the language model, achieving semantic alignment from image to text; finally, combined with a language decoder using preset prompts, guided by the visual context, the original answer text conforming to mathematical expression habits is generated. This multimodal deep model architecture overcomes the limitations of traditional OCR, which relies solely on local character recognition while ignoring the overall structural semantics, thereby achieving high-precision recognition in highly nested, non-linearly laid-out handwritten formula scenarios.

[0054] In step 140, the original answer text is semantically understood and normalized using a generative language model to generate a structured standard mathematical representation.

[0055] Among them, generative language models are large-scale pre-trained language models specifically optimized for mathematical semantic understanding and equivalent reasoning tasks. They are used to perform semantic parsing, ambiguity resolution, and format standardization on the original answer text output by multimodal deep models to generate structured standard mathematical representations.

[0056] In one specific implementation, the generative language model is selected from at least one of MathBERT, T5-Math, or a self-developed mathematical large language model. During the pre-training stage, it integrates mathematical formulas, textbook exercises, and error samples, and during the fine-tuning stage, it introduces an equivalence contrast learning objective to enhance robustness to common error patterns in educational assessment scenarios.

[0057] For example, the generative language model receives non-normalized free text such as "1 / 2", "0.5", "5cm" or "x^2+1" as input. Based on the semantic knowledge learned from large-scale mathematical corpora, it performs writing ambiguity resolution (unifying equivalent but different expressions into a preset standard form), mathematical equivalence modeling (identifying numerical equivalence relations, algebraic equivalence transformations and approximation tolerances), and structural normalization (calling the symbolic computation engine to perform algebraic simplification, unit separation or factorization to ensure that the output meets the formal constraints required by the problem).

[0058] For example, taking the unification of equivalent but formally different expressions into a pre-defined standard form as an example, "1 / 2" is converted to "\frac{1}{2}" in LaTeX format. For example, when identifying numerical equivalence relations, 0.5 = 1 / 2 = 50%. For example, when taking algebraic equivalence transformations, a + b = b + a. For example, when taking approximate tolerance, |3.14 - π| < 0.01 is considered correct. For example, when taking structural normalization, the identified "\frac12" is completed into the legal LaTeX expression "\frac{1}{2}", the identified "sqrt4" is corrected to "\sqrt{4}", and the implicit expression "sin2x" is explicitly defined as "\sin(2x)".

[0059] Standard mathematical representations include LaTeX format, normalized numerical values, unit identifiers, and semantic tags. The LaTeX format uses standard LaTeX syntax to encode mathematical expressions. Examples include "\frac{1}{2}", "x^{2}+y_{1}", and "\sqrt{a+b}". This ensures that the formula structure is complete, renderable, and parsable, avoiding common syntax omissions found in the original recognition results. For example, "\frac12" lacks curly braces.

[0060] Standardized values ​​involve converting all numerical values ​​in the answers to a standardized form. For example, fractions are forced to their simplest proper fractions. For instance, "4 / 8" is converted to "1 / 2". Decimals are retained to the specified number of decimal places as required by the question. For example, two decimal places are retained. Irrational numbers or constants are mapped to standard signs within a tolerance range. For example, "3.14" is converted to "π".

[0061] The unit identifier is as follows: when the answer involves physical quantities, the numerical value and unit are automatically separated, and the unit type and dimension are identified by a structured field. For example, the identified "5cm" is parsed as {value:5,unit:"cm",dimension:"length"}, supporting subsequent unit consistency checks and conversions.

[0062] Semantic labels are high-level semantic information extracted from the context of the question and the content of the answer, used to assist in scoring decisions. For example, labels such as "simplest fraction", "factored", "containing radicals", "approximate value", and "error pattern: missing negative sign" provide interpretable features for the weighted scoring model.

[0063] The four elements of LaTeX format, standardized numerical values, unit identifiers, and semantic tags together constitute a machine-computable, comparable, and interpretable structured representation, which serves as the basis for subsequent multi-dimensional comparisons of structural equivalence, numerical equivalence, and formal compliance.

[0064] Specifically, in some implementations, the step "using a generative language model to perform semantic understanding and normalization on the original answer text to generate a structured standard mathematical representation" may include the following steps:

[0065] (1) The original answer text is written to resolve ambiguity and unify the format, and equivalent but different expressions are converted into the preset intermediate standard form; equivalent expressions include the conversion between fractions, decimals and percentages, the equivalent forms of the commutative law of algebraic expressions, the equivalent forms of the associative law of algebraic expressions, and the approximate values ​​within the preset error range;

[0066] (2) Call the symbolic computation engine to perform algebraic simplification, unit conversion or numerical approximation on the intermediate standard form to generate a standard mathematical representation.

[0067] Identify and eliminate non-essential differences in expression caused by candidates' writing habits, input diversity, or recognition errors, and map semantically equivalent but formally different answers to a unified intermediate standard form so that subsequent steps can accurately compare and score them.

[0068] For example, Different expressions such as “1 / 2”, “0.5”, and “50%” are uniformly converted into preset standard forms, such as “\frac{1}{2}” or “0.50” with two decimal places. The choice depends on the requirements of the question or the system default strategy.

[0069] For example, we can use mathematical laws of operation (e.g., the commutative law of addition and the associative law of multiplication) to determine whether expressions are equivalent. For instance, we can consider "x+3" and "3+x" to be the same, or determine whether "2(x+1)" and "2x+2" are equivalent based on whether the question requires "expansion".

[0070] For example, answers to questions involving irrational numbers, measured values, or estimations are considered correct if they fall within a preset error range. For instance, if the reference answer is π, and a test taker answers "3.14" with a system tolerance of ±0.01, the answer is considered equivalent.

[0071] The original answer text is written to resolve ambiguity and unify format. Equivalent but different expressions are converted into a preset intermediate standard form. This is accomplished by a generative language model combined with a built-in mathematical equivalence rule library and contextual understanding ability. Although the output intermediate standard form has not yet undergone final simplification by the symbolic computation engine, it has eliminated the main ambiguities, laying the foundation for subsequent algebraic simplification, unit conversion and form compliance verification.

[0072] Calling a symbolic computation engine refers to passing the intermediate standard form output by the generative language model to a symbolic computation system with precise mathematical reasoning capabilities for structured processing, in order to generate a final standard mathematical representation that can be used for intelligent scoring. Calling a symbolic computation engine performs calculations based on deterministic mathematical rules, avoiding the illusions or errors that may occur in pure language models during complex mathematical transformations.

[0073] For example, the symbolic computation engine standardizes and simplifies algebraic expressions. For instance, it merges "2x+3x" into "5x", reduces "\frac{4}{8}" to "\frac{1}{2}", or expands "(x+1)^2" into "x^2+2x+1".

[0074] For example, when the answer contains physical quantities, it automatically converts them according to the International System of Units (SI) or the unit system specified in the question. For instance, it standardizes "500g" to "0.5kg", or converts "2h30min" to "2.5h".

[0075] For example, for answers containing irrational numbers or complex expressions, a numerical approximation can be generated at a preset precision. For instance, “\sqrt{2}” can be calculated as “1.414” (rounded to three decimal places), or “\pi / 2” can be approximated as “1.571”, so as to compare with the numerical reference answer.

[0076] After processing by the symbolic computation engine, it outputs a standard mathematical representation that is syntactically correct, structurally sound, and numerically consistent, serving as a reliable basis for multi-dimensional scoring comparison in subsequent steps.

[0077] Furthermore, to verify whether the standard mathematical representation accurately reproduces the structural semantics of the candidate's original handwritten content, this application also introduces a formula-based rendering verification mechanism. Specifically, in some embodiments, the handwritten fill-in-the-blank question grading method for AI-based marking further includes the following steps:

[0078] (1) Use a formula rendering tool to render the standard mathematical representation into an image, and compare its structural similarity with the corresponding answer area image to obtain the similarity score;

[0079] (2) If the similarity is lower than the preset threshold, it is marked as pending manual review.

[0080] Using a formula rendering tool (e.g., KaTeX, MathJax, or LaTeX engine), standard mathematical representations (e.g., "\frac{1}{2}" or "x^{2}+y_{1}") are rendered as high-fidelity vector images, generating a standard answer rendering image. Then, this rendering image is compared with its corresponding answer region image for structural similarity.

[0081] The comparison uses image quality assessment algorithms (e.g., Structural Similarity Index SSIM, Perceptual Hash, or Deep Feature Cosine Similarity) to calculate the visual consistency between the rendered image and its corresponding response region image in terms of character layout, symbol shape, spatial relationship, etc., and outputs a similarity score between 0 and 1.

[0082] If the similarity is lower than the preset threshold (e.g., SSIM < 0.75), it indicates that the standardization process may introduce structural biases (e.g., incorrect parsing of subscripts and superscripts, omission of square root range, misjudgment of score structure, etc.). The system will automatically mark the question as awaiting manual review and submit it to the examiner for final decision, thereby effectively intercepting the risk of misjudgment caused by identification or standardization errors.

[0083] Thus, leveraging the powerful semantic understanding and transformation capabilities of generative language models, the original answer texts provided by test takers underwent in-depth analysis and precise standardization. This process not only encompassed basic grammatical error correction and symbol standardization (e.g., unifying fraction formats and exponential expressions), but more importantly, it achieved a deep understanding and transformation of mathematical concepts, thereby generating a structured standard mathematical representation that strictly adheres to mathematical expression norms while retaining the complete semantic information of the original answers.

[0084] In step 150, the standard mathematical representation is compared with the preset reference answer in multiple dimensions, and the scoring result is output.

[0085] In order to accurately assess the quality of candidates' answers, the obtained standard mathematical representations are compared with the preset reference answers in a detailed multi-dimensional manner, so as to provide a comprehensive and objective scoring result for each candidate's answer.

[0086] Among them, the preset reference answer can be the standard correct answer that the question setter or the question bank system pre-sets for each fill-in-the-blank question before grading. It not only includes the final result, but also encodes all the semantic and format constraints required for grading in a structured form.

[0087] In some implementations, the preset reference answer includes a standard mathematical expression, formal constraint rules, equivalence tolerance parameters, and scoring weight configuration.

[0088] Taking a standard mathematical expression as an example, the correct answer is represented by a canonical LaTeX or symbolic expression tree, such as "\frac{1}{2}", "x^{2}+4x", or "{value:9.8,unit:"m / s^2"}.

[0089] For example, formal constraint rules specify the answer format required by the question, such as "simplest fraction", "round to two decimal places", "must be factored", and "no decimals allowed".

[0090] Taking the equivalence tolerance parameter as an example, it includes the allowable range of numerical errors (e.g., ±0.01), the acceptable set of algebraic equivalent transformations (e.g., commutative law, associative law), and the unit conversion strategy (e.g., g and kg are interchangeable).

[0091] Taking the scoring weight configuration as an example, the scoring weights for the three dimensions of structural equivalence, numerical equivalence and formal compliance are as follows (e.g., w1=0.5, w2=0.3, w3=0.2, and the deduction rules).

[0092] Specifically, in some implementations, the step of "comparing the standard mathematical representation with the preset reference answer in multiple dimensions and outputting the scoring result" may include the following steps:

[0093] (1) Construct mathematical expression syntax trees corresponding to standard mathematical representations and preset reference answers respectively;

[0094] (2) Based on the tree edit distance and the preset mathematical equivalence transformation rule base, whether the standard mathematical representation and the preset reference answer satisfy structural equivalence is obtained, and the structural equivalence score is obtained;

[0095] (3) Perform numerical analysis and calculation on the standard mathematical representation and the preset reference answer respectively, compare the results within the preset floating-point error tolerance, and obtain the numerical equivalence score;

[0096] (4) Verify whether the standard mathematical representation meets the formal constraints specified in the question and obtain a score for formal compliance; formal constraints include the simplest fraction, the number of decimal places to retain, unit integrity or factorization requirements;

[0097] (5) Based on the structural equivalence score, numerical equivalence score and formal compliance score, combined with the deduction items corresponding to the identified error patterns, a weighted fusion model is used to generate the scoring results.

[0098] To achieve accurate structural comparison between test takers' answers and standard answers, the system first parses the standard mathematical representation and the preset reference answer into Mathematical Expression Trees (METs). A Mathematical Expression Tree is a data structure that represents the semantic relationships of mathematical expressions in a tree structure.

[0099] In a mathematical expression syntax tree, leaf nodes represent operands (e.g., numbers, variables, or constants), internal nodes represent operators or functions (e.g., +, ×, \frac{}{}, \sqrt{}, \sin, etc.), and parent-child relationships reflect the precedence and nesting structure of operations (e.g., the left subtree of a fraction node is the numerator, and the right subtree is the denominator; the left child of an exponentiation node is the base, and the right child is the exponent).

[0100] Mathematical expression syntax trees can be automatically constructed from LaTeX or structured expressions using a recursive descent parser or a symbolic computation engine (e.g., SymPy's parse_expr module). For example, the expression \frac{x+1}{2} will be parsed into a tree structure with \frac{}{} as the root node, x+1 as the left subtree (which itself is a + node), and 2 as the right subtree.

[0101] After the mathematical expression syntax tree is constructed, a comprehensive evaluation is performed using a combination of tree edit distance (TED) and a pre-defined mathematical equivalence transformation rule base. Tree edit distance refers to the editing cost required to transform one syntax tree into another using the fewest node insertions, deletions, or replacements. The smaller the distance, the more similar the structures of the two expressions. However, relying solely on edit distance cannot identify semantically equivalent but structurally different expressions (e.g., a+b vs. b+a), therefore a mathematical equivalence transformation rule base is introduced.

[0102] The mathematical equivalence transformation rule base can include operational laws, identity transformations, functional identities, and rational / radical standardization rules. For example, operational laws include the commutative law of addition (x+y=y+x) and the associative law of multiplication ((xy)z=x(yz)); identity transformations include… Taking the functional identity as an example, sin 2 x+cos 2 x = 1, log(ab) = log(a) + log(b); Taking the standardization rule of fractions / radicals as an example,

[0103] To overcome the limitations of pure structural comparison in handling approximate values, irrational numbers, or numerical answers, the system further performs numerical equivalence verification. Standard mathematical representations and preset reference answers are input into the symbolic computation engine or numerical parser for automatic evaluation to obtain numerical equivalence scores.

[0104] For expressions containing variables (such as algebraic expressions), the system uses a multi-set random test point substitution method, selecting several combination values ​​(e.g., x = 1, 2, -0.5, etc.) within the domain of the question, and substituting them into the candidate's answer and the reference answer respectively to calculate the corresponding output value; for constant expressions without variables (e.g., π / 2, 2 or 3.14), its floating-point approximation value is directly calculated.

[0105] The system then compares the calculation results of the two systems point by point within a preset floating-point error tolerance (e.g., absolute error ≤ 0.01% or relative error ≤ 1%). If the difference between all test points is within the tolerance range, the results are considered numerically equivalent; otherwise, they are considered inequivalent. This tolerance can be dynamically configured according to the type of problem; for example, physics problems allow for larger errors, while pure mathematics problems require higher precision.

[0106] The numerical equivalence score is determined by a combination of the proportion of successfully matched test points and the magnitude of the error. For example, if all test points meet the tolerance requirements, the score is 1.0; if some points exceed the tolerance, the score decreases linearly according to the degree of exceedance; if key points (such as domain boundaries) are mismatched, the score is directly set to 0. This score serves as the core indicator for measuring the "correctness of the result" in the multi-dimensional scoring model, and is particularly suitable for approximate calculation problems that cannot be completely determined through structure.

[0107] Formal compliance verification is used to determine whether the candidate's answer conforms to the expression format requirements explicitly specified in the question. Even if the answer is correct in structure or value, if it does not follow the specified formal norms, the corresponding points should still be deducted to reflect the assessment of the standardization of mathematical writing and the rigor of problem-solving.

[0108] Specifically, the system extracts the formal constraints specified in the question from the question annotation information or metadata, and applies them as validation rules to the standard mathematical representation. Formal constraints may include: simplified fractions, number of decimal places, unit integrity, and factorization requirements.

[0109] Taking the simplest fraction as an example, the numerator and denominator must be coprime and the denominator must be positive. For example, “\frac{4}{8}” is considered non-compliant, while “frac{1}{2}” is not. Regarding the number of decimal places, the result must be rounded to two decimal places, so “3.14159” should be output as “3.14”. Answers like “3.1” or “3.142” are considered non-compliant. Regarding unit completeness, when the question involves physical quantities, the answer must include the correct unit. For example, “speed is __--__” requires an answer in the form of “{value:5,unit:"m / s"}”; simply filling in 5 is considered a missing unit. Regarding factorization requirements, if the question asks to “factor the polynomial”, the answer must be in product form (e.g., (x-1)(x+2)). If it is still an expansion of x^2+x-2, even if algebraically equivalent, it is considered a non-compliant form.

[0110] Formal compliance score is generated based on constraint satisfaction. If all formal requirements are met, the score is 1.0; if one or more requirements are not met, points are deducted according to preset weights (e.g., 0.3 for missing units, 0.2 for decimal place errors), with a minimum score of 0. This score serves as a key indicator for measuring "answer standardization" in the multi-dimensional scoring model, ensuring that the scoring focuses not only on whether the answer is correct but also on whether it is written in a standardized manner.

[0111] The system comprehensively scores the three dimensions of structure, numerical values, and format, and deducts points based on typical errors identified (such as missing units or incorrect symbols) to generate a final score. This ensures accuracy while reflecting teaching feedback on common errors.

[0112] Please see Figure 2 , Figure 2 This illustration shows a structural diagram of a handwritten fill-in-the-blank question grading device for AI-based exam marking, according to an embodiment of this application. The device 200 includes: a data acquisition module 210, a first execution module 220, a second execution module 230, a third execution module 240, and an output module 250. Specifically:

[0113] Acquisition module 210 is used to acquire images of the target answer sheet;

[0114] The first execution module 220 is used to match the target answer sheet image with the preset answer sheet template to obtain the answer area image corresponding to each fill-in-the-blank question;

[0115] The second execution module 230 is used to convert the handwritten answer content in the answer area map into the original answer text through a multimodal deep model; the multimodal deep model integrates a visual encoder, a cross-modal adapter, and a language decoder;

[0116] The third execution module 240 is used to perform semantic understanding and normalization processing on the original answer text through a generative language model, and generate a structured standard mathematical representation; the standard mathematical representation includes LaTeX format, normalized numerical values, unit identifiers and semantic tags;

[0117] Output module 250 is used to compare the standard mathematical representation with the preset reference answer in multiple dimensions and output the scoring results.

[0118] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0119] In the several embodiments provided in this application, the coupling or direct coupling or communication connection between the modules shown or discussed may be an indirect coupling or communication connection through some interface, device or module, and may be electrical, mechanical or other forms.

[0120] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0121] Please see Figure 3 , Figure 3 This illustration shows a structural diagram of a handwritten fill-in-the-blank question grading device for artificial intelligence-based marking, provided in an embodiment of this application. The handwritten fill-in-the-blank question grading device 300 for artificial intelligence-based marking in this application may include one or more of the following components: a processor 310, a memory 320, and one or more application programs. The one or more application programs may be stored in the memory 320 and configured to be executed by one or more processors 310. The one or more programs are configured to execute the handwritten fill-in-the-blank question grading method for artificial intelligence-based marking as described in the foregoing method embodiments.

[0122] The processor 310 may include one or more processing cores. The processor 310 connects to various parts of the handwritten fill-in-the-blank question grading device 300 for AI-based grading via various interfaces and lines. It executes various functions and processes data within the device 300 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 320, and by calling data stored in the memory 320. Optionally, the processor 310 may be implemented using at least one of the following hardware forms: Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 310 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and Modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understandable that the aforementioned modem may not be integrated into the processor 310, but may be implemented using a separate communication chip.

[0123] The memory 320 may include random access memory (RAM) or read-only memory (ROM). The memory 320 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 320 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created during use by the handwritten fill-in-the-blank question grading device 300 for AI-based grading.

[0124] Please see Figure 4 , Figure 4 The diagram illustrates the structure of a computer-readable storage medium 400 provided in an embodiment of this application. The computer-readable storage medium 400 stores program code, which can be called by a processor to execute the handwritten fill-in-the-blank grading method for artificial intelligence grading described in the above method embodiment.

[0125] The computer-readable storage medium 400 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has storage space for program code 410 that performs any of the method steps described above. This program code can be read from or written to one or more computer program devices. The program code 410 may be compressed, for example, in a suitable form.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A handwriting fill-in-the-blank test grading method for artificial intelligence marking, characterized in that, The method comprises: Collecting a target answer sheet image; Matching the target answer sheet image with a preset sheet template to obtain an answer area graph corresponding to each fill-in-the-blank question; Converting handwritten answer content in the answer area graph into original answer text through a multi-modal deep model; the multi-modal deep model is fused with a visual encoder, a cross-modal adapter, and a language decoder; Performing semantic understanding and normalization processing on the original answer text through a generative language model to generate a structured standard mathematical representation; the standard mathematical representation includes LaTeX format, standardized numerical values, unit identifiers, and semantic labels; Comparing the standard mathematical representation with a preset reference answer in multiple dimensions to output a scoring result.

2. The handwriting fill-in-the-blank test grading method for artificial intelligence grading according to claim 1, wherein, The conversion of the handwritten answer content in the answer area graph into the original answer text through the multi-modal deep model comprises: Performing global feature extraction on the answer area graph through the visual encoder to generate a high-dimensional semantic feature sequence; Performing semantic alignment and dimension mapping on the high-dimensional semantic feature sequence through the cross-modal adapter to convert it into a visual context vector sequence compatible with the embedding space of the language decoder; The language decoder generates the original answer text as conditional input of the visual context vector sequence combined with a preset prompt instruction.

3. The method for grading a handwriting fill-in-the-blank question for artificial intelligence scoring according to claim 1 or 2, wherein The visual encoder adopts a Vision Transformer-based architecture that divides the answer area graph into multiple image blocks and models the spatial dependency between the image blocks through a global self-attention mechanism; the Vision Transformer-based architecture is selected from at least one of SigLIP-ViT, CLIP-ViT, or MobileViT.

4. The handwriting fill-in-the-blank test grading method for artificial intelligence grading according to claim 1 or 2, characterized in that, The language decoder adopts a large-scale pre-trained language model; the large-scale pre-trained language model is selected from at least one of DeepSeek-MoE, Qwen3-0.5B, Qwen3-1.5B, or Qwen3-4B. 5.The method for handwriting fill-in-the-blank test grading for artificial intelligence scoring according to claim 1, wherein, The semantic understanding and normalization processing of the original answer text through the generative language model to generate the structured standard mathematical representation comprises: Performing writing ambiguity resolution and format unification on the original answer text to convert equivalent but different forms into a preset intermediate standard form; the equivalent expressions include conversion between fractions, decimals, and percentages, commutative law equivalent forms of algebraic expressions, associative law equivalent forms of algebraic expressions, and approximate values within a preset error range; Calling a symbol calculation engine to perform algebraic simplification, unit conversion, or numerical approximation processing on the intermediate standard form to generate the standard mathematical representation. 6.The method for grading a handwriting fill-in-the-blank question using artificial intelligence according to claim 5, wherein, The method further comprises: Rendering the standard mathematical representation into an image using a formula rendering tool and performing structure similarity comparison with the answer area graph corresponding thereto to obtain a similarity; If the similarity is lower than a preset threshold, marking it for manual review. 7.The method for handwriting fill-in-the-blank test grading for artificial intelligence scoring according to claim 1, wherein, The comparison of the standard mathematical representation with the preset reference answer in multiple dimensions to output the scoring result comprises: constructing a mathematical expression syntax tree corresponding to the standard mathematical expression and the preset reference answer respectively; based on the tree edit distance and the preset mathematical equivalence transformation rule library, whether the standard mathematical expression and the preset reference answer satisfy structural equivalence is obtained, and a structural equivalence score is obtained; the standard mathematical expression and the preset reference answer are respectively subjected to numerical analysis and calculation, and the results are compared within the preset floating point error tolerance, and a numerical equivalence score is obtained; checking whether the standard mathematical expression satisfies the form constraint specified by the question, and obtaining a form compliance score; the form constraint includes the simplest fraction, the number of decimal places, the integrity of the unit or the factor decomposition requirement; based on the structural equivalence score, the numerical equivalence score and the form compliance score, and combined with the deduction items corresponding to the identified error modes, the scoring result is generated through a weighted fusion model.

8. A handwriting fill-in-the-blank test grading device for artificial intelligence marking, characterized by, The device comprises: a collection module for collecting a target answer sheet image; a first execution module for matching the target answer sheet image with a preset sheet template to obtain an answer area image corresponding to each fill-in-the-blank question; a second execution module for converting the handwritten answer content in the answer area image into original answer text through a multi-modal deep model; the multi-modal deep model integrates a visual encoder, a cross-modal adapter and a language decoder; a third execution module for performing semantic understanding and standardization processing on the original answer text through a generative language model to generate a structured standard mathematical expression; the standard mathematical expression includes LaTeX format, standardized numerical value, unit identifier and semantic label; an output module for comparing the standard mathematical expression with a preset reference answer in multiple dimensions to output a scoring result.

9. A handwriting fill-in-the-blank grading device for artificial intelligence grading, comprising: comprise: one or more processors; a memory; one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the handwriting fill-in-the-blank question correction method for artificial intelligence marking as claimed in any one of claims 1-7.

10. A computer readable storage medium, characterized in that, The computer readable storage medium stores program code, and the program code can be called and executed by the processor to execute the handwriting fill-in-the-blank question correction method for artificial intelligence marking as claimed in any one of claims 1-7.