Intelligent paper marking method, device and equipment for open education drawing questions and storage medium

By combining a pre-set detection model with a multimodal human-like cognitive model, open-ended drawing questions are automatically identified and scored, solving the problem of time-consuming and labor-intensive manual review and achieving efficient, fair, and interpretable scoring results.

CN121661664APending Publication Date: 2026-03-13SHENZHEN SEA SKY LAND TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Scoring open-ended drawing questions relies on manual review, which is time-consuming, labor-intensive, and costly. Furthermore, the scoring is highly subjective and inconsistent, making it difficult to guarantee the fairness and efficiency of the assessment.

Method used

The system uses a pre-defined detection model to locate the bounding box coordinates and question type labels of drawing questions, crops the target answer sheet image, dynamically loads chain-like reasoning prompt templates, calls a multimodal human-like cognitive model to generate implicit reasoning chains, verifies based on implicit reasoning chains, and obtains scoring results.

Benefits of technology

It enables intelligent grading of open-ended drawing questions, improving the fairness and efficiency of the assessment, reducing the time and cost of manual grading, and ensuring the consistency and interpretability of the grading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661664A_ABST
    Figure CN121661664A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent paper marking method and device for open education drawing questions, equipment and a storage medium, and the method comprises the steps: inputting a collected target answer sheet image into a preset detection model, and determining bounding box coordinates and question type labels corresponding to the open education drawing questions included in the target answer sheet image; according to the bounding box coordinates, cutting in the target answer sheet image to obtain a target answer area sub-image marked with a question type label; a chain thinking reasoning prompt template corresponding to the question type label is dynamically loaded, a multi-mode-based human-like cognition model is called, reasoning is executed in a target mode, and an implicit reasoning chain is generated; and based on an implicit reasoning chain, verifying the target answer area sub-graph to obtain a scoring result. Therefore, through mutual cooperation of the preset detection model and the multi-modal human-like cognitive model, intelligent paper marking of open drawing questions is realized, and the fairness and efficiency of evaluation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent marking, and more specifically, to an intelligent marking method, apparatus, device, and storage medium for open-ended educational drawing questions. Background Technology

[0002] Open-ended drawing questions play a vital role in science and engineering education. Unlike multiple-choice and short-answer questions, these questions require students to express their understanding of scientific concepts through drawing. Examples include circuit diagrams, optical path diagrams, mechanical diagrams, or statistical charts. Open-ended drawing questions comprehensively assess students' spatial imagination, logical reasoning, and scientific literacy, and are important indicators for measuring higher-order cognition and creativity. Therefore, they are widely used in physics, electricity, and general science examinations.

[0003] In related technologies, the scoring of open-ended drawing questions has long relied on manual review. Manual scoring is not only time-consuming, labor-intensive, and costly, especially in large-scale standardized tests; moreover, faced with students' diverse writing styles, ambiguous markings, or atypical answers, different examiners are prone to misunderstandings, resulting in strong subjectivity and poor consistency in scoring, making it difficult to guarantee the fairness and efficiency of the assessment. Summary of the Invention

[0004] In view of the above problems, this application proposes an intelligent marking method, device, equipment and storage medium for open-ended educational drawing questions, which can solve the above problems.

[0005] In a first aspect, embodiments of this application provide an intelligent marking method for open-ended educational drawing questions. The method includes: inputting a collected target answer sheet image into a preset detection model to determine the bounding box coordinates and question type labels corresponding to the open-ended educational drawing questions included in the target answer sheet image; cropping a target answer area sub-image labeled with question type labels from the target answer sheet image based on the bounding box coordinates; dynamically loading a chain-like reasoning prompt template corresponding to the question type labels and calling a multimodal human-like cognitive model to perform reasoning in the target mode to generate an implicit reasoning chain; and verifying the target answer area sub-image based on the implicit reasoning chain to obtain a scoring result.

[0006] Secondly, embodiments of this application also provide an intelligent marking device for open-ended educational drawing questions. The device includes: a detection module, used to input a collected target answer sheet image into a preset detection model to determine the bounding box coordinates and question type labels corresponding to the open-ended educational drawing questions included in the target answer sheet image; a cropping module, used to crop a target answer area sub-image labeled with question type labels from the target answer sheet image based on the bounding box coordinates; a reasoning module, used to dynamically load chain-like reasoning prompt templates corresponding to the question type labels and call a multimodal human-like cognitive model to perform reasoning in the target mode, generating an implicit reasoning chain; and a scoring module, used to verify the target answer area sub-image based on the implicit reasoning chain to obtain a scoring result.

[0007] Thirdly, embodiments of this application also provide a cooking device, including a processor, a memory, and one or more applications; the one or more applications are stored in the memory and configured to be executed by the processor to implement the above-described intelligent marking method for open-ended educational drawing questions.

[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium storing program code, wherein the above-described intelligent grading method for open-ended educational drawing questions is executed when the program code is run by a processor.

[0009] The technical solution provided in this application includes the following method: inputting the collected target answer sheet image into a preset detection model to determine the bounding box coordinates and question type labels corresponding to the open-ended educational drawing questions included in the target answer sheet image; cropping a target answer area sub-image labeled with question type labels from the target answer sheet image based on the bounding box coordinates; dynamically loading the chain-like reasoning prompt template corresponding to the question type labels, and calling a multimodal human-like cognitive model to perform reasoning in the target mode to generate an implicit reasoning chain; and verifying the target answer area sub-image based on the implicit reasoning chain to obtain the scoring result. Thus, by combining the preset detection model and the multimodal human-like cognitive model, intelligent grading of open-ended drawing questions is achieved, improving the fairness and efficiency of the assessment. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments and drawings obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0011] Figure 1The illustration shows a flowchart of an intelligent grading method for open-ended educational drawing questions provided in an embodiment of this application.

[0012] Figure 2 This illustration shows a structural schematic diagram of an intelligent marking device for open-ended educational drawing questions provided in an embodiment of this application.

[0013] Figure 3 This illustration shows a structural schematic diagram of an intelligent marking device for open-ended educational drawing questions, provided in an embodiment of this application.

[0014] Figure 4 This illustration shows a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0016] Open-ended drawing questions play a vital role in science and engineering education. Unlike multiple-choice and short-answer questions, these questions require students to express their understanding of scientific concepts through drawing. Examples include circuit diagrams, optical path diagrams, mechanical diagrams, or statistical charts. Open-ended drawing questions comprehensively assess students' spatial imagination, logical reasoning, and scientific literacy, and are important indicators for measuring higher-order cognition and creativity. Therefore, they are widely used in physics, electricity, and general science examinations.

[0017] In related technologies, the scoring of open-ended drawing questions has long relied on manual review. Manual scoring is not only time-consuming, labor-intensive, and costly, especially in large-scale standardized tests; moreover, faced with students' diverse writing styles, ambiguous markings, or atypical answers, different examiners are prone to misunderstandings, resulting in strong subjectivity and poor consistency in scoring, making it difficult to guarantee the fairness and efficiency of the assessment.

[0018] To address the aforementioned issues, this application provides an intelligent marking method, apparatus, device, and storage medium for open-ended educational drawing questions. The method includes: inputting a collected target answer sheet image into a preset detection model to determine the bounding box coordinates and question type labels corresponding to the open-ended educational drawing questions included in the target answer sheet image; cropping a target answer area sub-image labeled with question type labels from the target answer sheet image based on the bounding box coordinates; dynamically loading a chain-like reasoning prompt template corresponding to the question type labels and calling a multimodal human-like cognitive model to perform reasoning in the target mode, generating an implicit reasoning chain; and verifying the target answer area sub-image based on the implicit reasoning chain to obtain the scoring result.

[0019] Therefore, by combining a pre-set detection model with a multimodal human-like cognitive model, intelligent grading of open-ended drawing questions can be achieved, improving the fairness and efficiency of the assessment.

[0020] Please see Figure 1 , Figure 1 This illustration shows a flowchart of an intelligent grading method for open-ended educational drawing questions, provided in an embodiment of this application. Figure 1 As shown, the method may include steps 110 to 140.

[0021] In step 110, the collected target answer sheet image is input into the preset detection model to determine the bounding box coordinates and question type labels corresponding to the open-ended educational drawing questions included in the target answer sheet image.

[0022] In some implementations, the target answer sheet image may include a digital image of the student's answers to open-ended educational drawing questions. In one specific implementation, a document scanner or high-speed scanner is used to digitize the paper answer sheet to capture the target answer sheet image. In another specific implementation, the student's hand-drawn answer area is photographed using the camera of a mobile terminal such as a smartphone or tablet to capture the target answer sheet image. In yet another specific implementation, the image is obtained from an image file (e.g., JPG, PNG format) uploaded by the user on an online education platform.

[0023] Preferably, the target answer sheet image should be clear, evenly lit, free from severe perspective distortion, and the drawing area should be fully visible.

[0024] In some implementations, each target answer sheet image is bound to a unique student ID metadata when it is entered into the system. This ID serves as an additional tag for the image and is continuously retained throughout the subsequent detection, recognition, cropping, reasoning, and scoring processes to ensure traceability of the results.

[0025] In some implementations, the preset detection model is a pre-trained and deployed target detection model used to automatically identify and locate the position and type of open-ended educational drawing questions from the entire target answer sheet image.

[0026] In one specific implementation, the preset detection model is built on a deep learning architecture (e.g., YOLOv8, Faster R-CNN, or DETR) and fine-tuned on an educational scene image dataset containing a large number of annotated drawing question regions and corresponding question type labels. The input to the preset detection model is the entire target answer sheet image, and the output is one or more detection results, each of which includes bounding box coordinates and question type labels.

[0027] In some implementations, bounding box coordinates are used to precisely describe the spatial location of the open-ended educational drawing question within the target answer sheet image. The target answer sheet image is then cropped based on the bounding box coordinates to extract a sub-image containing only the drawing answers.

[0028] Bounding box coordinates are typically represented by the four parameters of the rectangle. For example, the parameters are represented as (x... min y min x max y max ), (x min y min (x) represents the pixel coordinates of the top-left corner of the bounding box. max y max (x, y, w, h) represents the bottom right pixel coordinates of the bounding box. For example, the parameters can be represented as (x, y, w, h), where (x, y) is the top left corner coordinate of the bounding box, and (w, h) is the bottom right corner coordinate.

[0029] In some implementations, question type labels are semantic category identifiers associated with the detected drawing question area, indicating the subject or graphic type to which the question belongs. These labels are synchronously output by a pre-defined detection model during the detection process, serving as a key basis for subsequent steps (e.g., loading chained thinking prompt templates, activating subject rule knowledge bases, and selecting scoring strategies), thereby achieving intelligent grading that adapts to question types.

[0030] For example, question type tags include circuit diagrams, optical path diagrams, force analysis diagrams, statistical pie charts, and geometric construction questions.

[0031] By inputting the collected target answer sheet image into the preset detection model, the system can automatically and accurately locate the area containing open-ended educational drawing questions in the target answer sheet image, simultaneously identify the question type category, and output the corresponding bounding box coordinates and question type label.

[0032] Among them, the bounding box coordinates are used to accurately crop out the sub-image containing only the student's answer, effectively eliminating interference from irrelevant information, while the question type label serves as a key context for dynamically loading the chain-like thinking prompt template, subject rule knowledge base and scoring criteria that match the question, thereby achieving intelligent grading that adapts to question type.

[0033] This process significantly improves the system's automation level, processing robustness, and cross-question type generalization ability in real educational scenarios, providing a well-structured and semantically clear input foundation for multimodal human-like cognitive scoring.

[0034] In step 120, based on the bounding box coordinates, a target answer area sub-image labeled with question type is cropped from the target answer sheet image.

[0035] In some implementations, the target answer area sub-graph contains only the student's hand-drawn answer to a specific open-ended educational drawing question (e.g., circuit connections, light paths, force arrows, or statistical charts), and does not contain any distracting information from other parts of the test paper.

[0036] The target answer region sub-image is associated with its corresponding question type label (e.g., circuit diagram, optical path diagram, etc.) within the system, serving as a semantic context identifier for multimodal reasoning and scoring in subsequent steps. This pruning operation not only reduces the processing scope and improves computational efficiency but also provides the model with focused and clean visual input, significantly enhancing the accuracy of symbol recognition and rule verification.

[0037] In step 130, the chain-like reasoning prompt template corresponding to the question type label is dynamically loaded, and the reasoning based on the multimodal human-like cognitive model is invoked to perform reasoning in the target mode to generate an implicit reasoning chain.

[0038] In some implementations, the chain-thinking reasoning prompt template can be a structured text instruction template, pre-designed according to question type tags and stored in a prompt template library.

[0039] The chain-thinking reasoning prompt template guides the multimodal human-like cognitive model to think step by step according to the scoring logic of human experts, using natural language. For example: "First, check if the power supply is labeled with positive and negative terminals; second, determine if each branch forms a closed loop; then verify if the current direction is reasonable..."

[0040] During reasoning, the system dynamically loads the corresponding chain-like reasoning prompt template based on the detected question type label, and splices it with the target answer area sub-graph to form a multimodal input, thereby guiding the multimodal human-like cognitive model to focus on key scoring elements and improve the logic and interpretability of the reasoning.

[0041] In some implementations, the multimodal human-like cognitive model is a specialized artificial intelligence model built on a large-scale visual-language model (VLM). The multimodal human-like cognitive model is used to simulate the cognitive and scoring process of human teachers for open-ended drawing questions.

[0042] In one specific implementation, the multimodal human-like cognitive model is customized based on the Qwen3-VL-8B architecture and optimized by combining the GRPO algorithm with human feedback reinforcement learning (RLHF), enabling it to have symbol recognition, structural understanding, rule reasoning and semantic judgment capabilities in educational scenarios such as circuit diagrams, optical path diagrams and force diagrams.

[0043] In some implementations, the target pattern can be the thinking reasoning pattern of a multimodal human cognitive model.

[0044] Multimodal human-like cognitive models can simultaneously process image and text inputs and output scoring results and the internal reasoning representations supporting those results. Specifically, in some implementations, this intelligent grading method for open-ended educational drawing questions may include the following steps:

[0045] (1) Based on each sample in the educational drawing question scoring training set, call the multimodal human-like cognitive model to generate score prediction and implicit inference chain; the implicit inference chain is the internal intermediate representation on which the score prediction is generated.

[0046] (2) Calculate multi-dimensional reward signals based on internal intermediate representations; the multi-dimensional reward signals include visual structure matching scores, symbol semantic equivalence scores, and subject rule compliance scores.

[0047] (3) Construct a comprehensive reward function based on the multi-dimensional reward signal and the KL divergence regularization term;

[0048] (4) A human feedback reinforcement learning framework based on the GRPO algorithm is used to iteratively optimize the policy network of the multimodal human-like cognitive model using a comprehensive reward function in order to complete the training of the multimodal human-like cognitive model.

[0049] In some implementations, the training set for scoring educational drawing questions consists of a dataset of numerous labeled samples. Each sample in the training set includes a sub-image of the target answer area, standard reference information, and human expert scoring labels. The standard reference information includes the corresponding standard answer image, symbol annotations, subject-specific rules, and elements of a correct answer; the human expert scoring labels are binary (e.g., "correct / incorrect") or multi-level (e.g., 0–5 points) scores given by experienced teachers or scoring experts, serving as a supervisory signal.

[0050] The training set covers a variety of question types and typical error patterns (e.g., reversed polarity, unclosed light path, incorrect force direction, etc.) and is used to train and optimize the scoring ability of the multimodal human-like cognitive model.

[0051] In some implementations, the score prediction can be an automatic score output by a multimodal human-like cognitive model for the input target answer region subimage. The score prediction can be binary (e.g., "pass / fail"), multi-level scores (e.g., 0–5 points), or a probability distribution, representing the multimodal human-like cognitive model's comprehensive judgment on whether the answer conforms to subject-specific norms, logical correctness, and symbolic accuracy. The score prediction is not directly based on image pixel matching, but rather relies on logical deduction through an implicit inference chain generated by the model in the target pattern, thus possessing the scoring criteria of a human-like expert.

[0052] In some implementations, the implicit inference chain can be an intermediate representation sequence that is not explicitly output during the internal autoregressive inference stage of the multimodal human-like cognitive model in the process of generating rating predictions.

[0053] The implicit reasoning chain consists of a series of intermediate tokens related to the scoring, their corresponding hidden state vectors, and cross-modal attention weights. It encodes the stepwise analysis process of the multimodal human-like cognitive model on graph structure, symbolic semantics, and subject rules (e.g., "Power source detected → polarity is positive or negative → branch A is not closed → violating Kirchhoff's laws").

[0054] Although the implicit inference chain is not visible to the user, it serves as the internal basis for rating prediction and can be parsed by the system to calculate multi-dimensional reward signals, supporting reinforcement learning fine-tuning and interpretability verification.

[0055] In some implementations, the internal intermediate representation can be the internal semantic and structured information generated by the multimodal human-like cognitive model during the rating prediction process to support the scoring decision.

[0056] In one embodiment, the internal intermediate representations specifically include intermediate tokens in the implicit inference chain, their corresponding hidden state vectors, and spatial attention maps of these tokens on the target answer region subgraph. Although these representations are not output externally, they encode the model's understanding of key elements such as graph topology, symbol positions, and physical quantity relationships, and serve as the fundamental data source for calculating multi-dimensional reward signals.

[0057] In some implementations, visual structure matching scores can be used to measure the consistency between the target answer region subgraph and the standard answer in terms of overall graphic structure or topological relationship.

[0058] In some implementations, symbolic semantic equivalence scores can be used to assess the semantic equivalence of handwritten symbols or annotations in student responses within a target answer area subgraph to standard subject-specific terminology.

[0059] In some implementations, subject rule compliance scores can be used to determine whether a student's answer in a target answer area subgraph satisfies the formal physical or mathematical constraints of the corresponding subject.

[0060] In some implementations, the KL divergence regularization term can be a constraint introduced during Human Feedback Reinforcement Learning (RLHF) training to measure the difference in output distribution between the current multimodal human-like cognitive model and the reference policy model.

[0061] In this application, the KL divergence regularization term is incorporated into the comprehensive reward function and participates in the policy gradient update in the form of a forward negative sign. This prevents the multimodal human-like cognitive model from deviating excessively from its original semantic understanding ability during reinforcement learning, avoids overfitting the reward signal or generating semantic distortion and logical confusion in the scoring behavior, and thus ensures that the multimodal human-like cognitive model maintains language consistency, reasoning stability and generalization ability while improving the scoring accuracy.

[0062] In some implementations, the expression for the comprehensive reward function can be:

[0063] r total =r semantic +r visual +r reward-model -γD KL (π θ ||π efr )

[0064] Where, r total For the total reward, r semantic For symbolic semantic equivalence scores, r visual For visual structure matching score, r reward-model The score represents the compliance score with subject rules, γ represents the regularization strength, and D represents the score. KL This is the KL divergence regularization term.

[0065] Furthermore, during the deployment phase of the multimodal human-like cognitive model, in order to improve inference efficiency and reduce hardware resource requirements, the system adopts the vLLM high-performance inference framework for service-oriented deployment of the multimodal human-like cognitive model. The vLLM framework effectively reduces memory fragmentation of the key-value cache (KVCache) through memory optimization techniques such as PagedAttention, significantly improving batch processing throughput and response speed.

[0066] Furthermore, to reduce the memory usage and inference latency of the multimodal human-like cognitive model, 8-bit GPTQ (Generalized Post-Training Quantization) quantization is performed on the model weights before deployment. This quantization method, without relying on training data, compresses the original FP16 (16-bit floating-point) weights into INT8 (8-bit integers) while preserving the model's semantic understanding and inference capabilities to the maximum extent through a layer-by-layer error compensation mechanism.

[0067] After quantization, the model weight size was reduced from approximately 16GB to 8.9GB, and the memory usage was reduced by nearly 45%. As a result, the multimodal human-like cognitive model can be fully loaded and run efficiently on a single NVIDIA RTX 4090 GPU (24GB of memory). In contrast, the model before quantization required more memory than a single card and needed to be loaded collaboratively by two GPUs in parallel, which incurred communication overhead and scheduling complexity.

[0068] Thus, by constructing a reinforcement learning training framework for educational drawing problems, while calling a multimodal human-like cognitive model to generate score predictions, the implicit inference chain generated within it is used as an intermediate representation to calculate reward signals covering three dimensions: visual structure, symbolic semantics, and subject rules, and combined with the KL divergence regularization term to form a comprehensive reward function.

[0069] Based on this, the GRPO algorithm is used to iteratively optimize the model policy network, so that the multimodal human-like cognitive model can improve the scoring accuracy while maintaining the logical consistency and reasoning stability of human-like experts, effectively solving the problems of weak interpretability and difficulty in rule generalization in the automatic scoring of open-ended drawing questions.

[0070] Furthermore, in some implementations, the step of "dynamically loading the chain-like reasoning prompt template corresponding to the question type label, and calling the multimodal human-like cognitive model to perform reasoning in the target mode and generate an implicit reasoning chain" may include the following steps:

[0071] (1) Based on question type tags, retrieve and dynamically load from the pre-set prompt template library to guide the multimodal human-like cognitive model to focus on the key scoring elements in the target answer area sub-graph of the chain thinking reasoning prompt template;

[0072] (2) The target answer area sub-graph is spliced ​​with the chain thinking reasoning prompt template to form a multimodal input sequence, and the multimodal input sequence is input into the multimodal human-like cognitive model to trigger the multimodal human-like cognitive model to enter the target mode;

[0073] (3) In the target mode, the multimodal human-like cognitive model generates multiple intermediate tokens and corresponding spatial attention weights in the autoregressive generation process, forming an implicit inference chain.

[0074] Each template in the pre-set prompt template library is a structured natural language instruction designed to guide models in simulating the scoring logic of human experts.

[0075] For example, when the question type is labeled "circuit diagram", the corresponding template can be: "Please judge step by step: (1) whether the power supply is marked with positive and negative terminals; (2) whether each branch forms a closed loop; (3) whether the current direction conforms to the conventional flow direction...". For another example, when the question type is labeled "optical path diagram", the template is adjusted to focus on elements such as the angle of incidence, the refraction surface, and the normal.

[0076] During the reasoning phase, the system retrieves and loads matching chain-like reasoning prompt templates from the template library in real time based on the question type tags identified in the previous steps, and then splices the templates with the target answer area subgraph to form a multimodal input sequence.

[0077] The multimodal input sequence contains a preset control token, which is used to activate the Thinking inference mode (i.e., the target mode) of the multimodal human-like cognitive model, so that the model performs an internal implicit inference process before generating the final output.

[0078] In the Thinking inference mode, when the multimodal human-like cognitive model performs autoregressive inference, it outputs the hidden state vector of the last layer Transformer for each intermediate word generated, and calculates the cross-modal attention weight of the word to each spatial location of the input image.

[0079] The system performs semantic matching on intermediate words using a pre-defined logical keyword library (such as "closed", "polarity", "angle", etc.), retaining only the hidden states of semantically related words and their corresponding attention heatmaps. These retained representations are organized into a structured sequence according to their generation order, which is the implicit reasoning chain, used for subsequent symbol recognition, rule verification, and multi-dimensional reward calculation.

[0080] In step 140, the target answer region subgraph is verified based on the implicit inference chain to obtain the scoring result.

[0081] In some implementations, the scoring results can be automated evaluation outputs of a multimodal human-like cognitive model on the target answer region subgraph, used to characterize the correctness and standardization level of the student's answers.

[0082] In one specific implementation, the scoring result can be in binary form (e.g., "correct / incorrect" or "pass / fail"). In another specific implementation, the scoring result can be a multi-level score (e.g., a 0–5 scale, reflecting fine-grained differences such as partially correct, incorrect orientation, or complete absence). In yet another specific implementation, the scoring result can be structured feedback (e.g., "power polarity error, -1 point; branch not closed, -2 points").

[0083] It is worth noting that the scoring results are not directly output from the end-to-end network, but are generated based on a multi-dimensional rule verification process supported by an implicit reasoning chain, thus possessing stronger interpretability, logical consistency, and educational adaptability.

[0084] The system extracts structured scoring elements (such as whether the element exists, whether the connection is closed, whether the symbol polarity is correct, and whether the angle or value is reasonable) by parsing the intermediate words retained in the implicit inference chain, their hidden state vectors, and the corresponding spatial attention weights.

[0085] Subsequently, these elements are substituted into the subject rule knowledge base that matches the question type label (for example, circuit diagrams must satisfy Kirchhoff's laws, and light path diagrams must conform to Snell's law of refraction), and logical consistency checks and numerical compliance judgments are performed.

[0086] Finally, based on the number of dimensions that passed verification, the severity of errors, and the rule weights, a quantitative or categorical scoring result is generated. Specifically, in some implementations, the step "verifying the target answer region subgraph based on the implicit inference chain to obtain the scoring result" may include the following steps:

[0087] (1) A structured representation of the target answer region subgraph is generated based on the implicit inference chain, and the structured representation is matched with the standard reference structured representation to obtain the visual structure verification result;

[0088] (2) Based on the implicit inference chain, the semantic equivalence of the symbol content in the target answer area subgraph is judged to determine the symbol verification result;

[0089] (3) Extract semantic information related to physical quantities, directions, polarities, connections or numerical logic from the target answer area subgraph based on implicit reasoning chains, and deduce or verify the semantic information with the target subject rule knowledge base to obtain rule verification results;

[0090] (4) Determine the scoring results based on the visual structure verification results, symbol verification results, and rule verification results.

[0091] In some implementations, the structured representation of the target answer region subgraph can be graphical semantics and topological information extracted from the target answer region subgraph and expressed in a machine-parseable form. The structured representation is not the original pixel image, but a symbolic data structure reconstructed based on intermediate lexical units and their attention weights related to spatial layout and connectivity in the implicit inference chain.

[0092] For example, for a circuit diagram, the structured representation of the target answer area subgraph can be modeled as a directed graph, where nodes represent components (such as power sources, resistors, and switches), edges represent wire connections, and attributes (such as power source polarity and resistor values). As another example, for an optical path diagram, the structured representation of the target answer area subgraph can be represented as a sequence of geometric relationships consisting of a light source, incident rays, refracting surfaces, normals, and outgoing rays.

[0093] The structured representation of the target answer area sub-image retains the key scoring elements of the answer while filtering out irrelevant noise such as hand-drawn style and line thickness.

[0094] In some implementations, the standard reference structured representation can be a pre-constructed structured representation corresponding to the standard answer to the question. The standard reference structured representation is automatically generated by educational experts or the system based on the standard answer image, and adopts the same modeling format as the student's answer (such as a compositional structure or geometric relationship sequence) to ensure comparability between the two.

[0095] For example, if the standard answer is a circuit containing a battery, two parallel resistors, and a closed switch, then its standard reference structured representation is a graph structure with the same node types, connection topology, and attribute values. This standard reference structured representation serves as a validation benchmark, stored in a question bank or rule knowledge base, and is invoked during visual structure matching.

[0096] In some implementations, the visual structure verification result can be a similarity or consistency determination result obtained by matching the structured representation of the student's answer with the standard reference structured representation.

[0097] The matching process can employ algorithms such as graph edit distance, subgraph isomorphism detection, and GNN embedding vector cosine similarity to calculate the degree of similarity between the two in terms of component integrity, connection correctness, and topological logic. Specifically, in some implementations, the step "generating a structured representation of the target answer region subgraph based on the implicit inference chain, and matching the structured representation with the standard reference structured representation to obtain the visual structure verification result" may include the following steps:

[0098] (1) Extract the attention weight distribution and hidden state vector corresponding to the target answer region subgraph from the implicit reasoning chain; the attention weight distribution and hidden state vector are related to the position, adjacency relationship and topological connection of the graph elements;

[0099] (2) Use attention weight distribution to locate key component regions in the target answer region subgraph, and construct the structured representation of the target answer region subgraph based on the key component regions and hidden state vectors;

[0100] (3) Perform graph matching calculations between the structured representation and the standard reference structured representation to determine the similarity measurement results between the two;

[0101] (4) Determine the visual structure verification result based on the similarity measurement result and the first preset value.

[0102] As described above, an implicit inference chain can be an information flow that is automatically captured by a deep learning model (such as a neural network) when processing graphical information. This information includes, but is not limited to, features such as the position, adjacency, and topological connections of graphical elements.

[0103] The attention weight distribution reflects the degree of attention the model pays to different parts of the image when processing it; while the hidden state vector contains an abstract representation of the graphic elements and their relationships. Both components are important for understanding the semantics of graphics.

[0104] Furthermore, based on the obtained attention weight distribution, the locations of key elements in the target answer region subgraph can be identified. This typically involves the algorithm determining which regions are given higher attention and are therefore considered crucial for answering the question.

[0105] Next, by combining the deep information provided by the hidden state vectors, these key elements and their logical relationships are transformed into a structured form—a structured representation. This representation allows computers to understand and manipulate graphical data in a more intuitive and manageable way.

[0106] Once a structured representation of the student's response is obtained, the next step is to compare it with a predefined standard reference structured representation. The technique used here is graph matching computation, which aims to quantify the similarity between the two. Graph matching can be implemented using various algorithms, such as Graph Edit Distance (GED) or graph embedding-based methods. The goal is to find correspondences between two graphs and evaluate their degree of structural similarity.

[0107] The system then uses a first preset value to determine whether the similarity between the structured representation and the standard reference structured representation meets the requirements. If the similarity measurement result exceeds or equals the first preset value, the student's answer can be considered to meet the expected standard at the visual structure level; otherwise, it does not.

[0108] The above process directly determines the final visual structure verification result, providing an important reference for the scoring system.

[0109] In some implementations, the symbols in the target answer area subgraph can be graphic marks or text symbols with specific subject meanings, hand-drawn or written by students in the target answer area subgraph.

[0110] For example, for circuit diagrams, the symbols in the target answer area sub-diagram could be: "+", "-" polarity markers, power supply symbols, and resistor labels. For optics diagrams, the symbols in the target answer area sub-diagram could be: the angle of incidence label "θ1", the normal "N", and the refractive surface label. For mechanics diagrams, the symbols in the target answer area sub-diagram could be: force vector arrows and the label "F = 10N". For statistical graphs, the symbols in the target answer area sub-diagram could be: the numerical label "30%", and coordinate axis units, etc.

[0111] The symbols in the target answer area subgraph are usually located at key scoring positions in the graph, and their correctness directly affects the scientific and standardized nature of the answer. The system locates these areas through words with high attention weights in the implicit reasoning chain.

[0112] By comparing the symbol content in the target answer area subgraph with the standard terminology of the corresponding question type in a unified semantic space, it is determined whether the two are equivalent in meaning, rather than just literal, so as to obtain the symbol verification result.

[0113] In some implementations, the symbol verification result can be a symbol correctness judgment output after semantic equivalence judgment. For example, the symbol verification result can be in binary form (e.g., "symbol correct" or "symbol incorrect"). Another example is that the symbol verification result can be fine-grained feedback (e.g., "polarity symbol missing," "force direction label reversed," "angle unit not labeled," etc.). Yet another example is that the symbol verification result can be a quantitative score (e.g., each key symbol is assigned 0–1 points, and the sum is used as the symbol dimension score).

[0114] Specifically, in some implementations, the step "to perform semantic equivalence judgment on the symbol content in the target answer region subgraph based on the implicit inference chain and determine the symbol verification result" may include the following steps:

[0115] (1) Using the built-in visual-text cross-modal alignment mechanism of the multimodal human-like cognitive model, the token positions related to symbol semantics are identified in the text token sequence corresponding to the implicit inference chain;

[0116] (2) Based on the spatial attention mapping corresponding to the token position, the handwritten symbols and labeled text in the target answer area sub-image are focused and recognized to obtain candidate text sequences;

[0117] (3) Encode the candidate text sequence and the preset standard terms into semantic embedding vectors respectively, and calculate the cosine similarity between the semantic embedding vectors corresponding to the two;

[0118] (4) Determine the symbol semantic verification result based on the cosine similarity and the second preset value.

[0119] In the autoregressive generation process, the implicit inference chain of the multimodal human-like cognitive model is represented as a sequence of internal text tokens. The built-in visual-text cross-modal alignment mechanism (such as the cross-attention layer) of the multimodal human-like cognitive model enables semantic association between each text token and the spatial region of the input image.

[0120] The system analyzes the lexical sequence to identify key lexical positions related to symbol semantics (e.g., lexical positions containing semantics such as "positive pole", "30°", "F="). These positions are learned by the model during the training phase and correspond to key scoring symbols.

[0121] For the identified symbol-related word locations, the system extracts their corresponding spatial attention maps. These maps reflect which pixel regions in the image received high attention when the multimodal human-like cognitive model generated the word. Using this attention heatmap as a mask or guiding signal, the system performs refined OCR or visual-language alignment processing on local areas within the target answer region sub-image, thereby accurately locating and recognizing handwritten symbols or labeled text (such as "+", "θ=45°", "R1=10Ω"), and outputting a candidate text sequence.

[0122] The candidate text sequence and the preset standard terms corresponding to the question type (such as "positive power supply", "45-degree angle of incidence", and "10 ohms resistance" as defined in the standard answer) are respectively input into the text encoder of the multimodal human-like cognitive model to obtain their respective semantic embedding vectors in a unified semantic space.

[0123] Subsequently, the cosine similarity between the two vectors is calculated as a quantitative indicator to measure their semantic equivalence. This method can effectively handle writing variations (such as "+" vs. "positive pole", "V" vs. "voltage") and avoid misjudgment due to formal differences.

[0124] The calculated cosine similarity is compared with a preset second value (e.g., 0.85). If the similarity is greater than or equal to the second preset value, the symbol is considered semantically correct; otherwise, it is considered incorrect or missing. The system can perform the above process on multiple key symbols separately and summarize the verification results of each symbol to form the final symbol semantic verification result, which is used for subsequent comprehensive scoring decisions.

[0125] In some implementations, semantic information can be structured semantic elements parsed from implicit inference chains that are directly related to the subject-specific scoring logic. It is noteworthy that semantic information is not the original image pixels or natural language sentences, but rather computationally achievable and verifiable physical or logical attributes encoded by intermediate lexical units and their hidden states generated by the model in the target pattern.

[0126] Semantic information includes, but is not limited to, physical quantities, directions, polarities, connections, or numerical logic. For example, physical quantities might include the magnitude of a force (e.g., "F = 10N"), voltage ("U = 5V"), or angle ("θ = 30°"). Directions might include the direction of current flow (e.g., "from left to right"), the direction of light propagation, or the direction of a force-bearing arrow. Polarity might include the positive and negative terminals of a power supply (e.g., "+" or "-"), and the polarity of a capacitor. Connections might include whether components in a circuit are connected in series or parallel, or whether a light path passes through a specified interface. Numerical logic might include deductive statements such as "the resultant force is zero," "the angle of incidence equals the angle of reflection," and "the sum of branch currents equals the main current."

[0127] In some implementations, the rule verification result can be a compliance judgment obtained by formal deduction or constraint verification of the extracted semantic information and the target subject rule knowledge base. The system substitutes the extracted semantic information into the corresponding rules for automatic reasoning (such as symbolic calculation, logical judgment, or numerical verification). If all relevant rules are satisfied, the rule verification result is "passed"; if there are violations, the violation type and severity are marked (such as "polarity error" or "unclosed loop"). This result can be a binary judgment, a multi-level score, or a structured error report, serving as a key basis for the final comprehensive score.

[0128] Specifically, in some implementations, the step "extracting semantic information related to physical quantities, directions, polarities, connectivity, or numerical logic from the target answer region subgraph based on implicit inference chains, and deducing or verifying the semantic information with the target subject rule knowledge base to obtain rule verification results" may include the following steps:

[0129] (1) Activate the target subject rule knowledge base according to the question type tag; the target subject rule knowledge base has pre-stored formalized physical or mathematical constraint rules;

[0130] (2) Extract reasoning semantic fragments related to physical quantities, directions, polarities, connections, or numerical logic from the target answer area subgraph from the implicit reasoning chain;

[0131] (3) Parse the reasoning semantic fragments into structured parameters, and substitute the structured parameters into the target subject rule knowledge base for logical deduction or numerical verification to obtain the rule verification results.

[0132] In some implementations, the target subject rule knowledge base can be a collection of subject principles organized by question type tags and stored in the form of executable logic. For example, for circuit diagrams, it includes rules such as "all branches must form closed loops" and "power supply polarity must be consistent." For optics diagrams, it includes rules such as "incident rays, normal rays, and refracted rays are coplanar." For mechanics diagrams, it includes rules such as "the net force is zero when an object is at rest" and "the direction of friction is opposite to the tendency of motion." For statistical charts, it includes rules such as "the sum of the angles of each sector in a pie chart is 360°" and "the scale of the coordinate axes must increase monotonically."

[0133] The system traverses the intermediate word sequence in the implicit inference chain to identify inference semantic fragments related to the scoring. These inference semantic fragments are generated by the model in Thinking mode and contain explicit physical quantities, directions, polarities, connections, or numerical logic information, such as: "Current flows out from the positive terminal of the power source," "Branch A is not closed," "The angle of incidence is 45 degrees," and "The resultant force F..." x =0". By using keyword matching (such as "polarity", "equal to", "direction", "closed") or methods based on named entity recognition (NER), relevant lexical units and their contexts are extracted into original semantic fragments.

[0134] The above semantic fragments of reasoning are further parsed into structured parameters. For example, "current flows out from the positive terminal" is transformed into:<direction:from_positive_terminal> For example, converting "θ=30°" into...<angle:30,unit:degree> Then, these parameters are substituted into the activated target subject rule knowledge base to perform logical deduction or numerical verification:

[0135] For logical rules (e.g., "must be closed"), determine whether the condition is met;

[0136] For numerical rules (such as "n1sinθ1=n2sinθ2"), substitute specific values ​​to verify the equation.

[0137] If all relevant rules pass, the system outputs "Rule verification passed"; if violations are found, the violation type and location are recorded. The final judgment is the rule verification result, used for subsequent comprehensive scoring decisions.

[0138] After obtaining the visual structure verification results, symbol verification results, and rule verification results through the above methods, the visual structure verification results, symbol verification results, and rule verification results are then integrated to generate a scoring result.

[0139] For example, assuming the visual structure verification result is 2 points, the symbol verification result is 0 points, and the rule verification result is 0 points, then the score result is 2 = 2 + 0 + 0.

[0140] For example, suppose the visual structure verification result is: the arrangement of the incident ray, normal, and refracted ray is basically correct (pass); the symbol verification result is: the angle label "θ1=30°" is semantically equivalent to the standard terminology (pass); the rule verification result is: after substituting n1=1.0 and n2=1.5, sinθ2≈0.333 is calculated, but the angle of refraction in the figure is significantly greater than arcsin(0.333), violating Snell's law (fail). Since rule verification has a veto power, the final score is "fail".

[0141] Please see Figure 2 , Figure 2 The diagram illustrates the structure of an intelligent marking device for open-ended educational drawing questions according to an embodiment of this application. The intelligent marking device 200 includes: a detection module 210, a cropping module 220, a reasoning module 230, and a scoring module 240. Specifically:

[0142] The detection module 210 is used to input the collected target answer sheet image into the preset detection model to determine the bounding box coordinates and question type labels corresponding to the open-ended educational drawing questions included in the target answer sheet image;

[0143] The cropping module 220 is used to crop a target answer sheet sub-image labeled with question type tags from the target answer sheet image based on the bounding box coordinates;

[0144] The reasoning module 230 is used to dynamically load the chain-like reasoning prompt templates corresponding to the question type tags, and call the multimodal human-like cognitive model to perform reasoning in the target mode and generate implicit reasoning chains.

[0145] The scoring module 240 is used to verify the target answer region subgraph based on the implicit reasoning chain and obtain the scoring result.

[0146] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0147] In the several embodiments provided in this application, the coupling or direct coupling or communication connection between the modules shown or discussed may be an indirect coupling or communication connection through some interface, device or module, and may be electrical, mechanical or other forms.

[0148] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0149] Please see Figure 3 , Figure 3 This illustration shows a structural diagram of an intelligent marking device for open-ended educational drawing questions provided in an embodiment of this application. The intelligent marking device 300 for open-ended educational drawing questions in this application may include one or more of the following components: a processor 310, a memory 320, and one or more application programs. The one or more application programs may be stored in the memory 320 and configured to be executed by one or more processors 310. The one or more programs are configured to execute the intelligent marking method for open-ended educational drawing questions as described in the foregoing method embodiments.

[0150] The processor 310 may include one or more processing cores. The processor 310 connects to various parts of the intelligent marking device 300 for open-ended educational drawing questions using various interfaces and lines. It executes various functions and processes data within the intelligent marking device 300 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 320, and by calling data stored in the memory 320. Optionally, the processor 310 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 310 may integrate one or more of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understandable that the aforementioned modem may not be integrated into the processor 310, but may be implemented using a separate communication chip.

[0151] The memory 320 may include random access memory (RAM) or read-only memory (ROM). The memory 320 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 320 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created during use by the intelligent marking device 300 for open-ended educational drawing questions.

[0152] Please see Figure 4 , Figure 4 The diagram illustrates the structure of a computer-readable storage medium 400 provided in an embodiment of this application. The computer-readable storage medium 400 stores program code, which can be called by a processor to execute the intelligent marking method for open-ended educational drawing questions described in the above method embodiments.

[0153] The computer-readable storage medium 400 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has storage space for program code 410 that performs any of the method steps described above. This program code can be read from or written to one or more computer program devices. The program code 410 may be compressed, for example, in a suitable form.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An intelligent grading method for open-ended educational drawing questions, characterized in that: The method includes: The collected target answer sheet image is input into a preset detection model to determine the bounding box coordinates and question type labels corresponding to the open-ended educational drawing questions included in the target answer sheet image; Based on the bounding box coordinates, a target answer area sub-image labeled with the question type is cropped from the target answer sheet image; Dynamically load the chain-like reasoning prompt template corresponding to the question type tag, and call the multimodal human-like cognitive model to perform reasoning in the target mode to generate an implicit reasoning chain; Based on the implicit reasoning chain, the target answer region subgraph is verified to obtain the scoring result.

2. The intelligent marking method for open-ended educational drawing questions according to claim 1, characterized in that, The dynamic loading of the chain-like reasoning prompt template corresponding to the question type label, and the invocation of a multimodal human-like cognitive model to perform reasoning in the target mode, generating an implicit reasoning chain, includes: Based on the question type tags, the chain-like reasoning prompt templates are retrieved from the preset prompt template library and dynamically loaded to guide the multimodal human-like cognitive model to focus on the key scoring elements in the target answer area sub-graph. The target answer area sub-image is concatenated with the chain-like reasoning prompt template to form a multimodal input sequence, and the multimodal input sequence is input into the multimodal human-like cognitive model to trigger the multimodal human-like cognitive model to enter the target mode; The multimodal human-like cognitive model generates multiple intermediate tokens and corresponding spatial attention weights during the autoregressive generation process under the target mode, forming the implicit inference chain.

3. The intelligent marking method for open-ended educational drawing questions according to claim 1, characterized in that, The step of verifying the target answer region subgraph based on the implicit inference chain to obtain the scoring result includes: A structured representation of the target answer region subgraph is generated based on the implicit inference chain, and the structured representation is matched with the standard reference structured representation to obtain the visual structure verification result. Based on the implicit inference chain, semantic equivalence judgment is performed on the symbol content in the target answer region subgraph to determine the symbol verification result; Based on the implicit reasoning chain, semantic information related to physical quantities, directions, polarities, connectivity, or numerical logic is extracted from the target answer area subgraph, and the semantic information is deduced or verified with the target subject rule knowledge base to obtain rule verification results. The scoring result is determined based on the visual structure verification result, the symbol verification result, and the rule verification result.

4. The intelligent marking method for open-ended educational drawing questions according to claim 3, characterized in that, The process of generating a structured representation of the target answer region subgraph based on the implicit inference chain, and matching the structured representation with a standard reference structured representation to obtain a visual structure verification result includes: The attention weight distribution and hidden state vector corresponding to the target answer region subgraph are extracted from the implicit inference chain; the attention weight distribution and hidden state vector are related to the position, adjacency relationship and topological connection of the graphic elements; The attention weight distribution is used to locate key element regions in the target answer region subgraph, and a structured representation of the target answer region subgraph is constructed based on the key element regions and the hidden state vector; The structured representation is compared with the standard reference structured representation to perform graph matching calculation, and the similarity measurement result between the two is determined. The visual structure verification result is determined based on the similarity measurement result and the first preset value.

5. The intelligent marking method for open-ended educational drawing questions according to claim 3, characterized in that, The step of performing semantic equivalence judgment on the symbol content in the target answer region subgraph based on the implicit inference chain to determine the symbol verification result includes: Using the built-in visual-text cross-modal alignment mechanism of the multimodal human-like cognitive model, the token positions related to symbol semantics are identified in the text token sequence corresponding to the implicit inference chain; Based on the spatial attention mapping corresponding to the token position, the handwritten symbols and labeled text in the target answer area sub-image are focused and recognized to obtain a candidate text sequence; The candidate text sequence and the preset standard terms are encoded into semantic embedding vectors respectively, and the cosine similarity between the semantic embedding vectors corresponding to the two is calculated. The symbol semantic verification result is determined based on the cosine similarity and the second preset value.

6. The intelligent marking method for open-ended educational drawing questions according to claim 3, characterized in that, The semantic information related to physical quantities, directions, polarities, connectivity, or numerical logic in the target answer region subgraph is extracted based on the implicit inference chain, and the semantic information is deduced or verified against the target subject rule knowledge base to obtain rule verification results, including: The target subject rule knowledge base is activated based on the question type tag; the target subject rule knowledge base pre-stores formalized physical or mathematical constraint rules. Extract inference semantic fragments related to physical quantities, directions, polarities, connectivity, or numerical logic from the target answer region subgraph from the implicit inference chain; The reasoning semantic fragment is parsed into structured parameters, and the structured parameters are substituted into the target subject rule knowledge base for logical deduction or numerical verification to obtain the rule verification result.

7. The intelligent marking method for open-ended educational drawing questions according to claim 1, characterized in that, The method includes: For each sample in the educational drawing question scoring training set, the multimodal human-like cognitive model is invoked to generate a score prediction and an implicit inference chain; the implicit inference chain is the internal intermediate representation on which the score prediction depends. Based on the internal intermediate representation, a multi-dimensional reward signal is calculated; the multi-dimensional reward signal includes visual structure matching score, symbol semantic equivalence score, and subject rule compliance score. Based on the multi-dimensional reward signal and the KL divergence regularization term, a comprehensive reward function is constructed; The human feedback reinforcement learning framework based on the GRPO algorithm uses the comprehensive reward function to iteratively optimize the policy network of the multimodal human-like cognitive model in order to complete the training of the multimodal human-like cognitive model. The comprehensive reward function is: R total =R semantic +r visual +r reward-model -γD KL (p θ ||p efr ) Among them, R total For the total reward, r semantic For the semantic equivalence score of the symbols, r visual For the visual structure matching score, r reward-model The subject rule compliance score is represented by γ, where γ is the regularization strength and D is the regularization score. KL This is the KL divergence regularization term.

8. An intelligent marking device for open-ended educational drawing questions, characterized in that, The device includes: The detection module is used to input the collected target answer sheet image into a preset detection model to determine the bounding box coordinates and question type labels corresponding to the open-ended educational drawing questions included in the target answer sheet image; The cropping module is used to crop a target answer area sub-image labeled with the question type from the target answer sheet image according to the bounding box coordinates. The reasoning module is used to dynamically load the chain-like reasoning prompt templates corresponding to the question type tags, and call the multimodal human-like cognitive model to perform reasoning in the target mode and generate implicit reasoning chains. The scoring module is used to verify the target answer region subgraph based on the implicit reasoning chain and obtain the scoring result.

9. An intelligent marking device for open-ended educational drawing questions, characterized in that: include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the intelligent grading method for open-ended educational drawing questions as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code, which can be called by a processor to execute the intelligent marking method for open-ended educational drawing questions as described in any one of claims 1-7.