An intelligent homework correction method and device based on multi-modal recognition
Patent Information
- Application Number
- CN202611269114.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-20
- Publication Date
- 2026-09-25
AI Technical Summary
但在应对复杂的主观题时,大多基于单一OCR文字识别技术实现作业内容识别与批改,仅能针对作业最终答案的对错进行二元判断,无法满足精细化批改需求
[0038]本申请的基于多模态识别的智能作业批改方法,能够实现对复杂主观题的批改,提高作业批改效率,并提升批改结果准确性。
Smart Images

Figure CN122820408A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of intelligent education, and specifically relates to an intelligent homework correction method and device based on multimodal recognition. Background Technology
[0002] With the rapid popularization of online education and smart campus technologies, intelligent homework grading has become one of the core technologies for the digital transformation of education. It can effectively replace manual homework grading, reduce teachers' teaching burden, and improve homework grading efficiency.
[0003] Current intelligent homework grading systems are quite mature in handling objective and standardized subjective questions, boasting high accuracy and efficiency. However, when dealing with complex subjective questions, most rely on single OCR text recognition technology for content identification and grading, only able to make binary judgments on the correctness of the final answer, failing to meet the needs of refined grading.
[0004] Therefore, there is an urgent need for a technical solution to overcome or mitigate at least one of the aforementioned defects in the existing technology. Summary of the Invention
[0005] The purpose of this application is to provide an intelligent job grading method and apparatus based on multimodal recognition to solve at least one problem existing in the prior art.
[0006] The technical solution of this application is:
[0007] The first aspect of this application provides an intelligent homework grading method based on multimodal recognition, including:
[0008] S1. Obtain the answer content to be graded, and divide the answer content into multiple grading nodes;
[0009] S2. Identify the modified node using a multimodal recognition model to obtain the multimodal information of the modified node;
[0010] S3. Perform an initial score on the grading node based on the multimodal information to obtain the initial score of the grading node;
[0011] S4. Filter out the correction nodes whose initial scores are lower than the preset score threshold, perform error cause analysis on the correction nodes according to the multimodal information, and match error cause tags from the error cause tag library;
[0012] S5. Based on the severity level of the error cause label, correct the initial score of the corresponding correction node, and calculate the total score based on the corrected score.
[0013] In at least one embodiment of this application, in S2, the multimodal recognition model includes:
[0014] The OCR recognition sub-model is used to recognize the text information in the modified node;
[0015] The formula recognition sub-model is used to identify the formula information in the batching node;
[0016] A graphic recognition sub-model is used to identify graphic information in the modified nodes.
[0017] In at least one embodiment of this application, in step S3, the initial score of the correction node is obtained by performing an initial score on the correction node based on the multimodal information, including:
[0018] Obtain the base score of the revise node;
[0019] Calculate the first similarity between the text information in the grading node and the standard answer text, the second similarity between the formula information and the standard answer formula, and the third similarity between the graphic information and the standard answer graphic. Then, perform a weighted sum of the first similarity, the second similarity, and the third similarity to obtain the comprehensive similarity of the grading node.
[0020] The product of the base score and the comprehensive similarity is used as the initial score of the grading node.
[0021] In at least one embodiment of this application, in step S4, error cause analysis is performed on the correction node based on the multimodal information, and error cause tags are matched from the error cause tag library, including:
[0022] Extract the multidimensional error cause feature vector of the correction node based on the multimodal information;
[0023] Calculate the similarity between the multidimensional error cause feature vector and the feature vector of each error cause label in the error cause label library;
[0024] The error cause label with the highest similarity and above the preset similarity threshold is selected as the error cause label of the correction node.
[0025] In at least one embodiment of this application, the error cause label includes calculation error, logical deduction error, missing step, conceptual confusion, and misreading of the question.
[0026] In at least one embodiment of this application, in step S5, the initial score of the corresponding correction node is corrected according to the severity level of the error cause label, including:
[0027] Obtain the severity level corresponding to the error cause label;
[0028] The correction factor for the error cause label is determined based on the severity level.
[0029] The product of the initial score and the correction coefficient is used as the corrected score of the grading node.
[0030] In at least one embodiment of this application, the method further includes step S6: extracting and recommending exercises from the exercise bank based on the error cause labels.
[0031] A second aspect of this application provides an intelligent job grading device based on multimodal recognition, which, based on the intelligent job grading method based on multimodal recognition as described above, includes:
[0032] The node partitioning module is used to obtain the answer content to be graded and divide the answer content into multiple grading nodes;
[0033] The multimodal information acquisition module is used to identify the modification node through a multimodal recognition model and acquire the multimodal information of the modification node;
[0034] An initial score acquisition module is used to perform an initial score on the grading node based on the multimodal information and acquire the initial score of the grading node.
[0035] The error cause label matching module is used to filter out the correction nodes whose initial scores are lower than a preset score threshold, perform error cause analysis on the correction nodes based on the multimodal information, and match error cause labels from the error cause label library.
[0036] The total score calculation module is used to correct the initial score of the corresponding correction node according to the severity level of the error cause label, and calculate the total score based on the corrected score.
[0037] The invention has at least the following beneficial technical effects:
[0038] The intelligent homework grading method based on multimodal recognition proposed in this application can grade complex subjective questions, improve homework grading efficiency, and enhance the accuracy of grading results. Attached Figure Description
[0039] Figure 1 This is a flowchart of an intelligent job grading method based on multimodal recognition according to one embodiment of this application;
[0040] Figure 2 This is a schematic diagram of an intelligent job correction device based on multimodal recognition according to one embodiment of this application. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, embodiments of this application. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0042] The following is in conjunction with the appendix Figures 1 to 2 This application will be described in further detail.
[0043] The first aspect of this application provides an intelligent homework grading method based on multimodal recognition, comprising the following steps:
[0044] S1. Obtain the answer content to be graded and divide the answer content into multiple grading nodes;
[0045] S2. Identify the batching nodes using a multimodal recognition model to obtain multimodal information about the batching nodes;
[0046] S3. Perform initial scoring on the grading nodes based on multimodal information to obtain the initial scores of the grading nodes;
[0047] S4. Filter out the correction nodes whose initial scores are lower than the preset score threshold, perform error cause analysis on the correction nodes based on multimodal information, and match error cause labels from the error cause label library.
[0048] S5. Based on the severity level of the error cause label, correct the initial score of the corresponding correction node, and calculate the total score based on the corrected score.
[0049] The intelligent homework grading method based on multimodal recognition in this application, in step S1, firstly, places the student's answer sheet or homework book in a designated area, and then uses an image acquisition device on a smart terminal to photograph or scan the student's written answers to the exercises on the answer sheet or homework book to obtain image data of the answer content. Then, the answer content is divided into multiple grading nodes, each corresponding to a relatively independent answering step, so that each grading node can be independently identified, scored, and the error causes analyzed subsequently.
[0050] In a preferred embodiment of this application, a standard answer is pre-constructed for each exercise. The standard answer includes a standard answer path for the exercise, which consists of multiple standard answer steps arranged in sequence. Each standard answer step corresponds to a standard node, and each standard node is pre-marked with a corresponding score. When dividing the grading nodes, the answer content is aligned and matched with the standard answer to identify the answer areas corresponding to each standard step. The answer content in each answer area is then divided into a grading node. By dividing the grading nodes, a refined evaluation of each step in the answer process can be achieved, accurately pinpointing the specific point where the student's error occurred.
[0051] The intelligent homework grading method based on multimodal recognition in this application, in step S2, inputs each grading node into a pre-constructed multimodal recognition model to identify the multimodal information of the grading node. This multimodal information is structured data that integrates multiple information types, comprehensively representing the student's answer status at each grading node, and adapting to various question types. In this embodiment, the multimodal recognition model adopts a parallel sub-model architecture, including three parallel processing channels: an OCR recognition sub-model, a formula recognition sub-model, and an image recognition sub-model. Preferably, the OCR recognition sub-model uses an end-to-end text recognition network based on deep learning, the formula recognition sub-model uses a mathematical formula recognition network based on an encoder-decoder architecture, and the image recognition sub-model uses an image segmentation and object detection fusion architecture based on deep learning, capable of recognizing non-textual answer content such as geometric images, function graphs, icons, and force analysis diagrams in the grading nodes. Specifically, the OCR recognition sub-model identifies text information in the grading nodes; the formula recognition sub-model identifies formula information in the grading nodes; and the image recognition sub-model identifies graphic information in the grading nodes. The three sub-models independently identify the input grading node image data, output the feature information of their respective modalities, and then perform feature alignment and fusion to finally output a unified multimodal information data structure. This application can simultaneously extract three modalities of information: text, formulas, and graphics, significantly improving the ability to grade assignments, and is especially suitable for subjects such as mathematics, physics, and chemistry that contain a large number of formulas and graphics.
[0052] In the intelligent job grading method based on multimodal recognition of this application, the initial score of the grading node is obtained in S3 as follows:
[0053] Obtain the base score for each graded node;
[0054] Calculate the first similarity between the text information in the grading node and the standard answer text, the second similarity between the formula information and the standard answer formula, and the third similarity between the graphic information and the standard answer graphic. Then, perform a weighted sum of the first, second, and third similarities to obtain the comprehensive similarity of the grading node.
[0055] The product of the base score and the overall similarity is used as the initial score for the grading node.
[0056] In this embodiment, the score of the standard node in the standard answer is used as the base score of the grading node. The similarity of each modality is fused to obtain the comprehensive similarity. Then, the base score is multiplied by the comprehensive similarity to obtain the initial score of the grading node. In this way, the initial quantitative score result of each grading node is obtained.
[0057] In the intelligent homework correction method based on multimodal recognition of this application, the error cause label matching process in S4 includes:
[0058] Extract multidimensional error cause feature vectors from the grading nodes based on multimodal information;
[0059] Calculate the similarity between the feature vector of the multidimensional error cause and the feature vector of each error cause label in the error cause label library;
[0060] Select the error cause label with the highest similarity that is higher than the preset similarity threshold as the error cause label of the correction node.
[0061] In a preferred embodiment of this application, the multidimensional error cause feature vector includes result deviation features, logical deviation features, process coverage features, concept matching features, and question reading deviation features. Specifically, the calculation results of the grading node are compared with the calculation results of the standard answer to calculate the result deviation value, and then the result deviation value is mapped to the result deviation feature. The answer steps of the grading node are compared with the answer steps of the standard answer to calculate the logical deviation value, and then the logical deviation value is mapped to the logical deviation feature. The answer area coverage of the grading node is detected, and combined with its logical continuity with the preceding and succeeding nodes, the process completeness is calculated, and then the process completeness is mapped to the process coverage feature. The core concepts, formulas, theorems, etc. involved in the grading node are extracted, semantically matched with the concepts used in the standard answer, the concept deviation degree is calculated, and then the concept deviation degree is mapped to the concept matching feature. The answer content in the grading node is semantically matched with the key constraints such as the problem-solving method, accuracy requirements, unit requirements, and condition limitations in the question stem to calculate the question reading deviation value, and then the question reading deviation value is mapped to the question reading deviation feature. Finally, the five error cause features are combined into a multi-dimensional error cause feature vector. In this embodiment, the error cause label library includes error cause labels such as calculation error, logical derivation error, missing steps, conceptual confusion, and misreading the question. Cosine similarity is used to calculate the similarity between the multi-dimensional error cause vector and the feature vectors of each error cause label, and the error cause labels of the grading nodes are selected based on the similarity.
[0062] In the intelligent job grading method based on multimodal recognition of this application, the correction process of the initial score of the grading node in S5 includes:
[0063] Obtain the severity level corresponding to the error cause label;
[0064] The correction factor for the error cause label is determined based on the severity level;
[0065] The product of the initial score and the correction coefficient is used as the corrected score for the grading node.
[0066] The severity level is used to characterize the degree of impact of different error types on the overall performance of the answer. In this embodiment, conceptual confusion and misreading the question are defined as high severity, logical deduction errors and missing steps are defined as medium severity, and calculation errors are defined as low severity. A correction coefficient is used to adjust the initial score; the higher the severity, the smaller the correction coefficient. The corrected scores at each grading node are summed to obtain the total score for the answer.
[0067] The intelligent homework grading method based on multimodal recognition in this application further includes step S6: extracting and recommending exercises from the exercise bank based on error cause tags. A mapping relationship library between error cause tags and knowledge points is pre-built, with each error cause tag associated with one or more knowledge point identifiers. For example, conceptual confusion is associated with the knowledge point of similar triangle properties, and calculation errors are associated with the knowledge point of the difference of squares formula. The corresponding knowledge point identifier is queried based on the error cause tag. Using the knowledge point identifier as an index, exercises associated with the knowledge points corresponding to the error cause tags are extracted from the exercise bank, generating a recommended exercise set which is then pushed to the student's end.
[0068] This application presents an intelligent homework grading method based on multimodal recognition, refining the grading granularity from the overall to the specific details, enabling precise location of errors. By extracting multimodal information through a multimodal recognition model, it avoids information loss caused by single-modal information and improves adaptability to different subjects. It achieves automatic error diagnosis and differentiates initial scores based on the severity of the error, resulting in different final scores for the same initial score due to different errors, making the grading results more fair and reasonable. This forms a closed-loop teaching system of assessment-diagnosis-correction-reinforcement, providing students with personalized solutions for identifying and addressing learning gaps, effectively improving learning efficiency and the accuracy of instructional guidance.
[0069] Based on the above-described intelligent job grading method based on multimodal recognition, a second aspect of this application provides an intelligent job grading device based on multimodal recognition, comprising:
[0070] The node segmentation module is used to obtain the answer content to be graded and divide it into multiple grading nodes based on the answer steps in the answer content;
[0071] The multimodal information acquisition module is used to identify the batching nodes through a multimodal recognition model and acquire the multimodal information of the batching nodes;
[0072] The initial score acquisition module is used to perform initial scoring on the grading nodes based on multimodal information and obtain the initial score of the grading nodes.
[0073] The error cause label matching module is used to filter out the correction nodes whose initial scores are lower than the preset score threshold, perform error cause analysis on the correction nodes based on multimodal information, and match error cause labels from the error cause label library.
[0074] The total score calculation module is used to correct the initial score of the corresponding grading node according to the severity level of the error cause label, and calculate the total score based on the corrected score.
[0075] The intelligent job grading device based on multimodal recognition in this application is designed in detail for each module with reference to the intelligent job grading method based on multimodal recognition, which will not be described in detail here.
[0076] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for intelligent homework grading based on multimodal recognition, characterized in that, include: S1. Obtain the answer content to be graded, and divide the answer content into multiple grading nodes; S2. Identify the modified node using a multimodal recognition model to obtain the multimodal information of the modified node; S3. Perform an initial score on the grading node based on the multimodal information to obtain the initial score of the grading node; S4. Filter out the correction nodes whose initial scores are lower than the preset score threshold, perform error cause analysis on the correction nodes according to the multimodal information, and match error cause tags from the error cause tag library; S5. Based on the severity level of the error cause label, correct the initial score of the corresponding correction node, and calculate the total score based on the corrected score.
2. The intelligent job grading method based on multimodal recognition according to claim 1, characterized in that, In S2, the multimodal recognition model includes: The OCR recognition sub-model is used to recognize the text information in the modified node; The formula recognition sub-model is used to identify the formula information in the batching node; A graphic recognition sub-model is used to identify graphic information in the modified nodes.
3. The intelligent job grading method based on multimodal recognition according to claim 2, characterized in that, In S3, an initial score is given to the grading node based on the multimodal information to obtain the initial score of the grading node, including: Obtain the base score of the revise node; Calculate the first similarity between the text information in the grading node and the standard answer text, the second similarity between the formula information and the standard answer formula, and the third similarity between the graphic information and the standard answer graphic. Then, perform a weighted sum of the first similarity, the second similarity, and the third similarity to obtain the comprehensive similarity of the grading node. The product of the base score and the comprehensive similarity is used as the initial score of the grading node.
4. The intelligent job grading method based on multimodal recognition according to claim 3, characterized in that, In S4, error cause analysis is performed on the correction node based on the multimodal information, and error cause tags are matched from the error cause tag library, including: Extract the multidimensional error cause feature vector of the correction node based on the multimodal information; Calculate the similarity between the multidimensional error cause feature vector and the feature vector of each error cause label in the error cause label library; The error cause label with the highest similarity and above the preset similarity threshold is selected as the error cause label of the correction node.
5. The intelligent job grading method based on multimodal recognition according to claim 4, characterized in that, The error labels include calculation errors, logical deduction errors, missing steps, conceptual confusion, and misreading of the question.
6. The intelligent job grading method based on multimodal recognition according to claim 5, characterized in that, In S5, the initial score of the corresponding correction node is corrected according to the severity level of the error cause label, including: Obtain the severity level corresponding to the error cause label; The correction factor for the error cause label is determined based on the severity level. The product of the initial score and the correction coefficient is used as the corrected score of the grading node.
7. The intelligent job grading method based on multimodal recognition according to claim 6, characterized in that, It also includes S6, which involves extracting and recommending exercises from the exercise bank based on the error cause labels.
8. An intelligent job grading device based on multimodal recognition, based on the intelligent job grading method based on multimodal recognition as described in any one of claims 1 to 7, characterized in that, include: The node partitioning module is used to obtain the answer content to be graded and divide the answer content into multiple grading nodes; The multimodal information acquisition module is used to identify the modification node through a multimodal recognition model and acquire the multimodal information of the modification node; An initial score acquisition module is used to perform an initial score on the grading node based on the multimodal information and acquire the initial score of the grading node. The error cause label matching module is used to filter out the correction nodes whose initial scores are lower than a preset score threshold, perform error cause analysis on the correction nodes based on the multimodal information, and match error cause labels from the error cause label library. The total score calculation module is used to correct the initial score of the corresponding correction node according to the severity level of the error cause label, and calculate the total score based on the corrected score.