Automatic paper marking and scoring method and system based on subjective questions
By combining optical character recognition technology, deep learning and logical reasoning models, the subjective questions are automatically marked and scored, which solves the problem of insufficient evaluation of answer diversity and logic in the existing technology, and achieves a more accurate and transparent scoring process.
Patent Information
- Application Number
- CN202510108667.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
The existing automatic marking and scoring technology for subjective questions has problems such as insufficient ability to handle answer diversity, ignoring logical and reasoning processes, poor fault tolerance for spelling and grammatical errors, and lack of interpretability in the scoring process.
Optical character recognition technology and deep learning technology are used to initial recognition and feature extraction of answers, combined with logical reasoning models and knowledge graphs for in-depth understanding and analysis, output detailed scoring processes and results, and improve the transparency of scoring through multi-level feedback suggestions.
It improves the ability to accurately identify and deeply understand answers, enhances the evaluation of logic and reasoning processes, improves the accuracy and transparency of scores, and reduces scoring bias caused by differences in language expression.
Smart Images

Figure CN120047954A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic marking and scoring, and particularly to an automatic marking and scoring method and system for subjective questions. Background Art
[0002] In recent years, with the continuous advancement of educational informatization and the in-depth reform of the education evaluation system, the automated scoring technology for subjective questions has gradually become an important research direction in the field of educational technology. Compared with the automation and maturity of objective question scoring, the scoring of subjective questions has higher technical difficulty due to its complex characteristics such as semantic understanding, logical analysis, and language expression. The scoring requirements for this type of question not only evaluate the accuracy of the answer content, but also require comprehensive analysis of the logic, coherence, and innovation of language expression. However, the traditional manual scoring method has problems such as low efficiency, high cost, and strong subjectivity of scoring results, and cannot meet the needs of modern education for efficient and fair evaluation.
[0003] With the development of science and technology, the current automatic marking and scoring of subjective questions mainly relies on the following implementation technologies: keyword matching-based scoring methods, natural language processing (NLP) technology-based scoring methods, machine learning-based scoring methods, syntax and logic analysis-based scoring methods, etc. Although the existing technologies have made certain progress in the automatic scoring of subjective questions, there are still the following main defects: 1) The answer expressions of candidates are diverse, and the ability to handle answer diversity is insufficient; 2) Most of the existing technologies focus on content matching of answers, while ignoring the logic, coherence, and reasoning process of answers; 3) There are often spelling mistakes, grammar mistakes, etc. in candidates' answers, and the existing technologies have poor fault tolerance for these mistakes; 4) Although the machine learning-based scoring method has a high degree of intelligence, the scoring process is often "black box", lacking interpretability of the scoring process, etc. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide an automatic marking and scoring method and system for subjective questions, which can accurately identify and deeply understand the input answers, and can also analyze and demonstrate through powerful logical reasoning ability, output accurate scoring processes and scoring results, and increase the accuracy and transparency of automatic scoring.
[0005] To achieve the above object, the present invention provides the following solution: An automatic marking and scoring method for subjective questions, including:
[0006] Obtain handwritten or electronic answers, use optical character recognition technology and deep learning technology for feature extraction and sequence modeling, complete initial recognition, and obtain digital text;
[0007] Normalize, correct spelling and grammar of the digital text to obtain a standard digital text;
[0008] Use a classifier to identify the subject and question type of the standard digital text, obtain a text classification result, select a semantic model according to the classification result, and use the LoRA fine-tuning technique to perform domain fine-tuning and task adaptation on the selected semantic model to obtain a deep semantic understanding model;
[0009] Use the deep semantic understanding model to process the expression diversity of the standard digital text, extract key concepts and viewpoints in the standard digital text to obtain structured semantic information;
[0010] Combine the structured semantic information, use a logical reasoning model to analyze the logical framework and reasoning process of the standard digital text, introduce a knowledge graph to verify the validity of facts and arguments in the standard digital text, and then extract the arguments, evidence and conclusions in the standard digital text for argument evaluation to obtain a logical reasoning evaluation result;
[0011] Combine the logical reasoning evaluation result, calculate the scores of the content relevance dimension, logic dimension, language quality dimension and creativity dimension respectively, and then synthesize the scores of each dimension to obtain a total score and generate a score report;
[0012] Output the score report to display the scoring rules, provide multi-level feedback suggestions for the score report, and introduce the results of manual marking review to correct the total score and calculation process.
[0013] Optionally, obtain handwritten or electronic answers, use optical character recognition technology and deep learning technology for feature extraction and sequence modeling to complete initial recognition to obtain digital text, including:
[0014] Obtain a handwritten answer image, perform denoising, binarization, skew correction, cropping and normalization processing on the handwritten answer image, and then perform text line segmentation and character segmentation to obtain an initial answer image;
[0015] Use a convolutional neural network to extract a two-dimensional feature map of the initial answer image, expand the two-dimensional feature map along the width direction of the initial answer image to generate a one-dimensional feature sequence in the form of a time series to complete visual feature extraction;
[0016] Use a recurrent neural network and a Transformer model to model the one-dimensional feature sequence at the character level and word level respectively to extract local features and global dependencies to obtain the context relationship between characters, complete the initial recognition of the handwritten answer image, and obtain digital text.
[0017] Optionally, perform format normalization, spelling correction, and grammar correction on the digital text to obtain a standard digital text, including:
[0018] Combine regular expressions to perform case conversion, punctuation processing, and whitespace processing on the digital text to complete format normalization; wherein, the case of proper nouns is preserved during case conversion;
[0019] Construct a domain dictionary, detect misspelled words in the digital text, calculate the similarity between words and words in the domain dictionary using the edit distance algorithm, generate multiple correction candidates based on the calculation results, and then calculate the context probabilities of the multiple correction candidates using a pre-trained language model to select the optimal candidate to replace the misspelled word to complete spelling correction;
[0020] Use a grammar analysis model to detect grammar errors in the digital text, generate multiple correction candidate sentences using predefined rules or deep learning, and then calculate the context probabilities of the correction candidate sentences using the pre-trained language model to select the optimal candidate sentence to replace the grammar error sentence to complete grammar correction.
[0021] Optionally, use a classifier to identify the subject and question type of the standard digital text to obtain a text classification result, select a semantic model according to the classification result, and use the LoRA fine-tuning technique to perform domain fine-tuning and task adaptation on the selected semantic model to obtain a deep semantic understanding model, including:
[0022] Construct an annotated dataset containing multiple subjects and question types, select a Transformer-based text classification model, train the text classification model using the annotated dataset to obtain a classifier, splice the questions and answers of the standard digital text and input them into the classifier for feature extraction to obtain a classification result;
[0023] Collect subjective question answer data for each subject and question type, construct a fine-tuning dataset, select a semantic model according to the classification result, and insert LoRA modules into the attention layer and feed-forward layer of the selected semantic model using the LoRA fine-tuning technique;
[0024] Train the semantic model of the LoRA module using the fine-tuning dataset, optimize the low-rank matrix to maximize the semantic similarity between the answer and the reference answer to obtain a deep semantic understanding model.
[0025] Optionally, use the deep semantic understanding model to process the expression diversity of the standard digital text, extract key concepts and viewpoints in the standard digital text to obtain structured semantic information, including:
[0026] Calculate the word similarity of the standard digital text using the deep semantic understanding model, perform synonym replacement, and perform sentence normalization to identify different expressions;
[0027] Convert the standard digital text into an answer vector and match it with the reference answer vector, calculate the semantic similarity, and then decompose the semantic units in the answer vector to generate a structured representation; semantic units include concepts, viewpoints, and arguments;
[0028] Combine the structured representation to perform keyword extraction, concept recognition, and viewpoint extraction on the standard digital text, and convert the extracted key concepts and viewpoints into a graph structure to obtain structured semantic information.
[0029] Optionally, combine the structured semantic information and use a logical reasoning model to analyze the logical framework and reasoning process of the standard digital text, including:
[0030] Introduce a dependency parsing tool in the natural language inference model to construct a logical reasoning model;
[0031] Use the dependency parsing tool in the logical reasoning model to parse the sentence structure of the standard digital text, identify grammatical relationships, and construct a logical tree for representing the logical hierarchy;
[0032] Combine the logical tree and use the logical reasoning model to identify the logical relationships between sentences in the standard digital text, and concatenate the logical relationships to construct an inference chain for representing the reasoning process.
[0033] Optionally, introduce a knowledge graph to verify the validity of facts and arguments in the standard digital text, and then extract the arguments, evidence, and conclusions in the standard digital text for argument evaluation, including:
[0034] Obtain an existing open-source knowledge graph or domain knowledge base, and represent the knowledge in the form of triples to obtain a triple knowledge graph;
[0035] Identify the entities in the standard digital text to obtain answer entities, match the answer entities with the nodes in the triple knowledge graph, and establish entity links;
[0036] Identify the relationships in the standard digital text to obtain answer relationships, and verify whether the answer relationships are consistent with the relationships in the triple knowledge graph for relationship matching;
[0037] Check whether the facts in the standard digital text exist in the triple knowledge graph. If not, mark the non-existent facts as unverified to complete fact verification;
[0038] Identify and examine the argument structure in the standard digital text, mark the arguments, evidence, and conclusions, classify the evidence into factual evidence and logical evidence, and then evaluate whether the evidence supports the arguments and whether the arguments reasonably lead to the conclusions to obtain the evaluation results of the integrity and rationality of the argument.
[0039] Optionally, in combination with the logical reasoning evaluation results, calculate the scores for the content relevance dimension, logical dimension, language quality dimension, and creativity dimension respectively, and then synthesize the scores of each dimension to obtain the total score and generate a score report, including:
[0040] Use a semantic similarity model to calculate the semantic similarity between the standard digital text and the reference answer, extract the key concepts in the standard digital text and the reference answer, and calculate the key concept coverage rate of the standard digital text. Then multiply the semantic similarity and the key concept coverage rate according to the set weights to obtain the content relevance score;
[0041] Calculate the total coherence score according to the coherence and rationality of the reasoning chain, calculate the argument integrity score according to the argument structure of the standard digital text, and then synthesize the total coherence score and the argument integrity score according to the set weights to obtain the logical score;
[0042] Calculate the language quality score according to the grammar error rate and spelling error rate of the standard digital text, calculate the lexical richness score according to the lexical diversity and sentence pattern diversity of the standard digital text, calculate the sentence coherence score according to the semantic coherence of the standard digital text, and then synthesize the language quality score, the lexical richness score, and the sentence coherence score according to the set weights to obtain the language quality score;
[0043] Calculate the creativity score according to the semantic similarity between the standard digital text and the reference answer, calculate the new idea score according to the number of ideas in the standard digital text, and then synthesize the semantic similarity and the new idea score according to the set weights to obtain the creativity score;
[0044] Adjust the weights of the scores of each dimension according to the teaching objectives, and then synthesize the scores of each dimension to obtain the total score and generate a score report including the advantages of the answer, the disadvantages of the answer, and suggestions for improving the answer.
[0045] Optionally, the expression of the total score is:
[0046] W = a·W 1 + b·W 2 + c·W 3 + d·W 4
[0047] where W is the total score, W 1 、W2 , W 3 , W 4 They are the content relevance score, logicality score, language quality score, and creativity score respectively. a, b, c, and d are the weights of the content relevance score, logicality score, language quality score, and creativity score respectively.
[0048] The present invention also provides an automatic marking and scoring system based on subjective questions, including:
[0049] An initial recognition module, used to obtain handwritten or electronic answers, perform feature extraction and sequence modeling using optical character recognition technology and deep learning technology, complete initial recognition, and obtain digital text;
[0050] A preprocessing module, used to perform normalization, spelling correction, and grammar correction processing on the digital text to obtain standard digital text;
[0051] A model construction module, used to identify the subject and question type of the standard digital text using a classifier to obtain a text classification result, select a semantic model according to the classification result, and use the LoRA fine-tuning technology to perform domain fine-tuning and task adaptation on the selected semantic model to obtain a deep semantic understanding model;
[0052] A deep understanding module, used to process the expression diversity of the standard digital text using the deep semantic understanding model, extract the key concepts and viewpoints in the standard digital text, and obtain structured semantic information;
[0053] A logical reasoning module, used to combine the structured semantic information, analyze the logical framework and reasoning process of the standard digital text using a logical reasoning model, introduce a knowledge graph to verify the validity of the facts and arguments in the standard digital text, and then extract the arguments, evidence, and conclusions in the standard digital text to perform argument evaluation and obtain a logical reasoning evaluation result;
[0054] A scoring calculation module, used to combine the logical reasoning evaluation result, calculate the scores of the content relevance dimension, logicality dimension, language quality dimension, and creativity dimension respectively, and then synthesize the scores of each dimension to obtain a total score and generate a scoring report;
[0055] An output feedback module, used to output the scoring report to display the scoring rules, provide multi-level feedback suggestions for the scoring report, and introduce the result of manual marking review to correct the total score and calculation process.
[0056] By providing an automatic marking and scoring method and system based on subjective questions, the present invention discloses the following technical effects:
[0057] 1. Accurate answer recognition: 1) By using optical character recognition technology and deep learning technology to recognize answer images, the recognition ability for complex handwritten fonts can be increased, supporting different writing styles and language habits, applicable to various exam scenarios, ensuring that the system can quickly process large-scale handwritten answers, and reducing the impact of recognition errors on grading. 2) By normalizing, spell-checking, and grammar-checking the answers, the fault tolerance for language errors can be improved, ensuring the parsability of the text, while retaining the original meaning, avoiding grading biases caused by poor language expression ability, and providing high-quality input text for the subsequent grading module.
[0058] 2. Deep answer understanding: 1) By constructing a classifier, the subject (such as mathematics, physics, history) and question type (such as short answer questions, essay questions, calculation questions) to which the answer belongs can be determined according to the content of the answer, enabling the system to select more appropriate semantic models and grading criteria. 2) By constructing a deep semantic understanding model, the diversity of answer expressions can be processed, key concepts and viewpoints can be extracted, grading biases caused by language style differences can be avoided, accurate evaluation of the content and logic of the answer can be achieved, and powerful semantic analysis capabilities can be provided for the automatic grading of subjective questions.
[0059] 3. Strong logical reasoning ability: 1) By constructing a logical reasoning model, the logic and reasoning ability of the answer can be accurately evaluated; 2) By verifying the facts and the validity of the arguments in the answer through a knowledge graph, grading biases caused by knowledge background differences can be avoided; 3) By extracting the arguments, evidence, and conclusions in the answer, the integrity and rationality of the argument can be evaluated.
[0060] 4. Reliable grading: By comprehensively adjusting various grading dimensions and weights, the accuracy of grading can be improved, and by outputting detailed grading reports and improvement suggestions feedback, students can be helped to discover deficiencies and improve their abilities, increasing the transparency of the grading process and user trust.
[0061] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other accompanying drawings without creative efforts based on these drawings.
[0063] Figure 1 It is a schematic flowchart of the method provided by the embodiment of the present invention;
[0064] Figure 2 It is a schematic flowchart of deep understanding and logical reasoning provided by the embodiment of the present invention;
[0065] Figure 3 This is a schematic diagram of the system architecture provided by the embodiments of the present invention. Detailed implementation manners
[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0067] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0068] As Figure 1 shown, the present invention provides an automatic marking and scoring method based on subjective questions, including:
[0069] 1. Obtain handwritten or electronic answers, and use optical character recognition technology and deep learning technology for feature extraction and sequence modeling to complete initial recognition and obtain digital text. This includes:
[0070] 1.1 Obtain a handwritten answer image, perform denoising, binarization, skew correction, cropping, and normalization on the handwritten answer image, and then perform text line segmentation and character segmentation to obtain an initial answer image.
[0071] Denoising: Use Gaussian filtering or bilateral filtering to remove noise in the image (such as background noise during scanning or photographing).
[0072] Binarization: Convert the image into a black-and-white binary image, using the Otsu algorithm or an adaptive threshold segmentation method.
[0073] Skew correction: Detect the skew angle of the text line through the Hough transform and perform rotation correction.
[0074] Cropping and normalization: Crop the blank area at the edge of the image and normalize the image size to a fixed size (such as 32×128 pixels).
[0075] Text line segmentation: Use the projection method or connected component analysis (CCA) to segment the image into individual text lines.
[0076] Character segmentation: For handwritten characters with a large character spacing, use the vertical projection method to segment the characters; for cursive handwritten characters, use a deep learning-based segmentation model (such as U-Net) for character segmentation.
[0077] 1.2 Use a convolutional neural network to extract the two-dimensional feature map of the initial answer image, unfold the two-dimensional feature map along the width direction of the initial answer image to generate a one-dimensional feature sequence in the form of a time series, and complete the visual feature extraction.
[0078] Convolutional neural network (CNN): Extract local features of handwritten images, such as stroke shapes, character contours, etc.
[0079] Among them, the convolutional layer: Use multiple convolutional kernels (such as 3×3, 5×5, 7×7) to extract low-level features (such as edges, corners).
[0080] Pooling layer: Reduce the dimension of the feature map through max pooling or average pooling to reduce the computational amount.
[0081] Batch Normalization: Accelerate training and improve the stability of the model.
[0082] Activation function: Use the ReLU activation function to introduce non-linear characteristics.
[0083] 1.3 Use a recurrent neural network and a Transformer model to model the one-dimensional feature sequence at the character level and word level respectively to extract local features and global dependencies, obtain the context relationship between characters, complete the initial recognition of the handwritten answer image, and obtain digital text.
[0084] Recurrent neural network (RNN): Model the character sequence to capture short-term dependencies between characters. Among them, use bidirectional LSTM (Bi-LSTM): Capture sequence information from both the forward and backward directions simultaneously to enhance the understanding of the context. Attention mechanism (Attention Mechanism): Assign weights to each time step to focus on important characters or features.
[0085] Transformer model: Capture long-range dependencies in the character sequence and solve the performance bottleneck of RNN in long sequence modeling. Among them, use multi-head self-attention mechanism (Multi-Head Self-Attention): Parallelly calculate different attention patterns to capture multiple dependencies. Positional Encoding: Add position information to each character to preserve the sequence order. Feed-Forward Network (FFN): Perform non-linear transformation on the features of each character.
[0086] 2. Perform normalization, spelling correction, and grammar correction processing on the digital text to obtain standard digital text. Include:
[0087] 2.1 Combine regular expressions to perform case conversion, punctuation processing, and whitespace character processing on the digital text to complete format normalization; among them, the case of proper nouns is preserved during case conversion.
[0088] Case conversion: Convert all text to lowercase (or capitalize the first letter according to language rules). Use regular expressions (Regex) to detect proper nouns (such as names and place names) and preserve their original case.
[0089] Punctuation processing: Use regular expressions to clean up redundant punctuation marks (such as consecutive commas and periods). Replace non-standard punctuation marks (such as the mixed use of Chinese and English punctuation). Detect the absence of punctuation marks (such as periods and question marks) and complete them.
[0090] Whitespace character processing: Delete redundant spaces or line breaks. Insert spaces where necessary (such as after punctuation marks).
[0091] 2.2 Build a domain dictionary, detect misspelled words in the digital text, calculate the similarity between the words and the words in the domain dictionary using an edit distance algorithm (such as the Levenshtein distance), generate multiple correction candidate words based on the calculation results, and then use a pre-trained language model to calculate the context probabilities of the multiple correction candidate words to select the optimal candidate word to replace the misspelled word and complete the spelling correction process.
[0092] 2.3 Use a syntax analysis model (such as SpaCy, Stanza) to detect syntax errors in the digital text, generate multiple correction candidate sentences using predefined rules or deep learning, and then use the pre-trained language model to calculate the context probabilities of the correction candidate sentences to select the optimal candidate sentence to replace the syntax error sentence and complete the syntax correction process.
[0093] 3. As Figure 2 shown, use a classifier to identify the subject and question type of the standard digital text to obtain a text classification result, select a semantic model according to the classification result, and use the LoRA fine-tuning technique to perform domain fine-tuning and task adaptation on the selected semantic model to obtain a deep semantic understanding model. Including:
[0094] 3.1 Construct an annotated dataset that includes multiple subjects (such as mathematics, physics, history) and multiple question types (such as short-answer questions, essay questions, calculation questions). Select a text classification model based on Transformer (such as BERT, RoBERTa), and use the annotated dataset to train the text classification model to obtain a classifier. Concatenate the questions and answers of the standard digital text and input them into the classifier for feature extraction to obtain a classification result. For example, input: question text + answer text, output: subject label (such as "physics"), question type label (such as "short-answer question").
[0095] 3.2 Collect subjective question answer data for each subject and question type, and construct a fine-tuning dataset (format: question + answer). Select a semantic model according to the classification result, and insert LoRA modules into the attention layer and feed-forward layer of the selected semantic model using LoRA fine-tuning technology.
[0096] Selection of semantic models, for example:
[0097] BERT: Suitable for extracting key concepts and viewpoints, applicable to short-answer questions and essay questions.
[0098] GPT: Good at generative tasks, applicable to semantic analysis of open-ended questions.
[0099] T5: Supports multi-task learning and can handle semantic understanding and generation tasks simultaneously.
[0100] Dynamically adjust the insertion position of the LoRA module and the dimension of the low-rank matrix according to the classification result. For example:
[0101] Short-answer questions: The LoRA module is mainly inserted into the attention layer to optimize semantic matching ability.
[0102] Essay questions: The LoRA module is inserted into both the attention layer and the feed-forward layer to optimize logical reasoning ability.
[0103] 3.2 Use the fine-tuning dataset to train the semantic model of the LoRA module, optimize the low-rank matrix to maximize the semantic similarity between the answer and the reference answer, and obtain a deep semantic understanding model.
[0104] 4. As Figure 2 shown, use the deep semantic understanding model to process the expression diversity of the standard digital text, and extract the key concepts and viewpoints in the standard digital text to obtain structured semantic information.
[0105] Including:
[0106] 4.1 Use the deep semantic understanding model to calculate the word similarity of the standard digital text, perform synonym replacement, and perform sentence pattern normalization processing to identify different expressions.
[0107] 4.2 Convert the standard digital text into an answer vector, match it with the reference answer vector, calculate the semantic similarity, and then decompose the semantic units in the answer vector to generate a structured representation; the semantic units include concepts, viewpoints, and arguments.
[0108] 4.3 Combine the structured representation, perform keyword extraction, concept recognition, and viewpoint extraction on the standard digital text, and convert the extracted key concepts and viewpoints into a graph structure to obtain structured semantic information.
[0109] For example, keyword extraction: Use algorithms such as TF-IDF and TextRank to extract keywords in the answer.
[0110] Concept recognition: Use a named entity recognition (NER) model to identify the core concepts (such as person names, place names, terms) in the answer.
[0111] Viewpoint extraction: Use argument mining technology to extract the arguments, evidence, and conclusions in the answer.
[0112] For example, answer: "Newton's first law states that an object will remain at rest or in uniform motion in a straight line when no external force acts on it."
[0113] Extraction result: Concepts: Newton's first law, external force, rest, uniform motion in a straight line. Viewpoint: An object will remain at rest or in uniform motion in a straight line when no external force acts on it.
[0114] 5. As Figure 2 shown, combine the structured semantic information, use a logical reasoning model to analyze the logical framework and reasoning process of the standard digital text, introduce a knowledge graph to verify the validity of the facts and arguments in the standard digital text, and then extract the arguments, evidence, and conclusions in the standard digital text for argument evaluation to obtain a logical reasoning evaluation result.
[0115] 5.1 Introduce a dependency parsing tool (such as SpaCy, Stanza) into a natural language inference model (such as DeBERTa, RoBERTa) to build a logical reasoning model.
[0116] 5.2 Use the dependency parsing tool in the logical reasoning model to parse the sentence structure of the standard digital text, identify the syntactic relationships (subject-predicate-object, modifiers, etc.), and build a logical tree for representing the logical hierarchy.
[0117] 5.3 Combine the logical tree, use the logical reasoning model to identify the logical relationships between sentences in the standard digital text (such as causal relationship, contrast relationship, conditional relationship), and string together the logical relationships to build an inference chain for representing the reasoning process (from the argument to the conclusion).
[0118] 5.4 Obtain existing open-source knowledge graphs (such as DBpedia, Wikidata) or domain knowledge bases (such as mathematical formula libraries, historical event libraries), and represent the knowledge in the form of triples to obtain a triple knowledge graph; the form of triples is as follows: "Newton's First Law" → "description" → "An object remains at rest or in uniform motion in a straight line under the action of no external force".
[0119] 5.5 Identify entities (such as person names, place names, terms) in the standard digital text to obtain answer entities, match the answer entities with the nodes in the triple knowledge graph, and establish entity links.
[0120] 5.6 Identify relationships (such as "description", "causes") in the standard digital text to obtain answer relationships, verify whether the answer relationships are consistent with the relationships in the triple knowledge graph, and perform relationship matching.
[0121] 5.7 Check whether the facts in the standard digital text exist in the triple knowledge graph. If not, mark the non-existent facts as unverified to complete fact verification.
[0122] 5.8 Identify and check the argument structure in the standard digital text, mark the argument, evidence, and conclusion, classify the evidence into factual evidence and logical evidence, and then evaluate whether the evidence supports the argument and whether the argument reasonably leads to the conclusion to obtain the evaluation results of argument integrity and reasonableness.
[0123] 6. Combine the logical reasoning evaluation results, calculate the scores for the content relevance dimension, logical dimension, language quality dimension, and creativity dimension respectively, and then synthesize the scores of each dimension to obtain the total score and generate a score report. Including:
[0124] 6.1 Content relevance dimension (40%)
[0125] Use a semantic similarity model (such as SBERT, USE) to calculate the semantic similarity between the standard digital text and the reference answer; the semantic similarity score ranges from 0 to 1, and the higher the score, the stronger the content relevance.
[0126] Extract the key concepts (such as terms, formulas, events) of the standard digital text and the reference answer, and calculate the key concept coverage rate of the standard digital text; coverage rate = the number of matching key points in the standard digital text / the number of key points in the reference answer.
[0127] Then multiply the semantic similarity and the key concept coverage rate according to the set weights to obtain the content relevance score; for example, content relevance score = semantic similarity score × 0.6 + key point coverage rate × 0.4.
[0128] 6.2 Logical dimension (30%)
[0129] Calculate the total coherence score based on the coherence and rationality of the reasoning chain. Coherence: Whether the reasoning chain is complete and whether there are logical jumps. Rationality: Whether the reasoning conforms to common sense and domain knowledge. Total coherence score = coherence score × 0.5 + rationality score × 0.5.
[0130] Calculate the argument integrity score based on the argument structure (arguments, evidence, conclusions) of the standard digital text; Argument integrity score = argument coverage rate × 0.4 + evidence support rate × 0.4 + conclusion rationality × 0.2.
[0131] Then, based on the set weights, synthesize the total coherence score and the argument integrity score to obtain the logical score.
[0132] 6.3 Language quality dimension (20%)
[0133] Calculate the language quality score based on the grammar error rate and spelling error rate of the standard digital text; Grammar error rate = number of errors / total number of sentences; Spelling error rate = number of errors / total number of words; Language quality score = 1 - (grammar error rate × 0.6 + spelling error rate × 0.4)
[0134] Calculate the lexical richness score based on the lexical diversity (such as Type - Token Ratio, TTR) and sentence pattern diversity of the standard digital text; Lexical richness score = TTR × 0.5 + sentence pattern diversity × 0.5.
[0135] Calculate the sentence coherence score based on the semantic coherence of the standard digital text, and then, based on the set weights, synthesize the language quality score, the lexical richness score, and the sentence coherence score to obtain the language quality score.
[0136] 6.4 Creativity dimension (10%)
[0137] Calculate the creativity score based on the semantic similarity between the standard digital text and the reference answer; Creativity score = 1 - semantic similarity (the lower the similarity, the higher the creativity).
[0138] Calculate the new idea score based on the number of ideas in the standard digital text; New idea score = number of new ideas / total number of ideas.
[0139] Then, based on the set weights, synthesize the semantic similarity and the new idea score to obtain the creativity score.
[0140] 6.5 Adjust the weights of the scores for each dimension according to the teaching objectives, then synthesize the scores of each dimension to obtain the total score, and generate a scoring report including the advantages, disadvantages, and improvement suggestions of the answers.
[0141] The expression of the total score is:
[0142] W = a·W 1 + b·W 2 + c·W 3 + d·W 4
[0143] where W is the total score, and W 1 , W 2 , W 3 , W 4 are the content relevance score, logical score, language quality score, and creativity score respectively, and a, b, c, and d are the weights of the content relevance score, logical score, language quality score, and creativity score respectively.
[0144] 7. Output the scoring report to display the scoring rules, provide multi-level feedback suggestions for the scoring report, and introduce the results of manual marking review to correct the total score and calculation process.
[0145] As Figure 3 shown, the present invention also provides an automatic marking and scoring system for subjective questions, including:
[0146] An initial recognition module, which is used to obtain handwritten or electronic answers, perform feature extraction and sequence modeling by using optical character recognition technology and deep learning technology, complete the initial recognition, and obtain digital text;
[0147] A preprocessing module, which is used to perform normalization, spelling correction, and grammar correction processing on the digital text to obtain standard digital text;
[0148] A model construction module, which is used to identify the subject and question type of the standard digital text by using a classifier to obtain a text classification result, select a semantic model according to the classification result, and use the LoRA fine-tuning technology to perform domain fine-tuning and task adaptation on the selected semantic model to obtain a deep semantic understanding model;
[0149] A deep understanding module, which is used to process the expression diversity of the standard digital text by using the deep semantic understanding model, extract the key concepts and viewpoints in the standard digital text, and obtain structured semantic information;
[0150] A logical reasoning module, which is used to combine the structured semantic information, analyze the logical framework and reasoning process of the standard digital text by using a logical reasoning model, introduce a knowledge graph to verify the validity of facts and arguments in the standard digital text, and then extract the arguments, evidence and conclusions in the standard digital text, conduct argument evaluation, and obtain a logical reasoning evaluation result;
[0151] A scoring calculation module, which is used to combine the logical reasoning evaluation result, calculate the scores of the content relevance dimension, logicality dimension, language quality dimension and creativity dimension respectively, then synthesize the scores of each dimension to obtain a total score, and generate a scoring report;
[0152] An output feedback module, which is used to output the scoring report to display the scoring rules, provide multi-level feedback suggestions for the scoring report, and introduce the results of manual marking review to correct the total score and calculation process.
[0153] Therefore, by providing an automatic marking and scoring method and system based on subjective questions, the present invention can accurately identify and deeply understand the input answers, and can also analyze and demonstrate through powerful logical reasoning ability, output accurate scoring processes and results, increasing the accuracy and transparency of automatic scoring.
[0154] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0155] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An automatic marking and grading method based on subjective questions, characterized in that: include: Obtain handwritten or electronic answers, use optical character recognition technology and deep learning technology to perform feature extraction and sequence modeling, complete initial recognition, and obtain digital text; Normalizing, correcting spelling and grammar of the digital text to obtain a standard digital text; Using a classifier to identify the subject and title type of the standard digital text to obtain a text classification result, selecting a semantic model based on the classification result, and using LoRA fine-tuning technology to perform domain fine-tuning and task adaptation on the selected semantic model to obtain a deep semantic understanding model; Using the deep semantic understanding model to process the expression diversity of the standard digital text, extract key concepts and viewpoints in the standard digital text, and obtain structured semantic information; In combination with the structured semantic information, a logical reasoning model is used to analyze the logical framework and reasoning process of the standard digital text, a knowledge graph is introduced to verify the validity of facts and arguments in the standard digital text, and then the arguments, evidence and conclusions in the standard digital text are extracted to conduct argumentation evaluation and obtain a logical reasoning evaluation result; Combined with the logical reasoning evaluation results, the scores of the content relevance dimension, the logic dimension, the language quality dimension and the creativity dimension are calculated respectively, and then the scores of each dimension are combined to obtain a total score, and a score report is generated; The scoring report is output to display the scoring details, multi-level feedback suggestions are provided for the scoring report, and the results of manual marking and review are introduced to correct the total score and calculation process.
2. The automatic marking and grading method based on subjective questions according to claim 1 is characterized in that: Obtain handwritten or electronic answers, use optical character recognition technology and deep learning technology to perform feature extraction and sequence modeling, complete initial recognition, and obtain digital text, including: Acquire a handwritten answer image, perform denoising, binarization, tilt correction, cropping and normalization on the handwritten answer image, and then perform text line segmentation and character segmentation to obtain an initial answer image; A convolutional neural network is used to extract a two-dimensional feature map of the initial answer image, and the two-dimensional feature map is expanded along the width direction of the initial answer image to generate a one-dimensional feature sequence in the form of a time series, thereby completing visual feature extraction; The one-dimensional feature sequence is modeled at the character level and the word level respectively using a recurrent neural network and a Transformer model to extract local features and global dependencies, obtain the contextual relationship between characters, complete the initial recognition of the handwritten answer image, and obtain the digital text.
3. The automatic marking and grading method based on subjective questions according to claim 2 is characterized in that: The digital text is subjected to format normalization, spelling correction and grammar correction processing to obtain a standard digital text, including: Combined with regular expressions, the digital text is case-converted, punctuation marks and blank characters are processed to complete format normalization processing; wherein the case of proper nouns is retained during the case conversion process; Constructing a domain dictionary, detecting misspelled words in the digital text, calculating the similarity between the word and the words in the domain dictionary using an edit distance algorithm, generating multiple correction candidate words according to the calculation results, and then calculating the context probabilities of the multiple correction candidate words using a pre-trained language model to select the best candidate word to replace the misspelled word, thereby completing the spelling correction process; A grammatical analysis model is used to detect grammatical errors in the digital text, and a plurality of correction candidate sentences are generated using predefined rules or deep learning. The pre-trained language model is then used to calculate the context probability of the correction candidate sentences, so as to select the optimal candidate sentence to replace the grammatically incorrect sentence and complete the grammatical correction process.
4. The automatic marking and grading method based on subjective questions according to claim 3 is characterized in that: The classifier is used to identify the subject and title type of the standard digital text to obtain a text classification result, a semantic model is selected according to the classification result, and the LoRA fine-tuning technology is used to perform domain fine-tuning and task adaptation on the selected semantic model to obtain a deep semantic understanding model, including: Construct a labeled data set containing multiple subjects and multiple question types, select a Transformer-based text classification model, use the labeled data set to train the text classification model to obtain a classifier, splice the questions and answers of the standard digital text and input them into the classifier for feature extraction to obtain a classification result; Collect the subjective question answer data of each subject and question type, construct a fine-tuning dataset, select a semantic model based on the classification results, and use LoRA fine-tuning technology to insert LoRA modules in the attention layer and feedforward layer of the selected semantic model; The semantic model of the LoRA module is trained using the fine-tuning dataset, and the low-rank matrix is optimized to maximize the semantic similarity between the answer and the reference answer, thereby obtaining a deep semantic understanding model.
5. The automatic marking and grading method based on subjective questions according to claim 4 is characterized in that: The deep semantic understanding model is used to process the expression diversity of the standard digital text, extract key concepts and viewpoints in the standard digital text, and obtain structured semantic information, including: Utilizing the deep semantic understanding model to calculate the word similarity of the standard digital text, perform synonym replacement, and perform sentence normalization processing to identify different expressions; Converting the standard digital text into an answer vector and matching it with a reference answer vector, calculating semantic similarity, and then decomposing the semantic units in the answer vector to generate a structured representation; the semantic units include concepts, viewpoints, and arguments; In combination with the structured representation, keyword extraction, concept recognition and viewpoint extraction of the standard digital text are performed, and the extracted key concepts and viewpoints are converted into a graph structure to obtain structured semantic information.
6. The automatic marking and grading method based on subjective questions according to claim 5 is characterized in that: In combination with the structured semantic information, a logical reasoning model is used to analyze the logical framework and reasoning process of the standard digital text, including: Introduce dependency syntax analysis tools into the natural language reasoning model to build a logical reasoning model; Utilizing the dependency syntax analysis tool in the logic reasoning model to parse the sentence structure of the standard digital text, identify grammatical relations, and construct a logic tree for representing the logic hierarchy; In combination with the logic tree, the logic reasoning model is used to identify the logical relationship between sentences in the standard digital text, and the logical relationship is connected in series to construct a reasoning chain for representing the reasoning process.
7. The automatic marking and grading method based on subjective questions according to claim 6 is characterized in that: The knowledge graph is introduced to verify the validity of the facts and arguments in the standard digital text, and then the arguments, evidence and conclusions in the standard digital text are extracted to conduct argument evaluation, including: Obtain an existing open source knowledge graph or domain knowledge base, and represent the knowledge in the form of triples to obtain a triple knowledge graph; Identify entities in the standard digital text to obtain answer entities, match the answer entities with nodes in the triple knowledge graph, and establish entity links; Identify the relationship in the standard digital text, obtain the answer relationship, verify whether the answer relationship is consistent with the relationship in the triple knowledge graph, and perform relationship matching; Check whether the fact in the standard digital text exists in the triple knowledge graph. If not, mark the non-existent fact as unverified, and complete the fact verification; Identify and check the argument structure in the standard digital text, mark the arguments, evidence and conclusions, classify the evidence into factual arguments and logical arguments, and then evaluate whether the evidence supports the arguments and whether the arguments reasonably lead to conclusions, so as to obtain an evaluation result of the completeness and rationality of the argument.
8. The automatic marking and grading method based on subjective questions according to claim 7 is characterized in that: Combined with the logical reasoning evaluation results, the scores of the content relevance dimension, logic dimension, language quality dimension and creativity dimension are calculated respectively, and then the scores of each dimension are combined to obtain the total score, and a score report is generated, including: Calculating the semantic similarity between the standard digital text and the reference answer using a semantic similarity model, extracting key concepts in the standard digital text and the reference answer, and calculating the key concept coverage of the standard digital text, and then multiplying the semantic similarity and the key concept coverage according to a set weight to obtain a content relevance score; Calculating a total coherence score according to the coherence and rationality of the chain of reasoning, calculating an argument integrity score according to the argument structure of the standard digital text, and then combining the total coherence score and the argument integrity score according to a set weight to obtain a logic score; Calculate a language quality score according to the grammatical error rate and spelling error rate of the standard digital text, calculate a vocabulary richness score according to the vocabulary diversity and sentence diversity of the standard digital text, calculate a sentence coherence score according to the semantic coherence of the standard digital text, and then combine the language quality score, the vocabulary richness score and the sentence coherence score according to the set weights to obtain a language quality score; Calculate the creativity score according to the semantic similarity between the standard digital text and the reference answer, calculate the new viewpoint score according to the number of viewpoints in the standard digital text, and then combine the semantic similarity and the new viewpoint score according to the set weight to obtain the creativity score; The weights of each dimension are adjusted according to the teaching objectives, and the scores of each dimension are combined to obtain the total score, and a scoring report is generated including the advantages and disadvantages of the answer and suggestions for answer improvement.
9. The automatic marking and grading method based on subjective questions according to claim 8 is characterized in that: The expression of the total score is: W=a·W1+b·W2+c·W3+d·W4 Among them, W is the total score, W1, W2, W3, W4 are the content relevance score, logic score, language quality score and creativity score respectively, a, b, c, d are the weights of the content relevance score, logic score, language quality score and creativity score respectively.
10. An automatic marking and grading system based on subjective questions, characterized in that: include: The initial recognition module is used to obtain handwritten or electronic answers, and uses optical character recognition technology and deep learning technology to perform feature extraction and sequence modeling to complete initial recognition and obtain digital text; A preprocessing module, used for normalizing, correcting spelling and grammar of the digital text to obtain a standard digital text; A model building module is used to use a classifier to identify the subject and title type of the standard digital text, obtain a text classification result, select a semantic model according to the classification result, and use LoRA fine-tuning technology to perform domain fine-tuning and task adaptation on the selected semantic model to obtain a deep semantic understanding model; A deep understanding module, used to process the expression diversity of the standard digital text using the deep semantic understanding model, extract key concepts and viewpoints in the standard digital text, and obtain structured semantic information; A logical reasoning module is used to combine the structured semantic information, use a logical reasoning model to analyze the logical framework and reasoning process of the standard digital text, introduce a knowledge graph to verify the validity of the facts and arguments in the standard digital text, and then extract the arguments, evidence and conclusions in the standard digital text to conduct argumentation evaluation and obtain a logical reasoning evaluation result; A scoring calculation module is used to calculate the scores of the content relevance dimension, logic dimension, language quality dimension and creativity dimension respectively based on the logical reasoning evaluation results, and then combine the scores of each dimension to obtain a total score and generate a scoring report; The output feedback module is used to output the scoring report to display the scoring details, provide multi-level feedback suggestions for the scoring report, and introduce the results of manual marking and review to correct the total score and calculation process.
Citation Information
Patent Citations
Automatic subjective question scoring method based on multi-feature fusion
CN111310458A
Subjective question marking and scoring method and device, computer equipment and storage medium
CN113822040A
Method for constructing pre-training subject corpus of digital human teacher multi-mode large language model
CN117076693A
Answer scoring method and system based on knowledge graph
CN117648915A
Subjective question test paper marking method based on large model
CN118861522A
Cited By
Dynamic rule base construction method and device based on semantic contact calculation and medium
CN120409705A
Learner argumentation process effect evaluation method in critical thinking training mode
CN120470133A
Methods for evaluating the effectiveness of learners' argumentation process under critical thinking training models
CN120470133B
Multi-agent-based automatic subjective question scoring method and device and storage medium
CN120471596A
Online examination integrated intelligent paper marking method and system
CN120745643A