A method and system for automatic marking of examination papers based on generative artificial intelligence

By applying the generative artificial intelligence big model ChatGPT in the automatic marking system, the limitations of the existing system in content recognition, feature extraction and score generation are solved, the efficiency and accuracy of reviews are improved, and the interpretability of the scoring results is enhanced.

CN119540972BActive Publication Date: 2025-05-27SHANDONG JIANZHU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510103874.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing automatic marking system has limitations in content recognition, feature extraction, text similarity comparison and score generation, resulting in poor review efficiency and accuracy and insufficient ability to interpret scoring results.

Method used

The generative artificial intelligence big model ChatGPT is used in the review process throughout the cycle, combining natural language processing and machine learning technology to assist OCR recognition, enhance feature extraction and text similarity comparison, and introduce collaborative active learning mechanisms and attention mechanisms to optimize the scoring model.

Benefits of technology

It improves review efficiency and accuracy, enhances the interpretability of scoring results, provides efficient and accurate scoring and feedback, and improves the quality and efficiency of educational assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540972B_ABST
    Figure CN119540972B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and system for automatic marking of examination papers based on generative artificial intelligence, which relates to the technical fields of artificial intelligence and educational evaluation, and includes: obtaining an image of a paper answer sheet, performing content recognition on the processed image to obtain the content of the answer image; performing anomaly detection and correction on the content of the answer image through a generative artificial intelligence large model; extracting features from the content of the corrected answer image; using the generative artificial intelligence large model to extract key information from the content of the corrected answer image and embed it to form a feature vector of the new answer image content; calculating the cosine similarity score and the prediction score of the generative artificial intelligence large model; when there is a serious gap, introducing a collaborative active learning mechanism to evaluate the confidence of the candidate's answer; when the confidence is higher than a given threshold, outputting the predicted score. The present invention applies the generative artificial intelligence large model technology throughout the marking process to achieve automatic marking of examination papers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and educational assessment, and particularly to a method and system for automatically grading test papers based on generative artificial intelligence. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] The digital wave has promoted the transformation of the education field. The automated grading and rapid feedback functions of campus examination systems have become important driving forces for improving the efficiency of educational assessment, and there is a great demand for educational informatization. At the same time, the rapid development of artificial intelligence technology, especially the progress in the field of natural language processing (NLP), has led to an increasing application of machine learning technology in the field of educational assessment.

[0004] Traditional test paper grading methods mainly rely on teachers' subjective judgments. This method is not only inefficient but also easily affected by graders' fatigue and subjectivity. To improve the efficiency and accuracy of grading, researchers have begun to explore the use of machine learning technology to achieve automatic test paper grading.

[0005] Currently, although there have been many research works on automatic marking, there are still some challenges and limitations. First, at the information input end, the accuracy of automatic marking systems depends to a large extent on the accuracy of scanning and recognition technologies, and there are still challenges in accurately recognizing handwritten text and complex formulas. Second, to improve the accuracy of grading, automatic marking systems require a large amount of labeled data for training, which may be time-consuming and costly in actual operation. Third, compared with manual marking, self-service marking systems have deficiencies in explaining grading results, which may affect the acceptance of grading results by students and teachers.

[0006] Generative artificial intelligence, especially ChatGPT, as a generative artificial intelligence dialogue system based on large language models, understands and interprets user requests through massive data storage and an efficient design architecture, and generates "response texts with relatively high complexity" in a manner close to human natural language.

[0007] Currently, the applications of generative artificial intelligence in the education field are increasing, but in existing educational assessment systems, generative artificial intelligence large models have not been applied to the entire cycle process of the grading system, and still have the limitations of traditional systems in aspects such as content recognition, feature extraction, text similarity comparison, and score generation, resulting in poor efficiency and accuracy of grading. Summary of the Invention

[0008] To overcome the deficiencies of the above-mentioned existing technologies, the present invention provides a method and system for automatic marking of examination papers based on generative artificial intelligence, which applies the generative artificial intelligence large model, namely ChatGPT technology, throughout the marking process to achieve automatic marking of examination papers. By combining natural language processing and machine learning technologies, ChatGPT is used to assist OCR recognition, enhance feature extraction and text similarity comparison, and combined with a collaborative active learning mechanism to continuously optimize the scoring model, improve the marking efficiency and accuracy, and introduce an attention mechanism to enhance the interpretability of the scoring results, providing efficient and accurate scoring and feedback, thereby improving the quality and efficiency of educational evaluation.

[0009] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0010] The first aspect of the present invention provides a method for automatic marking of examination papers based on generative artificial intelligence.

[0011] A method for automatic marking of examination papers based on generative artificial intelligence includes:

[0012] Obtain an image of a paper answer sheet and perform preprocessing to obtain a processed answer image;

[0013] Perform content recognition on the processed answer image to obtain the content of the answer image; perform anomaly detection on the content of the answer image through a generative artificial intelligence large model, and correct the content of the answer image after obtaining the detection result to obtain the corrected content of the answer image;

[0014] Extract features from the corrected content of the answer image to obtain a feature vector of the content of the answer image; use a generative artificial intelligence large model to extract key information from the corrected content of the answer image and embed it into the feature vector of the content of the answer image to form a new feature vector of the content of the answer image;

[0015] Calculate the cosine similarity score for the feature vector of the content of the answer image; at the same time, also obtain a predicted score for the feature vector of the content of the answer image through a generative artificial intelligence large model;

[0016] When there is a serious gap between the cosine similarity score and the predicted score of the generative artificial intelligence large model, introduce a collaborative active learning mechanism and evaluate the confidence of the candidate's answer through an uncertainty sampling query method; when the confidence is higher than a given threshold, output the predicted score.

[0017] The second aspect of the present invention provides a system for automatic marking of examination papers based on generative artificial intelligence.

[0018] A system for automatic marking of examination papers based on generative artificial intelligence includes:

[0019] The test paper acquisition module is configured to: acquire an image of a paper answer sheet and perform preprocessing to obtain a processed answer image;

[0020] The content recognition module is configured to: perform content recognition on the processed answer image to obtain the content of the answer image; perform anomaly detection on the content of the answer image through a generative artificial intelligence large model, and correct the content of the answer image after obtaining the detection result to obtain the corrected content of the answer image;

[0021] The feature extraction module is configured to: extract features from the corrected content of the answer image to obtain a feature vector of the content of the answer image; use a generative artificial intelligence large model to extract key information from the corrected content of the answer image and embed it into the feature vector of the content of the answer image to form a new feature vector of the content of the answer image;

[0022] The prediction scoring module is configured to: calculate a cosine similarity score for the feature vector of the content of the answer image; at the same time, obtain a prediction score for the feature vector of the content of the answer image through a generative artificial intelligence large model;

[0023] The confidence evaluation module is configured to: when there is a serious gap between the cosine similarity score and the prediction score of the generative artificial intelligence large model, introduce a collaborative active learning mechanism and evaluate the confidence of the candidate's answer through an uncertainty sampling query method; when the confidence is higher than a given threshold, output the prediction score.

[0024] A third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in a method as described in the first aspect of the present invention are implemented.

[0025] A fourth aspect of the present invention provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the steps in a method as described in the first aspect of the present invention are implemented.

[0026] A fifth aspect of the present invention provides a computer program product containing instructions. When it runs on a computer, it causes the computer program to be executed by a processor to implement the steps in a method as described in the first aspect of the present invention.

[0027] The above one or more technical solutions have the following beneficial effects:

[0028] The present invention applies the generative artificial intelligence large model, i.e., the ChatGPT large model, throughout the entire review process, including content recognition, feature extraction, text similarity comparison, score generation, etc., overcoming the limitations of traditional systems. In addition, the present invention uses ChatGPT to assist in OCR recognition, enhance feature extraction and text similarity comparison, and combines a collaborative active learning mechanism to continuously optimize the scoring model, improving the review efficiency and accuracy, and introducing an attention mechanism to enhance the interpretability of the scoring results.

[0029] The present invention integrates the generative artificial intelligence large model, i.e., the ChatGPT large model, with various algorithmic techniques. It uses ChatGPT to assist the BPD and SideNet models in OCR recognition and combines ChatGPT for anomaly detection; ChatGPT optimizes algorithms such as Word2Vec, BERT, LaMDA, etc. for feature extraction; ChatGPT enhances text vector representation and calculates cosine similarity; and ChatGPT predicts scores and combines an active learning mechanism for scoring decisions. Through cosine similarity calculation and ChatGPT's predicted scores, and combined with a weight factor for weighted averaging, the final score is obtained, making the scoring result more transparent and interpretable, which helps students and teachers understand and accept the scoring result.

[0030] The automatic test paper / homework grading system of the present invention based on the generative artificial intelligence large model ChatGPT has remarkable practicality and can be widely applied in the education field. It overcomes the problem that traditional test paper grading methods highly rely on manual feature extraction or data collection, greatly improves the grading efficiency through an automated process, reduces labor costs, and uses ChatGPT's powerful language understanding ability to accurately understand the text content and perform semantic analysis, thereby improving the grading accuracy and reliability. And it overcomes the deficiency of the self-service grading system in explaining the grading results, thus realizing more efficient, accurate, and interpretable test paper grading. It can be widely applied in scenarios such as schools, training institutions, online education platforms, etc., and has important practical value.

[0031] Advantages of additional aspects of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0033] Figure 1 It is a flowchart of an automatic test paper grading method based on generative artificial intelligence according to Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention pertains.

[0035] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0036] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0037] Embodiment 1

[0038] This embodiment discloses an automatic test paper marking method based on generative artificial intelligence, and proposes an automatic marking method based on a generative artificial intelligence large model, with full-cycle cooperation between ChatGPT and expert knowledge. The ChatGPT large model is applied to multiple key links such as answer content recognition, feature extraction, text similarity comparison, and score generation. In this way, an attempt is made to use large model technology to penetrate and combine relevant methods in the fields of pattern recognition, multi-modal learning, natural language processing, deep learning, active learning, etc. to achieve automatic scoring of test papers, including:

[0039] Obtain the image of the paper answer sheet and perform preprocessing to obtain the processed answer image;

[0040] Perform content recognition on the processed answer image to obtain the content of the answer image; perform anomaly detection on the content of the answer image through a generative artificial intelligence large model, and after obtaining the detection result, correct the content of the answer image to obtain the corrected answer image content;

[0041] Extract the feature vectors of the content of the corrected answer image; use the generative artificial intelligence large model to extract key information from the content of the corrected answer image and embed it into the feature vectors of the content of the answer image to form new feature vectors of the content of the answer image;

[0042] Calculate the cosine similarity score for the feature vectors of the content of the answer image; at the same time, also obtain the predicted score for the feature vectors of the content of the answer image through the generative artificial intelligence large model;

[0043] When there is a serious gap between the cosine similarity score and the predicted score of the generative artificial intelligence large model, introduce a collaborative active learning mechanism, and evaluate the confidence of the candidate's answer through an uncertainty sampling query method; when the confidence is higher than a given threshold, output the predicted score.

[0044] In this embodiment, the generative artificial intelligence large model is specifically the ChatGPT large model.

[0045] To more clearly elaborate on this embodiment, as Figure 1 shown, the implementation process of a test paper automatic grading method based on generative artificial intelligence can be specifically described as follows:

[0046] Step 1, First stage: Content recognition of the answer image.

[0047] Step 1-1, Data collection and image segmentation: Obtain the image of the paper answer sheet and perform preprocessing to obtain the processed answer image.

[0048] (1) The specific method of obtaining the image of the paper answer sheet is as follows:

[0049] Image acquisition: Obtain the image of the paper answer sheet through devices such as scanners and digital cameras.

[0050] In this embodiment, obtaining the image of the paper answer sheet specifically includes: the student's handwritten answer image and the reference standard answer image.

[0051] (2) Perform preprocessing on the image of the paper answer sheet, specifically as follows:

[0052] Image denoising: Remove the noise in the answer sheet image, such as salt and pepper noise, Gaussian noise, etc., to improve the image quality.

[0053] Image binarization: Convert the color or grayscale image into a black and white binary image for subsequent question segmentation and character recognition.

[0054] Skew correction: Detect the skew angle of the image and perform correction to ensure accurate question segmentation and correct character recognition.

[0055] Image segmentation: Segment the preprocessed test paper image into individual question images according to the content, so as to score each question image one by one. Use methods such as the connected region method, key character segmentation method, or superpixel segmentation to segment the test paper image into multiple answer sub-images, and each sub-image corresponds to a question and the student's answer situation.

[0056] Step 1-2, Perform content recognition on the processed answer image to obtain the content of the answer image. Perform anomaly detection on the content of the answer image through the generative artificial intelligence large model, and after obtaining the detection result, correct the content of the answer image to obtain the corrected content of the answer image.

[0057] Step 1-2-1, Perform content recognition on the processed answer image through OCR technology.

[0058] In this embodiment, the OCR (Optical Character Recognition) optical character recognition technology driven by a large model is used to recognize the content of the processed answer image. The OCR technology aims to use a large-scale machine learning model to improve the accuracy and efficiency of text recognition in images.

[0059] In this embodiment, the machine learning models used by the OCR technology specifically include the BPD (Tree-based Model with Branch Parallel Decoding, BPD) model, the SideNet model, and various image segmentation models.

[0060] In this embodiment, the content of the processed answer image includes mathematical formulas, text, and subgraphs.

[0061] For example, the answer image of a math test question will contain subgraphs, mathematical formulas, text, etc. Therefore, using the OCR technology to achieve answer content recognition specifically includes:

[0062] (1) Mathematical formula recognition

[0063] For the mathematical formulas in the processed answer image, the BPD (Tree-based Model with Branch Parallel Decoding, BPD) model is used for mathematical formula recognition.

[0064] The BPD model can effectively recognize the structure of mathematical formulas and use a pre-trained Transformer model to predict symbols and relationships. The core idea of the model is to reduce the number of decoding time steps through a parallel decoding mechanism, thereby improving the recognition efficiency and accuracy.

[0065] By using the BPD model, various types of mathematical formulas can be accurately recognized. For example: basic arithmetic symbols such as addition, subtraction, multiplication, and division, parentheses, fractions, exponents, trigonometric functions, logarithms, etc. The model can recognize these operators and match them with the corresponding operands.

[0066] (2) Text recognition

[0067] For the text (non-mathematical formula) in the processed answer image, the SideNet model is used for OCR recognition to convert the text in the image into text data.

[0068] The SideNet model can effectively recognize Chinese characters by jointly learning radical and glyph information, solve the problem of repeated mapping, and improve the recognition accuracy.

[0069] The main components of the SideNet model include:

[0070] CNN backbone network: Extract visual features from input images and glyphs.

[0071] Dimension Decomposition Conversion Module (DDCM): Extract structural information from the radical tree and convert it into structure-aware embeddings. Through dimension decomposition, this module decomposes radical information into three dimensions: category, depth, and position, and uses learnable parameters and non-linear fusion strategies to enhance the embedding ability.

[0072] Count-based Spatial Conversion Module (CSCM): Learn the radical distribution in glyph images and generate spatial-aware embeddings. By predicting radical categories on the glyph feature map and using count loss for supervision, this module accurately encodes the position information of radicals and solves the problem of repeated mapping.

[0073] Two-stage classifier: Integrate radical and glyph information and classify characters. In the first stage, similarity map-guided feature fusion is used to convert the similarity map into a mask for scaling side information features, and a residual convolution is used to model the similarity map and add it to the fusion process. In the second stage, the Transformer decoder layer is used to calculate feature correlations and output the final probability matrix to identify unseen characters.

[0074] (3) Subgraph recognition and segmentation

[0075] For the subgraphs in the processed answer images (i.e., the test papers containing image elements), use image segmentation methods such as the connected region method, key character segmentation method, or superpixel segmentation to segment the image elements in the test paper images and extract image feature representations.

[0076] Step 1-2-2: Perform anomaly detection on the content of the answer image through a generative artificial intelligence large model.

[0077] Anomaly detection using the generative artificial intelligence large model (ChatGPT) is as follows:

[0078] For the text and formula information on the test paper, errors may occur during image conversion only through OCR technology. Therefore, before feature processing, it is necessary to handle the outliers that appear after OCR technology recognition.

[0079] First, for the content of the answer image, use the confidence scoring technology provided by the OCR software to obtain the uncertain text regions after OCR recognition and identify the low-confidence text.

[0080] Second, for the identified low-confidence text, use the ChatGPT large model to detect abnormal words or phrases in the OCR output and perform further manual review or automatic correction.

[0081] More specifically, to detect abnormal words or phrases in the OCR output using the ChatGPT large model, the steps are as follows:

[0082] Communicate using the API interface provided by ChatGPT.

[0083] Before sending the request, clean the text recognized by OCR, remove extra spaces, special characters, and ensure that the text encoding is compatible with the ChatGPT API.

[0084] The text data recognized by OCR is encapsulated in JSON format and sent to the ChatGPT server via an HTTPS request. The request includes the OCR text, context information, and any specific instructions to guide ChatGPT in performing anomaly detection.

[0085] ChatGPT analyzes the text recognized by OCR to detect whether there are grammar errors, incoherent sentence structures, or content unrelated to the topic. It also includes identifying words or phrases that OCR may misrecognize, which usually do not match the context.

[0086] The following is an example process:

[0087] Suppose the OCR outputs a text: "10 + 6 = / 6."

[0088] The system sends this text to the ChatGPT API and requests to detect possible anomalies.

[0089] ChatGPT returns the analysis result, indicating that " / " may be a misrecognition, and the correct symbol may be "1".

[0090] The system can then mark the symbol for further manual review or automatic correction, and finally obtain the corrected content of the answer image.

[0091] ChatGPT not only improves the quality of the OCR text but also enhances the accuracy and reliability of the entire test paper automatic grading system.

[0092] Step 2, Second part: Word segmentation and knowledge graph construction.

[0093] Knowledge graphs can uncover potential patterns and trends hidden in large amounts of information and provide support for decision-making. Taking the mathematics curriculum as an example, constructing a knowledge graph for it can systematically organize various knowledge points, theorems, formulas, and their interrelationships in the field of mathematics in a structured form. Teachers can clearly obtain the mathematics knowledge system, consult key knowledge points, and intuitively discover the associations between knowledge points. Introducing a knowledge graph into the automatic grading system can not only improve the quality of feature extraction but also improve the interpretability of automatic grading.

[0094] Step 2-1, Keyword Extraction.

[0095] Identify important words or phrases from the response text, which usually represent the theme or main content of the text.

[0096] In this embodiment, Jieba word segmentation and TF-IDF technology are used for keyword extraction, and the specific steps are as follows:

[0097] (1) First, calculate the term frequency (TF): For each word, calculate its frequency of occurrence in the text.

[0098] (2) Then calculate the inverse document frequency (IDF): For each word, calculate its rarity in the entire document set.

[0099] (3) After that, calculate the TF-IDF score: For each word, its TF-IDF score is the product of TF and IDF. Finally, select the word with the highest score as the keyword. This keyword may contain the knowledge points of these questions. We have performed text embedding before keyword extraction, so the obtained keywords are in the form of vectors with a fixed dimension.

[0100] (4) To identify keywords related to knowledge points, use the attention mechanism of the ChatGPT model to analyze the importance of each word in the sentence.

[0101] The ChatGPT model can judge the relevance of a word to knowledge points based on the context information of the word. Through its powerful language understanding ability, ChatGPT can identify key semantic elements that may be ignored by traditional TF-IDF methods, thereby improving the accuracy and comprehensiveness of the keyword vector representation. We perform semantic analysis on these keywords through ChatGPT to ensure that they are closely related to knowledge points and verify whether the keywords extracted by the TF-IDF algorithm are relevant to the theme of the text. The specific process is as follows:

[0102] ① Reconstruction of the input sentence

[0103] First, use the masking mechanism in self-supervised learning to randomly mask some words in the student's answer text to obtain a new sentence, and the masked words are used as the predicted output of ChatGPT.

[0104] ② Model fine-tuning

[0105] Next, the reconstructed sentence is passed to ChatGPT. Specifically, the API interface provided by ChatGPT is used for communication and sent to the ChatGPT server via an HTTPS request. The request includes the text of the reconstructed sentence and the masked words, guiding ChatGPT to perform self-supervised fine-tuning training.

[0106] Step 2-2: Knowledge Graph Construction.

[0107] In step 2-1, the keywords (i.e., entities) in the test questions are obtained. For example, for a math test paper, math concepts, formulas, theorem names, etc. can be extracted, and semantic analysis is performed on these keywords through ChatGPT to ensure their close relevance to knowledge points, identify the relationships between entities, such as mathematical relationships like "equal to", "greater than", "less than", etc., or "definition", "derivation", etc., as well as the logical relationships between concepts. With this information, a relevant knowledge graph can be constructed.

[0108] The specific construction process is as follows:

[0109] (1) Knowledge Representation and Fusion

[0110] The extracted entities and relationships are represented in the form of triples, and entity linking, data alignment, etc. are used to align and fuse the entities and relationships from different test questions, eliminating redundancy and conflicts.

[0111] (2) Text Generation

[0112] A generative model is used to assist in generating labeled data for training the information extraction model, or generating variants of math test questions to enrich the content of the knowledge graph. The generative model (such as GPT) is used to deeply understand math test questions and extract implicit knowledge points and logical relationships.

[0113] (3) Knowledge Graph Construction

[0114] Graph construction tools such as Neo4j and RDF4J are used to store and manage the extracted entities and relationships, and visualization tools are used to display the knowledge graph for teachers to understand and analyze.

[0115] Step 3: Third part: Feature extraction.

[0116] After obtaining the content information of the obtained response image, feature extraction needs to be performed to facilitate text analysis. Feature extraction is one of the key steps in this embodiment. In this embodiment, for the recognized formulas, strings are used to describe and represent them, and no additional feature extraction is performed. For the recognized text and subgraphs, combined with the knowledge graph constructed in step 2, various methods are used to extract their features, including but not limited to related technologies such as Word2Vec, Bert, ResNet, etc., and the generative artificial intelligence large model is used to refine and adjust the features of the text information through semantic analysis.

[0117] In this embodiment, the content of the corrected response image is divided into text and non-text elements such as charts, graphics, or handwritten notes in the test paper.

[0118] Step 3-1: Extract features from the text content of the corrected response image to obtain the feature vector of the text content of the response image.

[0119] The following are the steps for feature extraction of the text content:

[0120] (1) The first step to be carried out is tokenization: decompose the text content of the response image into words or phrases (tokens) and remove common but insignificant words in the text, such as "de", "he", etc.

[0121] (2) Then build a window: For each word, select the surrounding words to form a context window.

[0122] (3) After tokenization and building the context window, use ChatGPT to analyze the context of each word to generate a richer semantic representation, which helps the Word2Vec model capture deeper semantic relationships.

[0123] (4) Then train the Word2Vec model to capture the semantic relationship features between words and convert them into vector representations of a fixed dimension.

[0124] (5) Next, use the pre-trained BERT model to perform sentence encoding on the candidate answers and the standard answers to obtain the semantic relationship features at the sentence level and convert them into vector representations of a fixed dimension.

[0125] BERT (Bidirectional Encoder Representations from Transformers, BERT) is a pre-trained language model based on Transformer. The BERT model can effectively capture the semantic information in a sentence and convert it into a vector of a fixed dimension.

[0126] (6) According to the semantic relationship feature vector representation captured by the Word2Vec model and the sentence-level semantic relationship feature vector representation obtained by the pre-trained BERT model, the feature vector of the answer image text content is obtained.

[0127] In this embodiment, the Word2Vec model is trained to capture the semantic relationship between words. Specifically, the Word2Vec model is used for text embedding:

[0128] The original text is converted into a vector representation of a fixed dimension, retaining the semantic information of the text.

[0129] There are two models in the Word2Vec model, the Skip-Gram model and the CBOW model, and each model has its own objective function. The objective function defines the learning objective of the model, that is, by adjusting the model parameters (such as word vectors), the value of the objective function is minimized. In Word2Vec, the optimization process of the objective function is the process of learning word vectors, and the value of the objective function can be used to evaluate the quality of the model. During the training process, if the value of the objective function decreases, it usually means that the model performance is improving. The objective function of the Skip-Gram model is expressed as follows:

[0130] (1)

[0131] where is the conditional probability of the context word given the center word , is the total number of training samples, t is the subscript index variable, j is the subscript index variable, is the size of the context window. θ represents the model parameters to be learned, and the J(θ) function represents the conditional probability of the context word given the center word . The purpose of optimization is to maximize this probability. In actual operation, the conditional probability is usually calculated by the softmax function:

[0132] (2)

[0133] W is the size of the vocabulary, represents the subscript index variable, represents the th word vector representation. and are the vector representations of and respectively.

[0134] The objective function of the CBOW model is expressed as follows:

[0135] (3)

[0136] is the total number of training samples, is the central word, is the central word background word set of is the background word set lower central word conditional probability of

[0137] Similar to formula (2), the conditional probability is also calculated through the softmax function. These objective functions are minimized through optimization algorithms such as gradient descent during training, so as to learn the effective vector representation of words.

[0138] Step 3-2: Use the generative artificial intelligence large model to extract key information from the text content of the corrected answer image and embed it into the feature vector of the answer image text content to form a new feature vector of the answer image text content.

[0139] In this embodiment, in order to further improve the quality of semantic representation, ChatGPT is used to optimize the sentence encoding result of BERT obtained in step 3-1 (i.e., the feature vector of the answer image content). Through its powerful language understanding ability, ChatGPT can identify the key information in the sentence and integrate it into the semantic representation, so that the semantic representation can more accurately reflect the semantic information of the answer. The process of using ChatGPT to optimize the sentence encoding result of BERT is as follows:

[0140] ① Reconstruction of the input sentence

[0141] First, reconstruct the text content of the answer image as the input sentence, and convert the text content of the answer image (including the student's answer and the reference standard answer) into a more structured format.

[0142] This process is achieved through the following steps:

[0143] Combine the student's answer with the corresponding knowledge concepts (KnowledgeConcepts, KCs) and question text in the knowledge graph constructed in step 2 so that ChatGPT can understand the context. Rearrange the input to ensure that key information such as the question statement and the student's answer content is included.

[0144] ② Semantic information extraction

[0145] Next, the text content of the reconstructed response image is passed as input to ChatGPT. Through this process, ChatGPT utilizes its powerful natural language understanding ability to extract semantic information, specifically including:

[0146] Identify the key knowledge points mentioned in the answer;

[0147] Evaluate the rationality and accuracy of the student's answer based on the context of the question;

[0148] Assign weights to each knowledge concept being evaluated to reflect its importance in the answer.

[0149] ③ Embedding integration

[0150] After extracting the semantic information of the key information from ChatGPT, these semantic information are embedded into the feature vectors of the response image content (i.e., integrated with the sentence encoding results of BERT). The semantic embedding generated by ChatGPT is concatenated with the corresponding sentence representation of BERT, and these two representations are integrated together through a simple linear layer to finally form a more effective sentence embedding representation, forming a new vector that integrates rich semantic information.

[0151] ④ Model fine-tuning

[0152] Use the sentence features containing ChatGPT output for model fine-tuning to improve performance on specific grading tasks. Through the method of gradient descent, train on the new dataset and optimize the model parameters so that it can better capture the features of diverse student response content.

[0153] Through the above steps, the powerful language understanding ability of ChatGPT is combined with the sentence encoding ability of BERT, optimizing the semantic representation of the student response content in the automatic test paper grading system.

[0154] Step 3-3: Extract features from the non-text content of the corrected response image to obtain the feature vectors of the non-text content of the response image.

[0155] In this embodiment, the feature extraction of the non-text content is specifically reflected in the extraction of image features (ImageFeature Extraction) (if there are image elements in the test paper):

[0156] For non-text elements such as charts, graphs, or handwritten notes in the test papers, we use the pre-trained image recognition model ResNet (Residual Network) to extract visual features. The image data is input into the ResNet model, and through multiple convolutional layers and pooling layers, a series of high-dimensional feature vectors are extracted. The feature map corresponding to each image, usually the output of the last layer, is obtained as the visual representation of the image.

[0157] Step 3-4: Use the generative artificial intelligence large model to extract key information from the non-text content of the corrected answer image and embed it into the feature vector of the non-text content of the answer image to form a new feature vector of the non-text content of the answer image.

[0158] In this embodiment, the non-text content is specifically embodied as the image elements contained in the answer image.

[0159] To better understand the semantic information in the image, we use ChatGPT to perform semantic analysis on the image features, such as identifying the objects or symbols in the image. Through its powerful language understanding ability, ChatGPT can associate the image features with the text information, thereby better understanding the semantic information of the image. The process of using ChatGPT to optimize the visual features extracted by ResNet is as follows:

[0160] (1) Feature reconstruction

[0161] Since relying solely on visual features may not provide enough context information, we combine the image features with the answer text so that ChatGPT can fully understand the relationship between the image content and the student's answer. Therefore, the image features are reconstructed by integrating relevant text information, such as the description corresponding to the image, the question requirements, etc., to form a richer input format.

[0162] (2) Semantic information enhancement

[0163] The content of the answer image after reconstructing the image features is passed to ChatGPT. Through ChatGPT's language generation and understanding ability, high-level semantic descriptions are generated based on the image features and text information, and the improved semantic information is extracted. This helps to reveal the meaning of the key elements in the image, identify the features related to the question in the image, and evaluate the relevance and rationality of these features in the student's answer.

[0164] (3) Embedding integration

[0165] After extracting the improved semantic information from ChatGPT, the enhanced visual features are integrated with the features originally extracted by ResNet to utilize the advantages of both: the semantic embedding generated by ChatGPT is concatenated with the visual features extracted by ResNet to form a richer representation vector.

[0166] Linear transformation and normalization are performed to ensure that the dimensions of the integrated features are consistent, preparing for subsequent model training.

[0167] Step 4, Fourth part: Score prediction and output.

[0168] In this embodiment, in order to improve the accuracy and effectiveness of the generated scores, a score generation module combining collaborative active learning and ChatGPT is designed, including: score prediction based on text similarity, score prediction based on ChatGPT, consistency detection, and score correction.

[0169] Step 4-1: Calculate the cosine similarity score for the feature vector of the answer image content.

[0170] In this embodiment, the feature vector of the answer image content specifically includes: the standard answer text vector and the answer text vector.

[0171] In this embodiment, calculating the cosine similarity score for the feature vector of the answer image content specifically means calculating the cosine similarity between the standard answer text vector and the answer text vector to obtain the score prediction based on text similarity.

[0172] That is, after feature extraction of the student's answer and the standard answer, the features of the student's answer are compared with the features of the reference answer to measure the similarity between the two.

[0173] In this embodiment, we obtain the large model text similarity comparison and the large model knowledge point matching rate of this process by calculating the cosine similarity. The specific steps are as follows:

[0174] (1) Calculate the cosine similarity between the Embedding vector A of the standard answer text and the Embedding vector B of the answer text. The calculation formula for cosine similarity is:

[0175] (4)

[0176] Among them, A and B are the vector representations of the reference answer and the candidate's answer text respectively. And ||A|| and ||B|| are the Euclidean norms of vectors A and B, that is, the lengths of the vectors. The calculation formula for the Euclidean norm is as follows

[0177] (5)

[0178] where \(i\) is the subscript index variable, and \(n\) represents the length of vector \(A\), which are the respective components in vector \(A\).

[0179] After the calculation, a cosine similarity value will be obtained. The value range of cosine similarity is between \([-1, 1]\), where \(1\) indicates that the two vectors are exactly the same, that is, the student's answer is exactly the same as the standard answer; \(0\) indicates that the two vectors are orthogonal (i.e., not related), and \(-1\) indicates that the two vectors are exactly opposite, that is, the student's answer is not related to the standard answer (off-topic).

[0180] In this embodiment, considering that a higher cosine similarity score indicates that the candidate's answer is closer to the reference answer in terms of text content, the score based on cosine similarity is defined as follows:

[0181] If \(\cosine_similarity(A,B)>0\), then \(Score\_Cos = Total\_Score\times\cosine\_simila\), otherwise: \(Score\_Cos = 0\).

[0182] Step 4 - 2: Also obtain a predicted score for the feature vector of the answer image content through the generative artificial intelligence large model.

[0183] The score prediction based on ChatGPT is as follows:

[0184] ① Collect the feature vectors of the student answers and the feature vectors of the standard answers, communicate using the API interface provided by ChatGPT and send them to the ChatGPT server through an HTTPS request. Utilize ChatGPT's language understanding ability to analyze the relationship between them, and generate a predicted score result based on the context, existing language patterns, and knowledge related to this specific task (the knowledge graph constructed in step 2). This predicted score result is a specific numerical value between \(0 - 10\).

[0185] ② Supervised fine - tuning

[0186] Transfer the actual scores of a large number of student answer questions to ChatGPT, guide ChatGPT to perform self - supervised fine - tuning training, and adjust its parameters according to the labeled data (answer feature vectors, standard answer feature vectors, true scores) to obtain the trained ChatGPT score generation model.

[0187] ③ Test sample score generation

[0188] Input the embedded representation of the answer questions of the student sample to be tested and the corresponding standard answer representation into the trained ChatGPT score generation model to obtain the predicted score \(Score\_GPT\).

[0189] Step 4-3: When there is a significant gap between the cosine similarity score and the prediction score of the generative AI large model, introduce a collaborative active learning mechanism. Through the query method of uncertainty sampling, evaluate the confidence of the candidate's answer. When the confidence is higher than the given threshold, output the prediction score.

[0190] That is, consistency detection and score correction, and the specific steps are as follows:

[0191] In the case where there is a significant disagreement between the cosine similarity score and the prediction score of the generative AI large model (at this time, the confidence of the predicted score is low), introduce an active learning mechanism to further optimize the generation process of the scoring decision and improve the accuracy and reliability of scoring. For example, effectively judge the situation where "the student's answer content is a lot but the score is very low".

[0192] The following are the steps for generating the collaborative active learning scoring decision:

[0193] (1) Confidence evaluation

[0194] For the query strategy framework of active learning, we choose uncertainty sampling query (Uncertainty Sampling). The query method of uncertainty sampling is to extract the samples that are difficult to distinguish in the model (or there are significant differences in the prediction results obtained by different prediction methods) and provide them to business experts or annotators for annotation, so as to achieve the ability to improve the algorithm effect at a relatively fast speed. The key to the uncertainty sampling method is how to describe the uncertainty of the samples or data. The present invention selects the idea of whether the confidence meets the given threshold, that is, first calculate the confidence of the prediction score. If the confidence meets the given threshold (assumed to be 0.8), it is deterministic sampling, otherwise, uncertainty sampling will be used for the subsequent active learning module. The confidence of the scoring result can be calculated using the prediction scores obtained by two different methods, and the specific scheme is as follows:

[0195] Combine the cosine similarity score and the ChatGPT prediction score to evaluate the confidence of the score, which is represented by the symbol C here:

[0196] (6)

[0197] Among them, the Abs function represents taking the absolute value, Score_Cos is the cosine similarity score, Score_GPT is the prediction score of the generative large model ChatGPT, and Toal_Score is the full score value of the corresponding test question. The value range of the confidence C is [0,1]. The larger the value, the more consistent the scores generated by the two schemes and the more reliable the score. Then set the threshold T. If the confidence is not lower than the threshold T, calculate the average score under the two schemes for output. Otherwise, enter the active learning process for score correction.

[0198] (2)Score correction

[0199] ① Construction of active learning dataset:

[0200] For the test questions with answers below the confidence threshold T, invite human experts to make judgments and give expert scores. Further, use the expert scores as the true labels of the test questions, and construct a new dataset together with the answer content representation vectors.

[0201] ② Model update and iteration:

[0202] Use the active learning dataset to optimize the ChatGPT model, improve its ability to identify and evaluate low-confidence answers, and further determine whether the model accuracy meets the requirements. The model accuracy can be obtained by comparing the prediction of the active learning model with the expert scores. If not, still select the test question samples with predicted scores below the confidence threshold and put them back into the model for further training until the model accuracy meets the requirements.

[0203] The present invention proposes a full-cycle ChatGPT-driven automatic test paper grading method, which has the following advantages compared with the traditional test paper grading method:

[0204] Automated and efficient grading: The present invention applies the ChatGPT large model to multiple key links such as test paper content collection, test paper feature extraction, text similarity comparison, and score generation, realizing the automation of test paper grading, greatly improving the grading efficiency, and reducing the labor cost.

[0205] Improved accuracy and reliability: The powerful natural language processing ability of the ChatGPT large model can accurately understand the text content and perform semantic analysis, thus improving the accuracy and reliability of grading.

[0206] Strong interpretability: The present invention calculates the cosine similarity and the predicted score of ChatGPT, and performs weighted averaging in combination with the weight factor to obtain the final score, making the grading result more transparent and interpretable, which helps students and teachers understand and accept the grading result.

[0207] Overcoming the limitations of traditional methods: The present invention overcomes the problems of the traditional test paper grading method highly relying on manual feature extraction or data collection, and the deficiency of the self-grading system in explaining the grading result, thus realizing more efficient, accurate and interpretable test paper grading.

[0208] Embodiment 2

[0209] The purpose of this embodiment is to provide an automatic test paper grading system based on generative artificial intelligence, including:

[0210] The test paper acquisition module is configured to: acquire an image of a paper answer sheet and perform preprocessing to obtain a processed answer image;

[0211] The content recognition module is configured to: perform content recognition on the processed answer image to obtain the content of the answer image; perform anomaly detection on the content of the answer image through a generative artificial intelligence large model, and after obtaining the detection result, correct the content of the answer image to obtain the corrected content of the answer image;

[0212] The feature extraction module is configured to: extract features from the corrected content of the answer image to obtain a feature vector of the content of the answer image; use a generative artificial intelligence large model to extract key information from the corrected content of the answer image and embed it into the feature vector of the content of the answer image to form a new feature vector of the content of the answer image;

[0213] The prediction scoring module is configured to: calculate the cosine similarity score for the feature vector of the content of the answer image; at the same time, also obtain a prediction score for the feature vector of the content of the answer image through a generative artificial intelligence large model;

[0214] The confidence evaluation module is configured to: when there is a serious gap between the cosine similarity score and the prediction score of the generative artificial intelligence large model, introduce a collaborative active learning mechanism, and evaluate the confidence of the candidate's answer through an uncertainty sampling query method; when the confidence is higher than a given threshold, output the prediction score.

[0215] Embodiment III

[0216] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above method are implemented.

[0217] Embodiment IV

[0218] The purpose of this embodiment is to provide a computer-readable storage medium.

[0219] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are executed.

[0220] Embodiment V

[0221] The purpose of this embodiment is to provide a computer program product containing instructions, which when running on a computer, enables the computer to execute the methods and functions involved in any one of the above embodiments.

[0222] The steps involved in the devices of the above embodiments correspond to those of the first method embodiment, and for the specific implementation, reference may be made to the relevant description part of the first embodiment. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to execute any method in the present invention.

[0223] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device for execution by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0224] Although the specific implementation modes of the present invention have been described above in conjunction with the accompanying drawings, this is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.

Claims

1. A method for automatic examination paper review based on generative artificial intelligence, characterized in that: include: Acquire an image of a paper answer sheet, and perform preprocessing to obtain a processed answer image; Performing content recognition on the processed answer image to obtain the content of the answer image; The generative artificial intelligence model is used to detect abnormalities in the content of the answer image, and after obtaining the detection results, the content of the answer image is corrected to obtain the corrected answer image content; Perform feature extraction on the content of the corrected answer image to obtain a feature vector of the answer image content; use the generative artificial intelligence big model to extract key information from the content of the corrected answer image and embed it into the feature vector of the answer image content to form a new feature vector of the answer image content; The cosine similarity score is calculated for the feature vector of the answer image content; at the same time, the feature vector of the answer image content is also predicted and scored through the generative artificial intelligence model; When there is a serious gap between the cosine similarity score and the predicted score of the generative artificial intelligence large model, a collaborative active learning mechanism is introduced to evaluate the confidence of the candidate's answer through the query method of uncertainty sampling; when the confidence is higher than the given threshold, the predicted score is output.

2. The method for automatic examination paper review based on generative artificial intelligence as claimed in claim 1, characterized in that: The processed answer image content includes mathematical formulas, text and sub-images; The performing content recognition on the processed answer image to obtain the content of the answer image specifically includes: For the mathematical formulas in the processed answer images, the BPD model is used to identify the mathematical formulas; For the text in the processed answer image, use the SideNet model to perform OCR recognition and convert the text in the answer image into text data; For the sub-images in the processed answer image, the image elements in the test paper image are segmented using the image segmentation method to extract image feature representation; The method of performing abnormal detection on the content of the answer image by using the generative artificial intelligence large model, and correcting the content of the answer image after obtaining the detection result to obtain the corrected answer image content, specifically includes: For the content of the answer image, use the confidence scoring technology provided by the OCR software to identify low-confidence text; For texts with low confidence levels that are recognized, the ChatGPT large model is used to detect abnormal words or phrases in the OCR output, and further manual review or automatic correction is performed.

3. The method for automatic examination paper review based on generative artificial intelligence as claimed in claim 1, characterized in that: The corrected answer image content is divided into text and non-text elements; The feature extraction is performed on the corrected answer image content to obtain a feature vector of the answer image content. For the text content, the Word2Vec model is specifically used to embed the text, capture the semantic relationship features between text words, and convert them into a vector representation of fixed dimension; then the pre-trained BERT model is used to sentence encode the text of the answer image to obtain sentence-level semantic relationship features, and convert them into a vector of fixed dimension; For non-text content, a pre-trained image recognition model ResNet is used to extract visual features.

4. The method for automatic examination paper review based on generative artificial intelligence as claimed in claim 1, characterized in that: The method of using the generative artificial intelligence big model to extract key information from the content of the corrected answer image and embed it into the feature vector of the answer image content to form a new feature vector of the answer image content specifically includes: Reconstructing the content of the answer image as input, converting the content of the answer image into a more structured format, and obtaining the reconstructed content of the answer image; The content of the reconstructed answer image is input into ChatGPT, and the semantic information of the key information is extracted through ChatGPT. The semantic information is embedded into the feature vector of the answer image content to form a new feature vector of the answer image content.

5. The method for automatic examination paper review based on generative artificial intelligence as claimed in claim 1, characterized in that: The feature vector of the answer image content includes: a standard answer text vector and an answer text vector; The cosine similarity score is calculated for the feature vector of the answer image content, specifically, the cosine similarity between the standard answer text vector and the answer text vector is calculated to obtain a score prediction based on text similarity; The feature vector of the answer image content is also predicted and scored through the generative artificial intelligence model. Specifically, the standard answer text vector and the answer text vector are communicated using the API interface provided by ChatGPT and sent to the ChatGPT server via an HTTPS request to generate a prediction score result.

6. The method for automatic examination paper review based on generative artificial intelligence as claimed in claim 1, characterized in that: When there is a serious gap between the cosine similarity score and the predicted score of the generative artificial intelligence large model, a collaborative active learning mechanism is introduced to evaluate the confidence of the examinee's answer through the query method of uncertainty sampling, which is specifically expressed as follows: Among them, the Abs function represents the absolute value, Score_Cos is the cosine similarity score, Score_GPT is the predicted score of the generative artificial intelligence model ChatGPT, and Toal_Score is the full score of the corresponding test question; the confidence C value range is [0,1]; Then set a threshold T. If the confidence level is not lower than the threshold T, calculate the average score of the two prediction scores and output it. Otherwise, enter the active learning process to correct the score.

7. An automatic examination paper review system based on generative artificial intelligence, characterized in that: include: The test paper acquisition module is configured to: acquire an image of a paper answer paper, and perform preprocessing to obtain a processed answer image; The content recognition module is configured to: perform content recognition on the processed answer image to obtain the content of the answer image; perform abnormality detection on the content of the answer image through the generative artificial intelligence large model, and correct the content of the answer image after obtaining the detection result to obtain the corrected answer image content; The feature extraction module is configured to: extract features from the content of the corrected answer image to obtain a feature vector of the answer image content; extract key information from the content of the corrected answer image using a generative artificial intelligence large model, and embed it into the feature vector of the answer image content to form a new feature vector of the answer image content; The prediction and scoring module is configured to: calculate the cosine similarity score of the feature vector of the answer image content; at the same time, the feature vector of the answer image content is also predicted and scored through the generative artificial intelligence big model; The confidence assessment module is configured as follows: when there is a serious gap between the cosine similarity score and the predicted score of the generative artificial intelligence large model, a collaborative active learning mechanism is introduced to evaluate the confidence of the candidate's answer through the query method of uncertainty sampling; when the confidence is higher than a given threshold, the predicted score is output.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are performed.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Intelligent paper marking implementation method and system based on deep learning and computer program

    CN110110585A

  • Hybrid expert visual question-answering method and system based on strong visual semantics

    CN118070816A