Electronic paper marking method and device based on retrieval enhancement generation
Through the improvement of the intelligent marking system, text similarity and semantic comparison technology are used to solve the problems of existing systems in area division, handwritten text recognition and grade evaluation, and improve the accuracy and fairness of electronic marking.
Patent Information
- Application Number
- CN202411977080.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-09
AI Technical Summary
The existing intelligent marking system has major problems in the area division, handwritten text recognition and grade evaluation modules, resulting in low accuracy of electronic marking.
By obtaining the image of the test paper to be reviewed and the preset reference answer data, the image is converted into answer text data, the test area is determined by using text similarity comparison, and the objective and subjective questions scores are determined based on character comparison and semantic comparison, and the total score of the test paper is finally calculated.
It improves the accuracy of electronic marking, ensures objective and fair scoring, can better adapt to different types of test papers, and enhances the robustness of the system.
Smart Images

Figure CN119964186A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent paper marking, and in particular to an electronic paper marking method based on retrieval enhancement generation, an electronic paper marking device based on retrieval enhancement generation, an electronic device and a storage medium. Background Art
[0002] In the current related technologies, the intelligent marking system usually uses a pre-processing module to perform operations such as denoising, removing shadows, and tilt correction on the image obtained by scanning the test paper. The area division module divides the test paper into regions according to the fixed layout distribution of multiple-choice questions, fill-in-the-blank questions, and essay questions in the test paper. The text recognition module recognizes the test paper questions and the examinee's handwritten answers. The score evaluation module matches the examinee's answer results with the test paper answer text to determine whether the answer is correct, and accumulates the score of each question in the test paper to obtain the overall test paper score. The above-mentioned marking system method has simple, efficient, and relatively accurate results on certain specific types of test papers (such as test papers containing only multiple-choice questions and fill-in-the-blank questions), but this method still has major problems in the area division part, handwritten text recognition, and score evaluation module, resulting in low accuracy of electronic marking. Summary of the invention
[0003] In view of the above problems, embodiments of the present invention are proposed to provide an electronic examination paper method, an electronic examination paper device, an electronic device and a storage medium that overcome the above problems or at least partially solve the above problems.
[0004] In order to solve the above problems, in a first aspect of the present invention, an embodiment of the present invention discloses an electronic examination paper marking method, comprising:
[0005] Obtain the test paper image to be marked and the preset reference answer data;
[0006] Converting the test paper image to be marked into answer text data;
[0007] Comparing the text similarity between the answer text data and the preset reference answer data to determine the test question area; the test question area includes the objective question area and the subjective question area;
[0008] Performing a character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine the objective question score;
[0009] Performing retrieval enhancement and generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data;
[0010] Performing semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score;
[0011] The total score of the test paper is obtained by combining the objective question score and the subjective question score.
[0012] Optionally, it also includes:
[0013] Performing a character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine a first wrong question set;
[0014] Subjective question answer data semantically compares the answer text data corresponding to the subjective question area with the subjective question answer data to determine a second wrong question set;
[0015] Identify the question types of the first set of wrong questions and the question types of the second set of wrong questions, and determine wrong question feature data.
[0016] Optionally, it also includes:
[0017] Visualize the wrong question feature data.
[0018] Optionally, the step of converting the test paper image to be marked into answer text data includes:
[0019] Inputting the test paper image to be marked into a preset image-text multimodal model, wherein the preset image-text multimodal model is used to encode the test paper image to be marked and generate a first text feature vector;
[0020] The text feature vector is determined to be answer text data.
[0021] Optionally, the step of comparing the answer text data with the preset reference answer data for text similarity to determine the test question area includes:
[0022] Segmenting the preset reference answer data by using the preset graphic-text multimodal model to determine a second text feature vector;
[0023] Performing a dot product of the second text feature vector and the first text feature vector to determine a similarity matrix;
[0024] Determine the test question area based on the similarity matrix.
[0025] Optionally, the step of performing semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score includes:
[0026] Performing semantic recognition on the subjective question answer data through the preset graphic-text multimodal model to determine target test point data;
[0027] Performing semantic recognition on the answer text data corresponding to the subjective question area through the preset graphic-text multimodal model to determine the answer test point data;
[0028] The target test point data is compared with the answer test point data to determine the subjective question score.
[0029] Optionally, the preset image-text multimodal model is trained in the following manner:
[0030] Obtain the test paper multimodal training dataset and initial model;
[0031] The initial model is fine-tuned based on the test paper multimodal training data set to generate the preset graphic and text multimodal model.
[0032] In a second aspect of the present invention, an embodiment of the present invention discloses an electronic marking device, comprising:
[0033] An acquisition module is used to acquire the image of the test paper to be marked and the preset reference answer data;
[0034] A conversion module, used for converting the test paper image to be marked into answer text data;
[0035] A first comparison module is used to compare the text similarity between the answer text data and the preset reference answer data to determine the test question area; the test question area includes the objective question area and the subjective question area;
[0036] A second comparison module is used to compare the answer text data corresponding to the objective question area with the preset reference answer data corresponding to the objective question area to determine the objective question score;
[0037] A retrieval module, used to perform retrieval enhancement and generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data;
[0038] A third comparison module is used to perform semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score;
[0039] A combining module is used to combine the objective question scores and the subjective question scores to obtain a total score of the test paper.
[0040] According to a third aspect of the present invention, an embodiment of the present invention discloses an electronic device, comprising a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of the electronic marking method based on retrieval-enhanced generation as described above.
[0041] In a fourth aspect of the present invention, an embodiment of the present invention discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the electronic marking method based on retrieval-enhanced generation as described above are implemented.
[0042] The embodiments of the present invention include the following advantages:
[0043] The embodiment of the present invention obtains an image of a test paper to be marked and preset reference answer data; converts the image of the test paper to be marked into answer text data; performs text similarity comparison between the answer text data and the preset reference answer data to determine the test question area; the test question area includes an objective question area and a subjective question area; performs character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine the objective question score; performs retrieval enhancement generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data; performs semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score; obtains the total score of the test paper by combining the objective question score and the subjective question score; accurately extracts and identifies the test questions and answers in the test paper, performs different scoring methods based on different question types, directly determines the score of the objective question by judging whether the characters are consistent, and performs semantic analysis and judgment on the answer text data and the preset reference answer data to determine the score for the subjective test questions, and scores the test points based on the fit of the test points to ensure objective and fair scoring, thereby improving the accuracy of electronic marking. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of the steps of an embodiment of an electronic examination paper marking method based on retrieval enhancement generation of the present invention;
[0045] Figure 2 It is a flowchart of the steps of an embodiment of an electronic examination paper marking method based on retrieval enhancement generation of the present invention;
[0046] Figure 3 It is a schematic diagram of character recognition of an embodiment of an electronic examination paper marking method based on retrieval enhancement generation of the present invention;
[0047] Figure 4 It is a semantic recognition schematic diagram of an embodiment of an electronic examination paper marking method based on retrieval enhancement generation of the present invention;
[0048] Figure 5 It is a flowchart of an example of an electronic examination paper marking method based on retrieval enhancement generation of the present invention;
[0049] Figure 6It is a structural block diagram of an embodiment of an electronic examination paper marking device based on retrieval enhancement generation of the present invention;
[0050] Figure 7 is a structural block diagram of an electronic device provided by an embodiment of the present invention;
[0051] Figure 8 It is a structural block diagram of a storage medium provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0052] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0053] Intelligent marking is a system that uses artificial intelligence technology to automatically mark and analyze test papers. In traditional education, teachers need to spend a lot of time and energy marking test papers, which compresses their rest and preparation time, resulting in a significant reduction in teaching efficiency; and it is difficult for teachers to remain highly objective and consistent in the process of marking test papers without being affected by personal emotional bias or fatigue. However, the intelligent system can quickly, accurately and objectively score and provide feedback on a large number of test papers, ensuring that each candidate receives a fair evaluation, freeing teachers from the heavy task of marking, and conducting detailed analysis and collation of test results to locate students' knowledge gaps and weaknesses, so as to help teachers develop personalized teaching plans for different students.
[0054] At present, the intelligent marking system in the industry usually performs denoising, shadow removal, tilt correction and other operations on the image obtained by scanning the test paper through the preprocessing module, and the region division module divides the test paper into regions according to the fixed layout distribution of multiple-choice questions, fill-in-the-blank questions, and essay questions in the test paper. The text recognition module recognizes the test paper title and the examinee's handwritten answer. The score evaluation module matches the examinee's answer result with the test paper answer text to determine whether the answer is correct, and accumulates the score of each question in the test paper to obtain the overall test paper score. The above-mentioned marking system method has simple, efficient, and relatively accurate results on certain specific categories of test papers (such as test papers containing only multiple-choice questions and fill-in-the-blank questions), but the method still has major problems in the region division part, handwritten text recognition and score evaluation module. First of all, the region division needs to divide the region according to the fixed layout or key title, resulting in the test paper format must meet the template specification requirements, and cannot be well adapted to various layout test papers. Secondly, in the handwritten text recognition part, due to the different writing habits of each person, the fonts vary greatly. In addition, some candidates do not have clear ideas when answering questions and keep revising their answers, resulting in dirty papers, text alterations, and reversed sentence order. This leads to a large number of recognition errors and confusion in word order when using traditional OCR (Optical Character Recognition), which causes the algorithm to misunderstand the candidates' answers. Finally, the traditional text pattern matching method is used in the score evaluation module to compare the candidates' answers with the standard answers. This method works well for multiple-choice questions and fill-in-the-blank questions, but it cannot be well understood and matched for question-and-answer questions and answer questions that contain certain subjective judgments, resulting in inaccurate AI (artificial intelligence) grading in this part. In order to solve some of the above-mentioned problems, the present invention proposes an electronic marking method, which uses a multimodal model that has been fine-tuned and trained with a handwritten test paper dataset to accurately extract and identify test questions and examinees' handwritten answers in the test paper, uses retrieval enhancement technology to retrieve locally mounted test paper answer documents to generate benchmark answers, and inputs the examinees' answers and benchmark answers into the model at the same time for analysis and judgment. For objective questions, the answers are directly judged whether they are consistent. For subjective questions, the examinees' answers and benchmark answers are input into the model together for understanding and analysis, and it is judged whether the main idea of the answer is consistent with the benchmark answer. After extracting scoring points and error points, appropriate scores are output, and finally, the examinees' knowledge weaknesses are analyzed and counted according to the overall test paper answer results.
[0055] Reference Figure 1 , shows a flowchart of an embodiment of an electronic examination paper marking method based on retrieval enhancement generation of the present invention, the electronic examination paper marking method may specifically include the following steps:
[0056] Step 101, obtaining the image of the test paper to be marked and the preset reference answer data;
[0057] The image of the test paper to be marked and the preset reference answer data corresponding to the test paper to be marked can be obtained. The preset reference answer data can be manually generated in advance.
[0058] Step 102, converting the test paper image to be marked into answer text data;
[0059] The text in the test paper image to be marked can be recognized, and the test paper image to be marked can be converted into answer text data. The answer text data is the text data in the test paper image to be marked.
[0060] Step 103, comparing the text similarity between the answer text data and the preset reference answer data to determine the test question area; the test question area includes the objective question area and the subjective question area;
[0061] The answer text data is compared with the preset reference answer data for text similarity, the answer content and answer content corresponding to the same question are determined, and the type of question is identified to determine the test area. The test area includes the objective question area and the subjective question area.
[0062] Step 104, comparing the answer text data corresponding to the objective question area with the preset reference answer data corresponding to the objective question area to determine the objective question score;
[0063] The answer text data corresponding to the objective question area can be compared with the preset reference answer data corresponding to the objective question area, that is, in the objective question area, the answer text data and the preset reference answer data are compared to determine whether the characters are consistent and determine the objective question score.
[0064] Step 105, performing retrieval enhancement generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data;
[0065] The search enhancement generation technology can be used to search and analyze the answer data of the subjective questions, determine the scoring points in the subjective questions, and obtain the answer data of the subjective questions.
[0066] Step 106, performing semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score;
[0067] The answer text data corresponding to the subjective question area can be compared with the subjective question answer data by character, that is, in the subjective question area, the answer text data and the subjective question answer data can be semantically compared to determine whether the corresponding answer points are consistent and determine the subjective question score.
[0068] Step 107, combining the objective question score and the subjective question score to obtain the total score of the test paper.
[0069] Add the objective question scores and subjective question scores together to get the total score of the test.
[0070] The embodiment of the present invention obtains an image of a test paper to be marked and preset reference answer data; converts the image of the test paper to be marked into answer text data; performs text similarity comparison between the answer text data and the preset reference answer data to determine the test question area; the test question area includes an objective question area and a subjective question area; performs character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine the objective question score; performs retrieval enhancement generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data; performs semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score; obtains the total score of the test paper by combining the objective question score and the subjective question score; accurately extracts and identifies the test questions and answers in the test paper, performs different scoring methods based on different question types, directly determines the score of the objective question by judging whether the characters are consistent, and performs semantic analysis and judgment on the answer text data and the preset reference answer data to determine the score for the subjective test questions, and scores the test points based on the fit of the test points to ensure objective and fair scoring, thereby improving the accuracy of electronic marking.
[0071] Reference Figure 2 , shows a flowchart of another electronic marking method embodiment of the present invention, the electronic marking method may specifically include the following steps:
[0072] Step 201, obtaining the test paper image to be marked and the preset reference answer data;
[0073] The image of the test paper to be marked and the preset reference answer data can be obtained for marking analysis.
[0074] Step 202, converting the test paper image to be marked into answer text data;
[0075] The text data can be extracted from the test paper image to be marked and converted into answer text data.
[0076] In an optional embodiment of the present invention, the step of converting the test paper image to be marked into answer text data includes: inputting the test paper image to be marked into a preset graphic-text multimodal model, the preset graphic-text multimodal model is used to encode the test paper image to be marked and generate a first text feature vector; determine that the text feature vector is answer text data.
[0077] The image of the test paper to be marked can be input into a preset graphic multimodal model, and the preset graphic multimodal model can encode the image of the test paper to be marked to generate a first text feature vector. The first text feature vector is the text feature vector corresponding to the image of the test paper to be marked. The text feature vector is determined as the answer text data. Figure 3 , the image of the test paper to be marked can be encoded to obtain the answer text data.
[0078] For example, for image recognition, a suitable embedding model (such as the bge-base-zh model) can be selected for subsequent text feature extraction and retrieval.
[0079] The extracted test questions are used as query text and encoded using the embedding model to obtain the first text feature vector query_logit.
[0080] query_logit=Model(query_text)
[0081] Furthermore, the preset graphic-text multimodal model is trained in the following manner: obtaining a test paper multimodal training data set and an initial model; performing fine-tuning model training on the initial model based on the test paper multimodal training data set to generate the preset graphic-text multimodal model.
[0082] For the preset graphic-text multimodal model, the test paper multimodal training data set and the initial model can be obtained, and then the initial model is fine-tuned based on the test paper multimodal training data set. The model that has completed the training and meets the usage requirements is determined as the preset graphic-text multimodal model.
[0083] For example, the model training process may include the following steps:
[0084] 1) Construct a multimodal training dataset of test papers.
[0085] Collect batches of test paper images of examinees, obtain test question information and examinee answers through traditional area segmentation and OCR positioning and recognition extraction methods, and correct the text results extracted by OCR through manual review. Finally, store the image path, test question positioning bbox information, test question recognition text information, examinee answer positioning bbox information, and examinee answer text information in the annotation file json. Construct a multimodal training dataset by corresponding examinee test paper images to the annotation file json one by one.
[0086] 2) Fine-tuning the image-text multimodal model
[0087] The multimodal test paper dataset constructed in 1) is used to fine-tune the model training. In order to maintain the model's own image and text semantic understanding ability and generalization, the model is fine-tuned with reference to the LoRA fine-tuning method to enhance the model's understanding and recognition ability of handwritten test papers.
[0088] i. The multimodal large model used in this article consists of a visual encoder (Image Encoder), a text encoder (Text Encoder), a compression layer (Compression Layer) and a large language model (LLM).
[0089] ii. Insert the LoRA module into the self-attention layer of LLM. The LoRA module reduces the number of model parameters by introducing a low-rank matrix.
[0090] iii. The LoRA fine-tuning method in this paper freezes the main structural parts of the network, including the visual encoder, text encoder, and compression layer, during the model fine-tuning training process, and only fine-tunes some layers of the LLM module that contains the LoRA module. In this way, better model effects can be obtained with limited resource consumption and less training time.
[0091] 3) Model Inference Deployment
[0092] The large model is quantized and compiled and deployed using llama.cpp (large model reasoning framework). During reasoning, the test paper image is sent to the image-text multimodal model to extract the test paper information text, including the test paper question text and the examinee's answer text.
[0093] i. Model quantization and conversion: Use the tools provided in llama.cpp to convert the model weights into a gguf file, and use the llama-quantize tool to perform int8 quantization on the language module to improve the efficiency of the model during inference.
[0094] ii. Reasoning part: Use llama-cli to load the quantized model and output the corresponding parameters, such as the number of topk and the path of the test paper image, as well as the prompt word. The prompt used in this article is "Please extract the question text and the candidate's answer text in the image respectively."
[0095] Step 203, performing text similarity comparison between the answer text data and the preset reference answer data to determine the test question area; the test question area includes the objective question area and the subjective question area;
[0096] The answer text data may be compared with the preset reference answer data for text similarity to determine different test question areas, which may include objective question areas and subjective question areas.
[0097] In an optional embodiment of the present invention, the step of comparing the answer text data with the preset reference answer data for text similarity and determining the test question area includes: segmenting the preset reference answer data through the preset graphic and text multimodal model to determine a second text feature vector; performing a dot product of the second text feature vector and the first text feature vector to determine a similarity matrix; and determining the test question area based on the similarity matrix.
[0098] The preset reference answer data can be segmented according to question type, punctuation marks, etc. by a preset graphic multimodal model to determine the second text feature vector. The second text feature vector is the text feature corresponding to the preset reference answer data. Each second text feature vector can be multiplied by the first text feature vector to construct a similarity matrix; based on the similarity matrix, a pair of text feature vectors with greater similarity is determined, thereby determining the corresponding test question area.
[0099] For example, because different question types are usually divided, such as "I. Single-choice questions" and "II. Fill-in-the-blank questions" appear in the test paper, the locally mounted test answer text is segmented according to the question type number (such as "I, II, III") and segmentation symbols (such as period, exclamation mark, line break, etc.). The segmented text is sent to the embedding model selected in i for encoding features to obtain the answer knowledge vector library data_logit_store, which stores each segmented text feature vector.
[0100] data_logit_store={data_logit i ∶Model(answer_text i )}
[0101] Use the transpose of query_logit and data_logit_store to multiply, that is, dot product query_logit with each data_logit in the library to get the similarity between the two features, thereby constructing the feature similarity matrix of the entire answer text.
[0102] scores = torch.matmul(query logit ,data_logit_score.T)
[0103] The k texts with the highest similarity are selected from the feature similarity matrix scores to locate the question and the corresponding answer text information.
[0104] Step 204, comparing the answer text data corresponding to the objective question area with the preset reference answer data corresponding to the objective question area to determine the objective question score;
[0105] The corresponding answer text data in the objective question area and the preset reference answer data can be compared character by character. When the characters are consistent, the score of the corresponding question is obtained; when they are inconsistent, the corresponding question is not scored, and the objective question score is determined.
[0106] For example, for objective questions such as multiple-choice questions and fill-in-the-blank questions, you can directly compare the retrieved test paper answer text with the candidate's answer text to determine whether the answer is correct. If it is correct, you will be given a score for the question. If it is wrong, no score will be given for the question.
[0107] Step 205, performing retrieval enhancement and generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data;
[0108] The search enhancement generation technology can be used to search and analyze the answer data of the subjective questions, determine the scoring points in the subjective questions, and obtain the answer data of the subjective questions.
[0109] In addition, the preset reference answer data corresponding to the subjective question area can also be uploaded in PDF document format, which can be directly converted into specific text content based on the PDF document format, and then based on the retrieval enhancement and generation processing of the subjective question answer data, the subjective question answer data can be obtained.
[0110] Step 206, performing semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score;
[0111] The corresponding answer text data and subjective question answer data in the subjective question area can be semantically recognized to determine whether the response points involved in the answer text data match the test points in the preset reference answer data. If they match, the corresponding score will be obtained, and if they do not match, no score will be obtained, thereby obtaining the subjective question score.
[0112] In an optional embodiment of the present invention, the step of performing semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score includes: performing semantic recognition on the subjective question answer data through the preset graphic-text multimodal model to determine the target test point data; performing semantic recognition on the answer text data corresponding to the subjective question area through the preset graphic-text multimodal model to determine the answer test point data; comparing the target test point data with the answer test point data to determine the subjective question score.
[0113] The preset text-graphic multimodal model can be used to perform semantic recognition on the preset reference answer data corresponding to the subjective question area, determine the test point content involved in the reference answer, and generate the target test point data. Similarly, the preset text-graphic multimodal model can be used to perform semantic recognition on the answer text data corresponding to the subjective question area, determine the test point content answered in the current answer paper, and generate the answer test point data. Determine the match between the target test point data and the answer test point data, obtain the score corresponding to the matching test point, and thus obtain the subjective question score. For the recognition of the preset reference answer data, please refer to Figure 4 , you can determine the corresponding answer by identifying the specific question.
[0114] For example, for subjective questions such as essay questions or answer questions, it is necessary to understand and analyze the content of the test paper answers and the candidate's answers to determine whether the candidate's answers are consistent with the main idea or which key points they hit. The specific algorithm operation is to use the answer text information retrieved from the test as the known context information, and send it into the multimodal model together with the obtained candidate's answer text as prompt information. The semantic reasoning part obtains the candidate's answer analysis result, determines whether the main idea of the candidate's answer is correct, and determines which scoring points the candidate's answer answers based on the test point scoring points in the benchmark answer. A corresponding score is given for each correct answer to a scoring point, and finally the total score of the question is calculated.
[0115] Step 207, combining the objective question score and the subjective question score to obtain a total test score;
[0116] Add the objective question scores and subjective question scores together to get the total score of the test.
[0117] Step 208, comparing the answer text data corresponding to the objective question area with the preset reference answer data corresponding to the objective question area to determine a first wrong question set;
[0118] In the embodiment of the present invention, the corresponding answer text data in the objective question area and the preset reference answer data can also be compared to determine the wrong questions and obtain the first wrong question set. The first wrong question set is the wrong question set of the objective questions.
[0119] Step 209, the subjective question answer data semantically compares the answer text data corresponding to the subjective question area with the subjective question answer data to determine a second wrong question set;
[0120] The corresponding answer text data and subjective question answer data in the subjective question area can also be compared by characters to determine the wrong questions and obtain a second wrong question set. The second wrong question set is a wrong question set of subjective questions.
[0121] Step 210, identifying the question types of the first set of wrong questions and the question types of the second set of wrong questions, and determining wrong question feature data;
[0122] The question types of the first set of wrong questions and the question types of the second set of wrong questions can be identified, the test points and corresponding features corresponding to these wrong questions can be determined, and wrong question feature data can be generated.
[0123] Step 211, visualizing the wrong question feature data.
[0124] The characteristic data of wrong questions can be visualized so that students or teachers can clearly understand the summary of errors in the test paper. For example, the types of wrong questions and the direction of test points can be counted, and a line graph of weak knowledge points can be drawn based on the statistical results, which is convenient for the examiner to analyze the various aspects of the examinee's ability.
[0125] The embodiment of the present invention obtains an image of a test paper to be marked and preset reference answer data; converts the image of the test paper to be marked into answer text data; performs text similarity comparison between the answer text data and the preset reference answer data to determine the test question area; the test question area includes an objective question area and a subjective question area; performs character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine the objective question score; performs retrieval enhancement generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data; and performs retrieval enhancement generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data. The question text data is semantically compared with the subjective question answer data to determine the subjective question score; the objective question score and the subjective question score are combined to obtain the total score of the test paper; the answer text data corresponding to the objective question area is character-compared with the preset reference answer data corresponding to the objective question area to determine the first wrong question set; the subjective question answer data is semantically compared with the answer text data corresponding to the subjective question area to determine the second wrong question set; the question type of the first wrong question set and the question type of the second wrong question set are identified to determine the wrong question feature data; the wrong question feature data is visualized. By using a multimodal model optimized for handwritten font training to directly understand the test paper image semantically and extract the corresponding text information, many problems caused by traditional visual methods are avoided. Retrieval enhancement generation technology is used to retrieve the standard answer from the locally mounted test paper answer and combine it with the candidate's answer to send it to the multimodal model for reasoning and understanding. For subjective question type questions such as essay questions, the score is scored according to the similarity of the candidate's answer theme and the relevance to the test point to ensure objective and fair scoring. Conduct statistical analysis based on the test points and knowledge points involved in the wrong questions, and visualize the characteristics of the wrong questions, so that the examiners or the candidates can have a deeper understanding of their mastery of the knowledge points.
[0126] In order to make the implementation process of the embodiments of the present invention clear to those skilled in the art, the following reference is made to Figure 5 , use an example to illustrate: it can include four parts: multimodal test paper recognition and comprehension module, answer retrieval, answer analysis and comprehensive evaluation.
[0127] The multimodal visual recognition and understanding module includes: 1) building a multimodal training dataset for test papers; 2) fine-tuning the image-text multimodal model; and 3) model reasoning deployment.
[0128] Answer retrieval includes: 1) Test answer document retrieval.
[0129] Answer analysis includes: 1) Analysis and scoring of answer results.
[0130] Comprehensive evaluation includes: 1) Calculation of total test scores. 2) Analysis of weak points.
[0131] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0132] Reference Figure 6 , shows a structural block diagram of an electronic marking device embodiment of the present invention, and the electronic marking device may specifically include the following modules:
[0133] The acquisition module 601 is used to acquire the image of the test paper to be marked and the preset reference answer data;
[0134] A conversion module 602, used for converting the test paper image to be marked into answer text data;
[0135] The first comparison module 603 is used to compare the text similarity between the answer text data and the preset reference answer data to determine the test question area; the test question area includes the objective question area and the subjective question area;
[0136] The second comparison module 604 is used to compare the answer text data corresponding to the objective question area with the preset reference answer data corresponding to the objective question area to determine the objective question score;
[0137] A retrieval module 605 is used to perform retrieval enhancement and generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data;
[0138] The third comparison module 606 is used to perform semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score;
[0139] The combining module 607 is used to combine the objective question scores and the subjective question scores to obtain the total score of the test paper.
[0140] In an optional embodiment of the present invention, it also includes:
[0141] A fourth comparison module, used for comparing the answer text data corresponding to the objective question area with the preset reference answer data corresponding to the objective question area to determine a first wrong question set;
[0142] A fifth comparison module, for the subjective question answer data, performs semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine a second wrong question set;
[0143] The identification module is used to collectively identify the question types of the first wrong question set and the question types of the second wrong question set, and determine the wrong question feature data.
[0144] In an optional embodiment of the present invention, it also includes:
[0145] A visualization module is used to visualize the wrong question feature data.
[0146] In an optional embodiment of the present invention, the conversion module 602 includes:
[0147] A conversion submodule, used for inputting the test paper image to be marked into a preset image-text multimodal model, wherein the preset image-text multimodal model is used for encoding the test paper image to be marked to generate a first text feature vector;
[0148] The determination submodule is used to determine that the text feature vector is answer text data.
[0149] In an optional embodiment of the present invention, the first comparison module 603 includes:
[0150] A segmentation submodule, used for segmenting the preset reference answer data by using the preset graphic-text multimodal model to determine a second text feature vector;
[0151] A dot multiplication submodule, used for performing dot multiplication of the second text feature vector and the first text feature vector to determine a similarity matrix;
[0152] The region determination submodule is used to determine the test question region based on the similarity matrix.
[0153] In an optional embodiment of the present invention, the third comparison module 605 includes:
[0154] A first semantic analysis submodule is used to perform semantic recognition on the subjective question answer data through the preset graphic and text multimodal model to determine the target test point data;
[0155] A second semantic analysis submodule is used to perform semantic recognition on the answer text data corresponding to the subjective question area through the preset graphic-text multimodal model to determine the answer test point data;
[0156] The comparison submodule is used to compare the target test point data with the answer test point data to determine the subjective question score.
[0157] In an optional embodiment of the present invention, the preset graphic-text multimodal model is trained in the following manner:
[0158] Obtain the test paper multimodal training dataset and initial model;
[0159] The initial model is fine-tuned based on the test paper multimodal training data set to generate the preset graphic and text multimodal model.
[0160] The embodiment of the present invention obtains an image of a test paper to be marked and preset reference answer data; converts the image of the test paper to be marked into answer text data; performs text similarity comparison between the answer text data and the preset reference answer data to determine the test question area; the test question area includes an objective question area and a subjective question area; performs character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine the objective question score; performs retrieval enhancement generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data; performs semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score; obtains the total score of the test paper by combining the objective question score and the subjective question score; accurately extracts and identifies the test questions and answers in the test paper, performs different scoring methods based on different question types, directly determines the score of the objective question by judging whether the characters are consistent, and performs semantic analysis and judgment on the answer text data and the preset reference answer data to determine the score for the subjective test questions, and scores the test points based on the fit of the test points to ensure objective and fair scoring, thereby improving the accuracy of electronic marking.
[0161] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0162] Reference Figure 7 , an embodiment of the present invention further provides an electronic device, including:
[0163] A processor 701 and a storage medium 702, wherein the storage medium 702 stores a computer program executable by the processor 701. When the electronic device is running, the processor 701 executes the computer program to perform the electronic marking method based on retrieval enhancement generation as described in any one of the embodiments of the present invention.
[0164] The electronic examination paper marking method based on retrieval enhancement generation comprises:
[0165] Obtain the test paper image to be marked and the preset reference answer data;
[0166] Converting the test paper image to be marked into answer text data;
[0167] Comparing the text similarity between the answer text data and the preset reference answer data to determine the test question area; the test question area includes the objective question area and the subjective question area;
[0168] Performing a character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine the objective question score;
[0169] Performing retrieval enhancement and generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data;
[0170] Performing semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score;
[0171] The total score of the test paper is obtained by combining the objective question score and the subjective question score.
[0172] Optionally, it also includes:
[0173] Performing a character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine a first wrong question set;
[0174] Subjective question answer data semantically compares the answer text data corresponding to the subjective question area with the subjective question answer data to determine a second wrong question set;
[0175] Identify the question types of the first set of wrong questions and the question types of the second set of wrong questions, and determine wrong question feature data.
[0176] Optionally, it also includes:
[0177] Visualize the wrong question feature data.
[0178] Optionally, the step of converting the test paper image to be marked into answer text data includes:
[0179] Inputting the test paper image to be marked into a preset image-text multimodal model, wherein the preset image-text multimodal model is used to encode the test paper image to be marked and generate a first text feature vector;
[0180] The text feature vector is determined to be answer text data.
[0181] Optionally, the step of comparing the answer text data with the preset reference answer data for text similarity to determine the test question area includes:
[0182] Segmenting the preset reference answer data by using the preset graphic-text multimodal model to determine a second text feature vector;
[0183] Performing a dot product of the second text feature vector and the first text feature vector to determine a similarity matrix;
[0184] Determine the test question area based on the similarity matrix.
[0185] Optionally, the step of performing semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score includes:
[0186] Performing semantic recognition on the subjective question answer data through the preset graphic-text multimodal model to determine target test point data;
[0187] Performing semantic recognition on the answer text data corresponding to the subjective question area through the preset graphic-text multimodal model to determine the answer test point data;
[0188] The target test point data is compared with the answer test point data to determine the subjective question score.
[0189] Optionally, the preset image-text multimodal model is trained in the following manner:
[0190] Obtain the test paper multimodal training dataset and initial model;
[0191] The initial model is fine-tuned based on the test paper multimodal training data set to generate the preset graphic and text multimodal model.
[0192] The embodiment of the present invention obtains an image of a test paper to be marked and preset reference answer data; converts the image of the test paper to be marked into answer text data; performs text similarity comparison between the answer text data and the preset reference answer data to determine the test question area; the test question area includes an objective question area and a subjective question area; performs character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine the objective question score; performs retrieval enhancement generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data; performs semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score; obtains the total score of the test paper by combining the objective question score and the subjective question score; accurately extracts and identifies the test questions and answers in the test paper, performs different scoring methods based on different question types, directly determines the score of the objective question by judging whether the characters are consistent, and performs semantic analysis and judgment on the answer text data and the preset reference answer data to determine the score for the subjective test questions, and scores the test points based on the fit of the test points to ensure objective and fair scoring, thereby improving the accuracy of electronic marking.
[0193] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the processor.
[0194] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0195] Reference Figure 8 The embodiment of the present invention further provides a computer-readable storage medium 801, on which a computer program is stored. When the computer program is executed by a processor, the electronic marking method based on retrieval enhancement generation as described in any one of the embodiments of the present invention is executed.
[0196] The electronic examination paper marking method based on retrieval enhancement generation comprises:
[0197] Obtain the test paper image to be marked and the preset reference answer data;
[0198] Converting the test paper image to be marked into answer text data;
[0199] Comparing the text similarity between the answer text data and the preset reference answer data to determine the test question area; the test question area includes the objective question area and the subjective question area;
[0200] Performing a character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine the objective question score;
[0201] Performing retrieval enhancement and generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data;
[0202] Performing semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score;
[0203] The total score of the test paper is obtained by combining the objective question score and the subjective question score.
[0204] Optionally, it also includes:
[0205] Performing a character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine a first wrong question set;
[0206] Subjective question answer data semantically compares the answer text data corresponding to the subjective question area with the subjective question answer data to determine a second wrong question set;
[0207] Identify the question types of the first set of wrong questions and the question types of the second set of wrong questions, and determine wrong question feature data.
[0208] Optionally, it also includes:
[0209] Visualize the wrong question feature data.
[0210] Optionally, the step of converting the test paper image to be marked into answer text data includes:
[0211] Inputting the test paper image to be marked into a preset image-text multimodal model, wherein the preset image-text multimodal model is used to encode the test paper image to be marked and generate a first text feature vector;
[0212] The text feature vector is determined to be answer text data.
[0213] Optionally, the step of comparing the answer text data with the preset reference answer data for text similarity to determine the test question area includes:
[0214] Segmenting the preset reference answer data by using the preset graphic-text multimodal model to determine a second text feature vector;
[0215] Performing a dot product of the second text feature vector and the first text feature vector to determine a similarity matrix;
[0216] Determine the test question area based on the similarity matrix.
[0217] Optionally, the step of performing semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score includes:
[0218] Performing semantic recognition on the subjective question answer data through the preset graphic-text multimodal model to determine target test point data;
[0219] Performing semantic recognition on the answer text data corresponding to the subjective question area through the preset graphic-text multimodal model to determine the answer test point data;
[0220] The target test point data is compared with the answer test point data to determine the subjective question score.
[0221] Optionally, the preset image-text multimodal model is trained in the following manner:
[0222] Obtain the test paper multimodal training dataset and initial model;
[0223] The initial model is fine-tuned based on the test paper multimodal training data set to generate the preset graphic and text multimodal model.
[0224] The embodiment of the present invention obtains an image of a test paper to be marked and preset reference answer data; converts the image of the test paper to be marked into answer text data; performs text similarity comparison between the answer text data and the preset reference answer data to determine the test question area; the test question area includes an objective question area and a subjective question area; performs character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine the objective question score; performs retrieval enhancement generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data; performs semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score; obtains the total score of the test paper by combining the objective question score and the subjective question score; accurately extracts and identifies the test questions and answers in the test paper, performs different scoring methods based on different question types, directly determines the score of the objective question by judging whether the characters are consistent, and performs semantic analysis and judgment on the answer text data and the preset reference answer data to determine the score for the subjective test questions, and scores the test points based on the fit of the test points to ensure objective and fair scoring, thereby improving the accuracy of electronic marking.
[0225] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0226] It will be appreciated by those skilled in the art that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0227] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0228] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0229] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0230] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0231] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0232] The electronic marking method based on retrieval enhancement generation, the electronic marking device based on retrieval enhancement generation, the electronic device and the storage medium provided by the present invention are introduced in detail above. The principle and implementation mode of the present invention are explained in this article by using specific examples. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation mode and the scope of application. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. An electronic examination paper marking method based on retrieval enhancement generation, characterized in that: include: Obtain the test paper image to be marked and the preset reference answer data; Converting the test paper image to be marked into answer text data; Comparing the text similarity between the answer text data and the preset reference answer data to determine the test question area; the test question area includes the objective question area and the subjective question area; Performing a character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine the objective question score; Performing retrieval enhancement and generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data; Performing semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score; The total score of the test paper is obtained by combining the objective question score and the subjective question score.
2. The method according to claim 1, characterized in that Also includes: Performing a character comparison between the answer text data corresponding to the objective question area and the preset reference answer data corresponding to the objective question area to determine a first wrong question set; Subjective question answer data semantically compares the answer text data corresponding to the subjective question area with the subjective question answer data to determine a second wrong question set; Identify the question types of the first set of wrong questions and the question types of the second set of wrong questions, and determine wrong question feature data.
3. The method according to claim 2, characterized in that Also includes: Visualize the wrong question feature data.
4. The method according to claim 1, characterized in that The step of converting the test paper image to be marked into answer text data comprises: Inputting the test paper image to be marked into a preset image-text multimodal model, wherein the preset image-text multimodal model is used to encode the test paper image to be marked and generate a first text feature vector; The text feature vector is determined to be answer text data.
5. The method according to claim 4, characterized in that The step of comparing the text similarity between the answer text data and the preset reference answer data to determine the test question area includes: Segmenting the preset reference answer data by using the preset graphic-text multimodal model to determine a second text feature vector; Performing a dot product of the second text feature vector and the first text feature vector to determine a similarity matrix; Determine the test question area based on the similarity matrix.
6. The method according to claim 4, characterized in that The step of performing semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score comprises: Performing semantic recognition on the subjective question answer data through the preset graphic-text multimodal model to determine target test point data; Performing semantic recognition on the answer text data corresponding to the subjective question area through the preset graphic-text multimodal model to determine the answer test point data; The target test point data is compared with the answer test point data to determine the subjective question score.
7. The method according to any one of claims 4 to 6, characterized in that: The preset image-text multimodal model is trained in the following way: Obtain the test paper multimodal training dataset and initial model; The initial model is fine-tuned based on the test paper multimodal training data set to generate the preset graphic and text multimodal model.
8. An electronic examination paper marking device based on retrieval enhancement generation, characterized in that: include: An acquisition module is used to acquire the image of the test paper to be marked and the preset reference answer data; A conversion module, used for converting the test paper image to be marked into answer text data; A first comparison module is used to compare the text similarity between the answer text data and the preset reference answer data to determine the test question area; the test question area includes the objective question area and the subjective question area; A second comparison module is used to compare the answer text data corresponding to the objective question area with the preset reference answer data corresponding to the objective question area to determine the objective question score; A retrieval module, used to perform retrieval enhancement and generation processing on the preset reference answer data corresponding to the subjective question area to determine the subjective question answer data; A third comparison module is used to perform semantic comparison between the answer text data corresponding to the subjective question area and the subjective question answer data to determine the subjective question score; A combining module is used to combine the objective question scores and the subjective question scores to obtain a total score of the test paper.
9. An electronic device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of the electronic marking method based on retrieval-enhanced generation are implemented as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the electronic marking method based on retrieval-enhanced generation according to any one of claims 1 to 7 are implemented.