Positioning and recognition method, device, electronic device and storage medium

Through deep learning model and image processing technology, the templates and student answer sheets are positioned and cropped, and combined with the weighted voting algorithm, the error problem of machine marking and filling-in-the-blank identification is solved, and the accuracy of marking and the correctness of unified scores is improved.

CN118968534BActive Publication Date: 2025-06-13SHANDONG NEW KUNPENG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410645406.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-23
Publication Date
2025-06-13
Estimated Expiration
2044-05-23

AI Technical Summary

Technical Problem

In the prior art, there is an error in the identification of fill-in-the-blank questions by machine marking papers, resulting in incorrect statistical scores.

Method used

By inputting the template answer sheet images and the student answer sheet images respectively into the deep learning model, obtaining the target area image and the scoring area image, image processing and cropping, using a weighted voting algorithm to combine the positioning accuracy and search rate, and performing image classification and recognition to improve the accuracy of marking papers.

Benefits of technology

The accuracy of the marking of fill-in-the-blank questions and the accuracy of the unified scores are improved, and the score errors caused by OCR recognition errors or semantic mismatch are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968534B_ABST
    Figure CN118968534B_ABST
Patent Text Reader

Abstract

The present application relates to a positioning and recognition method. The template answer sheet image and the student answer sheet image are respectively input into the first deep learning model to obtain the template answer sheet target area image and the student answer sheet target area image. The obtained target area images are respectively input into the second deep learning model to obtain the respective target scoring area images of the template answer sheet and the student answer sheet. The coordinate set of the student answer sheet target scoring area image is denoted as the first set. The student answer sheet image is processed according to the positioning coordinate information of the template answer sheet image, so that the template answer sheet image and the student answer sheet image are mapped one by one in the same coordinate system. The student answer sheet image is cropped according to the coordinate set of the template answer sheet target scoring area image to obtain the second set. A preset algorithm is used to process the first set and the second set to obtain the target coordinate set. Image classification and recognition are performed according to the target coordinate set to improve the accuracy of machine marking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular, to a positioning and recognition method, device, electronic device, and storage medium. Background Art

[0002] To examine students' learning achievements, examinations are an essential form. In current offline examinations, students answer on paper test papers. After the examination is completed and the students hand in their test papers, teachers need to grade them and give scores. By looking at the scores and the scoring and deduction situations of the students, teachers can understand the students' learning situations at a certain stage.

[0003] With the popularization of artificial intelligence technology, introducing machine grading can reduce the workload of teachers. Fill-in-the-blank questions are a common type of question on test papers. Currently, for grading fill-in-the-blank questions, when introducing machine grading, a fully automatic recognition method is generally adopted. First, OCR (Optical Character Recognition) is used to recognize the students' answers on the test papers, then the students' answers are matched with the standard answers, and finally, automatic scoring is carried out according to the matching results.

[0004] However, this method has the problem that OCR character recognition may be incorrect, or even if the OCR character recognition is correct, but the recognition result is not the same as the standard answer in terms of characters, that is, the characters do not completely match, but the semantics are the same. In this case, machine grading will consider the answer incorrect, but in fact, it is the correct answer, resulting in grading errors and further incorrect statistical scores. Summary of the Invention

[0005] This application provides a positioning and recognition method, device, electronic device, and storage medium to solve the problem of incorrect statistical scores caused by errors in machine grading in the prior art.

[0006] In a first aspect, this application provides a positioning and recognition method, including:

[0007] Inputting the template answer sheet image and the student answer sheet image into a first deep learning model respectively to obtain a template answer sheet target region image and a student answer sheet target region image;

[0008] Inputting the template answer sheet target region image and the student answer sheet target region image into a second deep learning model respectively to obtain a template answer sheet target scoring region image and a student answer sheet target scoring region image, where the coordinate set of the student answer sheet target scoring region image is denoted as the first set;

[0009] Performing image processing on the student answer sheet image according to the positioning coordinate information of the template answer sheet image, so that the template answer sheet image and the student answer sheet image are mapped one-to-one in the same coordinate system;

[0010] Crop the student answer sheet image according to the coordinate set of the target scoring area image of the template answer sheet, and denote the coordinate set of the cropped image as the second set;

[0011] Process the first set and the second set using a preset algorithm to obtain a target coordinate set;

[0012] Perform image classification and recognition according to the target coordinate set.

[0013] In a second aspect, the present application provides a positioning and recognition device, including:

[0014] A first positioning module, configured to input the template answer sheet image and the student answer sheet image into a first deep learning model respectively to obtain a template answer sheet target area image and a student answer sheet target area image;

[0015] A second positioning module, configured to input the template answer sheet target area image and the student answer sheet target area image into a second deep learning model respectively to obtain a template answer sheet target scoring area image and a student answer sheet target scoring area image, wherein the coordinate set of the student answer sheet target scoring area image is denoted as the first set;

[0016] An alignment module, configured to perform image processing on the student answer sheet image according to the positioning coordinate information of the template answer sheet image, so that the template answer sheet image and the student answer sheet image are in one-to-one mapping in the same coordinate system;

[0017] A cropping module, configured to crop the student answer sheet image according to the coordinate set of the template answer sheet target scoring area image, and denote the coordinate set of the cropped image as the second set;

[0018] A fusion calculation module, configured to process the first set and the second set using a preset algorithm to obtain a target coordinate set;

[0019] A classification module, configured to perform image classification and recognition according to the target coordinate set.

[0020] In a third aspect, the present application provides an electronic device, including: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; at least one memory connected to the at least one bus, wherein the processor is configured to execute the positioning and recognition method described in the present application.

[0021] In a fourth aspect, the present application further provides a computer storage medium, storing computer-executable instructions, and the computer-executable instructions are used to execute the positioning and recognition method described in the present application.

[0022] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: In the method provided by the embodiments of the present application, the template answer sheet image and the student answer sheet image are respectively input into the first deep learning model to obtain the template answer sheet target region image and the student answer sheet target region image; the template answer sheet target region image and the student answer sheet target region image are respectively input into the second deep learning model to obtain the template answer sheet target scoring region image and the student answer sheet target scoring region image, wherein the coordinate set of the student answer sheet target scoring region image is denoted as the first set; image processing is performed on the student answer sheet image according to the positioning coordinate information of the template answer sheet image, so that the template answer sheet image and the student answer sheet image are in one-to-one mapping in the same coordinate system; the student answer sheet image is cropped according to the coordinate set of the template answer sheet target scoring region image, and the coordinate set of the cropped image is denoted as the second set; a preset algorithm is used to process the first set and the second set to obtain a target coordinate set; image classification and recognition are performed according to the target coordinate set. Since the model positioning of deep learning is relatively accurate but there are omissions, and there are no omissions in the mapping and positioning of the fixed feature points of the answer sheet on the template answer sheet and the student answer sheet in the same coordinate system, but the positioning may have deviations. The present application combines the two methods and uses a weighted voting algorithm to calculate the comprehensive value, integrating the precision rate of model positioning and the recall rate of mapping positioning, further improving the accuracy of answer sheet positioning. In addition, by identifying the artificial identification number, that is, the identification number made by the teacher when grading the written answers, the situation where machine text semantic recognition may go wrong is avoided, so as to achieve the effect of improving the accuracy of intelligent marking and at the same time improving the accuracy of intelligent score statistics. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] One or more embodiments are illustrated by way of example in the accompanying drawings, and these exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.

[0026] Figure 1Flowchart of a positioning and recognition method provided by an embodiment of the present application;

[0027] Figure 2 Schematic diagram of the fill-in-the-blank scoring area box provided by an embodiment of the present application;

[0028] Figure 3 Schematic diagram of preset feature points on the answer sheet provided by an embodiment of the present application;

[0029] Figure 4 Schematic diagram of the image category of the fill-in-the-blank scoring area provided by an embodiment of the present application;

[0030] Figure 5 Schematic diagram of the answer sheet applicable to text mapping positioning provided by an embodiment of the present application;

[0031] Figure 6 Flowchart of a positioning and recognition device module provided by an embodiment of the present application;

[0032] Figure 7 Schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0034] The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0035] To solve the technical problem of low accuracy in machine marking in the prior art, the present application provides a positioning and recognition method, which can improve the accuracy of marking through image positioning and recognition.

[0036] Figure 1 A positioning and recognition method provided by an embodiment of the present application, the method includes:

[0037] S101. Input the template answer sheet image and the student answer sheet image into the first deep learning model respectively to obtain the template answer sheet target area image and the student answer sheet target area image;

[0038] S102. Input the template answer sheet target area image and the student answer sheet target area image into the second deep learning model respectively to obtain the template answer sheet target scoring area image and the student answer sheet target scoring area image. Among them, denote the coordinate set of the student answer sheet target scoring area image as the first set;

[0039] In the embodiment of the present application, the marking of fill-in-the-blank questions is taken as an example for illustration. First, a scoring area box for fill-in-the-blank questions is set on the answer sheet. As Figure 2 shown, the small square box located behind the horizontal line can also be of other shapes. The present application does not limit the shape of the scoring area box. This scoring area box is used for teachers to score the answers to fill-in-the-blank questions. It can be understood that teachers draw marking numbers here to represent the marking results of fill-in-the-blank questions by teachers, which are subsequently used for machine intelligent recognition. The present application realizes machine marking of fill-in-the-blank questions by locating and recognizing the marking numbers marked by teachers on the answer sheet.

[0040] First, two deep learning location models are used to obtain the position coordinates of this scoring area box. These two deep learning location models are respectively the fill-in-the-blank question area location model and the fill-in-the-blank question scoring area location model. It can be understood that the entire fill-in-the-blank question area is first located on the answer sheet, and then each fill-in-the-blank question scoring area is located through this fill-in-the-blank question area. Exemplarily, in the field of computer vision, the fill-in-the-blank question area and the fill-in-the-blank question scoring area are identified by bbox, and the coordinate values (x, y, w, h) represent the specific positions. In addition, the template answer sheet is preferably an answer sheet on which students have not answered, and the student answer sheet is an answer sheet after students have answered.

[0041] The fill-in-the-blank question area positioning model and the fill-in-the-blank question scoring area positioning model in this embodiment are obtained through model training. Specifically, it includes the collection and processing of the dataset, the annotation of the dataset, and the model training process. For the training of the fill-in-the-blank question area positioning model, exemplarily, first collect 50,000 answer sheets as the training set, which includes blank answer sheets without students' answers (i.e., template answer sheets) and answer sheets after students' answers (i.e., students' answer sheets). The proportion of template answer sheets is maintained at about 10%. The requirements for the collected answer sheets are different styles, multiple sessions, multiple subjects, and different schools, so as to ensure the diversity of data. It should be noted that: different styles mean different styles of answer sheets, which can be understood as collecting answer sheets in different modes. Multiple sessions mean collecting multiple sessions of exams in different classes for the same answer sheet. In addition, a validation set and a test set need to be set. The training set is used to train the model, and the validation set is used to check and output the training effect every once in a while (training cycle) during the training process. The test set is used to conduct an overall test after the training is completed to verify the training effect. Exemplarily, the ratio of the training set, the validation set, and the test set can be set to 8:1:1, or it can also be set to 7:2:1. Then, data annotation is carried out. Using the labelme annotation tool, the fill-in-the-blank question area is marked in the form of a rectangular box, and the category is defined as TIANKONG_RECT. Finally, model training is carried out. Specifically, the yolov8 framework is used to complete the training. Considering the inference performance of the model in the later stage, the smallest model is used, and the preferred resolution is 1280*1280. It should be noted that when training the model, enhancement parameters need to be set to perform enhancement operations on the input images. The input for training is a color image, which will be converted into a grayscale image during internal data enhancement. In addition, the rotation probability is set to 0.2, and the rotation angle range is plus or minus 5 degrees. This is to improve the accuracy of model training. After the model training is completed, a fill-in-the-blank question area positioning model that can be actually applied is obtained. Specifically, when applying this model for fill-in-the-blank question area positioning, the template answer sheet image and the student answer sheet image are processed into grayscale images, and the resolution is adjusted to 1280*1280. Then, they are input into the fill-in-the-blank question area positioning model for prediction and positioning to obtain the position coordinates of the fill-in-the-blank question area of the template answer sheet and the position coordinates of the fill-in-the-blank question area of the student answer sheet.

[0042] For the training of the filling-in-the-blank question scoring area positioning model, the method is basically the same as that of the filling-in-the-blank question area positioning model. The difference is that the training set is the image output by the filling-in-the-blank question area positioning model, and then a suitable filling-in-the-blank question area data set is manually selected. The validation set and the test set are also filling-in-the-blank question area data sets. In addition, when annotating the data, the defined category is TIANKONG_DAFEN_RECT, and when setting the augmentation parameters during the model training process, there is no need to convert the image to grayscale. After the model training is completed, a filling-in-the-blank question scoring area positioning model that can be actually applied is obtained. Specifically, when applying this model to locate the filling-in-the-blank question scoring area, it is not necessary to process the filling-in-the-blank question area images of the template answer sheet and the student answer sheet into grayscale images. Only the resolution needs to be adjusted to 1280*1280, and then it is input into the filling-in-the-blank question scoring area positioning model for prediction and positioning to obtain the position coordinates of the filling-in-the-blank question scoring area of the template answer sheet and the position coordinates of the filling-in-the-blank question scoring area of the student answer sheet. Specifically, the first set is the set of position coordinates of the filling-in-the-blank question scoring area of the student answer sheet.

[0043] S103. Perform image processing on the student answer sheet image according to the positioning coordinate information of the template answer sheet image, so that the template answer sheet image and the student answer sheet image are mapped one-to-one in the same coordinate system;

[0044] In the embodiment of the present application, the positioning coordinate information of the template answer sheet image is the coordinate information of the preset feature points on the template answer sheet, which is used to assist the mapping and alignment positioning of the template answer sheet image and the student answer sheet image in the same coordinate system. It can be understood that in the field of computer vision, based on the preset feature points of the answer sheet, the template answer sheet and the student answer sheet are overlapped and aligned in the same coordinate system. The preset feature points are some fixed feature points or fixed recognition points on the answer sheet, which can be the positioning points, titles, page numbers, two-dimensional codes, barcodes, verification points, text content of a certain area or lines of a certain area of the answer sheet, such as Figure 3 shown, which shows some preset feature points. In the case of normal scanning of the paper blank answer sheet and the answer sheet after the student's answer, these preset feature points exist on both the template answer sheet and the student answer sheet, so they can be used for mapping and positioning. Moreover, this overlapping mapping and positioning in the same coordinate system can detect the target position more comprehensively without omission. There may be a situation where individual preset feature points are unclear or do not exist due to scanning reasons, but it is impossible for multiple preset feature points to be missing. The position coordinates of these preset feature points can be obtained through deep learning model training.

[0045] S104. Crop the student answer sheet image according to the coordinate set of the template answer sheet target scoring area image, and record the coordinate set of the cropped image as the second set;

[0046] In the embodiments of the present application, after the template answer sheet image and the student answer sheet image are mapped and aligned one by one in the same coordinate system, by determining the coordinate set of the fill-in-the-blank scoring area image of the template answer sheet, cropping can be performed on the student answer sheet image, that is, the sub-image of the fill-in-the-blank scoring area on the student answer sheet is obtained by means of image processing, that is, the position coordinate set of the fill-in-the-blank scoring area image of the student answer sheet, that is, the second set.

[0047] S105. Process the first set and the second set by using a preset algorithm to obtain a target coordinate set;

[0048] In the embodiments of the present application, S105 specifically includes the following steps:

[0049] S1051. Traverse the coordinate elements in the second set and perform a rectangle intersection-over-union ratio operation with the coordinate elements in the first set respectively;

[0050] S1052. When the intersection-over-union ratio operation value is greater than a preset value, obtain the position coordinate values of the coordinate elements in the current second set and the first set, and perform a weighted voting algorithm process on the two to obtain a first target coordinate;

[0051] S1053. When the intersection-over-union ratio operation values are all less than or equal to the preset value, determine the position coordinate values of the coordinate elements in the current second set as second target coordinates;

[0052] S1054. Combine the first target coordinates and the second target coordinates to obtain the target coordinate set.

[0053] S1051-S1054 provide a method for fusing the precision rate of model positioning and the recall rate of mapping positioning. Specifically, traverse the coordinate elements in the second set one by one, and perform a rectangle intersection-over-union ratio operation with each coordinate element in the first set respectively. If the rectangle intersection-over-union ratio operation value of the coordinate element and one of the coordinate elements in the first set is greater than 0.5, obtain the position coordinate value of the coordinate element in the second set and the position coordinate value of the coordinate element in the first set respectively, and perform a weighted voting algorithm process on the two. In the embodiments of the present application, for the weighting coefficients for performing the weighted voting algorithm process, the weighting coefficient in the first set is set to be greater than the weighting coefficient in the second set. Exemplarily, the position coordinate value in the first set is (X1, Y1, W1, H1), and the position coordinate value in the second set is (X2, Y2, W2, H2). The position coordinate value after the weighted voting algorithm process can be: 0.7(X1, Y1, W1, H1)+0.3(X2, Y2, W2, H2). The preferred weighting coefficients are 0.7 and 0.3, and can also be 0.6 and 0.4, or 0.8 and 0.2. Determine all the position coordinates after the weighted voting algorithm process as the first target coordinates.

[0054] In another case, if the rectangular intersection and union operation values ​​of the coordinate element and each coordinate element in the first set are all less than or equal to 0.5, it indicates that there is no coordinate value in the first set that is the same or close to the position coordinate value of the coordinate element, that is, the determination of the fill-in-the-blank question scoring area sub-image by the fill-in-the-blank question area positioning model and the fill-in-the-blank question scoring area positioning model is inaccurate, and there is a missed detection or false detection. At this time, the image obtained by mapping positioning and cropping is used as the basis, and its position coordinates are determined as the second target coordinates.

[0055] Combining the above two situations, and so on, until all coordinate elements in the second set have completed the rectangular intersection and union operation. Further, the position coordinates determined in the two situations are combined to obtain the target coordinate set for further image classification and recognition.

[0056] S106: Perform image classification and recognition according to the target coordinate set.

[0057] In the embodiment of the present application, S106 is specifically implemented as follows:

[0058] First, obtain the area image corresponding to the target coordinate set, that is, the fill-in-the-blank question scoring area sub-image. It should be noted that the image includes the identification number manually marked by the teacher.

[0059] Secondly, the sub-image of the fill-in-the-blank question scoring area is input into the classification model to obtain the corresponding image category and confidence.

[0060] Specifically, the classification model can be an EfficientNet model, and the image category can be divided into three categories: valid identification number, invalid identification number and empty (no identification number), such as Figure 4 As shown in the figure, a " / " or "|" in the small square box is a valid identification number, indicating that the student answers the corresponding fill-in-the-blank question correctly and the teacher gives points; there is an "×" in the small square box, or multiple lines are painted out, or the identification number is outside the box, which are invalid identification numbers, indicating that the student answers the corresponding fill-in-the-blank question incorrectly and the teacher does not give points; the small square box is empty, indicating that there are two situations, the teacher gives points or the teacher does not give points. The two situations are designed without identification numbers because when teachers review a large number of fill-in-the-blank questions, if most students answer the questions correctly, the teacher only needs to correct the wrong ones, and the blank indicates that the teacher gives points. If most students answer the questions incorrectly, the teacher only needs to correct the correct ones, and the blank indicates that the teacher does not give points. This can save teachers' workload and is also suitable for good schools and general schools, and school scenarios with different student levels.

[0061] It should be noted that due to the special situation of fill-in-the-blank questions, which are either full marks or zero marks, the score of the corresponding fill-in-the-blank question can be determined by identifying valid identification numbers and invalid identification numbers. In addition, the setting of " / " and "×" in this embodiment also has an advantage. If the teacher wants to modify the score after giving it and finds it incorrect, they only need to add a " / " in the opposite direction or cross it out. Of course, the valid identification number and the invalid identification number can also be other shaped identifications, and the present application does not make specific limitations.

[0062] In addition, for these 3 types of identification numbers, the classification model also outputs the corresponding confidence level, that is, the probability of belonging to this category. Exemplarily, if the confidence level is less than 0.5, this question requires manual intervention for grading. If the confidence level is greater than or equal to 0.5, if the image category is a valid identification number, the fill-in-the-blank question scores; if the image category is an invalid identification number, it can be set to require manual intervention or not. If it is set not to require manual intervention, the fill-in-the-blank question does not score; if the image category is empty, optional items are displayed for manual selection, and the teacher sets the empty to represent full marks or zero marks according to the objective situation of the student.

[0063] The training of the EfficientNet model in this application is basically the same as the method of the aforementioned fill-in-the-blank question area positioning model. The difference is that the training set needs to collect 3 types of sub-images of the fill-in-the-blank question scoring area, and the preferred ratio is 1:1:1. According to the actual collection difficulty, the ratio of valid identification number sub-images: invalid identification number sub-images: sub-images without identification numbers can also be 3:4:3. When performing data annotation, the 3 sub-images are placed in 3 folders respectively. When training the model, the image resolution is 64*64, and the online data augmentation methods are horizontal flipping and vertical flipping. The flipping probability is set to 0.5, the rotation probability is set to 0.2, and the rotation angle range is plus or minus 5 degrees. It should be noted that grayscale images cannot be used because this application needs to utilize the color information of the red line.

[0064] Furthermore, by identifying the identification number, the effect of quickly grading fill-in-the-blank questions can be achieved. For example: if the score of a single fill-in-the-blank question is 1.5 points, a data table can be constructed, with the preset scores being 1.5 and 0, establishing a mapping relationship table between various identification numbers and preset scores, and then counting the number of various identification numbers, and finally calculating the total score of the fill-in-the-blank question. And the setting of the confidence level also further improves the accuracy of machine grading, because when the confidence level is less than 0.5, manual inspection is set to avoid incorrect grading caused by misidentification, which affects the correctness of machine grading.

[0065] In this embodiment, by integrating the precision rate of the fusion model positioning and the recall rate of the mapping positioning, the accuracy and comprehensiveness of the answer sheet positioning are achieved, and the positioning accuracy is further improved. In addition, by identifying the artificial identification number, that is, the identification number made by the teacher when marking the written answers, the situation where the machine text semantic recognition may go wrong is avoided, so as to achieve the effect of improving the accuracy of intelligent marking and also improve the accuracy of intelligent score statistics.

[0066] Furthermore, in order to enable different styles of answer sheets to adopt the mapping positioning method and solve the problem of positioning difficulties caused by the poor image quality of the student answer sheet images (the poor image quality may cause the preset feature points to be incomplete or missing), three different positioning methods are provided in the embodiment of the present application, which are specifically as follows:

[0067] The first one: positioning point positioning

[0068] 1. Perform the first image processing on the student answer sheet image according to the positioning point coordinates of the template answer sheet image, where the first image processing includes image scaling, image translation or image rectification;

[0069] 2. After the first image processing, judge whether to rotate the student answer sheet image according to the identification point coordinates of the template answer sheet image, where the identification points include at least one of page number, check point, title, QR code and bar code.

[0070] In one embodiment, the positioning points of the template answer sheet image are the positioning graphics set on the answer sheet, generally set near the four vertices of the answer sheet, and there are 4 of them. Specifically, any 3 positioning points can be selected, and the coordinate values of these 3 positioning points in the template answer sheet image are obtained. Through the affine transformation algorithm, the student answer sheet image is processed to be exactly the same as the template answer sheet image. This image processing process involves image scaling, image translation, or image rectification of the student answer sheet image, so that the coordinate values of the respective positioning points of the student answer sheet image and the template answer sheet image are the same in the same coordinate system, that is, in a one-to-one mapping and overlapping state. Then, the coordinate values of the page numbers in the template answer sheet image are obtained, and it is judged whether the coordinate values of the page numbers in the student answer sheet image are the same or similar. If so, it means the positioning is correct. If not, the student answer sheet image is rotated 180° to complete the positioning. Since other identification points on the answer sheet, such as: verification points, titles, QR codes, and barcodes, can also be used to judge the front and back and the upright or upside-down state of the answer sheet like the page numbers, in other embodiments of the present application, after positioning using the positioning points, these identification points can also be used to replace the page numbers to further verify whether the answer sheet is reversed or upside-down, or continue to perform superimposed positioning verification, such as: first the positioning points, then the page numbers, then other identification points, or first the positioning points, then other identification points, then the page numbers, so as to improve the positioning accuracy. It should be noted that due to the quality of the student answer sheet image being affected by printing devices, scanning devices, as well as the size, color, paper damage, paper smearing, paper stains, etc. of the answer sheet paper, the image quality may be poor, which interferes with the positioning. Therefore, before the positioning process, it is necessary to preprocess the student answer sheet image, including image enhancement processing and denoising processing, to make the image clearer.

[0071] The second type: line identification positioning

[0072] 1. Judge whether to perform image scaling processing on the student answer sheet image according to the title font size of the template answer sheet image;

[0073] 2. Perform a second image processing on the student answer sheet image according to the identification line coordinates of the template answer sheet image, where the second image processing includes image rectification;

[0074] 3. After the second image processing, perform image translation on the student answer sheet image according to the title coordinates or the coordinates of the center point of the title in the template answer sheet image.

[0075] In one embodiment, if there are no positioning points on the answer sheet, a positioning method using lines and identification points is adopted at this time. Specifically, first, it is determined whether the student answer sheet image needs to be scaled according to the font size of the title area of the template answer sheet image. It can be understood that first, the two answer sheet images need to be processed to be basically the same size in the same coordinate system, so that the position deviation of other identification point elements on the answer sheet is relatively small, which is convenient for quick positioning. Exemplarily, if the sizes of the template answer sheet image and the student answer sheet image are the same, no image scaling processing is required. If they are different, the two images are scaled to be the same size through image scaling processing. After determining that the image sizes are the same, the coordinate values of the identification line in the template answer sheet image are obtained. The identification line is a line at a certain position on the answer sheet. The student answer sheet image is corrected through the coordinate values of the identification line, so that the coordinate values of the respective identification lines of the student answer sheet image and the template answer sheet image in the same coordinate system are the same, that is, in a one-to-one mapping and overlapping state. It should be noted that if the identification line on the answer sheet is not in a standard horizontal and vertical state, the identification line can be corrected to be horizontal and vertical, and then mapped and overlapped for positioning through the coordinate values. Further, the coordinate values of the title in the template answer sheet image or the coordinate values of the center point of the title are obtained, and the title area in the student answer sheet image is positioned through the coordinate values of the title or the center point of the title. This process involves translation operations of the student answer sheet image in the up, down, left, and right directions. Exemplarily, the number of characters selected for the title is 6 or more than 6. In addition, if the title area is not located, the student answer sheet image is rotated 180° and the foregoing steps are tried again. It should be noted that other identification points can also be used to replace the title in this method. For example, a characteristic text area on the answer sheet can be used to determine the image size and judge the front and back and the upright or inverted state of the answer sheet.

[0076] The third type: Feature area positioning

[0077] Perform third image processing on the student answer sheet image according to the identification text coordinates of the template answer sheet image, where the identification text includes at least 3 paragraphs of text, and the connecting lines of the center points of each paragraph of text form a plane, and the third image processing includes image scaling, image translation or image correction.

[0078] In one embodiment, if there are no obvious preset feature points such as positioning points, identification lines, and titles on the answer sheet, such as Figure 5The form of the answer sheet shown. At this time, according to the known principle of plane transformation, 3 feature points are found. Similar to the principle of positioning the positioning points mentioned above, the difference is that these 3 feature points are 3 segments of text. This method is applicable to the test papers with questions and answers combined on one sheet. There are a lot of texts on the answer sheet (which is also the test paper), and these texts can be used as positioning information. Exemplarily, any 3 segments of text are selected, and it is optimal if the number of words is 3 - 5. The area of the triangle formed by the center points of these 3 segments of text should be large, the larger the better, and the positions of the texts on different pages of the answer sheet are different.

[0079] Specifically, obtain the coordinate values of the identification text of the template answer sheet image, that is, the regional coordinate values of 3 segments of preset text. Through the affine transformation algorithm, process the student answer sheet image to be exactly the same as the template answer sheet image. In this image processing process, image scaling, image translation or image rectification are involved for the student answer sheet image, so that the regional coordinate values of these 3 segments of preset text of the student answer sheet image and the template answer sheet image are the same in the same coordinate system, that is, in a one-to-one mapping and coincidence state. It should be noted that for this feature region positioning method, in addition to selecting text features, pictures and grids can also be selected. Pictures and grids also have positioning features and, like text, are easy to identify and distinguish, so they can also be used for positioning.

[0080] Furthermore, other preset feature points can also be set on the template answer sheet in the embodiments of the present application as long as these feature points have positioning functions.

[0081] As Figure 6 shown, the embodiments of the present application provide a positioning and recognition device, including:

[0082] The first positioning module 61 is used to input the template answer sheet image and the student answer sheet image into the first deep learning model respectively to obtain the template answer sheet target region image and the student answer sheet target region image;

[0083] The second positioning module 62 is used to input the template answer sheet target region image and the student answer sheet target region image into the second deep learning model respectively to obtain the template answer sheet target scoring region image and the student answer sheet target scoring region image. Among them, the coordinate set of the student answer sheet target scoring region image is denoted as the first set;

[0084] The alignment module 63 is used to perform image processing on the student answer sheet image according to the positioning coordinate information of the template answer sheet image, so that the template answer sheet image and the student answer sheet image are in one-to-one mapping in the same coordinate system;

[0085] The cropping module 64 is used to crop the student answer sheet image according to the coordinate set of the template answer sheet target scoring region image, and denote the coordinate set of the cropped image as the second set;

[0086] A fusion calculation module 65, configured to process the first set and the second set by using a preset algorithm to obtain a target coordinate set;

[0087] A classification module 66, configured to perform image classification and recognition according to the target coordinate set.

[0088] As Figure 7 shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114.

[0089] The memory 113 is used to store a computer program;

[0090] In an embodiment of the present application, when the processor 111 executes the program stored on the memory 113, it implements the positioning and recognition method provided by any one of the foregoing method embodiments.

[0091] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the positioning and recognition method provided by any one of the foregoing method embodiments.

[0092] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0093] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0094] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless an execution order is explicitly stated. It should also be understood that additional or alternative steps may be used.

[0095] The foregoing are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A positioning and identification method, characterized in that: The method comprises: Input the template answer sheet image and the student answer sheet image into the first deep learning model respectively to obtain the template answer sheet target area image and the student answer sheet target area image; Input the target area image of the template answer sheet and the target area image of the student answer sheet into the second deep learning model respectively, to obtain the target scoring area image of the template answer sheet and the target scoring area image of the student answer sheet, wherein the coordinate set of the target scoring area image of the student answer sheet is recorded as the first set; Performing image processing on the student answer sheet image according to the positioning coordinate information of the template answer sheet image, so that the template answer sheet image and the student answer sheet image are mapped one by one in the same coordinate system; Cropping the student answer sheet image according to the coordinate set of the target scoring area image of the template answer sheet, and recording the coordinate set of the cropped image as a second set; Using a preset algorithm to process the first set and the second set to obtain a target coordinate set; Performing image classification and recognition according to the target coordinate set; The using a preset algorithm to process the first set and the second set to obtain a target coordinate set includes: Traversing the coordinate elements in the second set, and performing rectangular intersection and union operations with the coordinate elements in the first set respectively; When the intersection-and-union ratio operation value is greater than a preset value, the position coordinate values ​​of the coordinate elements in the second set and the first set are obtained, and a weighted voting algorithm is performed on the two to obtain the first target coordinates; When the intersection-and-union ratio operation values ​​are all less than or equal to the preset value, the position coordinate value of the coordinate element currently in the second set is determined as the second target coordinate; Merging the first target coordinates and the second target coordinates to obtain the target coordinate set; The performing image classification and recognition according to the target coordinate set includes: Acquire a region image corresponding to the target coordinate set, wherein the region image includes identification numbers manually marked by a teacher; The regional image is input into a classification model to obtain a corresponding image category and confidence level, wherein the image category is used to score the target area of ​​the answer sheet according to a preset score.

2. The method according to claim 1, characterized in that The first deep learning model is a fill-in-the-blank question area positioning model, and the second deep learning model is a fill-in-the-blank question scoring area positioning model; the template answer sheet target area image and the student answer sheet target area image are respectively the template answer sheet fill-in-the-blank question area image and the student answer sheet fill-in-the-blank question area image; the template answer sheet target scoring area image and the student answer sheet target scoring area image are respectively the template answer sheet fill-in-the-blank question scoring area image and the student answer sheet fill-in-the-blank question scoring area image.

3. The method according to claim 1 or 2, characterized in that: The positioning coordinate information of the template answer sheet image is the coordinate information of the preset feature points on the template answer sheet, which is used to assist the mapping, alignment and positioning of the template answer sheet image and the student answer sheet image in the same coordinate system.

4. The method according to claim 3, characterized in that The performing image processing on the student answer sheet image according to the positioning coordinate information of the template answer sheet image so that the template answer sheet image and the student answer sheet image are mapped one by one in the same coordinate system includes: Performing a first image processing on the student answer sheet image according to the coordinates of the positioning points of the template answer sheet image, wherein the first image processing includes image scaling, image translation or image correction; After the first image processing, it is determined whether to rotate the student answer sheet image according to the coordinates of the identification points of the template answer sheet image, wherein the identification points include at least one of a page number, a check point, a title, a QR code and a bar code.

5. The method according to claim 3, characterized in that: The step of performing image processing on the student answer sheet image according to the positioning coordinate information of the template answer sheet image so that the template answer sheet image and the student answer sheet image are mapped one by one in the same coordinate system further includes: Determining whether to perform image scaling processing on the student answer sheet image according to the title font size of the template answer sheet image; Performing a second image processing on the student answer sheet image according to the coordinates of the identification line of the template answer sheet image, wherein the second image processing includes image deflection correction; After the second image processing, the student answer sheet image is subjected to image translation according to the title coordinates or the title center point coordinates of the template answer sheet image.

6. The method according to claim 3, characterized in that The step of performing image processing on the student answer sheet image according to the positioning coordinate information of the template answer sheet image so that the template answer sheet image and the student answer sheet image are mapped one by one in the same coordinate system further includes: The student answer sheet image is subjected to a third image processing according to the coordinates of the identification text of the template answer sheet image, wherein the identification text includes at least 3 paragraphs of text, and the center points of the paragraphs of text are connected to form a surface, and the third image processing includes image scaling, image translation or image deflection correction.

7. A positioning and identification device, characterized in that: The device comprises: A first positioning module is used to input the template answer sheet image and the student answer sheet image into a first deep learning model respectively to obtain a template answer sheet target area image and a student answer sheet target area image; A second positioning module is used to input the target area image of the template answer sheet and the target area image of the student answer sheet into a second deep learning model respectively, to obtain a target scoring area image of the template answer sheet and a target scoring area image of the student answer sheet, wherein a coordinate set of the target scoring area image of the student answer sheet is recorded as a first set; An alignment module, used for performing image processing on the student answer sheet image according to the positioning coordinate information of the template answer sheet image, so that the template answer sheet image and the student answer sheet image are mapped one by one in the same coordinate system; A cropping module, used for cropping the student answer sheet image according to the coordinate set of the target scoring area image of the template answer sheet, and recording the coordinate set of the cropped image as a second set; A fusion calculation module, used to process the first set and the second set using a preset algorithm to obtain a target coordinate set; A classification module, used for performing image classification and recognition according to the target coordinate set; The using a preset algorithm to process the first set and the second set to obtain a target coordinate set includes: Traversing the coordinate elements in the second set, and performing rectangular intersection and union operations with the coordinate elements in the first set respectively; When the intersection-and-union ratio operation value is greater than a preset value, the position coordinate values ​​of the coordinate elements in the second set and the first set are obtained, and a weighted voting algorithm is performed on the two to obtain the first target coordinates; When the intersection-and-union ratio operation values ​​are all less than or equal to the preset value, the position coordinate value of the coordinate element currently in the second set is determined as the second target coordinate; Merging the first target coordinates and the second target coordinates to obtain the target coordinate set; The performing image classification and recognition according to the target coordinate set includes: Acquire a region image corresponding to the target coordinate set, wherein the region image includes identification numbers manually marked by a teacher; The regional image is input into a classification model to obtain a corresponding image category and confidence level, wherein the image category is used to score the target area of ​​the answer sheet according to a preset score.

8. An electronic device, characterized in that: include: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor coupled to the at least one bus; At least one memory connected to the at least one bus, wherein the processor is configured to execute the method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: Computer executable instructions are stored, and the computer executable instructions are used to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Universal intelligent scoring system and method

    CN110008933A

  • Test paper score unifying system and method

    CN112163529A