Answer sheet recognition method and device, storage medium and electronic equipment
By using the target detection box and relative position relationship in the answer sheet image, the problems of high cost and low accuracy in grading small batches of answer sheets are solved, and accurate grading of answer sheets of different formats is achieved.
Patent Information
- Application Number
- CN202210315723.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2042-03-28
AI Technical Summary
In the existing technology, the cost of correcting small batches of answer sheets is high and the accuracy of image recognition is affected by the quality of the photos and layout changes, resulting in inaccurate answer recognition.
The answer sheet image is detected through the preset target detection frame to obtain the answer sheet number, question and answer information. The target detection model is used to extract multiple target images, and they are accurately paired based on relative position relationships to achieve intelligent correction.
It achieves accurate grading of answer sheets of different formats, improves grading accuracy, and reduces dependence on photo quality and format changes.
Smart Images

Figure CN114708598B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image recognition technology, and in particular to an answer sheet recognition method, device, storage medium, and electronic device. Background Art
[0002] When correcting answer sheets, they can generally be corrected using supporting detection equipment. However, for small batches of answer sheets, the cost of using supporting detection equipment for correction is extremely high. Therefore, at present, for small batches of answer sheets, correction is generally done manually or through image recognition on mobile terminals. Among them, the accuracy of image recognition on mobile terminals is greatly affected by the quality of the photographed images, and due to the poor generalization ability of the image processing algorithm, it will cause misidentification of question positioning and answer content. For example, due to the influence of shooting angles and the various answer sheet formats, it will lead to the inability to accurately identify the answers on the answer sheets. Once the format of the answer sheet changes, it will lead to the inability to accurately identify the filled-in answers. Summary of the Invention
[0003] The purpose of the present disclosure is to provide a method, device, storage medium and electronic device for answer sheet recognition to improve the accuracy of answer sheet correction.
[0004] In a first aspect, an embodiment of the present disclosure provides an answer sheet recognition method, comprising:
[0005] Obtaining a captured image of the answer sheet;
[0006] Detecting the captured image based on a preset target detection frame to obtain a plurality of target images, wherein the target detection frame includes at least an answer sheet number detection frame, a question detection frame, and an answer detection frame;
[0007] Determining the number information of the answer sheet based on the first target image detected by the answer sheet number detection frame;
[0008] Determining the standard answer corresponding to the answer sheet based on the number information;
[0009] determining answer information for each question on the answer sheet based on the second target image detected by the question detection frame and the third target image detected by the answer detection frame;
[0010] The correction result of the answer sheet is obtained according to the answer information and the standard answer.
[0011] In a second aspect, an embodiment of the present disclosure provides an answer sheet recognition device, comprising:
[0012] an acquisition module configured to acquire a photographed image of the answer sheet;
[0013] a detection module configured to detect the captured image based on a preset target detection frame to obtain a plurality of target images, wherein the target detection frame includes at least an answer sheet number detection frame, a question detection frame, and an answer detection frame;
[0014] a number recognition module configured to determine the number information of the answer sheet based on the first target image detected by the answer sheet number detection frame;
[0015] A standard answer determination module is configured to determine the standard answer corresponding to the answer sheet based on the number information;
[0016] an answer recognition module configured to determine answer information for each question on the answer sheet based on the second target image detected by the question detection frame and the third target image detected by the answer detection frame;
[0017] The correction module is configured to obtain the correction result of the answer sheet according to the answer information and the standard answer.
[0018] In a third aspect, an embodiment of the present disclosure provides a non-temporary computer-readable storage medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when executed by a processor.
[0019] In a fourth aspect, an embodiment of the present disclosure provides an electronic device, including:
[0020] a memory having a computer program stored thereon;
[0021] A processor is used to execute the computer program in the memory to implement the steps of the method of the first aspect.
[0022] Based on the above technical solution, the captured image is detected through the answer sheet number detection frame, the question detection frame and the answer detection frame to obtain multiple target images, and based on the first target image detected by the answer sheet number detection frame, the number information of the answer sheet is determined, and based on the number information, the standard answer corresponding to the answer sheet is determined, and based on the second target image detected by the question detection frame and the third target image detected by the answer detection frame, the answer information of each question on the answer sheet is determined, and then the correction result of the answer sheet is obtained based on the answer information and the standard answer. Not only can the intelligent correction of the answer sheet be achieved by capturing the image, but the target image can be obtained by the answer sheet number detection frame, the question detection frame and the answer detection frame, so that the answer sheet can be accurately corrected without being affected by the layout of the answer sheet, the filling method, etc. For example, whether it is a single-column, double-column or other complex layout answer sheet, it can be corrected by the answer sheet recognition method proposed in the embodiment of the present disclosure.
[0023] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following detailed description, they are used to explain the present disclosure but do not constitute a limitation of the present disclosure. In the accompanying drawings:
[0025] Figure 1 The present invention is a flowchart of an answer sheet recognition method provided according to an exemplary embodiment.
[0026] Figure 2 FIG. 4 is a schematic diagram of an object detection frame according to an exemplary embodiment.
[0027] Figure 3 yes Figure 1 A detailed flow chart of step 150 is shown.
[0028] Figure 4 yes Figure 3 A detailed flow chart of step 152 is shown.
[0029] Figure 5 yes Figure 1 A detailed flow chart of step 110 is shown.
[0030] Figure 6 The present invention is a module connection diagram of an answer sheet recognition device provided according to an exemplary embodiment.
[0031] Figure 7 It is a schematic structural diagram of an electronic device provided according to an exemplary embodiment. DETAILED DESCRIPTION
[0032] The following describes the specific embodiments of the present disclosure in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present disclosure and are not intended to limit the present disclosure.
[0033] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the corresponding data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0034] Figure 1 FIG. 1 is a flow chart of a method for identifying an answer sheet according to an exemplary embodiment. Figure 1 As shown, the embodiment of the present disclosure provides an answer sheet recognition method, which can be executed by an electronic device, specifically by an answer sheet recognition device, which can be implemented by software and / or hardware and configured in an electronic device. Figure 1As shown, the answer sheet recognition method may include the following steps.
[0035] In step 110 , a photographed image of the answer sheet is acquired.
[0036] Here, the captured image of the answer sheet can be obtained by the examiner using a camera or mobile device to capture the answer sheet. Of course, in some embodiments, the captured image can be an image of the answer sheet to be graded received via a network. For example, the examiner can capture the answer sheet to be graded using a mobile terminal to obtain the captured image, and then upload the captured image to a server via the network so that the examiner can perform the grade on the captured image.
[0037] In step 120, the captured image is detected based on a preset target detection frame to obtain multiple target images, wherein the target detection frame at least includes an answer sheet number detection frame, a question detection frame, and an answer detection frame.
[0038] Here, after obtaining the captured image of the answer sheet, key target detection can be performed on the captured image based on a preset target detection frame, thereby obtaining multiple target images. Figure 2 is a schematic diagram of a target detection frame according to an exemplary embodiment. Figure 2 As shown, the target detection frame includes an answer sheet number detection frame for detecting the number information of the answer sheet, a question detection frame for detecting the question information of the answer sheet, and an answer detection frame for detecting the answer information of the answer sheet.
[0039] It should be understood that when the answer sheet includes other information, the target detection frame may also include other detection frames. For example, the target detection frame may also include a name detection frame for detecting the name on the answer sheet, a page number detection frame for detecting the page number of the answer sheet, and a student number detection frame for detecting the student number filled in on the answer sheet.
[0040] It is worth mentioning that Figure 2 The answer sheet images shown are only used to illustrate the target detection frame. Figure 2 The clarity of the text and symbols does not affect the understanding of the technical solution.
[0041] For example, a captured image can be used as input to a trained object detection model. The model then extracts multiple target images from the captured image using pre-set target detection frames. In this target detection model, the target detection frame is equivalent to at least one sub-model within the target detection model. For example, the title detection frame can be a network layer within the target detection model, such as a CNN network.
[0042] It is worth noting that the target detection model is obtained by machine learning training based on the first training sample, which is an answer sheet image with the corresponding areas marked by the answer sheet number detection box, the question detection box and the answer detection box.
[0043] In step 130, the number information of the answer sheet is determined based on the first target image detected by the answer sheet number detection frame.
[0044] Here, the first target image is a sub-image extracted from the captured image that carries the answer sheet number. After the first target image is extracted from the captured image using the answer sheet number detection frame, the answer sheet number information is determined based on the first target image. The answer sheet number information refers to the test paper number corresponding to the answer sheet. For example, the answer sheet number information used for each exam is unique.
[0045] Exemplarily, the number information of the answer sheet can be extracted from the first target image by performing image recognition processing on the first target image.
[0046] In step 140, based on the number information, the standard answer corresponding to the answer sheet is determined.
[0047] Here, after obtaining the number information of the answer sheet, the standard answer corresponding to the answer sheet can be obtained from the database based on the number information. The standard answer can be the answer to the test paper corresponding to the answer sheet uploaded to the database by the examiner.
[0048] In some embodiments, different number information and the standard answer corresponding to each number information are stored in a database. It should be understood that the database can be a local database or a cloud database.
[0049] In step 150, answer information for each question on the answer sheet is determined based on the second target image detected by the question detection frame and the third target image detected by the answer detection frame.
[0050] Here, the second target image is a sub-image containing question information extracted from the captured image via the question detection frame. The question information includes at least the question number. The third target image is a sub-image containing answer information extracted from the captured image via the answer detection frame. The answer information includes at least the user's answer.
[0051] After obtaining the second target image and the third target image, the answer information of each question in the answer sheet can be determined based on the second target image and the third target image, wherein the answer information includes the question number and the answer corresponding to each question number.
[0052] For example, the question number corresponding to each second target image and the answer filled in the third target image corresponding to the question number can be determined by performing image recognition processing on the second target image and the third target image.
[0053] In step 160, the correction result of the answer sheet is obtained based on the answer information and the standard answer.
[0054] Here, after obtaining the answer information for each question on the answer sheet, the answer information is compared with the standard answer to obtain the correction result of the answer sheet. The correction result includes at least the score of the answer sheet and the correctness of each question.
[0055] Thus, the captured image is detected by the answer sheet number detection frame, the question detection frame, and the answer detection frame to obtain multiple target images, and based on the first target image detected by the answer sheet number detection frame, the number information of the answer sheet is determined, and based on the number information, the standard answer corresponding to the answer sheet is determined, and based on the second target image detected by the question detection frame and the third target image detected by the answer detection frame, the answer information of each question on the answer sheet is determined, and then the correction result of the answer sheet is obtained based on the answer information and the standard answer. Not only can the intelligent correction of the answer sheet be achieved by capturing the image, but the target image can be obtained by the answer sheet number detection frame, the question detection frame, and the answer detection frame, so that the answer sheet can be accurately corrected without being affected by the layout of the answer sheet, the filling method, etc. For example, whether it is a single-column, double-column or other complex layout answer sheet, it can be corrected by the answer sheet recognition method proposed in the embodiment of the present disclosure.
[0056] Figure 3 yes Figure 1 Detailed flow chart of step 150 is shown in FIG. Figure 3 As shown, in some feasible implementations, in step 150, determining the answer information for each question on the answer sheet based on the second target image detected by the question detection box and the third target image detected by the answer detection box may include the following steps.
[0057] In step 151, based on the relative positional relationship between each second target image detected by the question detection frame and each third target image detected by the answer detection frame, a plurality of question block images are determined in the captured image, wherein one question block image includes one second target image and a third target image corresponding to the second target image.
[0058] Here, the number of second target images detected from the captured image by the question detection frame and the number of third target images detected from the captured image by the answer detection frame are both more than one, so it is necessary to match each second target image with the third target image corresponding to the question number corresponding to the second target image.
[0059] The position information of each second target image and each third target image in the captured image is fixed. Therefore, based on the relative positional relationship between the second target images and the third target images, multiple topic block images can be determined in the captured image. Each topic block image includes a second target image and the third target image corresponding to the second target image.
[0060] It should be understood that the relative position relationship reflects the distance and orientation information between the second target image and the third target image. Based on the relative position relationship, the extracted multiple second target images and third target objects can be paired to obtain multiple topic block images.
[0061] In step 152 , for each of the second target image and the third target image in the question block image, the question number and the answer included in the question block image are determined.
[0062] Here, after obtaining a plurality of question block images, image recognition is performed on the second target image and the third target image in each question block image to obtain a corresponding question number and an answer.
[0063] Therefore, by accurately matching the extracted multiple second target images and third target objects based on the relative positional relationship between the second target image and the third target image, the answer sheet can be matched to improve the recognition accuracy of the answer information.
[0064] Figure 4 yes Figure 3 Detailed flow chart of step 152 is shown. Figure 4 As shown, in some possible implementations, filling in the answers can be obtained through the following steps.
[0065] In step 1521, text recognition is performed on the first image in the title block image to obtain text information of the first image.
[0066] Here, the answer detection box may include an unwritten detection box and a written detection box, and the third target image may include a first image detected by the unwritten detection box and a second image detected by the written detection box. In the answer sheet, each answer area may be numbered to represent the options of the corresponding question, and the options that are checked or written are the answers selected by the user. For example, for question 1, the corresponding options include "A", "B", "C", and "D". If the user writes on the area where option "A" is located, it means that the answer selected by the user is "A". Therefore, the first image actually refers to the sub-image corresponding to the unwritten area in the answer area of the answer sheet, and the second image actually refers to the sub-image corresponding to the written area in the answer area of the answer sheet.
[0067] After obtaining the first image, text recognition can be performed on the first image to determine the text information included in the first image that has not been scribbled on by the user. For example, the first image may include the text information "A," "B," and "C" that has not been scribbled on. It should be understood that the corresponding text information may vary for different answer sheet types and may be in the form of text, symbols, graphics, numbers, etc.
[0068] Exemplarily, the specific process of performing text recognition on the first image may include: correcting the first image using a Spatial Transformer Network to correct the curved or tilted text in the first image. The corrected first image is then used as the input of the feature extraction network to obtain character features. The feature extraction network is used to extract features related to the characters from the first image, while filtering out some irrelevant features, such as color, font, size, background and other information. In some embodiments, the feature extraction network may be a Resnet50 network. The character features are used as the input of the encoding network to obtain a feature vector. The encoding network is used to capture contextual information from the character features and encode the character features into corresponding feature vectors. In some embodiments, the encoding network may be a BiLSTM network. After obtaining the feature vector, the feature vector may be processed based on the attention model to obtain the text information contained in the first image.
[0069] In step 1522, the answer is determined based on the position information of the second image in the question block image and the text information of the first image.
[0070] Here, the second image is an image scribbled by the user, so text recognition processing cannot be performed on the second image. The location information of the second image in the question block image and the text information of the first image can be used to determine the area where the user actually scribbled, thereby obtaining the filled-in answer.
[0071] The position information of the second image in the question block image actually refers to the coordinate information of the second image on the captured image, and the option written by the user can be accurately located through the position information.
[0072] For example, the options include "A", "B", "C", and "D". If the option actually written by the user is "D", the obtained location information is the location information of "D", and the text information is the letter information of "A", "B", and "C".
[0073] It should be understood that by combining location information and text information, the problem of incorrect answer recognition caused by inaccurate text recognition or inaccurate location positioning can be avoided, making the final answer more accurate.
[0074] In some feasible implementations, the captured image can be used as input to the target detection model to obtain multiple annotated images, and based on the relative positional relationship between the multiple annotated images, the annotated images can be tilt-corrected to obtain the target image.
[0075] Here, the target detection model extracts annotated images from the captured image through the answer sheet number detection box, question detection box, and answer detection box. After obtaining multiple annotated images, each annotated image is tilt corrected according to the relative position relationship between the multiple annotated images, and the corrected annotated images are determined as target images.
[0076] The relative positional relationship between multiple annotated images reflects the angular relationship between various elements in the captured image. For example, the relative positional relationship between the answer sheet number determined by the answer sheet number detection frame, the question number determined by the question detection frame, and the answer information determined by the answer detection frame in the captured image reflects the deformation of various elements in the captured image or changes in the shooting angle. This relative positional relationship can be used to perform tilt correction on the extracted annotated image, thereby avoiding the problem of inaccurate answer sheet recognition caused by shooting angle issues.
[0077] It is worth noting that the target detection model can use Cascade R-CNN (a target detection algorithm) to extract target images from captured images.
[0078] Figure 5 yes Figure 1 Detailed flow chart of step 110 is shown in FIG. Figure 5 As shown, in some feasible implementations, in step 110, obtaining a photographed image of the answer sheet may include the following steps.
[0079] In step 111, an answer sheet image is obtained.
[0080] Here, the answer sheet may be photographed to obtain an image of the answer sheet.
[0081] In step 112, the answer sheet image is segmented to obtain a foreground image containing the answer sheet text area.
[0082] Here, the answer sheet image includes a foreground image and a background image. The foreground image is the image containing the answer sheet text area, while the background image is the image excluding the answer sheet text area. By segmenting the answer sheet image, we can filter out useless background information and eliminate interference caused by the background image on answer sheet recognition.
[0083] In some embodiments, the answer sheet image can be segmented using the ShuffleNetV2 network to obtain a foreground image.
[0084] In step 113, edge contour points of the foreground image are calculated, and a minimum bounding rectangle is constructed based on the edge contour points.
[0085] Here, edge contour points refer to a plurality of preset points on the edge of the foreground image. The number of edge contour points can be preset, for example, 16. The shape of the foreground image is described by the 16 edge contour points. Constructing a minimum bounding rectangle based on the edge contour points can be accomplished by arbitrarily selecting four edge contour points from the edge contour points to construct a bounding rectangle, and determining the rectangle with the smallest area as the minimum bounding rectangle.
[0086] In step 114, the foreground image is segmented based on the minimum bounding rectangle to obtain a segmented image.
[0087] Here, the foreground image is segmented by the minimum bounding rectangle to obtain a segmented image. Particularly, segmenting the foreground image by the minimum bounding rectangle can obtain a standard rectangular segmented image.
[0088] In step 115, perspective transformation is performed on the segmented image to obtain a corrected image.
[0089] Here, by performing perspective transformation processing on the segmented image, the influence of perspective deformation caused by shooting on the final recognition result can be eliminated.
[0090] In step 116, the corrected image is preprocessed to obtain a photographed image of the answer sheet.
[0091] Here, the preprocessing may include performing image enhancement operations such as shadow removal and contrast enhancement on the corrected image, thereby obtaining a captured image.
[0092] Therefore, the captured images obtained through steps 111 to 116 can eliminate interference in answer sheet recognition caused by factors such as blurred captured images, light interference, background clutter, perspective deformation, and image quality in the subsequent recognition process, so that the recognition results can be more accurate.
[0093] In some feasible implementations, in step 130, the first target image can be used as input to a trained recognition model to obtain the numbering information of the answer sheet; wherein the recognition model includes a text correction layer, a feature extraction layer, a sequence modeling layer and a prediction layer connected in sequence, the text correction layer is used to perform tilt correction on the first target image, the feature extraction layer is used to extract text features, the sequence modeling layer is used to construct text sequence features based on the context information of the text features, and the prediction layer is used to obtain the numbering information of the answer sheet based on the text sequence features.
[0094] Here, the text correction layer can be a spatial transformation network, which performs tilt correction on the first target image through the spatial transformation network to obtain a corrected image. The corrected image is then used as input to the feature extraction layer to obtain text features. The feature extraction layer is used to extract text features related to characters from the first target image, while filtering out irrelevant features such as color, font, size, background, and other information. In some embodiments, the feature extraction layer can be a Resnet50 network. The text features are used as input to the sequence modeling layer to obtain text sequence features. The sequence modeling layer is used to capture contextual information from the text features and encode the text features into corresponding text sequence features based on this contextual information. In some embodiments, the sequence modeling layer can be a BiLSTM network. After obtaining the text sequence features, the text sequence features are used as input to the prediction layer to obtain the answer sheet number information. The prediction layer can process the text sequence features based on an attention model to obtain the number information contained in the first target image.
[0095] Figure 6 FIG. 1 is a schematic diagram of module connections of an answer sheet recognition device according to an exemplary embodiment. Figure 6 As shown, the embodiment of the present disclosure provides an answer sheet recognition device, the device 600 including:
[0096] An acquisition module 601 is configured to acquire a photographed image of the answer sheet;
[0097] A detection module 602 is configured to detect the captured image based on a preset target detection frame to obtain a plurality of target images, wherein the target detection frame includes at least an answer sheet number detection frame, a question detection frame, and an answer detection frame;
[0098] A number recognition module 603 is configured to determine the number information of the answer sheet based on the first target image detected by the answer sheet number detection frame;
[0099] A standard answer determination module 604 is configured to determine the standard answer corresponding to the answer sheet based on the number information;
[0100] an answer recognition module 605 configured to determine answer information for each question on the answer sheet based on the second target image detected by the question detection frame and the third target image detected by the answer detection frame;
[0101] The correction module 606 is configured to obtain the correction result of the answer sheet according to the answer information and the standard answer.
[0102] Optionally, the answer identification module 605 includes:
[0103] a question-connecting unit configured to determine a plurality of question block images in the captured image based on a relative positional relationship between each second target image detected by the question detection frame and each third target image detected by the answer detection frame, wherein each question block image includes a second target image and a third target image corresponding to the second target image;
[0104] The recognition unit is configured to determine the question number and the answer included in the question block image for the second target image and the third target image in each of the question block images.
[0105] Optionally, the answer detection frame includes an ungraffiti detection frame and a graffiti detection frame, the third target image includes a first image detected by the ungraffiti detection frame and a second image detected by the graffiti detection frame; and the recognition unit includes:
[0106] a text recognition unit configured to perform text recognition on a first image in the title block image to obtain text information of the first image;
[0107] The answer determination subunit is configured to determine the answer based on the position information of the second image in the question block image and the text information of the first image.
[0108] Optionally, the detection module 602 is specifically configured as follows:
[0109] Using the captured image as input to a target detection model to obtain multiple target images;
[0110] The target detection model is obtained by machine learning training based on a first training sample, and the first training sample is an answer sheet image with corresponding areas marked by the answer sheet number detection box, the question detection box and the answer detection box.
[0111] Optionally, the detection module 602 includes:
[0112] an image extraction unit configured to use the captured image as input to a target detection model to obtain a plurality of annotated images;
[0113] The correction unit is configured to perform tilt correction on the annotated image based on the relative position relationship between the plurality of annotated images to obtain the target image.
[0114] Optionally, the acquisition module 601 includes:
[0115] An image acquisition unit configured to acquire an image of the answer sheet;
[0116] a first image segmentation unit configured to segment the answer sheet image to obtain a foreground image containing a text area of the answer sheet;
[0117] a contour unit configured to calculate edge contour points of the foreground image and construct a minimum circumscribed rectangle based on the edge contour points; and
[0118] a second image segmentation unit configured to segment the foreground image based on the minimum circumscribed rectangle to obtain a segmented image;
[0119] a transformation unit configured to perform perspective transformation processing on the segmented image to obtain a corrected image;
[0120] The preprocessing unit is configured to preprocess the corrected image to obtain a photographed image of the answer sheet.
[0121] Optionally, the number identification module 603 is specifically configured as follows:
[0122] Using the first target image as input to the trained recognition model to obtain the number information of the answer sheet;
[0123] Among them, the recognition model includes a text correction layer, a feature extraction layer, a sequence modeling layer and a prediction layer connected in sequence. The text correction layer is used to perform tilt correction on the first target image, the feature extraction layer is used to extract text features, the sequence modeling layer is used to construct text sequence features based on the context information of the text features, and the prediction layer is used to obtain the number information of the answer sheet based on the text sequence features.
[0124] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0125] Figure 7 FIG. 7 is a block diagram of an electronic device 700 according to an exemplary embodiment. Figure 7 As shown, the electronic device 700 may include: a processor 701 , a memory 702 , and may further include one or more of a multimedia component 703 , an input / output (I / O) interface 704 , and a communication component 705 .
[0126] The processor 701 is used to control the overall operation of the electronic device 700 to complete all or part of the steps in the above-mentioned answer sheet recognition method. The memory 702 is used to store various types of data to support the operation of the electronic device 700. Such data may include, for example, instructions for any application or method operating on the electronic device 700, as well as application-related data, such as contact information, sent and received messages, pictures, audio, video, etc. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 703 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 702 or sent via the communication component 705. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 704 provides an interface between the processor 701 and other interface modules. The above-mentioned other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 705 is used for wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more thereof, is not limited here. Therefore, the corresponding communication component 705 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.
[0127] In an exemplary embodiment, the electronic device 700 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned answer sheet recognition method.
[0128] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When executed by a processor, the program instructions implement the steps of the above-mentioned answer sheet recognition method. For example, the computer-readable storage medium may be the aforementioned memory 702 including the program instructions. The program instructions may be executed by the processor 701 of the electronic device 700 to perform the above-mentioned answer sheet recognition method.
[0129] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details of the above embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the scope of protection of the present disclosure.
[0130] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. In order to avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.
[0131] In addition, the various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.
Claims
1. A method for identifying an answer sheet, characterized in that: include: Obtaining a captured image of the answer sheet; Detecting the captured image based on a preset target detection frame to obtain a plurality of target images, wherein the target detection frame includes at least an answer sheet number detection frame, a question detection frame, and an answer detection frame; Determining the number information of the answer sheet based on the first target image detected by the answer sheet number detection frame; Determining the standard answer corresponding to the answer sheet based on the number information; Determining a plurality of question block images in the captured image based on a relative positional relationship between each second target image detected by the question detection frame and each third target image detected by the answer detection frame, wherein each question block image includes a second target image and a third target image corresponding to the second target image, wherein the relative positional relationship is used to reflect distance information and orientation information between the second target image and the third target image in a question block image; For each of the second target image and the third target image in the question block image, determining the question number and the answer included in the question block image; Obtaining a correction result of the answer sheet according to the question number, the filled-in answer, and the standard answer; The answer detection frame includes an ungraffiti detection frame and a graffiti detection frame, and the third target image includes a first image detected by the ungraffiti detection frame and a second image detected by the graffiti detection frame; The answers are determined as follows: Performing text recognition on a first image in the title block image to obtain text information of the first image; The answer is determined based on the position information of the second image in the question block image and the text information of the first image.
2. The answer sheet recognition method according to claim 1, characterized in that: The detecting the captured image based on a preset target detection frame to obtain a plurality of target images includes: Using the captured image as input to a target detection model to obtain multiple target images; The target detection model is obtained by machine learning training based on a first training sample, and the first training sample is an answer sheet image with corresponding areas marked by the answer sheet number detection box, the question detection box and the answer detection box.
3. The answer sheet recognition method according to claim 2, characterized in that: The method of using the captured image as input to a target detection model to obtain a plurality of target images includes: Using the captured image as input to a target detection model to obtain a plurality of labeled images; Based on the relative positional relationship between the plurality of annotated images, the annotated images are tilt-corrected to obtain the target image.
4. The answer sheet recognition method according to claim 1, characterized in that: The step of obtaining the photographed image of the answer sheet includes: Get the answer sheet image; Segmenting the answer sheet image to obtain a foreground image containing the answer sheet text area; Calculating edge contour points of the foreground image, and constructing a minimum circumscribed rectangle based on the edge contour points; and Segmenting the foreground image based on the minimum circumscribed rectangle to obtain a segmented image; Performing perspective transformation on the segmented image to obtain a corrected image; The corrected image is preprocessed to obtain a photographed image of the answer sheet.
5. The answer sheet recognition method according to claim 1, characterized in that: The determining the number information of the answer sheet based on the first target image detected by the answer sheet number detection frame includes: Using the first target image as input to the trained recognition model to obtain the number information of the answer sheet; Among them, the recognition model includes a text correction layer, a feature extraction layer, a sequence modeling layer and a prediction layer connected in sequence. The text correction layer is used to perform tilt correction on the first target image, the feature extraction layer is used to extract text features, the sequence modeling layer is used to construct text sequence features based on the context information of the text features, and the prediction layer is used to obtain the number information of the answer sheet based on the text sequence features.
6. An answer sheet recognition device, characterized in that: include: an acquisition module configured to acquire a photographed image of the answer sheet; a detection module configured to detect the captured image based on a preset target detection frame to obtain a plurality of target images, wherein the target detection frame includes at least an answer sheet number detection frame, a question detection frame, and an answer detection frame; a number recognition module configured to determine the number information of the answer sheet based on the first target image detected by the answer sheet number detection frame; A standard answer determination module is configured to determine the standard answer corresponding to the answer sheet based on the number information; a question-connecting unit configured to determine, in the captured image, a plurality of question block images based on a relative positional relationship between each second target image detected by the question detection frame and each third target image detected by the answer detection frame, wherein each question block image includes a second target image and a third target image corresponding to the second target image, wherein the relative positional relationship is used to reflect distance information and orientation information between the second target image and the third target image in a question block image; an identification unit configured to determine the question number and the answer included in the question block image for each of the second target image and the third target image in the question block image; A correction module is configured to obtain a correction result of the answer sheet according to the answer information and the standard answer; The answer detection frame includes an ungraffiti detection frame and a graffiti detection frame, the third target image includes a first image detected by the ungraffiti detection frame and a second image detected by the graffiti detection frame; the recognition unit includes: a text recognition unit configured to perform text recognition on a first image in the title block image to obtain text information of the first image; The answer determination subunit is configured to determine the answer based on the position information of the second image in the question block image and the text information of the first image.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
8. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Test paper and paper inspection system capable of preventing examinees from cheating
CN103870911A
End-to-end printed Mongolian recognition translation method based on spatial transformation network
CN112329760A
Card filling question cloud reviewing system and method based on neural network
CN113920521A