Test paper electronic identification positioning control method and device, equipment and medium
By placing graphic and text feature points in the margin area of the exam paper, and combining a lightweight convolutional neural network and an OCR model, the problems of print quality dependence and background interference in the existing exam paper recognition technology are solved, and efficient, accurate recognition and automated processing of electronic exam papers are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG MOKEN EDUCATION TECH CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies such as visual positioning based on black squares have problems such as strong dependence on printing quality, poor resistance to severe occlusion, and susceptibility to background interference. QR code-based positioning and recording technologies have problems such as high cost and space occupation, high printing and recognition thresholds, and the need to change the layout of the original test paper.
A recognition method combining graphic and textual feature points is adopted. By placing multiple graphic and textual feature points in the margin area of the test paper, a lightweight convolutional neural network and OCR model are used for recognition, and a large language model is combined for localization verification to achieve accurate recognition and segmentation of test paper information.
Without altering the original layout of the exam paper, the accuracy and anti-interference capabilities of exam paper recognition were improved, automated processing of electronic exam papers was achieved, the cost of technology implementation and adaptation was reduced, and recognition speed and accuracy were enhanced.
Smart Images

Figure CN121963243A_ABST
Abstract
Description
Methods, devices, equipment and media for electronic identification and positioning control of exam papers Technical Field
[0001] This application relates to the field of electronic examination papers, and in particular to an electronic examination paper identification and positioning control method, a corresponding device, an electronic device, and a computer-readable storage medium. Background Technology
[0002] In an era of deep integration between educational informatization and artificial intelligence, the education industry is accelerating its transformation from traditional paper-based media to digital and intelligent models. The demand for rapid processing and accurate analysis of massive amounts of paper-based test papers and assignments is becoming increasingly prominent. Traditional methods of manual sorting and annotation for assignment grading and test paper management have drawbacks such as low efficiency, high error rates, and difficulty in providing personalized feedback, making it difficult to meet the dual requirements of teaching quality and efficiency in large-scale teaching scenarios.
[0003] Currently, the core supporting technologies for electronic exam paper processing are visual positioning technology based on black squares and positioning recording technology based on QR codes. However, these technologies have the following technical defects in practical applications: 1. Visual positioning technology based on black squares has the following technical defects: 1.1 High dependence on printing quality: If the black squares are printed blurry, have ink leakage, or are too light in color, recognition may fail; ink diffusion (such as inferior ink) will change the shape of the squares and affect edge detection; 1.2 Poor resistance to occlusion: If two or more squares are completely occluded (such as when the exam paper is folded or damaged), coordinate mapping cannot be established, resulting in positioning failure; 1.3 Susceptible to background interference: If there are non-text black patterns (such as illustrations or table borders) in the background of the exam paper, they may be misidentified as positioning points; uneven lighting during scanning (such as local shadows) will reduce contrast and cause incorrect square recognition.
[0004] 1.4 Limitation of directional uniqueness: In some scenarios (such as using only two diagonal blocks), there may be misjudgments of mirror flips (such as upside down), and additional markers (such as arrows) are required to assist in positioning.
[0005] 2. The location recording technology based on QR codes has the following technical defects, including: 2.1 High cost and space occupation: QR codes are usually larger than ordinary black squares (the smallest size is about 1cm×1cm), making them less adaptable to scenarios with tight layout space (such as exam papers with narrow margins); 2.2 High printing and recognition thresholds: QR codes have higher requirements for printing accuracy (the module spacing error must be ≤10%), and poor printing may lead to decoding failure; ordinary scanners need to support QR code decoding function; 2.3 It requires changes to the layout of the original exam paper.
[0006] In summary, existing technologies such as black square-based visual positioning technology suffer from issues such as strong dependence on printing quality, poor resistance to severe occlusion, and susceptibility to background interference. In contrast, QR code-based positioning and recording technology suffers from problems such as high cost and space occupation, high printing and recognition thresholds, and the need to change the layout of the original test paper. The applicant has made corresponding explorations to address these issues. Summary of the Invention
[0007] The purpose of this application is to solve the above-mentioned problems by providing a method for electronic identification and positioning control of test papers, a corresponding device, electronic equipment and computer-readable storage medium.
[0008] To achieve the various objectives of this application, the following technical solution is adopted: A test paper electronic identification and positioning control method, proposed to meet one of the objectives of this application, includes: acquiring an image of a test paper to be identified, uploaded by a target examinee, containing a target identification code, wherein the target identification code is located in the margin area of the test paper image to be identified, and includes multiple graphic feature points, multiple text feature points, and a test paper information identification code; the graphic feature points include main graphic feature points and auxiliary graphic feature points; identifying the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the test paper image to be identified based on a preset graphic feature point identification model; determining the third position coordinates corresponding to each of the text feature points in the test paper image to be identified based on a preset text feature point identification model and a large language model fusion; and determining the third position coordinates corresponding to each of the text feature points in the test paper image to be identified based on the first position coordinates and the second position coordinates. The coordinates and the third position coordinates are used to calculate the fourth position coordinate range corresponding to the candidate information area, the fifth position coordinate range corresponding to the objective question answer area, and the sixth position coordinate range corresponding to the subjective question answer area in the image of the test paper to be identified. Based on the preset image segmentation model, the image is segmented according to the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range to determine the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answer area, and the third mask image corresponding to the subjective question answer area. Based on the first mask image, the candidate information of the target candidate is identified. The second mask image, the third mask image, the test paper information identification code, and the candidate information are associated to determine the test paper answer data of the target candidate and pushed to the preset test paper grading system to complete the control of electronic identification and positioning of the test paper.
[0009] Optionally, the step of identifying the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the image of the test paper to be identified based on a preset graphic feature point recognition model includes: extracting graphic contour information in the image of the test paper to be identified that conforms to a preset size and a preset contour shape based on the preset graphic feature point recognition model, so as to match and locate the main graphic feature points at the four corners of the test paper to be identified, as well as multiple auxiliary graphic feature points on its sides; and calculating and outputting the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points, respectively, based on the image pixel coordinate system of the test paper to be identified.
[0010] Optionally, the step of determining the third position coordinates corresponding to each text feature point in the test paper image to be recognized based on a preset text feature point recognition model and a large language model fusion includes: obtaining text feature points located adjacent to the main graphic feature points; matching candidate regions of corresponding text feature points for each main graphic feature point based on the first position coordinates corresponding to the main graphic feature points at the four corners of the test paper to be recognized; calling the preset text feature point recognition model to perform character detection on the candidate regions respectively, so as to output the minimum bounding rectangle bounding box of the characters in the candidate regions and the preliminary recognized text; inputting the association relationship between the preliminary recognized text and the main graphic feature points corresponding to the candidate regions into the large language model constructed by the model. The bit-identifier semantic verification library uses a multi-layered mechanism, including distorted character restoration, similar-looking character and background filtering, and contextual logic verification, to filter out valid text feature points that conform to preset positioning identifiers. Using the image pixel coordinate system of the test paper image to be recognized as a reference, it calculates the center coordinates of the minimum bounding rectangle of the valid text feature points, which serve as the initial position coordinates of each valid text feature point. Based on the first position coordinates corresponding to the main graphic feature points and the theoretical distance between their corresponding text feature points, it calculates and determines the theoretical position coordinates of the text feature points. If the relative position distance between the initial position coordinates and the theoretical position coordinates is less than a preset distance threshold, it outputs the initial position coordinates of the valid text feature points as the third position coordinates. Optionally, the step of calculating the range of fourth position coordinates corresponding to the candidate information area, the range of fifth position coordinates corresponding to the objective question answering area, and the range of sixth position coordinates corresponding to the subjective question answering area in the candidate's information area of the candidate's information ... The system establishes a preset relative positional relationship between each pair of character feature points, and a mapping relationship between the main graphic feature points, the auxiliary graphic feature points, the text feature points, and each functional area of the test paper to be identified. It then performs cross-validation and fusion correction on the first, second, and third position coordinates to obtain a feature point coordinate set for the test paper to be identified. Based on the feature point coordinate set and a preset test paper functional area layout template, it calculates the fourth position coordinate range corresponding to the candidate information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the image of the test paper to be identified.
[0011] Optionally, the step of performing cross-validation and fusion correction on the first position coordinates, the second position coordinates, and the third position coordinates to obtain the feature point coordinate set of the test paper to be identified includes: performing cross-validation on the first position coordinates, the second position coordinates, and the third position coordinates based on the identification code layout rules, determining whether the relative position distance between each pair of the first position coordinates, the second position coordinates, and the third position coordinates is within a preset distance threshold; if not within the preset distance threshold, marking the abnormal position coordinates whose relative position distance exceeds the preset distance threshold; using a weighted average method to fuse and correct the abnormal position coordinates based on the remaining valid position coordinates to determine the corrected position coordinates; and generating the feature point coordinate set of the test paper to be identified based on the corrected position coordinates and the valid position coordinates.
[0012] Optionally, the step of performing image segmentation based on a preset image segmentation model according to the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range to determine the first mask image corresponding to the candidate information region, the second mask image corresponding to the objective question answering region, and the third mask image corresponding to the subjective question answering region includes: obtaining the fourth position coordinate range corresponding to the candidate information region, the fifth position coordinate range corresponding to the objective question answering region, and the sixth position coordinate range corresponding to the subjective question answering region in the test paper image to be identified; and delineating the candidate information region in the test paper image to be identified based on the fourth position coordinate range. The domain range is defined by delineating the objective question answering area in the test paper image to be identified based on the fifth position coordinate range, and the subjective question answering area in the test paper image to be identified based on the sixth position coordinate range. A pre-trained and converged image segmentation model is invoked to perform pixel-level segmentation on the defined candidate information area, objective question answering area, and subjective question answering area, respectively, to exclude irrelevant interference information outside each area, including question description text, headers and footers, paper background, and stains. The first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answering area, and the third mask image corresponding to the subjective question answering area are output.
[0013] Optionally, the main graphic feature points are located at the four corners of the test paper image to be identified, the auxiliary graphic feature points are located on the sides of the test paper image to be identified, and the text feature points are located adjacent to the main graphic feature points; the test paper information identification code includes the subject, grade, exam session, and question type distribution of the test paper image to be identified; the basic network architecture of the graphic feature point recognition model includes a lightweight convolutional neural network; the basic network architecture of the text feature point recognition model includes an OCR model; and the basic network architecture of the image segmentation model includes a U2net model.
[0014] A test paper electronic identification and positioning control device provided for another purpose of this application includes: a test paper image acquisition module, configured to acquire a test paper image to be identified uploaded by a target examinee, containing a target identification code, wherein the target identification code is located in the margin area of the test paper image to be identified, and includes multiple graphic feature points, multiple text feature points, and a test paper information identification code, wherein the graphic feature points include main graphic feature points and auxiliary graphic feature points; a feature point coordinate identification module, configured to identify the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the test paper image to be identified based on a preset graphic feature point identification model, and determine the third position coordinates corresponding to each of the text feature points in the test paper image to be identified based on a preset text feature point identification model and a large language model fusion; and a functional area coordinate determination module, configured to determine the coordinates of the first position coordinates, the second position coordinates, and the third position coordinates of the text feature points in the test paper image to be identified based on the first position coordinates, the second position coordinates, and the third position coordinates of the text feature points; and a functional area coordinate determination module, configured to determine the coordinates of the functional area based on the first position coordinates, the second position coordinates, and the third position coordinates of the text feature points. The third position coordinates are used to calculate the fourth position coordinate range corresponding to the candidate information area, the fifth position coordinate range corresponding to the objective question answer area, and the sixth position coordinate range corresponding to the subjective question answer area in the image of the test paper to be identified. The mask image segmentation module is configured to perform image segmentation based on the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range according to a preset image segmentation model to determine the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answer area, and the third mask image corresponding to the subjective question answer area. The answer data push module is configured to identify the candidate information of the target candidate based on the first mask image, associate the second mask image, the third mask image, the test paper information identification code, and the candidate information to determine the test paper answer data of the target candidate, and push it to a preset test paper grading system to complete the control of electronic identification and positioning of the test paper.
[0015] An electronic device provided for another purpose of this application includes a central processing unit and a memory, the central processing unit being configured to invoke and run a computer program stored in the memory to perform the steps of the electronic identification and positioning control method for examination papers described in this application.
[0016] A computer-readable storage medium is provided for another purpose of this application, which stores, in the form of computer-readable instructions, a computer program implemented according to the electronic identification and positioning control method for the test paper, which, when called by a computer, executes the steps included in the corresponding method.
[0017] Compared to existing technologies, this application addresses the shortcomings of existing visual positioning technologies based on black squares, such as strong dependence on print quality, poor resistance to severe occlusion, and susceptibility to background interference. It also addresses the issues of QR code-based positioning and recording technologies, such as high cost and space requirements, high printing and recognition barriers, and the need to alter the layout of existing exam papers. This application offers the following advantages, including but not limited to: First, this application uses only page margins to place the identification code on the existing exam paper format, without needing to adjust the main text layout, headers, footers, or reserve extra space. Compared to the drawbacks of QR codes requiring layout changes to accommodate graphics, and the complex design of black squares potentially requiring additional directional markings, this application perfectly aligns with traditional exam paper layout habits, reducing the adaptation costs for technology implementation.
[0018] Secondly, this application adopts a feature point design that combines graphic outlines and Chinese characters. The Chinese characters have high information content and strong outline recognition. Even if the printing is slightly blurry, the color is light, or the ink spreads slightly, recognition can still be completed through the dual features of character semantics and graphic outlines, solving the problem of the black squares being highly dependent on printing quality. Thirdly, this application uses a comprehensive positioning method that combines graphic feature points and text feature points. Even if the test paper is folded or missing corners, causing two or more corner graphic feature points to fail, the missing positioning point can still be calculated through the remaining diagonal feature points and text feature points, solving the pain point of positioning failure due to black square occlusion. Fourthly, for black illustrations and table borders in the test paper background, the combination of Chinese characters and graphic feature points can effectively distinguish positioning elements from interference patterns, avoiding misidentification. At the same time, through uniform illumination processing, large model semantic correction, and OCR algorithm optimization, it can still accurately identify even in the face of uneven illumination, character adhesion or breakage, and text deformation and distortion, breaking through the limitation of traditional technology where the recognition rate drops sharply in complex environments.
[0019] Fifth, this application uses 12 Chinese characters as the test paper identification code, 1 Chinese character as the page number, and 4 Chinese characters as the check code. The high information capacity of Chinese characters ensures that the identification code of each test paper is unique, which can accurately distinguish different test papers of the same subject and version. Compared with the limitations of black squares, which can only achieve positioning and cannot carry complex information, and the problems of limited information capacity and large space occupation of QR codes, this application greatly improves the accuracy of electronic identification and positioning of test papers, and provides complete data support for subsequent test paper grading.
[0020] Sixth, it prioritizes image recognition for rapid localization, and only starts OCR text recognition when image recognition fails, avoiding the inefficiency of relying entirely on OCR and significantly improving the recognition speed of a single test paper. By pre-setting the mapping relationship between feature points and functional areas (name, student ID, answer area), the coordinates of the target area can be directly calculated after locating the identification code, without the need for manual annotation. Combined with image segmentation and standardized mask output, it realizes full automation from test paper scanning to electronic data, solving the bottleneck of low efficiency in traditional manual sorting and annotation, and can support the batch test paper processing needs in large-scale teaching scenarios. Attached Figure Description
[0021] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 is a flowchart illustrating the electronic identification and positioning control method for examination papers in an embodiment of this application; Figure 2 is a schematic diagram illustrating the image of the examination paper to be identified in an embodiment of this application; Figure 3 is a block diagram illustrating the principle of the electronic identification and positioning control device for examination papers in an embodiment of this application; and Figure 4 is a structural schematic diagram of the computer equipment in an embodiment of this application. Detailed Implementation
[0022] Those skilled in the art will understand that although the various methods in this application are described based on the same concept and thus present commonality among them, they can be performed independently unless otherwise specified. Similarly, the various embodiments disclosed in this application are all based on the same inventive concept; therefore, concepts expressed in the same way, as well as concepts that are appropriately changed for convenience but are expressed differently, should be understood equivalently.
[0023] Unless otherwise expressly stated, the various embodiments disclosed in this application can be combined in a cross-cutting manner to flexibly construct new embodiments, as long as such combination does not depart from the inventive spirit of this application and can meet the needs of the prior art or solve a certain deficiency in the prior art. Those skilled in the art should be aware of such modifications.
[0024] Please refer to Figures 1 and 2. In one embodiment of the electronic test paper recognition and positioning control method of this application, the method includes: Step S10: Obtaining an image of a test paper to be recognized containing a target identification code uploaded by the target examinee, wherein the target identification code is located in the margin area of the image of the test paper to be recognized, and includes multiple graphic feature points, multiple text feature points, and a test paper information identification code; the graphic feature points include main graphic feature points and auxiliary graphic feature points; the electronic test paper recognition and positioning control system in the terminal device obtains the image of the test paper to be recognized containing the target identification code, wherein the target identification code is located in the margin area of the image of the test paper to be recognized, and includes multiple graphic feature points, multiple text feature points, and a test paper information identification code; the graphic feature points include main graphic feature points and auxiliary graphic feature points; the main graphic feature points are located in the margin area of the image of the test paper to be recognized. The auxiliary graphic feature points are located at the four corners of the image, and the text feature points are located adjacent to the main graphic feature points. The exam paper information identification code includes the subject, grade, exam session, and question type distribution of the exam paper image. In some embodiments, the main graphic feature points are located at the four corners of the image, the auxiliary graphic feature points are located at the sides of the image, and the text feature points are located adjacent to the main graphic feature points. The exam paper information identification code includes the subject, grade, exam session, and question type distribution of the exam paper image. The basic network architecture of the graphic feature point recognition model includes a lightweight convolutional neural network. The basic network architecture of the text feature point recognition model includes an OCR model. The basic network architecture of the image segmentation model includes a U2net model.
[0025] In some embodiments, a margin of at least 2cm is preset around the page when the test paper is created. The target identification code is only placed in this area and does not occupy the core areas such as the test text and answer spaces, so as to avoid being obscured by the candidate's answer content. Moreover, there is no need to adjust the original test paper's question layout, font size, spacing, and other formats. The graphic feature points are the core anchor points for locating the test paper image to be identified. The spatial coordinate system of the test paper to be identified is quickly established through geometric contour recognition. The main graphic feature points can be placed at the four corners of the test paper image to be identified, and their size is uniformly 21×21 pixels. For example, the position coordinates of the main graphic feature point l0 in the upper left corner are (25,35); the position coordinates of the main graphic feature point r1 in the lower right corner are (747,1066). The main graphic feature points serve as the reference vertices for establishing the global pixel coordinate system of the test paper, which can determine the overall boundary and orientation of the test paper to avoid misjudgments such as upside down.
[0026] The auxiliary graphic feature points are disposed on the side of the test paper image to be recognized. The text feature points are disposed at adjacent positions of the main graphic feature points, and their sizes are uniformly 21×21 pixels. For example, the position coordinates of the first auxiliary graphic feature point h1 are (25, 292); the position coordinates of the second auxiliary graphic feature point h2 are (25, 550); the position coordinates of the third auxiliary graphic feature point h3 are (25, 808), etc. They can calibrate the tilt and distortion of the test paper shooting (such as perspective deformation); complete the missing main graphic feature points. For example, when a corner of the main graphic of the test paper is missing and cannot be recognized, it can be deduced through the auxiliary graphic feature points.
[0027] The text feature points are disposed at adjacent positions of the main graphic feature points. The information content of Chinese characters is much higher than that of letters or numbers, the upper limit of accuracy is high, the distinction from the text and patterns of the test paper background is high, and it is not easily misrecognized. It can solve the problem that the black squares in the prior art are easily interfered by the background; the text feature points are used as the complement of the graphic feature points. When the main graphic feature points in a certain corner are blocked, the area can be located by recognizing the adjacent text feature points. Therefore, the text feature points are disposed at adjacent positions of the main graphic feature points. For example, to the right of the main graphic feature point l0 in the upper left corner is the text feature point "壹", to the left of the main graphic feature point r0 in the upper right corner is the text feature point "贰", to the right of the main graphic feature point l1 in the lower left corner is the text feature point "叁", and to the left of the main graphic feature point r1 in the lower right corner is "肆"; their sizes can be 17×17 pixels. For example, the position coordinates of the text feature point "壹" are (52, 37); the position coordinates of the text feature point "肆" are (724, 1068), etc. The test paper information identification code includes the subject, grade, exam session, and question type distribution of the test paper image to be recognized. For example, the subject is mathematics, Chinese, English, etc., and the grade is the first grade of junior high school, the third grade of senior high school, etc.; the exam session includes mid-term exam, final exam, mock exam, etc.; the question type distribution includes the question types and the position ranges of each question type. The question types include subjective questions and objective questions, and the objective questions include multiple-choice questions, true or false questions, etc.; the subjective questions include fill-in-the-blank questions or short-answer questions, etc. The test paper information identification code can be constructed by 12 Chinese characters (recording core information), 1 Chinese character (page number), and 4 Chinese characters (check code), and one set can be disposed at the header and footer of the test paper respectively.
[0028] The target candidate can upload the test paper image to be recognized containing the target identification code through a scanner or mobile phone camera. If the test paper image to be recognized does not contain a complete target identification code, the test paper electronic recognition and positioning control system will prompt "The image is invalid. Please re-upload the test paper image with complete page margins" to ensure that the subsequent positioning process can be started normally.
[0029] In some embodiments, barcodes are used to assist in recording test paper information. When the test paper information identification code fails to be recognized, the test paper content can also be identified by recognizing the barcode, thereby improving the recognition rate.
[0030] Step S20: Based on a preset graphic feature point recognition model, identify the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the test paper image to be recognized; based on a preset text feature point recognition model and a large language model fusion, determine the third position coordinates corresponding to each of the text feature points in the test paper image to be recognized; after obtaining the test paper image to be recognized uploaded by the target examinee containing the target identification code, based on the preset graphic feature point recognition model, identify the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the test paper image to be recognized; based on a preset text feature point recognition model and a large language model fusion, determine the third position coordinates corresponding to each of the text feature points in the test paper image to be recognized; wherein, the basic network architecture of the graphic feature point recognition model includes a lightweight convolutional neural network; the basic network architecture of the text feature point recognition model includes an OCR model.
[0031] In some embodiments, Convolutional Neural Networks (CNNs) extract abstract features through multiple convolutions. Even if the printed image is blurry or the ink has slightly diffused, such as shape deviations caused by inferior ink, it can still be identified through contour feature matching, reducing the dependence on printing accuracy and solving the drawback of "printing quality sensitivity" of black squares in the prior art. CNNs, through pooling layers and attention mechanisms, can distinguish between black patterns (illustrations, table borders) and positioned graphics in the background of the exam paper, avoiding misidentification and solving the problem of black squares being easily affected by background interference in the prior art. CNNs are robust to slight wrinkles and tilts in graphics, and can restore the essence of the graphics through feature normalization processing, avoiding directional misjudgment without additional markings (such as arrows), thus solving the limitation of directional uniqueness of black squares in the prior art. Therefore, a lightweight convolutional neural network is used as the graphic feature point recognition model in this application.
[0032] In some embodiments, the step of identifying the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the test paper image to be identified based on a preset graphic feature point recognition model includes: Step S201, extracting graphic contour information that conforms to a preset size and preset contour shape in the test paper image to be identified based on the preset graphic feature point recognition model, so as to match and locate the main graphic feature points at the four corners of the test paper to be identified, and multiple auxiliary graphic feature points on its sides; specifically, calling the preset graphic feature point recognition model, which is constructed based on a lightweight convolutional neural network, extracting graphic contour information that conforms to a preset size and preset contour shape in the test paper image to be identified based on the preset graphic feature point recognition model, and extracting the graphic feature points in the test paper image to be identified. The contour information is used to match and locate feature points through a dual-filtering rule. First, only graphics with dimensions that conform to the preset size are retained. Since the preset size of the main graphic feature point is 21×21 pixels, the preset size of the auxiliary graphic feature point is also 21×21 pixels, black patterns in the test paper background that do not match the size (such as illustrations, table borders, etc.) are excluded, thus solving the drawback of the existing black square positioning being easily affected by the background. Then, only graphics with contour shapes consistent with the preset contour shapes are retained. The contours of the main graphic feature points and the auxiliary graphic feature points are regular geometric shapes, such as squares and triangles. The lightweight convolutional neural network extracts abstract contour features through multi-layer convolution. Even if the graphic has slight printing blur or ink diffusion (caused by inferior ink), it can still be recognized through contour matching, reducing the dependence on printing accuracy and solving the drawback of the existing black square positioning being highly dependent on printing quality.
[0033] Furthermore, by extracting the graphic contour information from the image of the test paper to be identified that conforms to a preset size and preset contour shape, the main graphic feature points at the four corners of the test paper to be identified can be accurately matched, such as the main graphic feature point l0 at the top left corner, the main graphic feature point l1 at the bottom left corner, the main graphic feature point r0 at the top right corner, and the main graphic feature point r1 at the bottom left corner, as well as multiple auxiliary graphic feature points on the side of the test paper to be identified. For example, the first auxiliary graphic feature point h1, the second auxiliary graphic feature point h2, and the third auxiliary graphic feature point h3 are located on the left side of the test paper to be identified.
[0034] By employing dual screening based on size and outline, misidentification of irrelevant black patterns in the exam paper background is completely eliminated, demonstrating stronger anti-interference capabilities than traditional black square positioning. The lightweight CNN is robust to minor printing deviations and slight graphic deformations (such as incomplete outlines caused by wrinkles), allowing for identification without requiring strict printing precision.
[0035] Step S202: Using the image pixel coordinate system of the test paper to be identified as a reference, calculate and output the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points.
[0036] Using the image pixel coordinate system of the test paper to be identified as a reference, this image pixel coordinate system has the upper left corner of the image as the origin, the horizontal rightward direction as the X-axis, and the vertical downward direction as the Y-axis, with the coordinate unit being pixels. For each located graphic feature point, its corresponding feature center coordinates are calculated. The center coordinates of the smallest bounding rectangle of the graphic or the geometric center coordinates of the contour can be taken to ensure that the coordinates accurately represent the position of the graphic in the image. The first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points are calculated and output respectively. The first position coordinates corresponding to each of the main graphic feature points include the position coordinates of the upper left main graphic feature point l0 as (25, 35) and the coordinates of the lower right main graphic r1 as (747, 1066), etc.; the second position coordinates corresponding to each of the auxiliary graphic feature points include the position coordinates of the first auxiliary graphic feature point h1 on the left side of the test paper as (25, 292), the position coordinates of the second auxiliary graphic feature point h2 as (25, 550), and the position coordinates of the third auxiliary graphic feature point h3 as (25, 808), etc.
[0037] As described in steps S201 to S202 above, the visual graphic feature point positions are converted into calculable coordinate data, providing a core basis for subsequent steps to calculate the coordinates of the examinee's information area and answer area. A global coordinate system for the exam paper is established using the coordinates of the main graphic feature points and auxiliary graphic feature points, ensuring the accuracy of functional area calculations. Even if the exam paper is slightly tilted, the calculated coordinates still reflect the relative positions of the graphic feature points. Subsequent tilt calibration can be performed using the coordinates of the auxiliary graphic feature points, maintaining a high recognition rate even under extreme conditions such as folded corners and complex lighting.
[0038] In a further embodiment, the step of determining the third position coordinates corresponding to each of the text feature points in the to-be-recognized test paper image based on the fusion of a preset text feature point recognition model and a large language model includes: Step S2001, obtaining text feature points arranged at adjacent positions of the main graphic feature points, and based on the first position coordinates corresponding to the main graphic feature points at the four corners of the to-be-recognized test paper, matching a candidate region of the text feature point corresponding to each main graphic feature point; avoiding the inefficiency and misrecognition caused by the global scanning of the OCR model, solving the problem that traditional text positioning is easily interfered by background text, and loading the relative positions between the main graphic feature points and the text feature points. For example, the horizontal axis (X-axis) coordinate of the text feature point "壹" relative to the main graphic feature point l0 is offset by 27 pixels, and the vertical axis (Y-axis) is offset by 2 pixels, with a standard size of 17×17 pixels; matching a candidate region for each main graphic feature point: taking the position coordinates of the main graphic feature point l0 as (25, 35) as an example, the candidate region range of its corresponding text feature point "壹" is X∈[25 + 20, 25 + 44], Y∈[35 - 2, 35 + 19].
[0039] Specifically, starting from the X coordinate (25) of the main graphic feature point l0, adding the value obtained by superimposing the standard offset (27 pixels) minus the compensation value (7 pixels) gives the left boundary (25 + 20 = 45) of the candidate region of the text feature point "壹"; starting from the X coordinate (25) of the main graphic feature point l0 and adding the sum of the "standard offset (27 pixels) + 17 pixels" gives the right boundary (25 + 44 = 69) of the candidate region of the text feature point "壹". This candidate region range not only completely covers the 17×17 pixel standard width of the text feature point "壹", but also reserves left and right error tolerance spaces, adapting to scenarios of tilted shooting and slightly offset printing.
[0040] Starting from the Y coordinate (35) of the main graphic feature point l0, adding the standard offset (2 pixels) minus the compensation value (4 pixels) gives the upper boundary (35 - 2 = 33) of the candidate region of the text feature point "壹". Starting from the Y coordinate (35) of the main graphic feature point l0 and adding the sum of the standard offset (2 pixels) and 17 pixels gives the lower boundary (35 + 19 = 54) of the candidate region of the text feature point "壹", which also covers the standard height of "壹" while reserving upper and lower error tolerance spaces.
[0041] By the above method, the candidate regions of the text feature points "壹、贰、叁、肆" corresponding to each main graphic feature point one by one are determined, only covering the potential positions of "壹、贰、叁、肆", and excluding the interference of background regions such as the test paper text and illustrations.
[0042] Step S2002: Call the preset text feature point recognition model to perform character detection on the candidate regions respectively, so as to output the minimum bounding rectangle of the characters in the candidate regions and the preliminary recognition text; call the preset OCR model to perform character detection on the candidate regions corresponding to each text feature point respectively, so as to determine the minimum bounding rectangle of the characters in the candidate regions corresponding to each text feature point and the preliminary recognition text, such as "壹", "一", "壱" or fragmented characters, etc.; exclude the bounding boxes with an area much smaller than 17×17 pixels or much larger than 17×17 pixels.
[0043] Step S2003: Input the association relationship between the preliminary recognition text and the main graphic feature points corresponding to the candidate regions into the positioning identifier semantic verification library constructed by the large language model, and screen out the valid text feature points that meet the preset positioning identifier through a multi-layer mechanism including distorted character restoration, similar character and background filtering, and context logic verification; the positioning identifier semantic verification library includes "壹, 贰, 叁, 肆" and their corresponding common distorted forms; the distorted character restoration means that for broken (such as "壹" broken into "士" + "冖" + "豆", etc.), adhered, and distorted characters, through semantic completion and glyph association analysis of the large language model, they are restored to standard positioning characters; the similar character and background filtering means excluding similar characters such as "一, 二, 壱" and the text on the test paper background (such as "一, 二, 三" in the test paper questions), and only retaining the preset positioning identifier; the context logic verification follows that the text feature point corresponding to the main graphic feature point in the upper left corner is "壹", the text feature point corresponding to the main graphic feature point in the upper right corner is "贰", the text feature point corresponding to the main graphic feature point in the lower left corner is "叁", and the text feature point corresponding to the main graphic feature point in the lower right corner is "肆". If the recognition result of a certain region conflicts with the global logic, such as the text feature point corresponding to the main graphic feature point in the lower left corner is recognized as "贰", it is determined as a misrecognition and excluded; screen out the valid text feature points that pass all verifications, and clarify their corresponding candidate regions and minimum bounding rectangles. For problems that traditional OCR models cannot handle, such as characters adhered or broken due to light, and text deformed due to wrinkles, inputting the association relationship between the preliminary recognition text and the main graphic feature points corresponding to the candidate regions into the positioning identifier semantic verification library constructed by the large language model, and screening out the valid text feature points that meet the preset positioning identifier through a multi-layer mechanism including distorted character restoration, similar character and background filtering, and context logic verification can greatly improve the recognition error tolerance rate in extreme scenarios. <>
[0044] Step S2004: Based on the image pixel coordinate system of the test paper image to be recognized, calculate the center coordinates of the minimum bounding rectangle of the effective text feature points as the preliminary position coordinates of each of the effective text feature points; based on the image pixel coordinate system of the test paper image to be recognized, for each effective text feature point, take the center coordinates (X, Y) of its minimum bounding rectangle as the preliminary position coordinates of each of the effective text feature points.
[0045] Step S2005: Calculate and determine the theoretical position coordinates of the text feature points according to the first position coordinates corresponding to the main graphic feature points and the theoretical spacing of the corresponding text feature points. If the relative position distance between the preliminary position coordinates and the theoretical position coordinates is less than the preset distance threshold, output the preliminary position coordinates of the effective text feature points as the third position coordinates.
[0046] Load the first position coordinates corresponding to the main graphic feature points. For example, the position coordinates of the main graphic feature point l0 in the upper left corner are (24.8, 34.9). Combine the theoretical spacing between the main graphic feature point l0 in the upper left corner and the text feature point "壹" to calculate the theoretical position coordinates of the effective text feature point "壹". For example, the theoretical horizontal axis coordinate of the effective text feature point "壹" is 24.8 + 27 = 51.8, and the theoretical vertical axis coordinate is 34.9 + 2 = 36.9. Then the theoretical position coordinates of the effective text feature point "壹" are (51.8, 36.9); if the relative position distance between the preliminary position coordinates (X, Y) of the effective text feature point "壹" and the theoretical position coordinates (51.8, 36.9) of the effective text feature point "壹" is less than the preset distance threshold, output the preliminary position coordinates (X, Y) of the effective text feature point "壹" as the third position coordinates; if the relative position distance between the preliminary position coordinates (X, Y) of the effective text feature point "壹" and the theoretical position coordinates (51.8, 36.9) of the effective text feature point "壹" is greater than the preset distance threshold, where the preset distance threshold can be 3 pixels, etc.
[0047] Step S30: Calculate the range of the fourth position coordinates corresponding to the candidate information area, the range of the fifth position coordinates corresponding to the objective question answering area, and the range of the sixth position coordinates corresponding to the subjective question answering area in the to-be-recognized test paper image based on the first position coordinates, the second position coordinates, and the third position coordinates; After identifying the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the to-be-recognized test paper image based on a preset graphic feature point recognition model, and determining the third position coordinates corresponding to each of the text feature points in the to-be-recognized test paper image based on the fusion of a preset text feature point recognition model and a large language model, calculate the range of the fourth position coordinates corresponding to the candidate information area, the range of the fifth position coordinates corresponding to the objective question answering area, and the range of the sixth position coordinates corresponding to the subjective question answering area in the to-be-recognized test paper image based on the first position coordinates, the second position coordinates, and the third position coordinates; In some embodiments, the step of calculating the range of the fourth position coordinates corresponding to the candidate information area, the range of the fifth position coordinates corresponding to the objective question answering area, and the range of the sixth position coordinates corresponding to the subjective question answering area in the to-be-recognized test paper image based on the first position coordinates, the second position coordinates, and the third position coordinates includes: Step S301: Obtain the first position coordinates corresponding to the main graphic feature points, the second position coordinates corresponding to the auxiliary graphic feature points, and the third position coordinates corresponding to the text feature points in the to-be-recognized test paper image; The first position coordinates corresponding to the main graphic feature points include that the position coordinates of the main graphic feature point l0 at the upper left corner of the to-be-recognized test paper are (25, 35); the position coordinates of the main graphic feature point r1 at the lower right corner are (747, 1066), etc.; The second position coordinates corresponding to the auxiliary graphic feature points include that the position coordinates of the first auxiliary graphic feature point h1 are (25, 292); the position coordinates of the second auxiliary graphic feature point h2 are (25, 550); the position coordinates of the third auxiliary graphic feature point h3 are (25, 808); The third position coordinates corresponding to the text feature points include that the position coordinates of the text feature point "壹" are (52, 37); the position coordinates of the text feature point "肆" are (724, 1068), etc.
[0048] Step S302: Based on the preset identification code layout rule, determine the preset relative position relationships between the main graphic feature points, the auxiliary graphic feature points, and the text feature points in the to-be-recognized test paper image pairwise, as well as the mapping relationships between the main graphic feature points, the auxiliary graphic feature points, the text feature points, and each functional area of the to-be-recognized test paper; The preset relative position relationships between the main graphic feature points, the auxiliary graphic feature points, and the text feature points in the to-be-recognized test paper image pairwise, for example, the horizontal axis (X-axis) offset between the main graphic feature point l0 at the upper left corner and its corresponding text feature point "One" is 27 pixels, and the vertical axis (Y-axis) offset is 2 pixels, and the vertical axis (Y-axis) distance between the main graphic feature point l0 at the upper left corner and the first auxiliary graphic feature point h1 is 257 pixels, etc.; The mapping relationships between the main graphic feature points, the auxiliary graphic feature points, the text feature points, and each functional area of the to-be-recognized test paper can be based on the fixed relative positions of each functional area of the to-be-recognized test paper and the main graphic feature points, the auxiliary graphic feature points, and the text feature points. For example, the position coordinates of "Name" in the candidate information area are (126, 37). Taking the X-axis coordinate of the text feature point "One" as the reference, and offsetting 74 pixels to the right, the starting X-axis position of "Name" in the candidate information area is determined. The ranges of the objective question area and the subjective question area can also be fixed by the distances from each auxiliary graphic feature points h1, h2, h3.
[0049] Step S303: Perform cross-checking and fusion correction on the first position coordinate, the second position coordinate, and the third position coordinate to obtain the feature point coordinate set of the to-be-recognized test paper; In a specific embodiment, the step of performing cross-checking and fusion correction on the first position coordinate, the second position coordinate, and the third position coordinate to obtain the feature point coordinate set of the to-be-recognized test paper includes: Step S3001: Based on the identification code layout rule, perform cross-checking on the first position coordinate, the second position coordinate, and the third position coordinate, and judge whether the relative position distances between the first position coordinate, the second position coordinate, and the third position coordinate pairwise are within the preset distance threshold. If not within the preset distance threshold, mark the abnormal position coordinates where the relative position distances exceed the preset distance threshold; Perform pairwise cross-verification on the first, second, and third position coordinates. For example, verify whether the actual distance between the main graphic feature point l0(25, 35) and the text feature point "One"(52, 37) is within the 3-pixel threshold, and verify whether the actual Y-axis distance between the auxiliary graphic h1(25, 292) and the main graphic feature point l0 conforms to the preset value of 257 pixels; If the actual relative position distance between a certain two groups of feature points exceeds the preset distance threshold, for example, the actual coordinates of the text feature point "One" are (55, 38), and the X-axis distance from the main graphic feature point l0 is 30 pixels, which exceeds the 3-pixel threshold, then mark the text feature point "One" as an abnormal position coordinate.
[0050] Step S3002: Use the weighted average method to fuse and correct the abnormal position coordinates based on the remaining valid position coordinates to determine the corrected position coordinates; for the marked abnormal coordinates, screen all valid coordinates that have a preset relative position relationship with them. For example, when the coordinates of "One" are abnormal, the associated valid coordinates include the main graphic l0(25, 35), the auxiliary graphic h1(25, 292), the text feature point "Two" (724, 37), etc.; assign weights based on the positioning reliability of the feature points. The positioning accuracy of the first position coordinates corresponding to the main graphic feature points is the highest, and the weight can be set to 0.4. The second position coordinates corresponding to the auxiliary graphic feature points are the second, and the weight can be set to 0.3. The weight of the third position coordinates corresponding to the text feature points can be set to 0.3. Those skilled in the art can adjust according to the actual scenario and no specific limitation is made here; according to the preset relative position relationship, combine the weights to inversely calculate the correction value of the abnormal coordinates. For example, the abnormal coordinates of "One" are (55, 38). Through the valid coordinates of the main graphic feature point l0(25, 35) in the upper left corner and the valid coordinates of the text feature point "Two" (724, 37), substitute them into the weighted average formula to calculate the corrected abnormal position coordinates (such as the abnormal coordinates of the text feature point "One" are corrected to (52.1, 37.2)), ensuring that the distances between the corrected abnormal position coordinates and the main graphic feature point l0 in the upper left corner and the text feature point "Two" all meet the preset threshold values.
[0051] Step S3003: Based on the corrected position coordinates and the valid position coordinates, generate the feature point coordinate set of the to-be-recognized test paper.
[0052] Integrate all valid coordinates without marked abnormalities (such as l0(25, 35), h1(25, 292), "Two" (724, 37), etc.) and the abnormal coordinates corrected in Step S3002 (such as "One" (52.1, 37.2)); classify and label the coordinates according to the feature point types (main graphic feature points, auxiliary graphic feature points, text feature points), and clarify the name of the feature point corresponding to each coordinate (such as "Main graphic feature point - l0", "Text feature point - One", etc.) to form a standardized feature point coordinate set; Step S304: According to the feature point coordinate set and the preset layout template of the test paper function area, respectively calculate the range of the fourth position coordinates corresponding to the candidate information area, the range of the fifth position coordinates corresponding to the objective question answering area, and the range of the sixth position coordinates corresponding to the subjective question answering area in the to-be-recognized test paper image.
[0053] The candidate information area includes a name area, a student ID area, etc. For the "name" area in the candidate information area, first obtain the accurate X-axis coordinate of the text feature point "One" after cross-checking and fusion correction. On this basis, offset 74 pixels to the right to obtain the starting X-axis coordinate of the "name" area; the Y-axis coordinate of the "name" area is the same as the Y-axis coordinate of "One". Combining the preset standard width (e.g., 136 pixels) and height (e.g., 17 pixels) of the "name" area, taking the determined starting X-axis coordinate and Y-axis coordinate as the reference, extend 136 pixels to the right to define the X-axis range, and extend 17 pixels downward to define the Y-axis range. Finally, the coordinate range of the "name" area is X∈[126,262], Y∈[37,54].
[0054] For the "student ID" area in the candidate information area, use the same calculation logic as the "name" area. Based on the preset mapping relationship with the corresponding feature points, determine the starting X-axis coordinate (determined based on the associated feature point coordinate and the fixed offset) and the Y-axis coordinate (the same as the Y-axis coordinate of the associated feature point). Then, combine the preset standard width (e.g., 136 pixels) and height (e.g., 17 pixels) of the "student ID" area to calculate the coordinate range of the "student ID" area as X∈[564,700], Y∈[37,54].
[0055] For the objective question answering area, based on the corrected feature point coordinate set, combine the preset template rules to determine the coordinate range. The upper boundary is the Y-axis coordinate of the text feature point "One" offset 50 pixels upward; the lower boundary is the Y-axis coordinate of the second auxiliary graphic feature point h2 offset 30 pixels downward; the left boundary is the X-axis coordinate of the main graphic feature point l0 offset 50 pixels to the right; the right boundary is the X-axis coordinate of the main graphic feature point r0 offset 50 pixels to the left. Through the conversion of the above feature point coordinates, the fifth position coordinate range of the objective question answering area is X∈[75,697], Y∈[87,520].
[0056] For the subjective question answering area, according to the preset template rules, calculate based on the corrected feature point coordinates. The upper boundary is the Y-axis coordinate of the second auxiliary graphic feature point h2 offset 30 pixels upward; the lower boundary is the Y-axis coordinate of the text feature point "Three" offset 50 pixels downward; the left and right boundaries are the same as the left and right boundaries of the objective question answering area. Through the conversion of the corresponding feature point coordinates, the sixth position coordinate range of the subjective question answering area is finally determined as X∈[75,697], Y∈[580,1018].
[0057] Step S40: Based on a preset image segmentation model, image segmentation is performed according to the fourth, fifth, and sixth position coordinate ranges to determine the first mask image corresponding to the candidate information region, the second mask image corresponding to the objective question answering region, and the third mask image corresponding to the subjective question answering region. After calculating the fourth, fifth, and sixth position coordinate ranges corresponding to the candidate information region, the objective question answering region, and the subjective question answering region in the test paper image to be identified based on the first, second, and third position coordinates, image segmentation is performed again based on the preset image segmentation model according to the fourth, fifth, and sixth position coordinate ranges to determine the first mask image corresponding to the candidate information region, the second mask image corresponding to the objective question answering region, and the third mask image corresponding to the subjective question answering region. The basic network architecture of the image segmentation model includes the U2net model.
[0058] In some embodiments, the step of performing image segmentation based on a preset image segmentation model according to the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range to determine the first mask image corresponding to the candidate information region, the second mask image corresponding to the objective question answering region, and the third mask image corresponding to the subjective question answering region includes: step S401, obtaining the fourth position coordinate range corresponding to the candidate information region, the fifth position coordinate range corresponding to the objective question answering region, and the sixth position coordinate range in the candidate information region of the image to be identified. The sixth coordinate range corresponding to the question answering area; as can be seen from step 304 above, the fourth coordinate range corresponding to the candidate information area includes the "name" area (X∈[126,262], Y∈[37,54]) and the "student ID" area (X∈[564,700], Y∈[37,54]); the fifth coordinate range corresponding to the objective question answering area is X∈[75,697], Y∈[87,520]; the sixth coordinate range corresponding to the subjective question answering area is X∈[75,697], Y∈[580,1018].
[0059] Step S402: Define the range of the candidate information area in the to-be-recognized test paper image based on the fourth position coordinate range, define the range of the objective question answering area in the to-be-recognized test paper image based on the fifth position coordinate range, and define the range of the subjective question answering area in the to-be-recognized test paper image based on the sixth position coordinate range; Based on the fourth position coordinate range corresponding to the candidate information area, respectively enclose the rectangular areas of "Name" and "Student ID" in the to-be-recognized test paper image, ensuring complete coverage of the candidate filling area, excluding feature points within the margin (such as the text feature point "One" and the text feature point "Two") and the test paper information identification code; Based on the fifth position coordinate range, enclose a rectangular range covering all objective question answering areas (such as the multiple-choice question filling area, etc.), avoiding the stem description text and margin feature points, and only focusing on the blank answering area; According to the sixth position coordinate range, enclose the subjective question answering area in the lower half of the test paper, excluding irrelevant content above and below (such as residues in the objective question area and the test question information identification code at the footer), and accurately covering the handwritten answering area.
[0060] Isolate the functional areas of each to-be-recognized test paper through coordinate enclosure, enabling subsequent segmentation to operate only on the effective range, which not only improves the segmentation efficiency (reducing invalid pixel operations) but also avoids external interference affecting the segmentation result.
[0061] Step S403: Invoke the image segmentation model trained to convergence to perform pixel-level segmentation on the defined candidate information area, objective question answering area, and subjective question answering area respectively, to exclude irrelevant interference information including question description text, header and footer, paper background, and stains outside each area; The basic network architecture of the image segmentation model includes the U2net model, and the trained-to-convergence U2net model can be invoked. This U2net model has been trained with a large number of test paper images, can accurately identify the differences between handwritten text, filling marks and printed text, background, and has strong resistance to light changes and stain interference; Perform pixel-level segmentation on the defined candidate information area, objective question answering area, and subjective question answering area respectively. For the candidate information area, filter the printed text of the "Name" and "Student ID" labels, and only retain the pixels of the handwritten or printed content filled by the candidate, excluding the paper background and slight stains; For the objective question area, accurately identify the filling pixels of pencils or signature pens, filter the paper background color and residual stem text, and ensure clear separation of the filling marks; For the subjective question area, retain all valid pixels of the candidate's handwritten answers, filter the paper fold shadows, edge stains, and residual header and footer, and achieve complete separation of the handwritten content from the background.
[0062] Step S404: Output the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answering area, and the third mask image corresponding to the subjective question answering area.
[0063] The designated candidate information area, objective question answer area, and subjective question answer area are segmented at the pixel level to eliminate irrelevant interference information outside each area, including question description text, headers and footers, paper background, and stains. Then, a first mask image corresponding to the candidate information area, a second mask image corresponding to the objective question answer area, and a third mask image corresponding to the subjective question answer area are output. The first mask image retains only the pixels of the "name" and "student ID" fields; the second mask image only shows the pixels of the marking marks, with clear outlines, facilitating rapid identification of the marking location and intensity by the grading system; and the third mask image fully retains the pixels of the handwritten answers, providing high-quality material for subsequent subjective question grading.
[0064] Step S50: Based on the first mask image, identify the candidate information of the target candidate, associate the second mask image, the third mask image, the test paper information identification code with the candidate information to determine the test paper answer data of the target candidate, and push it to the preset test paper grading system to complete the control of electronic identification and positioning of the test paper.
[0065] Based on a preset image segmentation model, image segmentation is performed according to the fourth, fifth, and sixth position coordinate ranges to determine the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answer area, and the third mask image corresponding to the subjective question answer area. Then, the candidate information of the target candidate is identified based on the first mask image. The second mask image, the third mask image, the test paper information identification code, and the candidate information are associated with the candidate information to determine the test paper answer data of the target candidate, and pushed to a preset test paper grading system to complete the control of electronic identification and positioning of the test paper.
[0066] Specifically, the first mask image is a "cleaned version" of the candidate's information area after pixel-level segmentation. It retains only the pixels of the candidate's handwritten or printed content, such as the characters for "name" and "student ID," completely eliminating interference from paper backgrounds, stains, printed labels, etc. A preset OCR model can be called to extract information from the first mask image. For example, from the mask image of the "name" area, the candidate's name (e.g., "Zhang San") can be identified; from the mask image of the "student ID" area, the candidate's student ID (e.g., "2024001") can be identified; and the candidate information containing the name and student ID of the target candidate can be output, ensuring unique and traceable identity.
[0067] Using the candidate information, which includes the candidate's name and student ID, as a unique identifier, the system associates the second mask image of the objective question answering area, the third mask image of the subjective question answering area, and the test paper information identification code to generate an electronic answer data package, which is then pushed to a preset test paper grading system to complete the control of electronic identification and positioning of the test paper.
[0068] As can be seen from the above embodiments, compared with the prior art, this application addresses the problems of existing technologies such as strong dependence on printing quality, poor resistance to severe occlusion, and susceptibility to background interference in visual positioning technology based on black squares, and the problems of high cost and space occupation, high printing and recognition thresholds, and the need to change the original layout of the test paper in positioning and recording technology based on QR codes. This application has the following beneficial effects, including but not limited to: First, this application only uses the page margins to place the identification code on the basis of the original format of the test paper, without adjusting the layout of the main text, headers and footers, or reserving extra space. Compared with the drawbacks of QR codes that require changes in layout to accommodate graphics, and the complex design of black squares that may require additional directional markings, this application is completely in line with the layout habits of traditional test papers, reducing the adaptation cost of technology implementation.
[0069] Secondly, this application adopts a feature point design that combines graphic outlines and Chinese characters. The Chinese characters have high information content and strong outline recognition. Even if the printing is slightly blurry, the color is light, or the ink spreads slightly, recognition can still be completed through the dual features of character semantics and graphic outlines, solving the problem of the black squares being highly dependent on printing quality. Thirdly, this application uses a comprehensive positioning method that combines graphic feature points and text feature points. Even if the test paper is folded or missing corners, causing two or more corner graphic feature points to fail, the missing positioning point can still be calculated through the remaining diagonal feature points and text feature points, solving the pain point of positioning failure due to black square occlusion. Fourthly, for black illustrations and table borders in the test paper background, the combination of Chinese characters and graphic feature points can effectively distinguish positioning elements from interference patterns, avoiding misidentification. At the same time, through uniform illumination processing, large model semantic correction, and OCR algorithm optimization, it can still accurately identify even in the face of uneven illumination, character adhesion or breakage, and text deformation and distortion, breaking through the limitation of traditional technology where the recognition rate drops sharply in complex environments.
[0070] Fifth, this application uses 12 Chinese characters as the test paper identification code, 1 Chinese character as the page number, and 4 Chinese characters as the check code. The high information capacity of Chinese characters ensures that the identification code of each test paper is unique, which can accurately distinguish different test papers of the same subject and version. Compared with the limitations of black squares, which can only achieve positioning and cannot carry complex information, and the problems of limited information capacity and large space occupation of QR codes, this application greatly improves the accuracy of electronic identification and positioning of test papers, and provides complete data support for subsequent test paper grading.
[0071] Sixth, it prioritizes image recognition for rapid localization, and only starts OCR text recognition when image recognition fails, avoiding the inefficiency of relying entirely on OCR and significantly improving the recognition speed of a single test paper. By pre-setting the mapping relationship between feature points and functional areas (name, student ID, answer area), the coordinates of the target area can be directly calculated after locating the identification code, without the need for manual annotation. Combined with image segmentation and standardized mask output, it realizes full automation from test paper scanning to electronic data, solving the bottleneck of low efficiency in traditional manual sorting and annotation, and can support the batch test paper processing needs in large-scale teaching scenarios.
[0072] Please refer to Figure 3. An electronic test paper recognition and positioning control device provided for one of the purposes of this application includes a test paper image acquisition module 1100, a feature point coordinate recognition module 1200, a functional area coordinate determination module 1300, a mask image segmentation module 1400, and an answer data push module 1500. The test paper image acquisition module 1100 is configured to acquire a test paper image uploaded by the target examinee, containing a target identification code. The target identification code is located in the margin area of the test paper image and includes multiple graphic feature points, multiple text feature points, and a test paper information identification code. The graphic feature points include main graphic feature points and auxiliary graphic feature points. The feature point coordinate recognition module 1200 is configured to identify the first position coordinates corresponding to each main graphic feature point and the second position coordinates corresponding to each auxiliary graphic feature point in the test paper image based on a preset graphic feature point recognition model, and determine the third position coordinates corresponding to each text feature point in the test paper image based on a preset text feature point recognition model and a large language model fusion. The functional area coordinate determination module 1300 is configured to calculate the coordinates of the functional area based on the first position coordinates, the second position coordinates, and the third position coordinates. The image to be identified includes a fourth coordinate range corresponding to the candidate information area, a fifth coordinate range corresponding to the objective question answer area, and a sixth coordinate range corresponding to the subjective question answer area. A mask image segmentation module 1400 is configured to perform image segmentation based on a preset image segmentation model according to the fourth, fifth, and sixth coordinate ranges to determine a first mask image corresponding to the candidate information area, a second mask image corresponding to the objective question answer area, and a third mask image corresponding to the subjective question answer area. An answer data push module 1500 is configured to identify the candidate information of the target candidate based on the first mask image, associate the second mask image, the third mask image, and the exam paper information identification code with the candidate information to determine the target candidate's exam paper answer data, and push it to a preset exam paper grading system to complete the control of electronic identification and positioning of the exam paper.
[0073] Based on any embodiment of this application, referring to Figure 4, another embodiment of this application also provides an electronic device, which can be implemented by a computer device. As shown in Figure 4, this is a schematic diagram of the internal structure of the computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database may store a sequence of control information. When the computer-readable instructions are executed by the processor, the processor can implement a method for electronic identification and positioning control of examination papers. The processor of the computer device provides computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the method for electronic identification and positioning control of examination papers of this application. The network interface of the computer device is used for communication with a terminal. Those skilled in the art will understand that the structure shown in Figure 4 is merely a block diagram of a portion of the structure related to the solution of this application and does not constitute a limitation on the computer device to which the solution of this application is applied. Specific computer devices may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0074] In this embodiment, the processor executes the specific functions of each module in Figure 3, and the memory stores the program code and various types of data required to execute the above modules. The network interface is used for data transmission between the user terminal and the server. In this embodiment, the memory stores the program code and data required to execute all modules in the electronic test paper recognition and positioning control device of this application, and the server can call the server's program code and data to execute the functions of all modules.
[0075] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the electronic test paper identification and positioning control method described in any embodiment of this application.
[0076] This application also provides a computer program product, including a computer program / instructions, which, when executed by one or more processors, implement the steps of the electronic test paper identification and positioning control method described in any embodiment of this application.
[0077] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0078] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for electronic identification and positioning control of examination papers, characterized in that, include: The process involves acquiring an image of a test paper uploaded by a target examinee, containing a target identification code. The target identification code is located in the margin area of the test paper image and includes multiple graphic feature points, multiple text feature points, and a test paper information identification code. The graphic feature points include main graphic feature points and auxiliary graphic feature points. Based on a preset graphic feature point recognition model, the first position coordinates corresponding to each main graphic feature point and the second position coordinates corresponding to each auxiliary graphic feature point in the test paper image are identified. Based on a preset text feature point recognition model and a large language model fusion, the third position coordinates corresponding to each text feature point in the test paper image are determined. The examinee in the test paper image is then deduced based on the first, second, and third position coordinates. The system defines a fourth coordinate range corresponding to the information area, a fifth coordinate range corresponding to the objective question answering area, and a sixth coordinate range corresponding to the subjective question answering area. Based on a preset image segmentation model, image segmentation is performed according to these coordinate ranges to determine a first mask image corresponding to the candidate information area, a second mask image corresponding to the objective question answering area, and a third mask image corresponding to the subjective question answering area. Based on the first mask image, the candidate information of the target candidate is identified. The second mask image, the third mask image, and the test paper information identification code are associated with the candidate information to determine the target candidate's test paper answer data, which is then pushed to a preset test paper grading system to complete the electronic identification and positioning control of the test paper.
2. The electronic identification and positioning control method for examination papers according to claim 1, characterized in that, The steps of identifying the first position coordinates of each main graphic feature point and the second position coordinates of each auxiliary graphic feature point in the image of the test paper to be identified based on a preset graphic feature point recognition model include: extracting graphic contour information in the image of the test paper to be identified that conforms to a preset size and preset contour shape based on the preset graphic feature point recognition model, so as to match and locate the main graphic feature points at the four corners of the test paper to be identified, as well as multiple auxiliary graphic feature points on its sides; and calculating and outputting the first position coordinates of each main graphic feature point and the second position coordinates of each auxiliary graphic feature point based on the image pixel coordinate system of the test paper to be identified.
3. The electronic identification and positioning control method for examination papers according to claim 2, characterized in that, The step of determining the third position coordinates corresponding to each text feature point in the test paper image to be recognized based on a preset text feature point recognition model and a large language model fusion includes: obtaining text feature points located adjacent to the main graphic feature points; matching candidate regions of the corresponding text feature points for each main graphic feature point based on the first position coordinates corresponding to the main graphic feature points at the four corners of the test paper to be recognized; calling the preset text feature point recognition model to perform character detection on the candidate regions respectively, so as to output the minimum bounding rectangle bounding box of the characters in the candidate regions and the preliminary recognized text; and inputting the association relationship between the preliminary recognized text and the main graphic feature points corresponding to the candidate regions into the positioning marker constructed by the large language model. The semantic verification library uses a multi-layered mechanism, including distorted character restoration, similar-looking character and background filtering, and contextual logic verification, to filter out valid text feature points that conform to preset positioning identifiers. Using the image pixel coordinate system of the test paper image to be recognized as a reference, the center coordinates of the minimum bounding rectangle of the valid text feature points are calculated as the initial position coordinates of each valid text feature point. Based on the first position coordinates corresponding to the main graphic feature points and the theoretical distance between their corresponding text feature points, the theoretical position coordinates of the text feature points are calculated and determined. If the relative position distance between the initial position coordinates and the theoretical position coordinates is less than a preset distance threshold, the initial position coordinates of the valid text feature points are output as the third position coordinates.
4. The electronic identification and positioning control method for examination papers according to claims 2 to 3, characterized in that, The steps of calculating the range of fourth position coordinates corresponding to the candidate information area, the range of fifth position coordinates corresponding to the objective question answering area, and the range of sixth position coordinates corresponding to the subjective question answering area in the candidate's information area of the candidate's information ... The system establishes a preset relative positional relationship between each pair of feature points, and a mapping relationship between the main graphic feature points, the auxiliary graphic feature points, the text feature points, and each functional area of the test paper to be identified. It then performs cross-validation and fusion correction on the first, second, and third position coordinates to obtain the feature point coordinate set of the test paper to be identified. Based on the feature point coordinate set and a preset test paper functional area layout template, it calculates the fourth position coordinate range corresponding to the candidate information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the image of the test paper to be identified.
5. The electronic identification and positioning control method for examination papers according to claim 1, characterized in that, The step of performing cross-validation and fusion correction on the first position coordinates, the second position coordinates, and the third position coordinates to obtain the feature point coordinate set of the test paper to be identified includes: performing cross-validation on the first position coordinates, the second position coordinates, and the third position coordinates based on the identification code layout rules, determining whether the relative position distance between each pair of the first position coordinates, the second position coordinates, and the third position coordinates is within a preset distance threshold; if not within the preset distance threshold, marking the abnormal position coordinates whose relative position distance exceeds the preset distance threshold; using a weighted average method to fuse and correct the abnormal position coordinates based on the remaining valid position coordinates to determine the corrected position coordinates; and generating the feature point coordinate set of the test paper to be identified based on the corrected position coordinates and the valid position coordinates.
6. The electronic identification and positioning control method for examination papers according to claim 1, characterized in that, The steps of segmenting the image based on a preset image segmentation model according to the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range to determine the first mask image corresponding to the candidate information region, the second mask image corresponding to the objective question answering region, and the third mask image corresponding to the subjective question answering region include: obtaining the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range corresponding to the candidate information region in the image to be identified; and delineating the candidate information region in the image to be identified based on the fourth position coordinate range. The objective question answering area in the test paper image to be identified is defined based on the fifth position coordinate range, and the subjective question answering area in the test paper image to be identified is defined based on the sixth position coordinate range. The trained and converged image segmentation model is called to perform pixel-level segmentation on the defined candidate information area, objective question answering area, and subjective question answering area to exclude irrelevant interference information outside each area, including question description text, headers and footers, paper background, and stains. The first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answering area, and the third mask image corresponding to the subjective question answering area are output.
7. The electronic identification and positioning control method for examination papers according to any one of claims 1 to 6, characterized in that, The main graphic feature points are located at the four corners of the test paper image to be identified, the auxiliary graphic feature points are located on the sides of the test paper image to be identified, and the text feature points are located adjacent to the main graphic feature points; the test paper information identification code includes the subject, grade, exam session, and question type distribution of the test paper image to be identified; the basic network architecture of the graphic feature point recognition model includes a lightweight convolutional neural network; the basic network architecture of the text feature point recognition model includes an OCR model; and the basic network architecture of the image segmentation model includes a U2net model.
8. A test paper electronic identification and positioning control device, characterized in that, The system includes: a test paper image acquisition module, configured to acquire a test paper image uploaded by a target examinee containing a target identification code, wherein the target identification code is located in the margin area of the test paper image and includes multiple graphic feature points, multiple text feature points, and a test paper information identification code; the graphic feature points include main graphic feature points and auxiliary graphic feature points; a feature point coordinate recognition module, configured to identify the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the test paper image based on a preset graphic feature point recognition model, and determine the third position coordinates corresponding to each of the text feature points in the test paper image based on a preset text feature point recognition model and a large language model fusion; and a functional area coordinate determination module, configured to calculate the coordinates of the test paper image to be recognized based on the first position coordinates, the second position coordinates, and the third position coordinates. The system identifies the fourth coordinate range corresponding to the candidate information area, the fifth coordinate range corresponding to the objective question answer area, and the sixth coordinate range corresponding to the subjective question answer area in the test paper image. A mask image segmentation module is configured to perform image segmentation based on a preset image segmentation model according to the fourth, fifth, and sixth coordinate ranges to determine the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answer area, and the third mask image corresponding to the subjective question answer area. An answer data push module is configured to identify the candidate information of the target candidate based on the first mask image, associate the second mask image, the third mask image, and the test paper information identification code with the candidate information to determine the target candidate's test paper answer data, and push it to a preset test paper grading system to complete the control of electronic identification and positioning of the test paper.
9. An electronic device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 7, which, when invoked by a computer, executes the steps included in the corresponding method.