Methods, devices, equipment and media for electronic identification and positioning control of exam papers

CN121963243BActive Publication Date: 2026-09-01GUANGDONG MOKEN EDUCATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202512039697.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-09-01
Estimated Expiration
2045-12-31

AI Technical Summary

Technical Problem

2.1、成本与空间占用大:二维码尺寸通常大于普通黑色方块(最小尺寸约 1cm×1cm),对排版空间紧张的场景(如窄边距试卷)适应性较差;

Benefits of technology

其一、本申请在试卷原有格式基础上,仅利用页边距布设识别码,无需调整试卷正文布局、页眉页脚或预留额外空间。相较于二维码需要改变排版以容纳图形的弊端,以及黑色方块可能需额外添加方向标记的复杂设计,本申请完全契合传统试卷的排版习惯,降低了技术落地的适配成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963243B_ABST
    Figure CN121963243B_ABST
Patent Text Reader

Abstract

This application relates to a method, device, equipment, and medium for electronic identification and positioning control of examination papers. The method includes: identifying the first position coordinates of each main graphic feature point and the second position coordinates of each auxiliary graphic feature point in the examination paper image to be identified based on a preset graphic feature point recognition model; determining the third position coordinates of each text feature point in the examination paper image to be identified based on a preset text feature point recognition model and a large language model fusion; calculating the position coordinate range of the candidate information area, objective question answer area, and subjective question answer area in the examination paper image to be identified based on the first, second, and third position coordinates; performing image segmentation based on the position coordinate range to determine the mask images of the candidate information area, objective question answer area, and subjective question answer area, and pushing them to a preset examination paper grading system in association with the candidate information; this application greatly improves the accuracy of electronic identification and positioning of examination papers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic examination papers, and in particular to an electronic examination paper identification and positioning control method, a corresponding device, an electronic device, and a computer-readable storage medium. Background Technology

[0002] In an era of deep integration between educational informatization and artificial intelligence, the education industry is accelerating its transformation from traditional paper-based media to digital and intelligent models. The demand for rapid processing and accurate analysis of massive amounts of paper-based test papers and assignments is becoming increasingly prominent. Traditional methods of manual sorting and annotation for assignment grading and test paper management have drawbacks such as low efficiency, high error rates, and difficulty in providing personalized feedback, making it difficult to meet the dual requirements of teaching quality and efficiency in large-scale teaching scenarios.

[0003] Currently, the core supporting technologies for electronic processing of exam papers are visual positioning technology based on black squares and positioning and recording technology based on QR codes. However, these technologies have the following technical shortcomings in practical applications: 1. The visual positioning technology based on black squares has the following technical defects, including: 1.1 High dependence on printing quality: If the black square is printed blurry, has ink leakage, or is too light in color, it may lead to recognition failure; ink diffusion (such as inferior ink) will change the shape of the square and affect edge detection; 1.2 Poor resistance to occlusion: If two or more blocks are completely occluded (e.g., the exam paper is folded or damaged), coordinate mapping cannot be established, resulting in positioning failure; 1.3. Susceptible to background interference: If the background of the test paper contains non-text black patterns (such as illustrations or table borders), they may be misidentified as positioning points; uneven lighting during scanning (such as local shadows) will reduce contrast and cause incorrect recognition of squares.

[0004] 1.4 Limitation of directional uniqueness: In some scenarios (such as using only two diagonal blocks), there may be misjudgments of mirror flips (such as upside down), and additional markers (such as arrows) are required to assist in positioning.

[0005] 2. The location recording technology based on QR codes has the following technical defects, including: 2.1 High cost and space occupation: QR codes are usually larger than ordinary black squares (the smallest size is about 1cm×1cm), which makes them less suitable for scenarios with tight layout space (such as exam papers with narrow margins); 2.2 High barriers to entry in printing and recognition: QR codes have higher requirements for printing accuracy (module spacing error must be ≤10%), and poor printing may lead to decoding failure; ordinary scanners need to support QR code decoding function; 2.3 The layout of the original test paper needs to be changed.

[0006] In summary, existing technologies such as black square-based visual positioning technology suffer from issues such as strong dependence on printing quality, poor resistance to severe occlusion, and susceptibility to background interference. In contrast, QR code-based positioning and recording technology suffers from problems such as high cost and space occupation, high printing and recognition thresholds, and the need to change the layout of the original test paper. The applicant has made corresponding explorations to address these issues. Summary of the Invention

[0007] The purpose of this application is to solve the above-mentioned problems by providing a method for electronic identification and positioning control of test papers, a corresponding device, electronic equipment and computer-readable storage medium.

[0008] To achieve the various objectives of this application, the following technical solution is adopted: A method for electronic identification and positioning control of examination papers, proposed to meet one of the purposes of this application, includes: The method involves acquiring an image of a test paper uploaded by a target candidate, which contains a target identification code. The target identification code is located in the margin area of ​​the test paper image and includes multiple graphic feature points, multiple text feature points, and a test paper information identification code. The graphic feature points include main graphic feature points and auxiliary graphic feature points. Based on a preset graphic feature point recognition model, the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the test paper image to be recognized are identified. Based on a preset text feature point recognition model and a large language model fusion, the third position coordinates corresponding to each of the text feature points in the test paper image to be recognized are determined. Based on the first position coordinates, the second position coordinates, and the third position coordinates, the fourth position coordinate range corresponding to the candidate information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the candidate's test paper image to be identified are calculated. Based on a preset image segmentation model, image segmentation is performed according to the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range to determine the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answering area, and the third mask image corresponding to the subjective question answering area. Based on the first mask image, the candidate information of the target candidate is identified. The second mask image, the third mask image, the test paper information identification code and the candidate information are associated to determine the test paper answer data of the target candidate and pushed to the preset test paper grading system to complete the control of electronic identification and positioning of the test paper.

[0009] Optionally, the step of identifying the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the image of the test paper to be identified based on a preset graphic feature point recognition model includes: Based on the preset graphic feature point recognition model, the graphic contour information that conforms to the preset size and preset contour shape in the image of the test paper to be recognized is extracted, so as to match and locate the main graphic feature points at the four corners of the test paper to be recognized, as well as multiple auxiliary graphic feature points on its sides. Based on the image pixel coordinate system of the test paper to be identified, the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points are calculated and output respectively.

[0010] Optionally, the step of determining the third position coordinates corresponding to each text feature point in the test paper image to be recognized based on a preset text feature point recognition model and a large language model fusion includes: Obtain text feature points located adjacent to the main graphic feature points, and based on the first position coordinates corresponding to the main graphic feature points at the four corners of the test paper to be identified, match the candidate regions of the corresponding text feature points for each main graphic feature point. The preset text feature point recognition model is invoked to perform character detection on the candidate regions respectively, so as to output the minimum bounding rectangle of the characters in the candidate regions and the preliminary recognized text. The association between the preliminarily identified text and the main graphic feature points corresponding to the candidate region is input into the positioning identifier semantic verification library constructed by the large language model. The effective text feature points that meet the preset positioning identifier are filtered out through a multi-layer mechanism including distorted character restoration, similar character and background filtering, and context logic verification. Using the image pixel coordinate system of the test paper image to be recognized as a reference, the center coordinates of the minimum bounding rectangle of the effective text feature points are calculated as the preliminary position coordinates of each effective text feature point. Based on the first position coordinates corresponding to the main graphic feature points and the theoretical spacing between their corresponding text feature points, the theoretical position coordinates of the text feature points are calculated and determined. If the relative position distance between the preliminary position coordinates and the theoretical position coordinates is less than a preset distance threshold, the preliminary position coordinates of the effective text feature points are output as the third position coordinates. Optionally, the step of calculating the fourth position coordinate range corresponding to the candidate information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the test paper image to be recognized based on the first position coordinates, the second position coordinates, and the third position coordinates includes: Obtain the first position coordinates of the main graphic feature points, the second position coordinates of the auxiliary graphic feature points, and the third position coordinates of the text feature points in the test paper image to be identified; Based on the preset identification code layout rules, the preset relative positional relationships between the main graphic feature points, the auxiliary graphic feature points and the text feature points in the image of the test paper to be identified are determined, as well as the mapping relationship between the main graphic feature points, the auxiliary graphic feature points, the text feature points and each functional area of ​​the test paper to be identified. Cross-validation and fusion correction are performed on the first position coordinates, the second position coordinates, and the third position coordinates to obtain the feature point coordinate set of the test paper to be identified; Based on the feature point coordinate set and the preset test paper functional area layout template, the coordinate range of the fourth position corresponding to the candidate information area, the coordinate range of the fifth position corresponding to the objective question answering area, and the coordinate range of the sixth position corresponding to the subjective question answering area in the test paper image to be identified are calculated respectively.

[0011] Optionally, the step of performing cross-validation and fusion correction on the first position coordinates, the second position coordinates, and the third position coordinates to obtain the feature point coordinate set of the test paper to be identified includes: Based on the identification code layout rules, cross-validation is performed on the first position coordinates, the second position coordinates, and the third position coordinates to determine whether the relative position distance between each pair of the first position coordinates, the second position coordinates, and the third position coordinates is within the preset distance threshold. If it is not within the preset distance threshold, abnormal position coordinates whose relative position distance exceeds the preset distance threshold are marked. The weighted average method is used to fuse and correct the abnormal position coordinates based on the remaining valid position coordinates to determine the corrected position coordinates. Based on the corrected position coordinates and the valid position coordinates, a set of feature point coordinates for the test paper to be identified is generated.

[0012] Optionally, the step of performing image segmentation based on a preset image segmentation model according to the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range to determine the first mask image corresponding to the candidate information region, the second mask image corresponding to the objective question answering region, and the third mask image corresponding to the subjective question answering region includes: Obtain the fourth position coordinate range corresponding to the candidate information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the image of the test paper to be identified; The candidate information area in the test paper image to be identified is defined based on the fourth position coordinate range, the objective question answering area in the test paper image to be identified is defined based on the fifth position coordinate range, and the subjective question answering area in the test paper image to be identified is defined based on the sixth position coordinate range. The trained and converged image segmentation model is invoked to perform pixel-level segmentation on the defined candidate information region, objective question answer region, and subjective question answer region, so as to exclude irrelevant interference information outside each region, including question description text, headers and footers, paper background, and stains. Output the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answer area, and the third mask image corresponding to the subjective question answer area.

[0013] Optionally, the main graphic feature points are arranged at the four corners of the test paper image to be identified, the auxiliary graphic feature points are arranged on the sides of the test paper image to be identified, and the text feature points are arranged at the adjacent positions of the main graphic feature points. The test paper information identification code includes the subject, grade, exam session, and question type distribution of the test paper image to be identified; The basic network architecture of the graphic feature point recognition model includes a lightweight convolutional neural network; the basic network architecture of the text feature point recognition model includes an OCR model; and the basic network architecture of the image segmentation model includes a U2net model.

[0014] A test paper electronic identification and positioning control device provided for another purpose of this application includes: The test paper image acquisition module is configured to acquire the test paper image uploaded by the target examinee, which contains the target identification code. The target identification code is located in the margin area of ​​the test paper image and includes multiple graphic feature points, multiple text feature points, and test paper information identification code. The graphic feature points include main graphic feature points and auxiliary graphic feature points. The feature point coordinate recognition module is configured to identify the first position coordinates of each main graphic feature point and the second position coordinates of each auxiliary graphic feature point in the test paper image to be recognized based on a preset graphic feature point recognition model, and determine the third position coordinates of each text feature point in the test paper image to be recognized based on a preset text feature point recognition model and a large language model fusion. The functional area coordinate determination module is configured to calculate, based on the first position coordinate, the second position coordinate, and the third position coordinate, the fourth position coordinate range corresponding to the candidate information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the candidate's information area of ​​the test paper image to be identified; The mask image segmentation module is configured to perform image segmentation based on a preset image segmentation model according to the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range, so as to determine the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answering area, and the third mask image corresponding to the subjective question answering area. The answer data push module is configured to identify the candidate information of the target candidate based on the first mask image, associate the second mask image, the third mask image, the test paper information identification code with the candidate information to determine the test paper answer data of the target candidate, and push it to the preset test paper grading system to complete the control of electronic identification and positioning of the test paper.

[0015] An electronic device provided for another purpose of this application includes a central processing unit and a memory, the central processing unit being configured to invoke and run a computer program stored in the memory to perform the steps of the electronic identification and positioning control method for examination papers described in this application.

[0016] A computer-readable storage medium is provided for another purpose of this application, which stores, in the form of computer-readable instructions, a computer program implemented according to the electronic identification and positioning control method for the test paper, which, when called by a computer, executes the steps included in the corresponding method.

[0017] Compared to existing technologies, this application addresses the shortcomings of existing visual positioning technologies based on black squares, such as strong dependence on print quality, poor resistance to severe occlusion, and susceptibility to background interference. It also addresses the problems of QR code-based positioning and recording technologies, such as high cost and space requirements, high printing and recognition barriers, and the need to modify the layout of existing exam papers. This application offers the following advantages, including but not limited to: Firstly, this application, based on the original format of the exam paper, only utilizes the page margins to place the identification code, without requiring adjustments to the main text layout, headers and footers, or the provision of additional space. Compared to the drawbacks of QR codes requiring changes to the layout to accommodate graphics, and the complex design that may require additional directional markings for black squares, this application perfectly conforms to the traditional exam paper layout habits, reducing the adaptation costs for technology implementation.

[0018] Secondly, this application adopts a feature point design that combines graphic outlines and Chinese characters. The Chinese characters have high information content and strong outline recognition. Even if the printing is slightly blurry, the color is light, or the ink spreads slightly, it can still be recognized through the dual features of character semantics and graphic outlines, thus solving the problem of the strong dependence of black squares on printing quality. Thirdly, this application uses a combination of graphic and text feature points for positioning. Even if the test paper is folded or missing corners, causing two or more corner graphic feature points to fail, the missing positioning point can still be calculated through the remaining diagonal feature points and text feature points, thus solving the problem of positioning failure due to black square obstruction. Fourth, this application effectively distinguishes between positioning elements and interference patterns in the black illustrations and table borders in the background of the test paper, and the combination of Chinese characters and graphic feature points, thus avoiding misidentification. At the same time, through uniform illumination processing, large model semantic correction, and OCR algorithm optimization, it can still accurately identify even in the face of uneven illumination, character adhesion or breakage, and text deformation and distortion, thus breaking through the limitation of traditional technology in the sharp drop in recognition rate in complex environments.

[0019] Fifth, this application uses 12 Chinese characters as the test paper identification code, 1 Chinese character as the page number, and 4 Chinese characters as the check code. The high information capacity of Chinese characters ensures that the identification code of each test paper is unique, which can accurately distinguish different test papers of the same subject and version. Compared with the limitations of black squares, which can only achieve positioning and cannot carry complex information, and the problems of limited information capacity and large space occupation of QR codes, this application greatly improves the accuracy of electronic identification and positioning of test papers, and provides complete data support for subsequent test paper grading.

[0020] Sixth, it prioritizes image recognition for rapid localization, and only starts OCR text recognition when image recognition fails, avoiding the inefficiency of relying entirely on OCR and significantly improving the recognition speed of a single test paper. By pre-setting the mapping relationship between feature points and functional areas (name, student ID, answer area), the coordinates of the target area can be directly calculated after locating the identification code, without the need for manual annotation. Combined with image segmentation and standardized mask output, it realizes full automation from test paper scanning to electronic data, solving the bottleneck of low efficiency in traditional manual sorting and annotation, and can support the batch test paper processing needs in large-scale teaching scenarios. Attached Figure Description

[0021] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart illustrating the electronic identification and positioning control method for exam papers in this application. Figure 2 This is a schematic diagram of the test paper image to be identified in an embodiment of this application; Figure 3 This is a schematic diagram of the electronic identification and positioning control device for exam papers in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the computer device in the embodiments of this application. Detailed Implementation

[0022] Those skilled in the art will understand that although the various methods in this application are described based on the same concept and thus present commonality among them, they can be performed independently unless otherwise specified. Similarly, the various embodiments disclosed in this application are all based on the same inventive concept; therefore, concepts expressed in the same way, as well as concepts that are appropriately changed for convenience but are expressed differently, should be understood equivalently.

[0023] Unless otherwise expressly stated, the various embodiments disclosed in this application can be combined in a cross-cutting manner to flexibly construct new embodiments, as long as such combination does not depart from the inventive spirit of this application and can meet the needs of the prior art or solve a certain deficiency in the prior art. Those skilled in the art should be aware of such modifications.

[0024] Please see Figure 1 and Figure 2 In one embodiment of the electronic test paper identification and positioning control method of this application, the method includes: Step S10: Obtain the image of the test paper to be identified uploaded by the target candidate, which contains the target identification code. The target identification code is placed in the margin area of ​​the test paper image to be identified. It includes multiple graphic feature points, multiple text feature points and test paper information identification code. The graphic feature points include main graphic feature points and auxiliary graphic feature points. The electronic test paper recognition and positioning control system in the terminal device acquires an image of a test paper to be recognized, which contains a target identification code. The target identification code is located in the margin area of ​​the test paper image and includes multiple graphic feature points, multiple text feature points, and a test paper information identification code. The graphic feature points include main graphic feature points and auxiliary graphic feature points. The main graphic feature points are located at the four corners of the test paper image, the auxiliary graphic feature points are located on the sides of the test paper image, and the text feature points are located adjacent to the main graphic feature points. The test paper information identification code includes the subject, grade, exam session, and question type distribution of the test paper image. In some embodiments, the main graphic feature points are located at the four corners of the test paper image to be recognized, the auxiliary graphic feature points are located on the sides of the test paper image to be recognized, and the text feature points are located adjacent to the main graphic feature points; the test paper information identification code includes the subject, grade, exam session, and question type distribution of the test paper image to be recognized; the basic network architecture of the graphic feature point recognition model includes a lightweight convolutional neural network; the basic network architecture of the text feature point recognition model includes an OCR model; and the basic network architecture of the image segmentation model includes a U2net model.

[0025] In some embodiments, when producing the test paper, a margin of not less than 2 cm is preset around the test paper, and the target identification codes are only arranged in this area, which does not occupy core areas such as test questions, answer spaces and the like, so as to avoid being blocked by the examinee's answer content; and it is not necessary to adjust the original test paper formats such as question layout, font size, spacing, etc. Graphic feature points are the core anchor points for positioning the image of the test paper to be recognized. The spatial coordinate system of the test paper to be recognized is quickly established through geometric contour recognition. The main graphic feature points can be arranged at the four corners of the test paper image to be recognized, and their size is uniformly 21×21 pixels. For example, the position coordinate of the main graphic feature point l0 at the top-left corner is (25, 35); the position coordinate of the main graphic feature point r1 at the bottom-right corner is (747, 1066); as the reference vertices for establishing the global pixel coordinate system of the test paper, the main graphic feature points can determine the overall boundary and orientation of the test paper, so as to avoid misjudgments such as upside-down.

[0026] The auxiliary graphic feature points are arranged on the side edges of the image of the test paper to be recognized, and the text feature points are arranged at adjacent positions of the main graphic feature points, with a uniform size of 21×21 pixels. For example, the position coordinate of the first auxiliary graphic feature point h1 is (25, 292); the position coordinate of the second auxiliary graphic feature point h2 is (25, 550); the position coordinate of the third auxiliary graphic feature point h3 is (25, 808), etc.; they can calibrate the tilt and distortion (such as perspective distortion) of the captured test paper; and complete the missing main graphic feature points. For example, when a main graphic at a corner cannot be recognized due to a missing corner of the test paper, it can be calculated through the auxiliary graphic feature points.

[0027] The text feature points are arranged at adjacent positions of the main graphic feature points. Chinese characters have much higher information content than letters or numbers, with a high upper limit of precision, and have high distinguishing degree from the text and patterns of the test paper background, so they are not easy to be misrecognized, and can solve the problem that black squares are easily interfered by the background in the prior art; text feature points serve as supplementary positions for graphic feature points. When the main graphic feature point at a certain corner is blocked, the area can be positioned by identifying the adjacent text feature points. Therefore, the text feature points are arranged at adjacent positions of the main graphic feature points. For example, the text feature point "Yi (One)" is on the right side of the main graphic feature point l0 at the top-left corner, the text feature point "Er (Two)" is on the left side of the main graphic feature point r0 at the top-right corner, the text feature point "San (Three)" is on the right side of the main graphic feature point l1 at the bottom-left corner, and the text feature point "Si (Four)" is on the left side of the main graphic feature point r1 at the bottom-right corner; the size can be 17×17 pixels. For example, the position coordinate of the text feature point "Yi (One)" is (52, 37); the position coordinate of the text feature point "Si (Four)" is (724, 1068), etc. The test paper information identification code includes the subject, grade, exam session, and question type distribution of the test paper image to be identified. For example, the subject is mathematics, Chinese, or English, and the grade is junior high school grade 7 or senior high school grade 9. The exam session includes mid-term exams, final exams, or mock exams, etc. The question type distribution includes the question type and the location range of each question type. The question types include subjective questions and objective questions, etc. Objective questions include multiple choice questions and true / false questions, etc.; subjective questions include fill-in-the-blank questions or short answer questions, etc. The test paper information identification code can be constructed from 12 Chinese characters (recording core information), 1 Chinese character (page number), and 4 Chinese characters (check code), and one set can be placed at the top and bottom of the test paper.

[0028] The target candidate can scan or take a photo with their mobile phone to upload an image of the test paper containing the target identification code. If the image of the test paper does not contain the complete target identification code, the electronic test paper identification and positioning control system will prompt "Image invalid, please re-upload an image of the test paper containing the complete page margins" to ensure that the subsequent positioning process can start normally.

[0029] In some embodiments, barcodes are used to assist in recording test paper information. When the test paper information identification code fails to be recognized, the test paper content can also be identified by recognizing the barcode, thereby improving the recognition rate.

[0030] Step S20: Based on the preset graphic feature point recognition model, identify the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the test paper image to be recognized; and based on the preset text feature point recognition model and the fusion of the large language model, determine the third position coordinates corresponding to each of the text feature points in the test paper image to be recognized. After obtaining the image of the test paper to be identified, which contains the target identification code uploaded by the target candidate, the system identifies the first position coordinates of each main graphic feature point and the second position coordinates of each auxiliary graphic feature point in the image of the test paper based on a preset graphic feature point recognition model. The system then determines the third position coordinates of each text feature point in the image of the test paper based on a preset text feature point recognition model and a large language model fusion. The basic network architecture of the graphic feature point recognition model includes a lightweight convolutional neural network, and the basic network architecture of the text feature point recognition model includes an OCR model.

[0031] In some embodiments, Convolutional Neural Networks (CNNs) extract abstract features through multiple convolutions. Even if the printed image is blurry or the ink has slightly diffused, such as shape deviations caused by inferior ink, it can still be identified through contour feature matching, reducing the dependence on printing accuracy and solving the drawback of "printing quality sensitivity" of black squares in the prior art. CNNs, through pooling layers and attention mechanisms, can distinguish between black patterns (illustrations, table borders) and positioned graphics in the background of the exam paper, avoiding misidentification and solving the problem of black squares being easily affected by background interference in the prior art. CNNs are robust to slight wrinkles and tilts in graphics, and can restore the essence of the graphics through feature normalization processing, avoiding directional misjudgment without additional markings (such as arrows), thus solving the limitation of directional uniqueness of black squares in the prior art. Therefore, a lightweight convolutional neural network is used as the graphic feature point recognition model in this application.

[0032] In some embodiments, the step of identifying the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the image of the test paper to be identified based on a preset graphic feature point recognition model includes: Step S201: Based on the preset graphic feature point recognition model, extract the graphic contour information in the test paper image to be recognized that conforms to the preset size and preset contour shape, so as to match and locate the main graphic feature points at the four corners of the test paper to be recognized, as well as multiple auxiliary graphic feature points on its sides. Specifically, a preset graphic feature point recognition model is invoked. This model is constructed based on a lightweight convolutional neural network. Based on this model, graphic contour information conforming to preset sizes and shapes is extracted from the test paper image to be recognized. The graphic contour information is then matched and located using a dual-filtering rule. First, only graphics conforming to the preset size are retained. Since the preset size of the main graphic feature point is 21×21 pixels, and the preset size of the auxiliary graphic feature point is also 21×21 pixels, black patterns in the test paper background that do not match the size (such as illustrations, table borders, etc.) are excluded, thus solving the drawback of existing black square positioning being easily affected by background interference. Then, only graphics with contour shapes consistent with the preset contour shape are retained. The contours of the main graphic feature point and the auxiliary graphic feature point are regular geometric shapes, such as squares and triangles. The lightweight convolutional neural network extracts abstract contour features through multiple convolutions. Even if the graphic has slight printing blur or ink diffusion (caused by inferior ink), it can still be recognized through contour matching, reducing the dependence on printing accuracy and solving the drawback of existing black square positioning being highly dependent on printing quality.

[0033] Furthermore, by extracting the graphic contour information from the image of the test paper to be identified that conforms to a preset size and preset contour shape, the main graphic feature points at the four corners of the test paper to be identified can be accurately matched, such as the main graphic feature point l0 at the top left corner, the main graphic feature point l1 at the bottom left corner, the main graphic feature point r0 at the top right corner, and the main graphic feature point r1 at the bottom left corner, as well as multiple auxiliary graphic feature points on the side of the test paper to be identified. For example, the first auxiliary graphic feature point h1, the second auxiliary graphic feature point h2, and the third auxiliary graphic feature point h3 are located on the left side of the test paper to be identified.

[0034] By employing dual screening based on size and outline, misidentification of irrelevant black patterns in the exam paper background is completely eliminated, demonstrating stronger anti-interference capabilities than traditional black square positioning. The lightweight CNN is robust to minor printing deviations and slight graphic deformations (such as incomplete outlines caused by wrinkles), allowing for identification without requiring strict printing precision.

[0035] Step S202: Using the image pixel coordinate system of the test paper to be identified as a reference, calculate and output the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points.

[0036] Using the image pixel coordinate system of the test paper to be identified as a reference, this image pixel coordinate system has the upper left corner of the image as the origin, the horizontal direction to the right as the X-axis, and the vertical direction downward as the Y-axis, with the coordinate unit being pixels. For each located graphic feature point, its corresponding feature center coordinates are calculated. The center coordinates of the smallest bounding rectangle of the graphic or the geometric center coordinates of the contour can be taken to ensure that the coordinates can accurately represent the position of the graphic in the image. The first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points are calculated and output respectively. The first position coordinates corresponding to each of the main graphic feature points include the position coordinates of the upper left main graphic feature point l0 as (25, 35) and the coordinates of the lower right main graphic r1 as (747, 1066), etc. The second position coordinates corresponding to each auxiliary graphic feature point include the position coordinates of the first auxiliary graphic feature point h1 on the left side of the test paper as (25, 292), the position coordinates of the second auxiliary graphic feature point h2 as (25, 550), and the position coordinates of the third auxiliary graphic feature point h3 as (25, 808), etc.

[0037] It can be known from the above steps S201 to S202 that converting visually graphic feature point positions into computable coordinate data provides a core basis for subsequent steps to calculate the examinee information area, answer area coordinates, etc. A global coordinate system of the test paper is established through the coordinates of main graphic feature points and auxiliary graphic feature points, which ensures the accuracy of functional area calculation. Even if the test paper is slightly tilted, the calculated coordinates can still reflect the relative positions of the graphic feature points, and tilt calibration can be performed subsequently through the coordinates of the auxiliary graphic feature points, maintaining a high recognition rate under extreme conditions such as folded missing corners and complex illumination.

[0038] In a further embodiment, the step of determining the third position coordinates corresponding to each text feature point in the test paper image to be recognized based on the fusion of a preset text feature point recognition model and a large language model comprises: Step S2001, acquiring text feature points arranged at adjacent positions of the main graphic feature points, and matching candidate regions of the corresponding text feature points for each main graphic feature point based on the first position coordinates corresponding to the main graphic feature points at the four corners of the test paper to be recognized; It avoids low efficiency and false recognition caused by full-domain scanning of the OCR model, solves the problem that traditional text positioning is easily interfered by background text, and loads the relative positions between main graphic feature points and text feature points. For example, the text feature point "Yi" is offset by 27 pixels on the horizontal axis (X axis) and 2 pixels on the vertical axis (Y axis) relative to the main graphic feature point l0, with a standard size of 17×17 pixels; for matching candidate regions for each main graphic feature point, taking the position coordinate of the main graphic feature point l0 as (25, 35) as an example, the candidate region range of the corresponding text feature point "Yi" is X∈[25+20, 25+44], Y∈[35-2, 35+19].

[0039] Specifically, starting from the X coordinate (25) of the main graphic feature point l0, superimposing the standard offset (27 pixels) minus the compensation value (7 pixels) to obtain the left boundary (25+20=45) of the candidate region of the text feature point "Yi"; starting from the X coordinate (25) of the main graphic feature point l0, superimposing "standard offset (27 pixels) +17 pixels" to obtain the right boundary (25+44=69) of the candidate region of the text feature point "Yi". This candidate region range not only completely covers the 17×17 pixel standard width of the text feature point "Yi", but also reserves left and right tolerance space, which is suitable for scenes with shooting tilt and slight printing offset.

[0040] Starting from the Y-coordinate (35) of the main graphic feature point l0, add the standard offset (2 pixels) and subtract the compensation value (4 pixels) to obtain the upper boundary of the candidate region of the text feature point "Yi (One)" (35-2=33). Starting from the Y-coordinate (35) of the main graphic feature point l0, add the sum of the standard offset (2 pixels) and 17 pixels to obtain the lower boundary of the candidate region of the text feature point "Yi (One)" (35+19=54), which also covers the standard height of "Yi (One)" and reserves upper and lower fault-tolerant spaces.

[0041] Through the above method, the candidate regions of the text feature points "Yi (One), Er (Two), San (Three), Si (Four)" that correspond one-to-one with the main graphic feature points are determined, which only cover the potential positions of "Yi (One), Er (Two), San (Three), Si (Four)" and eliminate the interference of background regions such as the main text of the test paper and illustrations.

[0042] Step S2002: Invoke a pre-set text feature point recognition model to perform character detection on each of the candidate regions respectively, so as to output the minimum bounding rectangle bounding box of the character in the candidate region and the preliminary recognition text; Invoke a preset OCR model to perform character detection on the candidate regions corresponding to each text feature point respectively, so as to determine the minimum bounding rectangle bounding box of the character in the candidate region corresponding to each text feature point and the preliminary recognition text, such as "壹", "一", "壱" or broken character fragments, etc.; remove bounding boxes whose area is far smaller than 17×17 pixels or far larger than 17×17 pixels.

[0043] Step S2003: Input the association relationship between the preliminary recognition text and the main graphic feature point corresponding to the candidate region into a positioning identifier semantic verification library constructed by a large language model, and screen out effective text feature points that conform to the preset positioning identifier through a multi-layer mechanism including distorted character restoration, similar character and background filtering, and context logic verification; The positioning identifier semantic verification library includes "Yi (One), Er (Two), San (Three), Si (Four)" and their common distorted forms; the distorted character restoration refers to restoring broken (e.g., "Yi (One)" is broken into "士" + "冖" + "豆"), adhered, and distorted characters into standard positioning characters through semantic completion and glyph correlation analysis of the large language model; the similar character and background filtering refers to excluding similar characters such as "一, 二, 壱" and test paper background text (e.g., "一, 二, 三" in test paper questions), and only retaining the preset positioning identifiers; The context logic check follows that the text feature point corresponding to the main graphic feature point at the top-left corner is "Yi", the text feature point corresponding to the main graphic feature point at the top-right corner is "Er", the text feature point corresponding to the main graphic feature point at the bottom-left corner is "San", and the text feature point corresponding to the main graphic feature point at the bottom-right corner is "Si". If the recognition result of a certain area conflicts with the global logic, for example, the text feature point corresponding to the main graphic feature point at the bottom-left corner is recognized as "Er", it is determined as misrecognition and excluded; valid text feature points that pass all checks are screened out, and their corresponding candidate regions and minimum bounding rectangle boundary boxes are determined. For problems that cannot be handled by traditional OCR models, such as character adhesion or fracture caused by illumination, text deformation caused by wrinkles, the association relationship between the preliminarily recognized text and the main graphic feature points corresponding to the candidate regions is input into a positioning identifier semantic check library constructed by a large language model, and valid text feature points conforming to the preset positioning identifier are screened out through a multi-layer mechanism including distorted character restoration, similar character and background filtering, and context logic check, which can greatly improve the recognition fault tolerance in extreme scenarios.

[0044] Step S2004: based on the image pixel coordinate system of the to-be-identified test paper image, calculate the center coordinates of the minimum bounding rectangle boundary box of the valid text feature points, which are used as the preliminary position coordinates of each of the valid text feature points; based on the image pixel coordinate system of the to-be-identified test paper image, for each valid text feature point, the center coordinates (X, Y) of its minimum bounding rectangle boundary box are taken as the preliminary position coordinates of each of the valid text feature points.

[0045] Step S2005: according to the first position coordinates corresponding to the main graphic feature points and the theoretical spacing of the corresponding text feature points, calculate and determine the theoretical position coordinates of the text feature points, if the relative position distance between the preliminary position coordinates and the theoretical position coordinates is less than a preset distance threshold, output the preliminary position coordinates of the valid text feature points as the third position coordinates.

[0046] Loading first position coordinates corresponding to main graphic feature points, for example, the position coordinate of the main graphic feature point l0 at the top-left corner is (24.8, 34.9). The theoretical position coordinates of the valid text feature point "Yi (One)" are calculated in combination with the theoretical spacing between the top-left main graphic feature point l0 and the text feature point "Yi (One)". For example, the theoretical horizontal coordinate of the valid text feature point "Yi (One)" is 24.8+27=51.8, and the theoretical vertical coordinate is 34.9+2=36.9, so the theoretical position coordinate of the valid text feature point "Yi (One)" is (51.8, 36.9); if the relative position distance between the preliminary position coordinates (X, Y) of the valid text feature point "Yi (One)" and the theoretical position coordinates (51.8, 36.9) of the valid text feature point "Yi (One)" is less than a preset distance threshold, outputting the preliminary position coordinates (X, Y) of the valid text feature point "Yi (One)" as the third position coordinates; if the relative position distance between the preliminary position coordinates (X, Y) of the valid text feature point "Yi (One)" and the theoretical position coordinates (51.8, 36.9) of the valid text feature point "Yi (One)" is greater than the preset distance threshold, wherein the preset distance threshold can be 3 pixels, etc.

[0047] Step S30:推算 out, according to the first position coordinates, the second position coordinates and the third position coordinates, a fourth position coordinate range corresponding to the examinee information area, a fifth position coordinate range corresponding to the objective question answering area, and a sixth position coordinate range corresponding to the subjective question answering area in the to-be-identified test paper image; After identifying, based on a preset graphic feature point recognition model, the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the to-be-identified test paper image, and determining through fusion based on a preset text feature point recognition model and a large language model the third position coordinates corresponding to each of the text feature points in the to-be-identified test paper image,推算 out, according to the first position coordinates, the second position coordinates and the third position coordinates, the fourth position coordinate range corresponding to the examinee information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the to-be-identified test paper image; In some embodiments, the step of推算 out, according to the first position coordinates, the second position coordinates and the third position coordinates, the fourth position coordinate range corresponding to the examinee information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the to-be-identified test paper image comprises: Step S301: acquiring the first position coordinates corresponding to the main graphic feature points, the second position coordinates corresponding to the auxiliary graphic feature points, and the third position coordinates corresponding to the text feature points in the to-be-identified test paper image; The first position coordinates corresponding to the main graphic feature points include: the position coordinate of the main graphic feature point l0 at the top-left corner of the to-be-identified test paper is (25, 35); the position coordinate of the main graphic feature point r1 at the bottom-right corner is (747, 1066), etc. The second position coordinates corresponding to the auxiliary graphic feature points include: the position coordinate of the first auxiliary graphic feature point h1 is (25, 292); the position coordinate of the second auxiliary graphic feature point h2 is (25, 550); the position coordinate of the third auxiliary graphic feature point h3 is (25, 808). The third position coordinates corresponding to the text feature points include: the position coordinate of the text feature point "Yi (numeral one)" is (52, 37); the position coordinate of the text feature point "Si (numeral four)" is (724, 1068), etc.

[0048] Step S302: Based on a preset identification code layout rule, determine a preset relative positional relationship between every two of the main graphic feature points, the auxiliary graphic feature points and the text feature points in the to-be-identified test paper image, and a mapping relationship between the main graphic feature points, the auxiliary graphic feature points, the text feature points and each functional area of the to-be-identified test paper; The preset relative positional relationship between every two of the main graphic feature points, the auxiliary graphic feature points and the text feature points in the to-be-identified test paper image is, for example, that the main graphic feature point l0 at the top-left corner is offset by 27 pixels on the horizontal axis (X axis) and 2 pixels on the vertical axis (Y axis) from the corresponding text feature point "Yi (numeral one)", and the vertical axis (Y axis) spacing between the main graphic feature point l0 at the top-left corner and the first auxiliary graphic feature point h1 is 257 pixels, etc. The mapping relationship between the main graphic feature points, the auxiliary graphic feature points, the text feature points and each functional area of the to-be-identified test paper can be based on the fixed relative positions between each functional area of the to-be-identified test paper and the main graphic feature points, the auxiliary graphic feature points and the text feature points. For example, the position coordinate of "Name" in the candidate information area is (126, 37), which takes the X-axis coordinate of the text feature point "Yi (numeral one)" as a reference and is offset 74 pixels to the right to determine the X-axis starting position of "Name" in the candidate information area. The ranges of the objective question area and the subjective question area can also be fixed by the spacing from each of the auxiliary graphic feature points h1, h2 and h3.

[0049] Step S303: Perform cross-checking and fusion correction on the first position coordinates, the second position coordinates and the third position coordinates to obtain a feature point coordinate set of the to-be-identified test paper; In a specific embodiment, the step of performing cross-checking and fusion correction on the first position coordinates, the second position coordinates and the third position coordinates to obtain the feature point coordinate set of the to-be-identified test paper comprises: Step S3001: performing cross-check on the first position coordinate, the second position coordinate and the third position coordinate based on the identification code layout rule, determining whether a pairwise relative position distance between the first position coordinate, the second position coordinate and the third position coordinate is within the preset distance threshold, and if not, marking the abnormal position coordinate of which the relative position distance exceeds the preset distance threshold; performing pairwise cross-check on the first, second and third position coordinates, for example, checking whether an actual spacing between a main graphic feature point l0(25,35) and a text feature point "one" (52,37) is within a 3-pixel threshold, and checking whether an actual Y-axis spacing between an auxiliary graphic h1(25,292) and the main graphic feature point l0 conforms to a preset value of 257 pixels; if an actual relative position distance of two groups of feature points exceeds the preset distance threshold, for example, an actual coordinate of the text feature point "one" is (55,38), an X-axis spacing between the text feature point "one" and the main graphic feature point l0 is 30 pixels, which exceeds the 3-pixel threshold, the text feature point "one" is marked as an abnormal position coordinate.

[0050] Step S3002: performing fusion correction on the abnormal position coordinate according to other valid position coordinates by using a weighted average method to determine a corrected position coordinate; for the marked abnormal coordinate, screening all valid coordinates having a preset relative position relationship with the abnormal coordinate. For example, when the coordinate of "one" is abnormal, associated valid coordinates comprise the main graphic l0(25,35), the auxiliary graphic h1(25,292), the text feature point "two" (724,37) and the like; assigning weights based on positioning reliability of feature points: the first position coordinate corresponding to a main graphic feature point has the highest positioning accuracy, which can be set with a weight of 0.4; the second position coordinate corresponding to an auxiliary graphic feature point takes the second place, which can be set with a weight of 0.3; the weight of the third position coordinate corresponding to a text feature point can be set as 0.3; those skilled in the art can adjust the weights according to actual scenarios, which is not limited herein; calculating a correction value of the abnormal coordinate by reverse推算 in combination with the weights according to the preset relative position relationship. For example, the abnormal coordinate of "one" is (55,38), the corrected abnormal position coordinate is calculated by substituting a weighted average formula through the valid coordinate (25,35) of the main graphic feature point l0 at the upper left corner and the valid coordinate (724,37) of the text feature point "two" (for example, the abnormal coordinate of the text feature point "one" is corrected to (52.1, 37.2)), so as to ensure that spacings between the corrected abnormal position coordinate and the main graphic feature point l0 at the upper left corner as well as the text feature point "two" both conform to the preset threshold.

[0051] Step S3003: generating a feature point coordinate set of the to-be-identified test paper based on the corrected position coordinates and the valid position coordinates.

[0052] Integrate all valid coordinates of unmarked anomalies (such as l0(25,35), h1(25,292), "Er"(724,37), etc.) with the abnormal coordinates corrected in step S3002 (such as "Yi"(52.1,37.2)); classify and label the coordinates according to feature point types (main graphic feature points, auxiliary graphic feature points, text feature points), clarify the feature point name corresponding to each coordinate (such as "main graphic feature point-l0", "text feature point-Yi", etc.), and form a standardized feature point coordinate set; Step S304: According to the feature point coordinate set and a preset test paper functional area layout template, calculate a fourth position coordinate range corresponding to the examinee information area, a fifth position coordinate range corresponding to the objective question answering area, and a sixth position coordinate range corresponding to the subjective question answering area in the to-be-identified test paper image respectively.

[0053] The examinee information area includes a name area and a student number area, etc. For the "name" area in the examinee information area, first obtain the precise X-axis coordinate of the text feature point "Yi" after cross-verification and fusion correction, and offset 74 pixels to the right on this basis to obtain the X-axis starting coordinate of the "name" area; the Y-axis coordinate of the "name" area remains consistent with the Y-axis coordinate of "Yi". Combined with the preset standard width (for example, 136 pixels) and height (for example, 17 pixels) of the "name" area, based on the determined X-axis starting coordinate and Y-axis coordinate, extend 136 pixels to the right to delimit the X-axis range, and extend 17 pixels downward to delimit the Y-axis range, and finally obtain the coordinate range of the "name" area as X∈[126,262], Y∈[37,54].

[0054] For the "student number" area in the examinee information area, the same calculation logic as that for the "name" area is adopted, and according to the preset mapping relationship between the "student number" area and the corresponding feature point, the X-axis starting coordinate (determined based on the associated feature point coordinate and a fixed offset) and the Y-axis coordinate (consistent with the Y-axis coordinate of the associated feature point) are determined, and then combined with the preset standard width (for example, 136 pixels) and height (for example, 17 pixels) of the "student number" area, the coordinate range of the "student number" area is calculated as X∈[564,700], Y∈[37,54].

[0055] For the objective question answering area, based on the corrected feature point coordinate set, the coordinate range is determined by combining preset template rules: the upper boundary is obtained by shifting the Y-axis coordinate of the text feature point "Yi (one)" upward by 50 pixels; the lower boundary is obtained by shifting the Y-axis coordinate of the second auxiliary graphic feature point h2 downward by 30 pixels; the left boundary is obtained by shifting the X-axis coordinate of the main graphic feature point l0 rightward by 50 pixels; the right boundary is obtained by shifting the X-axis coordinate of the main graphic feature point r0 leftward by 50 pixels. Through conversion of the above feature point coordinates, the fifth position coordinate range of the objective question answering area is obtained as X∈[75,697], Y∈[87,520].

[0056] For the subjective question answering area, according to the preset template rules, calculation is performed based on the corrected feature point coordinates: the upper boundary is obtained by shifting the Y-axis coordinate of the second auxiliary graphic feature point h2 upward by 30 pixels; the lower boundary is obtained by shifting the Y-axis coordinate of the text feature point "San (three)" downward by 50 pixels; the left and right boundaries are kept consistent with those of the objective question answering area. Through conversion of corresponding feature point coordinates, the sixth position coordinate range of the subjective question answering area is finally determined as X∈[75,697], Y∈[580,1018].

[0057] Step S40: performing image segmentation based on a preset image segmentation model according to the fourth position coordinate range, the fifth position coordinate range and the sixth position coordinate range, so as to determine a first mask image corresponding to the examinee information area, a second mask image corresponding to the objective question answering area and a third mask image corresponding to the subjective question answering area; After calculating the fourth position coordinate range corresponding to the examinee information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the to-be-identified test paper image according to the first position coordinate, the second position coordinate and the third position coordinate, performing image segmentation based on a preset image segmentation model according to the fourth position coordinate range, the fifth position coordinate range and the sixth position coordinate range, so as to determine a first mask image corresponding to the examinee information area, a second mask image corresponding to the objective question answering area and a third mask image corresponding to the subjective question answering area; wherein the basic network architecture of the image segmentation model comprises a U2net model.

[0058] In some embodiments, the step of performing image segmentation based on a preset image segmentation model according to the fourth position coordinate range, the fifth position coordinate range and the sixth position coordinate range, so as to determine the first mask image corresponding to the examinee information area, the second mask image corresponding to the objective question answering area and the third mask image corresponding to the subjective question answering area comprises: Step S401: Acquire a fourth position coordinate range corresponding to the examinee information area, a fifth position coordinate range corresponding to the objective question answer area, and a sixth position coordinate range corresponding to the subjective question answer area in the to-be-identified examination paper image; It can be known from the foregoing step 304 that the fourth position coordinate range corresponding to the examinee information area includes a "name" area (X∈[126,262], Y∈[37,54]) and a "student ID" area (X∈[564,700], Y∈[37,54]); the fifth position coordinate range corresponding to the objective question answer area is X∈[75,697], Y∈[87,520]; the sixth position coordinate range corresponding to the subjective question answer area is X∈[75,697], Y∈[580,1018].

[0059] Step S402: Delimit the examinee information area range in the to-be-identified examination paper image based on the fourth position coordinate range, delimit the objective question answer area range in the to-be-identified examination paper image based on the fifth position coordinate range, and delimit the subjective question answer area range in the to-be-identified examination paper image based on the sixth position coordinate range; Based on the fourth position coordinate range corresponding to the examinee information area, circumscribe the rectangular areas of "name" and "student ID" respectively in the to-be-identified examination paper image, ensuring that the filling area of the examinee is completely covered, and feature points within the page margins (such as the text feature point "Yi", the text feature point "Er") and the examination paper information identification code are not included; based on the fifth position coordinate range, circumscribe a rectangular range covering all objective question answer areas (such as multiple-choice question shading areas), avoid question stem description text and page margin feature points, and only focus on the blank answer area; according to the sixth position coordinate range, circumscribe the subjective question answer area in the lower half of the examination paper, exclude irrelevant content above and below (such as residues in the objective question area, the test question information identification code in the page footer, etc.), and accurately cover the handwritten answer area.

[0060] Isolation of functional areas of each to-be-identified examination paper is realized through coordinate circumscription, so that subsequent segmentation only operates on the effective range, which not only improves segmentation efficiency (reduces invalid pixel operations), but also avoids external interference affecting the segmentation result.

[0061] Step S403: Invoke the image segmentation model that has been trained to convergence, and perform pixel-level segmentation on the delimited examinee information area, the delimited objective question answer area and the delimited subjective question answer area respectively, so as to exclude irrelevant interference information outside each area including question description text, header and footer, paper background and stains; The basic network architecture of the image segmentation model comprises the U2net model, and the U2net model trained to convergence can be invoked. The U2net model has been trained through a large number of examination paper images, can accurately identify the differences between handwritten text, shading marks and printed text as well as the background, and has strong resistance to illumination changes and stain interference; The designated candidate information area, objective question answer area, and subjective question answer area are segmented at the pixel level. For the candidate information area, printed text on the "name" and "student ID" labels is filtered out, retaining only the pixels of the candidate's handwritten or printed content, excluding paper background and minor stains. For the objective question area, the pixels of pencil or pen markings are accurately identified, filtering out paper background color and residual text from the question stem to ensure clear separation of marking marks. For the subjective question area, all valid pixels of the candidate's handwritten answers are retained, filtering out paper wrinkles, shadows, edge stains, and header / footer residue to achieve complete separation of handwritten content from the background.

[0062] Step S404: Output the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answering area, and the third mask image corresponding to the subjective question answering area.

[0063] The designated candidate information area, objective question answer area, and subjective question answer area are segmented at the pixel level to eliminate irrelevant interference information outside each area, including question description text, headers and footers, paper background, and stains. Then, a first mask image corresponding to the candidate information area, a second mask image corresponding to the objective question answer area, and a third mask image corresponding to the subjective question answer area are output. The first mask image retains only the pixels of the "name" and "student ID" fields; the second mask image only shows the pixels of the marking marks, with clear outlines, facilitating rapid identification of the marking location and intensity by the grading system; and the third mask image fully retains the pixels of the handwritten answers, providing high-quality material for subsequent subjective question grading.

[0064] Step S50: Based on the first mask image, identify the candidate information of the target candidate, associate the second mask image, the third mask image, the test paper information identification code with the candidate information to determine the test paper answer data of the target candidate, and push it to the preset test paper grading system to complete the control of electronic identification and positioning of the test paper.

[0065] Based on a preset image segmentation model, image segmentation is performed according to the fourth, fifth, and sixth position coordinate ranges to determine the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answer area, and the third mask image corresponding to the subjective question answer area. Then, the candidate information of the target candidate is identified based on the first mask image. The second mask image, the third mask image, the test paper information identification code, and the candidate information are associated with the candidate information to determine the test paper answer data of the target candidate, and pushed to a preset test paper grading system to complete the control of electronic identification and positioning of the test paper.

[0066] Specifically, the first mask image is a "cleaned version" of the candidate's information area after pixel-level segmentation. It retains only the pixels of the candidate's handwritten or printed content, such as the characters for "name" and "student ID," completely eliminating interference from paper backgrounds, stains, printed labels, etc. A preset OCR model can be called to extract information from the first mask image. For example, from the mask image of the "name" area, the candidate's name (e.g., "Zhang San") can be identified; from the mask image of the "student ID" area, the candidate's student ID (e.g., "2024001") can be identified; and the candidate information containing the name and student ID of the target candidate can be output, ensuring unique and traceable identity.

[0067] Using the candidate information, which includes the candidate's name and student ID, as a unique identifier, the system associates the second mask image of the objective question answering area, the third mask image of the subjective question answering area, and the test paper information identification code to generate an electronic answer data package, which is then pushed to a preset test paper grading system to complete the control of electronic identification and positioning of the test paper.

[0068] As can be seen from the above embodiments, compared with the prior art, the present application addresses the problems of existing technologies such as strong dependence on printing quality, poor resistance to severe occlusion, and susceptibility to background interference in visual positioning technology based on black squares, and the problems of high cost and space occupation, high printing and recognition thresholds, and the need to change the layout of the original test paper in positioning and recording technology based on QR codes. The present application has, but is not limited to, the following beneficial effects: Firstly, this application, based on the original format of the exam paper, only utilizes the page margins to place the identification code, without requiring adjustments to the main text layout, headers and footers, or the provision of additional space. Compared to the drawbacks of QR codes requiring changes to the layout to accommodate graphics, and the complex design that may require additional directional markings for black squares, this application perfectly conforms to the traditional exam paper layout habits, reducing the adaptation costs for technology implementation.

[0069] Secondly, this application adopts a feature point design that combines graphic outlines and Chinese characters. The Chinese characters have high information content and strong outline recognition. Even if the printing is slightly blurry, the color is light, or the ink spreads slightly, it can still be recognized through the dual features of character semantics and graphic outlines, thus solving the problem of the strong dependence of black squares on printing quality. Thirdly, this application uses a combination of graphic and text feature points for positioning. Even if the test paper is folded or missing corners, causing two or more corner graphic feature points to fail, the missing positioning point can still be calculated through the remaining diagonal feature points and text feature points, thus solving the problem of positioning failure due to black square obstruction. Fourth, this application effectively distinguishes between positioning elements and interference patterns in the black illustrations and table borders in the background of the test paper, and the combination of Chinese characters and graphic feature points, thus avoiding misidentification. At the same time, through uniform illumination processing, large model semantic correction, and OCR algorithm optimization, it can still accurately identify even in the face of uneven illumination, character adhesion or breakage, and text deformation and distortion, thus breaking through the limitation of traditional technology in the sharp drop in recognition rate in complex environments.

[0070] Fifth, this application uses 12 Chinese characters as the test paper identification code, 1 Chinese character as the page number, and 4 Chinese characters as the check code. The high information capacity of Chinese characters ensures that the identification code of each test paper is unique, which can accurately distinguish different test papers of the same subject and version. Compared with the limitations of black squares, which can only achieve positioning and cannot carry complex information, and the problems of limited information capacity and large space occupation of QR codes, this application greatly improves the accuracy of electronic identification and positioning of test papers, and provides complete data support for subsequent test paper grading.

[0071] Sixth, it prioritizes image recognition for rapid localization, and only starts OCR text recognition when image recognition fails, avoiding the inefficiency of relying entirely on OCR and significantly improving the recognition speed of a single test paper. By pre-setting the mapping relationship between feature points and functional areas (name, student ID, answer area), the coordinates of the target area can be directly calculated after locating the identification code, without the need for manual annotation. Combined with image segmentation and standardized mask output, it realizes full automation from test paper scanning to electronic data, solving the bottleneck of low efficiency in traditional manual sorting and annotation, and can support the batch test paper processing needs in large-scale teaching scenarios.

[0072] Please see Figure 3An electronic test paper identification and positioning control device provided for one of the purposes of this application includes a test paper image acquisition module 1100, a feature point coordinate identification module 1200, a functional area coordinate determination module 1300, a mask image segmentation module 1400, and an answer data push module 1500. The test paper image acquisition module 1100 is configured to acquire a test paper image uploaded by the target examinee, containing a target identification code. The target identification code is located in the margin area of ​​the test paper image and includes multiple graphic feature points, multiple text feature points, and a test paper information identification code. The graphic feature points include main graphic feature points and auxiliary graphic feature points. The feature point coordinate recognition module 1200 is configured to identify the first position coordinates corresponding to each main graphic feature point and the second position coordinates corresponding to each auxiliary graphic feature point in the test paper image based on a preset graphic feature point recognition model, and determine the third position coordinates corresponding to each text feature point in the test paper image based on a preset text feature point recognition model and a large language model fusion. The functional area coordinate determination module 1300 is configured to calculate the coordinates of the functional area based on the first position coordinates, the second position coordinates, and the third position coordinates. The image to be identified includes a fourth coordinate range corresponding to the candidate information area, a fifth coordinate range corresponding to the objective question answer area, and a sixth coordinate range corresponding to the subjective question answer area. A mask image segmentation module 1400 is configured to perform image segmentation based on a preset image segmentation model according to the fourth, fifth, and sixth coordinate ranges to determine a first mask image corresponding to the candidate information area, a second mask image corresponding to the objective question answer area, and a third mask image corresponding to the subjective question answer area. An answer data push module 1500 is configured to identify the candidate information of the target candidate based on the first mask image, associate the second mask image, the third mask image, and the exam paper information identification code with the candidate information to determine the target candidate's exam paper answer data, and push it to a preset exam paper grading system to complete the control of electronic identification and positioning of the exam paper.

[0073] Based on any embodiment of this application, please refer to Figure 4 Another embodiment of this application also provides an electronic device, which can be implemented by a computer device, such as... Figure 4The diagram shows the internal structure of a computer device. This computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database may store control information sequences. When the processor executes the computer-readable instructions, it enables the processor to implement a method for electronic identification and positioning control of examination papers. The processor provides computing and control capabilities to support the operation of the entire computer device. The memory stores computer-readable instructions, which, when executed by the processor, enable the processor to execute the electronic identification and positioning control method for examination papers described in this application. The network interface of the computer device is used for communication with a terminal. Those skilled in the art will understand that… Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0074] In this embodiment, the processor is used to execute... Figure 3 The specific functions of each module are defined within the device, and the memory stores the program code and various data required to execute these modules. A network interface is used for data transmission between the user terminal and the server. In this embodiment, the memory stores the program code and data required to execute all modules in the electronic test paper identification and positioning control device of this application, and the server can call the server's program code and data to execute the functions of all modules.

[0075] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the electronic test paper identification and positioning control method described in any embodiment of this application.

[0076] This application also provides a computer program product, including a computer program / instructions that, when executed by one or more processors, implement the steps of the electronic test paper identification and positioning control method described in any embodiment of this application.

[0077] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0078] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for electronic identification and positioning control of examination papers, characterized in that, include: The method involves acquiring an image of a test paper uploaded by a target candidate, which contains a target identification code. The target identification code is located in the margin area of ​​the test paper image and includes multiple graphic feature points, multiple text feature points, and a test paper information identification code. The graphic feature points include main graphic feature points and auxiliary graphic feature points. Based on a preset graphic feature point recognition model, the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the test paper image to be recognized are identified. Based on a preset text feature point recognition model and a large language model fusion, the third position coordinates corresponding to each of the text feature points in the test paper image to be recognized are determined, which includes: Obtain text feature points located adjacent to the main graphic feature points, and based on the first position coordinates corresponding to the main graphic feature points at the four corners of the test paper to be identified, match the candidate regions of the corresponding text feature points for each main graphic feature point. The preset text feature point recognition model is invoked to perform character detection on the candidate regions respectively, so as to output the minimum bounding rectangle of the characters in the candidate regions and the preliminary recognized text. The association between the preliminarily identified text and the main graphic feature points corresponding to the candidate region is input into the positioning identifier semantic verification library constructed by the large language model. The effective text feature points that meet the preset positioning identifier are filtered out through a multi-layer mechanism including distorted character restoration, similar character and background filtering, and context logic verification. Using the image pixel coordinate system of the test paper image to be recognized as a reference, the center coordinates of the minimum bounding rectangle of the effective text feature points are calculated as the preliminary position coordinates of each effective text feature point; based on the first position coordinates corresponding to the main graphic feature points and the theoretical spacing between their corresponding text feature points, the theoretical position coordinates of the text feature points are calculated and determined; if the relative position distance between the preliminary position coordinates and the theoretical position coordinates is less than a preset distance threshold, the preliminary position coordinates of the effective text feature points are output as the third position coordinates; Based on the first position coordinates, the second position coordinates, and the third position coordinates, the fourth position coordinate range corresponding to the candidate information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the candidate's test paper image to be identified are calculated. Based on a preset image segmentation model, image segmentation is performed according to the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range to determine the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answering area, and the third mask image corresponding to the subjective question answering area. Based on the first mask image, the candidate information of the target candidate is identified. The second mask image, the third mask image, the test paper information identification code and the candidate information are associated to determine the test paper answer data of the target candidate and pushed to the preset test paper grading system to complete the control of electronic identification and positioning of the test paper.

2. The electronic identification and positioning control method for examination papers according to claim 1, characterized in that, The step of identifying the first position coordinates of each main graphic feature point and the second position coordinates of each auxiliary graphic feature point in the image of the test paper to be identified based on a preset graphic feature point recognition model includes: Based on the preset graphic feature point recognition model, the graphic contour information that conforms to the preset size and preset contour shape in the image of the test paper to be recognized is extracted, so as to match and locate the main graphic feature points at the four corners of the test paper to be recognized, as well as multiple auxiliary graphic feature points on its sides. Based on the image pixel coordinate system of the test paper to be identified, the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points are calculated and output respectively.

3. The electronic identification and positioning control method for examination papers according to claim 2, characterized in that, The steps of calculating the range of fourth coordinates corresponding to the candidate information area, the range of fifth coordinates corresponding to the objective question answering area, and the range of sixth coordinates corresponding to the subjective question answering area in the image of the test paper to be identified based on the first, second, and third position coordinates include: Obtain the first position coordinates of the main graphic feature points, the second position coordinates of the auxiliary graphic feature points, and the third position coordinates of the text feature points in the test paper image to be identified; Based on the preset identification code layout rules, the preset relative positional relationships between the main graphic feature points, the auxiliary graphic feature points and the text feature points in the image of the test paper to be identified are determined, as well as the mapping relationship between the main graphic feature points, the auxiliary graphic feature points, the text feature points and each functional area of ​​the test paper to be identified. Cross-validation and fusion correction are performed on the first position coordinates, the second position coordinates, and the third position coordinates to obtain the feature point coordinate set of the test paper to be identified; Based on the feature point coordinate set and the preset test paper functional area layout template, the coordinate range of the fourth position corresponding to the candidate information area, the coordinate range of the fifth position corresponding to the objective question answering area, and the coordinate range of the sixth position corresponding to the subjective question answering area in the test paper image to be identified are calculated respectively.

4. The electronic identification and positioning control method for examination papers according to claim 1, characterized in that, The step of performing cross-validation and fusion correction on the first position coordinates, the second position coordinates, and the third position coordinates to obtain the feature point coordinate set of the test paper to be identified includes: Based on the identification code layout rules, cross-validation is performed on the first position coordinates, the second position coordinates, and the third position coordinates to determine whether the relative position distance between each pair of the first position coordinates, the second position coordinates, and the third position coordinates is within a preset distance threshold. If it is not within the preset distance threshold, abnormal position coordinates whose relative position distance exceeds the preset distance threshold are marked. The weighted average method is used to fuse and correct the abnormal position coordinates based on the remaining valid position coordinates to determine the corrected position coordinates. Based on the corrected position coordinates and the valid position coordinates, a set of feature point coordinates for the test paper to be identified is generated.

5. The electronic identification and positioning control method for examination papers according to claim 1, characterized in that, The steps of performing image segmentation based on a preset image segmentation model according to the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range to determine the first mask image corresponding to the candidate information region, the second mask image corresponding to the objective question answering region, and the third mask image corresponding to the subjective question answering region include: Obtain the fourth position coordinate range corresponding to the candidate information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the image of the test paper to be identified; The candidate information area in the test paper image to be identified is defined based on the fourth position coordinate range, the objective question answering area in the test paper image to be identified is defined based on the fifth position coordinate range, and the subjective question answering area in the test paper image to be identified is defined based on the sixth position coordinate range. The trained and converged image segmentation model is invoked to perform pixel-level segmentation on the defined candidate information region, objective question answer region, and subjective question answer region, so as to exclude irrelevant interference information outside each region, including question description text, headers and footers, paper background, and stains. Output the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answer area, and the third mask image corresponding to the subjective question answer area.

6. The electronic identification and positioning control method for examination papers according to any one of claims 1 to 5, characterized in that, The main graphic feature points are placed at the four corners of the test paper image to be identified, the auxiliary graphic feature points are placed on the sides of the test paper image to be identified, and the text feature points are placed at positions adjacent to the main graphic feature points. The test paper information identification code includes the subject, grade, exam session, and question type distribution of the test paper image to be identified; The basic network architecture of the graphic feature point recognition model includes a lightweight convolutional neural network; the basic network architecture of the text feature point recognition model includes an OCR model; and the basic network architecture of the image segmentation model includes a U2net model.

7. A test paper electronic identification and positioning control device, characterized in that, include: The test paper image acquisition module is configured to acquire the test paper image uploaded by the target examinee, which contains the target identification code. The target identification code is located in the margin area of ​​the test paper image and includes multiple graphic feature points, multiple text feature points, and test paper information identification code. The graphic feature points include main graphic feature points and auxiliary graphic feature points. The feature point coordinate recognition module is configured to identify the first position coordinates corresponding to each of the main graphic feature points and the second position coordinates corresponding to each of the auxiliary graphic feature points in the test paper image to be recognized based on a preset graphic feature point recognition model, and to determine the third position coordinates corresponding to each of the text feature points in the test paper image to be recognized based on a preset text feature point recognition model and a large language model fusion. This module includes: Obtain text feature points located adjacent to the main graphic feature points, and based on the first position coordinates corresponding to the main graphic feature points at the four corners of the test paper to be identified, match the candidate regions of the corresponding text feature points for each main graphic feature point. The preset text feature point recognition model is invoked to perform character detection on the candidate regions respectively, so as to output the minimum bounding rectangle of the characters in the candidate regions and the preliminary recognized text. The association between the preliminarily identified text and the main graphic feature points corresponding to the candidate region is input into the positioning identifier semantic verification library constructed by the large language model. The effective text feature points that meet the preset positioning identifier are filtered out through a multi-layer mechanism including distorted character restoration, similar character and background filtering, and context logic verification. Using the image pixel coordinate system of the test paper image to be recognized as a reference, the center coordinates of the minimum bounding rectangle of the effective text feature points are calculated as the preliminary position coordinates of each effective text feature point; based on the first position coordinates corresponding to the main graphic feature points and the theoretical spacing between their corresponding text feature points, the theoretical position coordinates of the text feature points are calculated and determined; if the relative position distance between the preliminary position coordinates and the theoretical position coordinates is less than a preset distance threshold, the preliminary position coordinates of the effective text feature points are output as the third position coordinates; The functional area coordinate determination module is configured to calculate, based on the first position coordinate, the second position coordinate, and the third position coordinate, the fourth position coordinate range corresponding to the candidate information area, the fifth position coordinate range corresponding to the objective question answering area, and the sixth position coordinate range corresponding to the subjective question answering area in the candidate's information area of ​​the test paper image to be identified; The mask image segmentation module is configured to perform image segmentation based on a preset image segmentation model according to the fourth position coordinate range, the fifth position coordinate range, and the sixth position coordinate range, so as to determine the first mask image corresponding to the candidate information area, the second mask image corresponding to the objective question answering area, and the third mask image corresponding to the subjective question answering area. The answer data push module is configured to identify the candidate information of the target candidate based on the first mask image, associate the second mask image, the third mask image, the test paper information identification code with the candidate information to determine the test paper answer data of the target candidate, and push it to the preset test paper grading system to complete the control of electronic identification and positioning of the test paper.

8. An electronic device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 6, which, when invoked by a computer, executes the steps included in the corresponding method.

Citation Information

Patent Citations

  • Answer sheet processing method and answer sheet processing system

    CN116012864A

  • Scanning checking system and method for identification checking and electronic equipment

    CN117275027A