A method for detecting and segmenting homework and test papers based on neural network

Through the neural network-based homework test paper detection and segmentation method, combined with the text area detection network of YOLOv7 and MobileNetV3, image correction and direction detection are carried out, and two-stage detection methods and post-processing optimization are adopted to solve the problems of small and medium-sized formulas, complex layout detection errors and image tilt direction errors, and efficient and accurate homework and test paper automation detection are achieved.

CN119478980BActive Publication Date: 2025-05-09BEIJING AIXUESI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411588730.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-05-09
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

The existing deep learning algorithms have large errors in detecting information such as small formulas and complex layouts, and when faced with complex backgrounds, image tilt or direction errors, the accuracy is low, which affects the widespread application of homework and test paper detection and slicing technology.

Method used

A neural network-based operation test paper detection and segmentation method is adopted, including text area detection, image correction, direction detection and correction, content detection classification and post-processing. The specific steps include using a text area detection network combined with YOLOv7 and MobileNetV3, correcting images through perspective transformation and direction detection, using a two-stage detection method to refine content information extraction, and optimizing the detection results through post-processing.

Benefits of technology

It significantly improves the ability to process non-standardized images, improves the accuracy of detection small targets, reduces missed detection, improves the speed and accuracy of detection, is highly adaptable, and can achieve efficient and accurate homework and automatic test paper detection in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478980B_ABST
    Figure CN119478980B_ABST
Patent Text Reader

Abstract

The invention discloses a homework and test paper detection and segmentation method based on a neural network, and relates to the field of teaching technology. The invention combines YOLOv7 with MobileNetV3, modifies the output structure of YOLOv7, so that it can detect any quadrilateral area, eliminates the dependence on the precise scanner, and significantly improves the ability to process non-standardized images. At the same time, a two-stage detection method is adopted to improve the overall detection accuracy, and the missed detection situation is reduced when processing small targets. In addition, MobileNetV3 is introduced as a backbone network to reduce the amount of model parameters and the requirement for image resolution, so that better detection effect can be obtained under medium and low resolution images. The invention performs text area detection based on the combination of YOLOv7 and MobileNetV3, uses OpenCV to perform image correction and rotation, adopts a two-stage detection method to extract refined content information, and finally optimizes the detection result through post-processing. It not only performs well in complex photo-taking scenes, but also significantly improves the speed and accuracy of detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of teaching technology, and in particular to a homework and test paper detection and segmentation method based on a neural network. Background Art

[0002] During the teaching process, teachers need to mark a large number of homework and test papers, which is extremely time-consuming and repetitive. In order to reduce the burden on teachers and improve teaching efficiency, some intelligent marking systems have emerged. These systems detect and segment the scanned test paper images, automatically identify question information, and realize subsequent automatic scoring and analysis, greatly reducing the workload of teachers.

[0003] With the rapid development of artificial intelligence technology, the efficiency and accuracy of educational assessment systems have attracted much attention. They are not only used in examination scenarios, but also in daily homework and exercise books. Automatic detection and segmentation of questions have become an important part of improving efficiency. With the development of deep learning technology, target detection algorithms such as YOLO, CTPN, and EAST have significantly improved the accuracy of question detection and segmentation in complex scenarios. With the support of these technologies, questions can be detected more accurately even in interference conditions such as poor lighting and image stains. In addition, the application of layout analysis algorithms such as Layoutparser and table recognition algorithms such as TableMaster also provides the possibility for the realization of a comprehensive automatic analysis system for test papers or homework.

[0004] The current marking scheme relies on the coordinates of the positioning blocks on the test paper template for image recognition and uses a scanner to obtain regular image input. The premise for this is that all input images are of the same size and have no obvious tilt, stains, and other problems. Otherwise, the system will have large errors in the position and boundary recognition of the test paper. At the same time, although deep learning algorithms have made significant progress in object detection and text detection, these algorithms rely on large-scale and diverse annotated data sets, and building these data sets requires a lot of manual work. Secondly, there are still large errors in the detection of small formulas, complex layouts and other information. In addition, the existing detection algorithms have low accuracy when facing complex backgrounds, image tilts or wrong directions. These problems have significantly affected the widespread application of homework and test paper detection and segmentation technology based on deep learning. Therefore, a neural network-based homework and test paper detection and segmentation method is urgently needed to solve such problems. Summary of the invention

[0005] In view of the above existing problems, the present invention is proposed.

[0006] The present invention provides a homework paper detection and segmentation method based on a neural network to solve the problem that deep learning algorithms still have large errors in detecting small formulas, complex layouts and other information, and have low accuracy when facing complex backgrounds, image tilts or wrong directions.

[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0008] The present invention provides a method for detecting and segmenting homework and test papers based on a neural network, which comprises:

[0009] Step S1, text area detection, the original input image is passed through the text area detection network to obtain the quadrilateral boundary of the target area,

[0010] Step S2: image correction, perspective transformation, correction of the detected quadrilateral boundary of the target area, and obtaining the corrected image to be detected.

[0011] Step S3: direction detection: the direction detection network is used to determine whether the direction of the current image needs to be rotated.

[0012] Step S4, direction correction, for the image that needs to be rotated, a rotation transformation operation is performed, and a unified positive image is obtained after the rotation transformation operation.

[0013] Step S5: Content detection and classification: the final forward image is passed through the content detection and classification network to obtain the coordinates and category information of the title, graphics, table, user answer and other information in the image.

[0014] Step S6, post-processing, corrects the detection result of step S5 through post-processing, corrects the erroneous boundaries, and supplements the missing detection information.

[0015] Furthermore, in step S1, a text region detection network is used to identify text regions in the input image. The text region detection network is based on YOLOv7, but in order to improve efficiency and speed, the backbone of the network is replaced with MobileNetV3 to reduce the number of parameters. Since YOLOv7 can only detect rectangular frames, here, in order to adapt to the shooting angle and perspective distortion problems, the output of the network is modified to detect quadrilateral regions, and a loss function is used to support quadrilateral output. The loss function is modified:

[0016] L=L class +λ1L quad , where L class is the classification loss, L quad is the coordinate loss of the quadrilateral detection box. The quadrilateral detection box is represented by four vertices (x1, y1), (x2, y2), (x3, y3), and (x4, y4). The coordinates of these four points are used to calculate the error loss. λ1 is the weight parameter used to balance the classification loss L class and quadrilateral detection box loss L quad Time influence.

[0017] Furthermore, in step S2, after the quadrilateral region is detected, the perspective transformation function of OpenCV is used to correct the irregular quadrilateral region into a regular rectangular region. The transformation matrix H is calculated from four points and satisfies: P′=H·P, where P is the coordinate of the original quadrilateral, P′ is the coordinate of the transformed rectangle, and H is the perspective transformation matrix.

[0018] Furthermore, in step S3, in order to ensure the correct direction of the image, the direction detection network PP-LCNet is used to determine whether the image needs to be rotated. Since part of the direction has been corrected in the previous step, the network only performs lightweight rotation detection on the image. The network divides the image orientation into four categories: 0 degrees, 90 degrees, 180 degrees and 270 degrees.

[0019] Furthermore, in step S4, based on the direction detection result of step S3, the image is rotated to the correct direction using the rotation operation of OpenCV, and the rotation angle is output by the direction detection network. The rotation operation is expressed as: Q′=R(θ)·Q, wherein R(θ) is the rotation matrix, θ is the rotation angle, Q is the image that needs to be rotated as determined in step S3, and Q′ is the forward image after direction correction.

[0020] Furthermore, in step S5, detailed information in the homework or test paper is detected, including questions, graphics, tables and formulas.

[0021] Furthermore, in step S5, both stages of the content detection classification network use the original yolov7 network as the reference network.

[0022] Furthermore, in step S5, the content detection classification adopts a two-stage detection method:

[0023] In the first stage, the approximate location of the questions and areas is detected;

[0024] In the second stage, each question is segmented and further information is tested.

[0025] The two-stage detection method not only improves the overall detection accuracy, but also reduces the requirements on the pixel resolution of the input image.

[0026] Furthermore, in step S6, based on the refined information detection result of step S5, the detection result is corrected for possible false detection and missed detection, and the correction rules include:

[0027] The detection boxes between different text lines cannot contain each other.

[0028] The formula area must be completely contained within the corresponding text line.

[0029] There cannot be irrelevant detection boxes in the same area.

[0030] Furthermore, the post-processing correction rule in step S6 is specifically:

[0031] Assume the coordinates of the upper left corner of the detection box A are (x1 A ,y1 A ), the coordinate of the lower right corner is (x2 A ,y2 A ), the coordinates of the upper left corner of the detection box B are (x1 B ,y1 B ), the coordinate of the lower right corner is (x2 B ,x2 B ), the coordinates of the upper left corner of the current question C are (x1 C ,y1 C ), the coordinate of the lower right corner is (x2 C ,y2 C ),

[0032] Assuming that detection boxes A and B belong to the text line category respectively, if:

[0033] x1 A ≥x1 B ,y1 A ≥y1 B ,x2 A ≤x2 B ,y2 A ≤y2 B , then delete the detection box A,

[0034] Regarding the relationship between the formula and the text line, if the formula box A is not completely contained by the text line box B, the coordinates of B are adjusted so that it can contain A. The adjustment rule is:

[0035] x1 B =min(x1 A ,x1 B ),y1 B =min(y1 A ,y1 B ),x2 B =max(x2 A ,x2 B ),y2 B =max(y2 A ,y2 B ), where min means taking the minimum function and max means taking the maximum function.

[0036] If A belongs to the formula category, but there is no text line related to A, the text line detection is considered lost, and the text line detection box D is supplemented according to the coordinates of the question. The coordinates are defined as x1 D =x1 C ,y1 D =y1A , x2 D =x2 A , x2 C , y2 D =y2 A .

[0037] The beneficial effects of the present invention are:

[0038] The present invention combines YOLOv7 with MobileNetV3 and modifies the output structure of YOLOv7 so that it can detect any quadrilateral area, solve the distortion problem caused by shooting angle, process tilted and irregular working images, eliminate the dependence on precise scanners, and significantly improve the ability to process non-standardized images.

[0039] The present invention adopts a two-stage detection method, first detecting the overall question area and then refining it into various types of information. The hierarchical detection method not only improves the overall detection accuracy, but also is more sensitive to detail detection, reducing missed detection when processing small targets.

[0040] The present invention introduces MobileNetV3 as the backbone network, which reduces the number of model parameters and lowers the requirements for image resolution. It not only improves the detection speed but also ensures the detection accuracy. It can obtain better detection effects under medium and low resolution images and significantly improve the user experience.

[0041] The present invention adds direction detection and correction, judges the direction of the image through the PPLCNet network, and implements image rotation correction in combination with OpenCV, ensuring that the input image will eventually be normalized to a standard direction regardless of the rotation angle, greatly improving the robustness of the detection.

[0042] The present invention adopts a post-processing step to correct the detection results of boundary detection errors or content classification errors that often occur in deep learning detection algorithms, correct erroneous boundaries, supplement missed detection areas, and make adjustments according to the logical relationship of the content. The correction strategy based on layout rules effectively improves the rationality of the detection results.

[0043] In summary, the present invention provides an end-to-end homework and test paper detection and segmentation method, which performs text area detection based on the combination of YOLOv7 and MobileNetV3, uses OpenCV for image correction and rotation, adopts a two-stage detection method to extract refined content information, and finally optimizes the detection results through post-processing. This method not only performs well in complex photo-taking scenarios, but also significantly improves the speed and accuracy of detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0045] Figure 1 It is a schematic diagram of the flow of the homework test paper detection and segmentation method based on neural network of the present invention. DETAILED DESCRIPTION

[0046] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0047] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0048] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0049] Example 1, reference Figure 1 This embodiment provides a method for detecting and segmenting homework papers based on a neural network, comprising the following steps:

[0050] Step S1: The original input image passes through the text region detection network to obtain the quadrilateral boundary of the target region.

[0051] For the text area detection network, the present invention adopts the target detection network yolov7. Since the text area usually represents the entire question position area of ​​the test paper or exercise book, the characteristics are relatively obvious. In order to improve the performance, the present invention adopts the yolov7 network, which has a smaller parameter amount. On the other hand, considering that further improvement in speed will bring better user experience, the present invention modifies the backbone network of the yolov7 network to mobilenetV3, and adjusts the head network structure of yolov7 according to the output of mobilenetV3, and reduces the number of convolution channels to further reduce the parameters of the model. Since the yolov7 network can only detect rectangular frames, and in actual scenes, due to the uncertainty of the shooting angle, it is difficult to obtain detection results close to the target text area in the rectangular area, therefore, the present invention modifies the output and loss function of the yolov7 network so that it can detect any quadrilateral area. Specifically, the predicted output of the network is modified from the center coordinates and width and height information of the rectangular box to the coordinates of the four points of the quadrilateral.

[0052] Specifically, replacing the backbone network of YOLOv7 with MobileNetV3 greatly reduces the number of model parameters, which not only speeds up the detection speed of the model, but also reduces the occupation of hardware resources. It is very suitable for deployment on mobile devices or resource-constrained environments. At the same time, the optimized YOLOv7 network can not only detect conventional rectangular boxes, but also realize the detection of arbitrary quadrilaterals by modifying the output structure and loss function, thereby adapting to scenes with uncertain shooting angles and perspective distortion, making the detection of text areas more accurate, and significantly improving the detection effect on irregular input images.

[0053] Step S2, correcting the detected target area through perspective transformation to obtain a corrected image to be detected.

[0054] For image correction, after obtaining the quadrilateral boundary, the present invention uses the image transformation operation of OpenCV to correct any quadrilateral area into a rectangular area, providing a standard output for subsequent detection. Table 1 shows the modified network parameter quantity and accuracy of the present invention.

[0055] Step S3: Determine whether the direction of the current image needs to be rotated through the direction detection network.

[0056] For the direction detection network, the present invention adopts a lightweight PP-LCNet network as the basic network. The prediction category of the network is modified according to the corresponding usage scenario of the present invention. Based on the previous two steps, the image between 0-90 degrees has been corrected. At this time, the direction detection network only needs to distinguish the orientation of the image, a total of 4 categories corresponding to: 0 degrees, 90 degrees, 180 degrees, and 270 degrees.

[0057] Step S4: for the image to be rotated, a rotation transformation operation is performed to obtain a unified forward image.

[0058] For direction correction, according to the classification result of the direction detection network, the present invention corrects the direction of the image through the rotation operation of OpenCV.

[0059] Specifically, adding image correction and direction detection steps can adapt to various irregular inputs. Any quadrilateral area can be corrected into a regular rectangular area through perspective transformation, which ensures the consistency and accuracy of the detection area. The direction detection network is introduced to quickly judge the image direction through the lightweight PPLCNet. Combined with rotation correction, the image at any rotation angle of the input can be automatically adjusted to a standardized direction, which effectively solves the problems of complex shooting environment and inconsistent image direction.

[0060] Step S5, the final image is passed through a content detection and classification network to obtain the coordinates and category information of the title, graphics, table, user answer and other information in the image.

[0061] The main purpose of the content detection classification network is to detect the position of different information in homework or exercise books. The category information set in the present invention includes graphics, figure captions, tables, table captions, multiple-choice question answer areas, fill-in-the-blank question answer areas, formulas, printed text and handwriting. Since the entire image is used as input, some smaller formulas may be missed due to unclear features. Therefore, this network adopts a two-stage method to first detect the position of the question, then segment the image area of ​​a single question, and then detect more detailed information for each question. This not only improves the overall detection accuracy, but also reduces the pixel requirements for the input image. After the previous steps S1 to S5, the current image content is basically in the horizontal direction. Considering accuracy and speed, the two stages of the content detection classification network use the original yolov7 network as the reference network.

[0062] Specifically, the content detection classification network adopts a two-stage detection method, which improves the detection accuracy and reduces the requirements for the input image resolution. The first stage detects the approximate area, and the second stage refines each question, making the detection of complex information such as graphics, tables, and formulas more accurate. In the detection of small targets such as formulas, it significantly reduces the number of missed detections and improves the overall detection reliability. This method can also adaptively adjust the detection strategy according to different assignments and test paper structures.

[0063] Step S6: Correct the detection results through post-processing to correct some erroneous boundaries and supplement the missing detection information.

[0064] After obtaining the preliminary detection results, since there will inevitably be false detections and missed detections, based on the content layout and the dependencies between different categories, the present invention designs a post-processing correction sorting scheme to correct and complete the network detection results, and to standardize and sort the final results according to the layout information, so as to facilitate subsequent recognition, correction and other tasks.

[0065] The present invention designs the following correction rules:

[0066] The positions of the rows cannot contain each other;

[0067] The formula must be completely contained within a line of text;

[0068] Two unrelated detection frames cannot appear in the same area.

[0069] The specific implementation is as follows: Let the coordinates of the upper left corner of the detection box A be (x1 A ,y1 A ), the coordinate of the lower right corner is (x2 A ,y2 A ), the coordinates of the upper left corner of the detection box B are (x1 B ,y1 B ), the coordinate of the lower right corner is (x2 B ,y2 B ), the coordinates of the upper left corner of the current question C are (x1 C ,y1 C ), the coordinate of the lower right corner is (x2 C ,y2 C ),

[0070] If both A and B belong to the text line category, if x1 A ≥x1 B And x1 A ≥y1 B And x2 A ≤x2 B And y2 A ≤y2 B , then delete the detection box A;

[0071] If A belongs to the Formula category and B belongs to the Text Line category, then if x1 A >x1 B or y1 A >y1 B or x2 A <x2 B or y2 A <y2 B , then adjust the coordinate of B to x1 B =min(x1 A ,x1 B ), y1B =min(y1 A ,y1 B ), x2 B =max(x2 A ,x2 B ), y2 B =max(y2 A ,y2 B ), where min means taking the minimum function and max means taking the maximum function.

[0072] If A belongs to the formula category, but there is no text line related to A, the text line detection is considered lost, and the text line detection box D needs to be supplemented according to the coordinates of the question. The coordinates are defined as x1 D =x1 C ,y1 D =y1 A , x2 D =x2 A , x2 C , y2 D =y2 A ,

[0073] The present invention establishes a data set corresponding to this task, including scanned images of different homework and test papers, network screenshots and photographed images. Table 2 introduces the accuracy of the present invention in various categories on the homework and test paper data set. Compared with the single-stage yolov7 network, the two-stage detection method proposed in the present invention has obvious improvement in accuracy in various categories, and the accuracy can be further improved after correction.

[0074] Specifically, post-processing effectively solves the problems of missed detection and false detection. It makes targeted corrections to the detection results based on the boundary relationship of the detection box, the inclusion relationship between the formula and the text line, etc., so that the output results conform to the layout logic in the actual scene, thereby improving the final detection accuracy.

[0075] Table 1 Parameters and accuracy comparison of text region detection models

[0076] Model Parameter quantity Accuracy yolov7-tiny 6020400 0.963 The present invention 1718033 0.958

[0077] Table 2 Comparison of detection accuracy of the present invention

[0078]

[0079]

[0080] In summary, combined with Table 1 and Table 2, it can be seen that the model parameter amount of the present invention is significantly reduced, which is only about 28.5% of YOLOv7tiny. The model is lighter and has lower computing resource requirements. The accuracy is almost not significantly reduced compared with YOLOv7tiny. The accuracy of YOLOv7tiny is 0.963, while the accuracy of the present invention is 0.958, with almost no sacrifice of accuracy. At the same time, by comparing the accuracy of the two-stage detection method of YOLOv7 and the present invention and the addition of the post-processing step, it can be seen that the detection accuracy is significantly better than that of YOLOv7 in multiple categories (formulas, printed lines, handwritten lines, graphics, etc.). After the introduction of the post-processing step, the detection accuracy of all categories is further improved, indicating that the scheme of the present invention reduces the parameter amount of the model and maintains efficient detection accuracy. After the introduction of post-processing, the detection capability of complex scenes is greatly improved. In practical applications, efficient and accurate automatic detection of homework and test papers can be achieved, and it has strong adaptability.

[0081] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for detecting and segmenting homework papers based on a neural network, characterized in that: include, Step S1, text region detection, the original input image is passed through the text region detection network to obtain the quadrilateral boundary of the target region; In step S1, a text region detection network is used to identify text regions in the input image. The text region detection network is based on YOLOv7. The backbone of the network is replaced with MobileNetV3. The output of the network is modified to detect quadrilateral regions, and a loss function is used to support quadrilateral output. The loss function is modified as follows: L=L class +λ1L quad , where L class is the classification loss, L quad is the coordinate loss of the quadrilateral detection box. The quadrilateral detection box is represented by four vertices (x1, y1), (x2, y2), (x3, y3), and (x4, y4). The coordinates of these four points are used to calculate the error loss. λ1 is the weight parameter used to balance the classification loss L class and quadrilateral detection box loss L quad Inter-influence; Step S2: image correction, perspective transformation, correction of the detected quadrilateral boundary of the target area, and obtaining the corrected image to be detected. Step S3: direction detection: the direction detection network is used to determine whether the direction of the current image needs to be rotated. Step S4, direction correction, for the image that needs to be rotated, a rotation transformation operation is performed, and a unified positive image is obtained after the rotation transformation operation. Step S5: Content detection and classification: the final forward image is passed through the content detection and classification network to obtain the coordinates and category information of the title, graphics, table, and user answer information in the image. Step S6, post-processing, correcting the detection result of step S5 by post-processing, correcting the wrong boundary, and supplementing the missing detection information; The post-processing correction rule in step S6 is specifically: Assume the coordinates of the upper left corner of the detection box A are (x1 A ,y1 A ), the coordinate of the lower right corner is (x2 A ,y2 A ), the coordinates of the upper left corner of the detection box B are (x1 B ,y1 B ), the coordinate of the lower right corner is (x2 B ,y2 B ), the coordinates of the upper left corner of the current question C are (x1 C ,y1 C ), the coordinate of the lower right corner is (x2 C ,y2 C ), Assuming that detection boxes A and B belong to the text line category respectively, if: x1 A ≥x1 B ,y1 A ≥y1 B ,x2 A ≤x2 B ,y2 A ≤y2 B , then delete the detection box A, Regarding the relationship between the formula and the text line, if the formula box A is not completely contained by the text line box B, the coordinates of B are adjusted so that it can contain A. The adjustment rule is: x1 B =min(x1 A ,x1 B ),y1 B =min(y1 A ,y1 B ),x2 B =max(x2 A ,x2 B ),y2 B =max(y2 A ,y2 B ), where min means taking the minimum function and max means taking the maximum function. If A belongs to the formula category, but there is no text line related to A, the text line detection is considered lost, and the text line detection box D is supplemented according to the coordinates of the question. The coordinates are defined as x1 D =x1 C , y1 D =y1 A , x2 D =x2 A , x2 C , y2 D =y2 A .

2. According to the neural network-based homework test paper detection and segmentation method of claim 1, it is characterized in that: In step S2, after the quadrilateral area is detected, the perspective transformation function of OpenCV is used to correct the irregular quadrilateral area into a regular rectangular area. The transformation matrix H is calculated by four points and satisfies: P'=H·P, where P is the coordinate of the original quadrilateral, P' is the coordinate of the transformed rectangle, and H is the perspective transformation matrix.

3. According to the neural network-based homework test paper detection and segmentation method of claim 2, it is characterized in that: In step S3, the direction detection network PP-LCNet is used to determine whether the image needs to be rotated. Since part of the direction has been corrected in the previous step, the network only performs lightweight rotation detection on the image. The network divides the image orientation into four categories: 0 degrees, 90 degrees, 180 degrees and 270 degrees.

4. A method for detecting and segmenting homework papers based on a neural network according to claim 3, characterized in that: In step S4, based on the direction detection result of step S3, the image is rotated to the correct direction using the rotation operation of OpenCV. The rotation angle is output by the direction detection network. The rotation operation is expressed as: Q'=R(θ)·Q, where R(θ) is the rotation matrix, θ is the rotation angle, Q is the image that needs to be rotated as determined in step S3, and Q' is the forward image after direction correction.

5. A method for detecting and segmenting homework papers based on a neural network according to claim 4, characterized in that: In step S5, detailed information in the homework or test paper is detected, including questions, graphics, tables and formulas.

6. A method for detecting and segmenting homework papers based on a neural network according to claim 5, characterized in that: In step S5, both stages of the content detection classification network use the original YOLOv7 network as the reference network.

7. A method for detecting and segmenting homework papers based on a neural network according to claim 6, characterized in that: In step S5, content detection and classification adopts a two-stage detection method: In the first stage, the location of the questions and areas is detected; In the second stage, each question is segmented for further information detection.

8. A method for detecting and segmenting homework papers based on a neural network according to claim 7, characterized in that: In step S6, based on the refined information detection result of step S5, the detection result is corrected for possible false detection and missed detection, and the correction rules include: The detection boxes between different text lines cannot contain each other. The formula area must be completely contained within the corresponding text line. There cannot be irrelevant detection boxes in the same area.

Citation Information

Patent Citations

  • Image character recognition method and system based on deep learning and medium

    CN112016547A

  • Electronic file image intelligent correction method based on text detection and table detection

    CN117496518A