Paper retrieval matching method and equipment and storage medium
By building a paper library and performing handwriting recognition and text matching calculation, the time-consuming problem of disassembly of template test papers is solved, the speed and accuracy of test paper correction are improved, and it is suitable for multi-subject test paper information retrieval.
Patent Information
- Application Number
- CN202510478753.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-01
AI Technical Summary
In the existing technology, in the full intelligent correction of student homework/examination papers, the disassembly and data extraction of template test papers takes a long time, and the large differences in the answers of students in Chinese test papers lead to low accuracy in image similarity matching.
Build a paper library, store the paper information and use handwritten text recognition and erase, search the template test paper information through text matching, and calculate the final matching degree based on the number of words on the front and back text content.
It improves the speed and accuracy of template test paper information retrieval, and is suitable for the correction of test papers in multiple subjects, especially the matching of Chinese composition paper information.
Smart Images

Figure BDA0005362066100000031 
Figure BDA0005362066100000061 
Figure BDA0005362066100000071
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent education technology, and particularly relates to a method, device and storage medium for retrieving and matching examination paper surfaces. Background Art
[0002] Currently, the large model technology is developing rapidly and has been widely applied in the field of intelligent education. At present, in the full-intelligent marking of students' homework / examination papers, a common method is to first scan the template examination paper, perform page layout splitting and question data extraction on the template examination paper, and then mark the scanned answer sheet images of students based on this. Generally speaking, various large model-related technologies are used for disassembling and data extraction of the template examination paper, which takes a relatively long time.
[0003] In actual scenarios, each grade is divided into multiple classes. Therefore, the same examination paper may be used by multiple classes and then scanned and marked separately in each class. Therefore, it is not necessary to make template examination papers and extract data for each class, and the same template examination paper data can be shared for the same assignment. Therefore, how to quickly, effectively and accurately retrieve and utilize the existing template examination paper information is very important. On the one hand, it saves computing power, and on the other hand, it can improve the speed of marking examination papers. Especially for Chinese examination papers, the content of students' answers varies greatly. If the template examination paper data is directly matched through simple image similarity, the matching accuracy is relatively low. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention aims to provide a method, device and storage medium for retrieving and matching examination paper surfaces.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for retrieving and matching examination paper surfaces includes the following steps:
[0007] S1. Construct an examination paper surface library, in which multiple examination paper surface information is stored. Each examination paper surface information includes a unique identifier of the examination paper surface, front data information of the examination paper surface and back data information of the examination paper surface. The front data and back data of the examination paper surface are both mapped to the unique identifier of the examination paper surface; the front data of the examination paper surface includes the front text content of the examination paper surface, and the back data of the examination paper surface includes the back text content of the examination paper surface;
[0008] S2. Obtain a first examination paper surface image, which includes a first front image and a first back image;
[0009] S3. Perform handwritten character recognition and erasure on the first front image and the first back image to obtain a second front image and a second back image, and the second front image and the second back image form a second examination paper surface image;
[0010] S4. Extract the text content of the second front image and the second back image, and perform matching retrievals in the front-side data and back-side data of the examination paper library respectively. If valid results are matched, reuse the matched examination paper information.
[0011] Further, the specific process of step S4 is as follows:
[0012] S4.1. Perform a text matching degree retrieval on the text content of the second front image in the front-side data of the examination paper library, and perform a text matching degree retrieval on the text content of the second back image in the back-side data of the examination paper library, and screen out the results with a text matching degree greater than a preset threshold;
[0013] S4.2. The matching results of the front-side data of the examination paper form the front-side matching result set of the examination paper, and the matching results of the back-side data of the examination paper form the back-side matching result set of the examination paper. If either the front-side matching result set or the back-side matching result set of the examination paper is empty, mark the front-side matching result set and the back-side matching result set as the first retrieval result and end the retrieval, otherwise go to step S4.3;
[0014] S4.3. Confirm the unique identification of the examination paper mapped by each result in the front-side matching result set and the back-side matching result set of the examination paper, and screen out the intersection of the unique identifications of the examination paper mapped by the front-side matching result set and the back-side matching result set of the examination paper, and then calculate the final matching degree corresponding to each unique identification of the examination paper in the intersection;
[0015] S4.4. Use the unique identification of the examination paper with the highest final matching degree as the retrieval result.
[0016] Even further, the process of calculating the final matching degree corresponding to the unique identification of the examination paper in step S4.3 is as follows:
[0017] S4.3.1. Respectively obtain the number of words Pn1 of the front-side text content and the number of words Pn2 of the back-side text content mapped by the unique identification of the examination paper;
[0018] S4.3.2. Respectively obtain the text matching degree Zn1 and Zn2 obtained from the retrieval matching of the front-side data and the back-side data of the examination paper mapped by the unique identification of the examination paper in step S4.1;
[0019] S4.3.3. Calculate the final matching degree of the unique identification of the examination paper according to the following formula:
[0020]
[0021] The present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of claims 1-3 is implemented.
[0022] The present invention also provides a computer device, including a processor and a memory, where the memory is used to store a computer program; when the processor is used to execute the computer program, the above method is implemented.
[0023] The beneficial effects of the present invention are as follows:
[0024] 1. The present invention retrieves the examination paper surface information of the template by means of text matching. Compared with simple image similarity matching retrieval, it not only has a faster speed but also a higher accuracy rate.
[0025] 2. The method of the present invention separately retrieves the front and back sides of the examination paper surface, and calculates the confidence level of the retrieval result according to the number of words in the text on the front and back sides of the retrieval result, which can improve the accuracy and credibility of the retrieval result, and thus can be applied to the retrieval and matching of examination paper surface information of more subjects. Specific Embodiments
[0026] The following will further describe the present invention. It should be noted that this embodiment is based on the present technical solution, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to this embodiment.
[0027] This embodiment provides an examination paper surface retrieval and matching method, including the following steps:
[0028] S1. Construct an examination paper surface library, where multiple examination paper surface information is stored in the examination paper surface library. Each examination paper surface information includes an examination paper surface unique identifier, examination paper surface front data information, and examination paper surface back data information. The examination paper surface front data and the examination paper surface back data are both mapped to the examination paper surface unique identifier; the examination paper surface front data includes the examination paper surface front text content, and the examination paper surface back data includes the examination paper surface back text content.
[0029] For example, if the examination paper surface unique identifier is 123456, the corresponding examination paper surface front text content is abcde, and the corresponding examination paper surface back text content is ABCDE, then when the text content on a page of the image is abcde, it can be determined that it corresponds to the front of the examination paper surface with the unique identifier of 123456.
[0030] According to actual application needs, each examination paper surface information may further include various examination question data of the examination paper surface, including question stems, standard answers, answer analyses, etc., to realize further use after successfully matching the examination paper surface information.
[0031] S2. Obtain a first examination paper surface image, where the first examination paper surface image includes a first front image and a first back image. The first examination paper surface image is the original examination paper surface image to be retrieved and matched, and can be obtained by double-sided scanning.
[0032] S3. Perform handwriting recognition and erasure on the first front image and the first back image to obtain a second front image and a second back image. The second front image and the second back image constitute a second test paper image. The handwriting differences among students are very large. Especially in subjects with a large amount of written answers, such as Chinese, the handwriting differences will seriously affect the accuracy of retrieval and matching through the test paper. The method of this embodiment can reduce the impact of handwriting differences on test paper retrieval and matching through handwriting recognition and erasure.
[0033] S4. Extract the text content of the second front image and the second back image, and perform matching retrieval in the front data and back data of the test papers in the test paper library respectively. If a valid result is matched, reuse the matched test paper information.
[0034] Further, in this embodiment, the specific process of step S4 is as follows:
[0035] S4.1. Perform text matching degree retrieval on the text content of the second front image in the front data of the test papers in the test paper library, and perform text matching degree retrieval on the text content of the second back image in the back data of the test papers in the test paper library, and screen out the results with a text matching degree greater than a preset threshold;
[0036] S4.2. The matching results of the front data of the test papers form a front matching result set of the test papers, and the matching results of the back data of the test papers form a back matching result set of the test papers. If either the front matching result set or the back matching result set of the test papers is empty, mark the front matching result set and the back matching result set of the test papers as the first retrieval result and end the retrieval. Otherwise, go to step S4.3;
[0037] S4.3. Confirm the unique test paper identifiers mapped by the results in the front matching result set and the back matching result set of the test papers, and screen out the intersection of the unique test paper identifiers mapped by the front matching result set and the back matching result set of the test papers, and then calculate the final matching degrees corresponding to the unique test paper identifiers in the intersection;
[0038] S4.4. Use the unique test paper identifier with the highest final matching degree as the retrieval result.
[0039] Even further, in this embodiment, the process of calculating the final matching degree corresponding to the unique test paper identifier in step S4.3 is as follows:
[0040] S4.3.1. Obtain the number of words Pn1 of the front text content and the number of words Pn2 of the back text content mapped by the unique test paper identifier respectively;
[0041] S4.3.2. Obtain the text matching degrees Zn1 and Zn2 obtained from the front - side data and back - side data of the test paper mapped by the unique test - paper identifier respectively in the retrieval and matching in step S4.1;
[0042] S4.3.3. Calculate the final matching degree of the unique test - paper identifier according to the following formula:
[0043]
[0044] Retrieval in the test - paper library is based on text information. Therefore, for a page with more words in the text content, the confidence level corresponding to its retrieval result is higher. Taking the Chinese composition test paper as an example, the front side generally records the requirements of the composition title, and the printed text content has more words, so the confidence level is high. The back side is generally a grid page for writing the composition, and the printed text content has fewer words. The confidence level of the result retrieved based on the text will decrease accordingly. In this embodiment, by combining the number of words in the text content on the front and back sides of the test paper to calculate the final matching degree corresponding to each unique test - paper identifier, the accuracy of the retrieval result can be improved.
[0045] For example, in the test - paper library, there are stored unique test - paper identifiers 123 and 124 and their mapped front - side data of the test paper (a large amount of printed text, for the sake of simplifying the calculation, it is assumed that the number of words in the text content is 90 for both here) and back - side data of the test paper (very little printed text, for the sake of simplifying the calculation, it is assumed that the number of words in the text content is 10 for both here).
[0046] After removing the handwritten characters from the front - side image and back - side image of the test - paper image to be retrieved and matched, extract the text content, and use the extracted text content to perform retrieval in the test - paper library. Set the text matching degree threshold to 70. The retrieval results are shown in Table 1.
[0047] Table 1
[0048]
[0049] It can be seen from the results in Table 1 that the intersection of the unique test - paper identifiers mapped by the matching results on the front and back sides is 123 and 124.
[0050] Then the final matching degree of the unique test - paper identifier 123 is 95.4:
[0051]
[0052] The final matching degree of the unique test - paper identifier 124 is 94.7:
[0053]
[0054] Therefore, the final retrieval result is the test - paper information of the unique test - paper identifier 123.
[0055] For those skilled in the art, various corresponding changes and modifications can be made based on the above technical solutions and concepts, and all such changes and modifications should be included within the protection scope of the claims of this invention.
Claims
1. A method for retrieving and matching examination papers, characterized in that, It includes the following steps: S1. Construct an examination paper library, in which multiple examination paper information is stored. Each piece of examination paper information includes a unique identifier of the examination paper, front-side data information of the examination paper, and back-side data information of the examination paper. Both the front-side data and the back-side data of the examination paper are mapped to the unique identifier of the examination paper; the front-side data of the examination paper includes the front-side text content of the examination paper, and the back-side data of the examination paper includes the back-side text content of the examination paper; S2. Obtain a first examination paper image, which includes a first front image and a first back image; S3. Perform handwritten character recognition and erasure on the first front image and the first back image to obtain a second front image and a second back image, and the second front image and the second back image form a second examination paper image; S4. Extract the text content of the second front image and the second back image, and perform matching retrieval in the front-side data and the back-side data of the examination paper in the examination paper library respectively. If a valid result is matched, reuse the matched examination paper information.
2. The method according to claim 1, wherein The specific process of step S4 is as follows: S4.
1. Perform text matching degree retrieval on the text content of the second front image in the front-side data of the examination paper library, and perform text matching degree retrieval on the text content of the second back image in the back-side data of the examination paper library, and screen out the results with a text matching degree greater than a preset threshold; S4.
2. The matching results of the front-side data of the examination paper form a front-side matching result set, and the matching results of the back-side data of the examination paper form a back-side matching result set. If either the front-side matching result set or the back-side matching result set is empty, mark the front-side matching result set and the back-side matching result set as the first retrieval result and end the retrieval, otherwise go to step S4.3; S4.
3. Confirm the unique identifiers of the examination papers mapped by each result in the front-side matching result set and the back-side matching result set, and screen out the intersection of the unique identifiers of the examination papers mapped by the front-side matching result set and the back-side matching result set, and then calculate the final matching degree corresponding to each unique identifier in the intersection; S4.
4. Use the unique identifier of the examination paper with the highest final matching degree as the retrieval result.
3. The method according to claim 2, wherein The process of calculating the final matching degree corresponding to the unique identifier of the examination paper in step S4.3 is: S4.3.
1. Respectively obtain the number of words Pn1 of the front-side text content and the number of words Pn2 of the back-side text content mapped by the unique identifier of the examination paper; S4.3.
2. Respectively obtain the text matching degrees Zn1 and Zn2 obtained from the retrieval matching of the front-side data and the back-side data mapped by the unique identifier of the examination paper in step S4.1; S4.3.
3. Calculate the final matching degree of the unique identifier of the examination paper according to the following formula:
4. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1-3 is implemented.
5. A computer device, characterized in that, It includes a processor and a memory. The memory is used to store a computer program; when the processor executes the computer program, the method described in any one of claims 1-3 is implemented.
Citation Information
Cited By
Paper sheet double-side correction judgment method, device and equipment and readable medium
CN120976937A