Paper teaching aid identification method and device based on Chinese feature code, equipment and medium
By adding feature codes to the headers and footers of printed teaching materials, combined with OCR technology and a server database, the problem of positioning and recognition in the connection between printed teaching materials and digital systems was solved, achieving efficient and accurate identification and correction of teaching materials, and reducing the cost of transformation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG MOKEN EDUCATION TECH CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-26
AI Technical Summary
The integration of paper-based teaching materials with digital systems presents challenges such as difficulties in page positioning, low accuracy in identifying question areas, insufficient error tolerance, and high transformation costs, resulting in low grading efficiency and poor accuracy.
This paper-based teaching aid recognition method is designed based on Chinese feature codes. By adding feature codes to the header and footer of the teaching aid pages, and combining OCR technology with a server database, a seamless integration is achieved. This includes feature code design, embedding, data association, image recognition and matching, and region segmentation, supporting high error tolerance and efficient positioning.
It achieves the goal of not changing the original layout and paper material of teaching materials, improving recognition efficiency and accuracy, supporting quick viewing of grading results and data analysis, and reducing transformation costs.
Smart Images

Figure CN122090464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic examination papers, and in particular to a method, device, equipment and medium for identifying paper-based teaching aids based on Chinese feature codes. Background Art
[0002] In the current process of educational informatization, paper-based teaching aids are still an important carrier for teaching activities of teachers and students. However, there are obvious pain points in the connection between paper-based teaching aids and digital marking and statistical systems: 1. Difficult page positioning: Traditional paper-based teaching aids have no exclusive identifiers. After being scanned or photographed and uploaded, it is difficult for the system to quickly match the corresponding teaching aid name, page number and specific question content, and manual intervention is required for confirmation, resulting in low efficiency. 2. Low recognition accuracy of question areas: The layout designs of different teaching aids vary greatly, and the system cannot accurately cut out the question areas and answer areas, resulting in recognition deviations easily occurring during subsequent marking, affecting the accuracy of marking. 3. Insufficient error tolerance: Existing image recognition technologies are easily interfered by blurred text and similar-looking characters. When characters such as "主", "玉", "王" are confused, it will directly lead to incorrect recognition results, thus affecting subsequent data matching. 4. High transformation cost: Some connection technologies require changing the original layout of paper-based teaching aids (such as adding QR codes and barcodes), which not only increases the printing cost of teaching aids but may also affect the reading experience of teachers and students.
[0003] In view of the above problems, there is an urgent need for a technology that does not require changing the original layout and paper material of teaching aids, has a high error tolerance and high recognition efficiency, and can achieve seamless connection between paper-based teaching aids and digital systems. Summary of the Invention
[0004] In view of the problems of the prior art, the present invention provides a method, device, equipment and medium for identifying paper-based teaching aids based on Chinese feature codes, which is ingeniously designed to achieve without modifying the basic attributes of teaching aids, without changing the original layout and paper material of paper-based teaching aids, without affecting the reading experience of teachers and students, and without increasing the printing cost; it has a high error tolerance, effectively avoids character recognition errors, and improves the accuracy of system data matching; it can achieve efficient positioning and recognition, greatly improving the recognition efficiency; it can be directly connected to the teacher-student end web page to quickly view marking results and perform data analysis.
[0005] To solve the above technical problems, the present invention adopts the following technical solutions: The present invention provides a method for identifying paper-based teaching aids based on Chinese feature codes, which includes the following steps: Step S1, Feature Code Design: Select N basic recognition characters from a preset set of Chinese characters that have been filtered for similar-looking characters and excluded from sensitive words; generate a basic feature code consisting of M of the aforementioned basic recognition characters for each page of the printed teaching aid, and assign a page number identifier character to identify the page number; Step S2, Feature Code Embedding: Without changing the original layout of the printed teaching aid, a first combination code containing the basic feature code and the page number identifier character is added to the header area of the printed teaching aid page with a fixed font size; and a second combination code containing the basic feature code, the page number identifier character, and a check code composed of K basic identification characters is added to the footer area of the printed teaching aid page with a fixed font size. Step S3, Data Association: Establish a database in the server backend to store structured data that uniquely corresponds to the basic feature code. The structured data includes at least teaching material metadata, question area coordinate information, and the page number identifier character. Step S4, Image Recognition and Matching: Acquire images of paper teaching aid pages containing the first and second combination codes through a terminal device; use OCR technology to recognize the first and second combination codes in the image; use the check code in the second combination code to verify the integrity of the recognition result; if the verification passes, match the corresponding structured data in the database according to the recognized basic feature code and page number identifier character; Step S5, Region Segmentation: Based on the coordinate information of the question region in the matched structured data, the acquired image is geometrically corrected and segmented to separate the question region from the answer region.
[0006] The preset Chinese character set contains 256 characters, the length M of the basic feature code is 12 digits, the page number identifier character uses uppercase Chinese numerals, and the length K of the check code is 4 digits.
[0007] In step S4, a similar-looking character correction step is also included: when a similar-looking character not included in the preset Chinese character set appears in the OCR recognition result, it is automatically corrected to the standard character corresponding to the preset Chinese character set according to the preset similar-looking character mapping rules.
[0008] In step S2, the font color of the first combination code and the second combination code is light gray or a low visual interference color similar to the background color, and the font size does not exceed 16px.
[0009] In step S4, if the second combination code cannot be fully recognized due to image damage, obstruction, or blurring, the system will switch to using the first combination code for preliminary matching and combine it with page layout structure features for auxiliary verification; the layout structure features include at least one of title position, question type distribution pattern, and white space ratio.
[0010] The paper-based teaching aid identification method includes a test question preparation step before step S1: the teaching aid content to be processed is made into a Word format file, ensuring that the question content and answer area layout in the file are complete and without formatting errors; the Word format file is imported into a test question creation tool for test question creation, which has the functions of "teaching aid content cutting", "automatic feature code generation" and "feature code addition".
[0011] After completing the region cutting in step S5, the following steps are also included: Step S6, Automated Grading: The cut-out answer area is intelligently analyzed, and grading results are generated according to the preset answer or scoring rules; Step S7, Data Visualization: The grading results are linked with student information and question information, and data query, statistical analysis and visualization services are provided to teachers and students through a web page.
[0012] The present invention also provides a recognition device based on the aforementioned paper-based teaching aid recognition method, comprising: Teaching aid production terminal: used to execute steps S1 and S2 to generate teaching aid files with Chinese character codes; Server-side: Includes a database module and a processing module; the database module is used to execute step S3 and store structured data; the processing module is used to execute steps S4 and S5 to perform image recognition, verification, matching, and region segmentation; User terminal: Used to perform image acquisition in step S4 and upload the image to the server.
[0013] The present invention also provides an electronic device, including a central processing unit and a memory, wherein the central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the paper-based teaching aid recognition method.
[0014] The present invention also provides a computer-readable storage medium that stores a computer program implemented according to the paper-based teaching aid recognition method in the form of computer-readable instructions. When the computer program is called and executed by a computer, it performs the steps included in the corresponding paper-based teaching aid recognition method.
[0015] The beneficial effects of this invention are: The present invention is ingeniously designed. Under the above settings, it is possible to achieve the goal without modifying the basic attributes of teaching aids, without changing the original layout and paper material of paper-based teaching aids. Only Chinese feature codes with a font size of 16px are added at fixed positions, which does not affect the reading experience of teachers and students and does not increase the printing cost. The present invention has a high error tolerance rate: by screening similar Chinese characters, establishing correction rules, and adding a 4-bit check code, it effectively avoids character recognition errors and improves the accuracy of system data matching. The present invention can achieve efficient positioning and recognition: the combination of 12-bit basic feature codes + 1-bit page number identification character can achieve the unique identification of teaching aid pages. Combining with the background positioning technology, the matching of images and teaching aid information can be completed within seconds, greatly improving the recognition efficiency. The present invention supports subsequent digital processes: the accurate cutting of the question area provides a basis for automated marking and answering data statistics, and can be directly connected to the teacher-student web page to achieve quick viewing of marking results and data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flowchart of the paper-based teaching aid recognition method based on Chinese feature codes of the present invention.
[0017] Figure 2 It is a schematic diagram of the image of the paper-based teaching aid of the present invention. SPECIFIC EMBODIMENTS
[0018] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with embodiments and drawings. The content mentioned in the embodiments is not a limitation to the present invention. The present invention will be described in detail below with reference to the drawings.
[0019] Embodiment 1 As shown in the embodiment of the present application Figure 1 A paper-based teaching aid recognition method based on Chinese feature codes includes the following steps: Step S1, Feature Code Design: Select N basic recognition characters from a preset Chinese character set that has been filtered for similar Chinese characters and excluded sensitive words; Generate a basic feature code composed of M of the basic recognition characters for each page of paper-based teaching aids, and assign a page number identification character for identifying the page number; The system pre-constructs a strictly screened 256-character Chinese character set, excludes sensitive words such as politics and religion, and normalizes easily confused similar Chinese characters such as "主 / 玉 / 王" and "日 / 曰" (only one standard character is retained for each group); On this basis, generate a unique 12-bit basic feature code (as a "digital ID card") for each page of teaching aids, and append 1 capital Chinese numeral as the page number identification character.
[0020] Step S2, Feature Code Embedding: Without changing the original layout of the paper-based teaching aid, add the first combined code containing the basic feature code and the page number identification character to the header area of the paper-based teaching aid page in a fixed font size; and add the second combined code containing the basic feature code, the page number identification character, and the check code composed of K basic recognition characters to the footer area of the paper-based teaching aid page in a fixed font size. This step S2 realizes dual-position redundant embedding. To balance aesthetics and robustness, the first combined code is embedded in the center of the header, and the second combined code is embedded on the left side of the footer. The entire process does not change the original layout, ensuring that the user reading experience is not affected.
[0021] Among them, the preset Chinese character set contains 256 characters, the length M of the basic feature code is 12 bits, the page number identification character uses capital Chinese numerals, and the length K of the check code is 4 bits. Among them, in step S2, the font colors of the first combined code and the second combined code are light gray or low-visual interference colors similar to the background color, and the font size does not exceed 16px.
[0022] Step S3, Data Association: Establish a database in the server background to store the structured data uniquely corresponding to the basic feature code. The structured data at least includes teaching aid metadata, question area coordinate information, and the page number identification character. In this step, structured data binding is realized. Before the teaching aid is printed, the precise coordinates (X, Y, Width, Height) of key elements such as the question area and the answer area of each page are extracted through a dedicated editing tool, together with metadata such as the teaching aid name, subject, and grade, and are packaged in JSON format and uniquely bound to the basic feature code of this page and stored in the server database.
[0023] Step S4, Image Recognition and Matching: Collect the paper-based teaching aid page image containing the first combined code and the second combined code through a terminal device; use OCR technology to recognize the first combined code and the second combined code in the image; use the check code in the second combined code to perform integrity verification on the recognition result; if the verification passes, match the corresponding structured data in the database according to the recognized basic feature code and page number identification character. In this step, after the user takes a photo of the teaching aid page with a mobile phone or other terminal, the system uses OCR technology to recognize the combined codes in the header and footer; if the recognition result contains non-standard similar-shaped characters (such as "主"), they are automatically corrected to standard characters (such as "王") according to the preset mapping rules; subsequently, the system preferentially analyzes the complete code in the footer and uses the check code (such as CRC or custom hash) to verify the integrity of the entire code string; only when the verification passes, the database is queried with the "basic code + page number" as the key value to obtain the corresponding structured data.
[0024] Step S5, Region Cutting: According to the question area coordinate information in the matched structured data, perform geometric correction and region cutting on the collected image to separate the question area and the answer area; specifically, this step achieves precise region cutting; after obtaining the structured data, perform geometric correction such as perspective transformation on the original image to eliminate shooting distortion, and then accurately crop the question area and the answer area of each question according to the pre-stored coordinate information, providing high-quality input for subsequent intelligent marking, content analysis, or AR interaction.
[0025] Specifically, the embodiments of this application are ingeniously designed. Under the above settings, it is possible to achieve without modifying the basic attributes of teaching aids, without changing the original layout and paper material of paper teaching aids, and only adding Chinese feature codes with a font size of 16px at fixed positions, without affecting the reading experience of teachers and students and without increasing the printing cost; the embodiments of this application have a high error tolerance rate: by screening similar-shaped characters, establishing correction rules, and adding a 4-bit check code, it effectively avoids character recognition errors and improves the accuracy of system data matching; the embodiments of this application can achieve efficient positioning and recognition: the combination of 12-bit basic feature codes + 1-bit page number identification character can uniquely identify the teaching aid page. Combining with the background positioning technology, the matching of the image and the teaching aid information can be completed in seconds, greatly improving the recognition efficiency; the embodiments of this application support subsequent digital processes: the precise cutting of the question area provides a basis for automatic marking and answering data statistics, and can be directly connected to the teacher-student end web page to achieve quick viewing of marking results and data analysis.
[0026] In the embodiments of this application, the specific principles of steps S1 to S5 are as follows: Feature Code Design (S1): Its advancement lies in the active design. Instead of random selection, it selects basic characters from a pre-filtered Chinese character set (containing 256 characters); the generated feature code consists of two parts: a 12-bit basic feature code that uniquely identifies the page, and a 1-bit page number identification character (such as "壹") for sorting and positioning, which avoids sensitive words and easily confused words from the source.
[0027] Feature Code Embedding (S2): Its non-invasive and robust design; the embedding follows two principles: Do not change the original format: This ensures the reading experience of teaching aids and low transformation costs; Dual-position redundant embedding: Place the "core identification code" (basic code + page number) in the header and the "complete verification code" (basic code + page number + 4-bit check code) in the footer respectively; this separates quick recognition and integrity verification.
[0028] The header number is located at the top center of the page, and its core function is visual indexing and quick rough positioning, allowing users or systems to see it immediately. The footer number is located at the bottom of the page, making it less likely to obstruct the view, and its core function is precise verification and fault recovery. The two are physically separate but logically strongly related. The setting of header and footer numbers is similar to the "preamble + data frame + checksum" structure in communication protocols. The header number is "preamble + partial information," used for waking up and pointing; the footer number is "complete data frame + strong checksum," used for precise reception and verification.
[0029] Data association (S3): Establishing a digital twin; on the server side, a structured data package containing teaching material metadata and precise question area coordinate information is pre-bound for each feature code; this achieves a unique mapping from physical pages to digital information.
[0030] Image Recognition and Matching (S4): A precise recognition and verification process. This step clarifies that data is collected using a terminal (such as a mobile phone). After recognizing the dual-position code using OCR, the checksum from the footer is first used to verify data integrity. After successful verification, the "basic feature code + page number" is used to perform precise matching in the database. This "verify first, then match" sequence ensures the reliability of the data source.
[0031] Region Segmentation (S5): The foundation for accurate output; based on the coordinate information in the structured data matched in the previous step, after geometric correction of the image, the target question area and answer area can be accurately segmented, preparing precise image materials for grading.
[0032] In this embodiment of the application, after completing the region cutting in step S5, the following steps are also included: Step S6, Automated Grading: The cut-out answer area is intelligently analyzed, and grading results are generated according to preset answer or scoring rules; In step S6, the intelligent analysis includes at least one of handwriting recognition, semantic understanding, formula analysis and graphic comparison, and supports mixed grading of subjective and objective questions; Step S7, Data Visualization: The grading results are linked with student information and question information, and data query, statistical analysis and visualization services are provided to teachers and students through a web page.
[0033] Furthermore, the system integrates the grading results from step S6 with students' historical answer data to construct a personal knowledge graph. Based on knowledge weaknesses, the system automatically pushes micro-lesson videos, variation exercises, or error notebooks, and feeds back new answer data to update the graph, forming a personalized learning loop of "identification—grading—diagnosis—consolidation".
[0034] In the embodiment of the present application, in step S4, it further includes a correction step for similar Chinese characters: when a similar Chinese character that is not included in the preset Chinese character set appears in the OCR recognition result, according to the preset mapping rule for similar Chinese characters, it is automatically corrected to the corresponding standard character in the preset Chinese character set.
[0035] In the embodiment of the present application, in step S4, if the second combined code cannot be completely recognized due to image damage, occlusion or blurring, the system will instead use the first combined code for preliminary matching and perform auxiliary verification in combination with the page layout structure features; the layout structure features include at least one of the title position, question type distribution pattern, and blank area ratio.
[0036] In the embodiment of the present application, before step S1 of the paper-based teaching aid recognition method, it further includes a test question preparation step: making the teaching aid content to be processed into a Word format file, ensuring that the question content and answer area layout in the file are complete and there is no format disorder; importing the above Word format file into a test question making tool for test question making, and this test question making tool has functions of "teaching aid content cutting", "automatic generation of feature codes", and "adding feature codes".
[0037] In the embodiment of the present application, in step S1, the verification code is generated by a hash algorithm based on the basic feature code and page number identification characters, and the hash algorithm includes an intercepted segment of CRC32, custom weighted exclusive OR checksum or national secret SM3 digest.
[0038] In the embodiment of the present application, in step S1, part of the bits of the basic feature code are jointly generated by a static teaching aid identifier and dynamic factors, and the dynamic factors include the printing batch number, time stamp or school institution code; the server side synchronously verifies the validity of the dynamic factors during matching to achieve the authenticity identification of teaching aids and piracy tracking.
[0039] In the embodiment of the present application, in step S1, a dedicated Chinese character set is constructed: N non-similar, semantically neutral and non-sensitive Chinese characters are selected from the first-level Chinese characters in GB2312 to form a dedicated character set C; where, for any two Chinese characters C1, C2 ∈ C, the Euclidean distance of the glyphs D(C1, C2) ≥ τ, and τ is a preset threshold; the size N of the dedicated character set C = 256; Generate a page-level combined code: Assign a unique 12-bit basic feature code B to each page of the paper-based teaching aid, where B[i] ∈ C, i = 1 to 12; obtain the page number p of this page in the affiliated teaching aid and map it to a capital Chinese numeral character P (such as p = 3 → P = "叁"); based on B and P, generate a 4-bit verification code V through a verification algorithm A, where V[j] ∈ C, j = 1 to 4; construct the first combined code M1 = B + P, and the second combined code M2 = B + P + V; Embedding the combination code into the teaching aid layout: Embed M1 in the center of the header area of the teaching aid page with a font size of 16px and black font; embed M2 in the left side of the footer area of the teaching aid page with a font size of 16px and black font; the embedding operation does not change the original teaching aid's question layout, white space area and answer box layout. Establish a structured data mapping. In the server-side database, create a record for each basic feature code B. This record contains: Teaching aid metadata fields: Teaching aid ID, subject, grade, chapter; Page metadata fields: Page number (p), Total number of pages; List of coordinates for the question area: {Q1:x1,y1,w1,h1),Q2:(x2,y2,w2,h2),...,Qk:(xk,yk,wk,hk)}, where each Qi represents the normalized coordinates of the answer area for the i-th question in the page image; Terminal image acquisition and OCR recognition: The user takes a picture of the teaching aid page I containing M1 and M2 through a mobile terminal; Image I is preprocessed, including noise reduction, perspective correction and text region localization; Using an OCR engine optimized for character set C, the text content of the regions where M1 and M2 are located is recognized respectively, and the recognition results R1 and R2 are obtained. Similar-looking character correction and integrity verification: For each character r in R1 and R2, if r If C is a character, then query the pre-set mapping table T for similar-looking characters and replace r with T(r)∈C; Using check algorithm A, recalculate the check code V' for the first 13 bits (i.e. B+P) of the corrected R2; If V' and the last 4 bits of R2 are completely identical, the verification is considered successful, and the basic feature code B and the page number character P are extracted; otherwise, the recognition fails. Structured data matching and region segmentation: Search the database for records that satisfy the following conditions: basic feature code = B* and page number corresponding to Chinese numerals = P*; if the match is successful, obtain the list of question region coordinates in the record; Based on the coordinate list, perform ROI (Region of Interest) cropping on the original image I, and output k independent question-answer image blocks {I1,I2,...,Ik}.
[0040] The verification algorithm A is a CRC-16 variant based on polynomial division, and its generator polynomial is G(x) = x. 4 +x+1, the operation domain is defined on a finite symbol space with character set C as the alphabet, and each bit of the check code V is generated by table lookup or modular arithmetic.
[0041] The similar-looking character mapping table T is constructed as follows: GB2312 Level 1 Chinese characters are clustered, and ResNet-18 is used to extract the deep feature vectors of the character images; in the feature space, for each character c∈C, all non-C set Chinese characters whose L2 distance is less than δ are found, forming its similar-looking character set N(c); T(n) = c is defined. n∈N(c), thus establishing a one-way mapping relationship.
[0042] Each answer image block I i The input is fed into a pre-trained question-answering content understanding model, which outputs the answer text T. i and confidence score i According to the pre-stored standard answer A in the database i and scoring rules R i , for T i Perform semantic comparison or pattern matching; generate grading results, including: correct / incorrect judgment, score, error type label, and associate the results with student ID, question ID and timestamp, and store them in the teaching analysis database.
[0043] Example 2 In Embodiment 2 of this application, a recognition device based on the aforementioned paper-based teaching aid recognition method is provided, comprising: Teaching aid production terminal: used to execute steps S1 and S2 to generate teaching aid files with Chinese character codes; Server-side: Includes a database module and a processing module; the database module is used to execute step S3 and store structured data; the processing module is used to execute steps S4 and S5 to perform image recognition, verification, matching, and region segmentation; User terminal: Used to perform image acquisition in step S4 and upload the image to the server.
[0044] Furthermore, the processing module is further configured with a similar-looking character correction unit and a multimodal fusion recognition unit. When the feature code recognition fails, the multimodal fusion recognition unit calls a layout analysis model based on deep learning to extract the semantic features of the page layout and performs similarity matching with the template layout stored in the database to achieve assisted positioning.
[0045] Furthermore, the functions of step S4 are deployed locally on the user terminal, including: integrating a lightweight OCR engine and a locally cached feature code-structured data index table on the terminal device; when the network is unavailable, the local cache is used first for feature code matching and region segmentation, and the batch correction results are synchronized to the server after the network is restored.
[0046] Example 3 In Embodiment 3 of this application, an electronic device is provided, including a central processing unit and a memory. The central processing unit is used to call and run a computer program stored in the memory to execute the steps of the paper-based teaching aid recognition method.
[0047] Example 4 In Embodiment 4 of this application, a computer-readable storage medium is provided, which stores a computer program implemented according to the paper-based teaching aid recognition method in the form of computer-readable instructions. When the computer program is called and executed by a computer, it performs the steps included in the corresponding paper-based teaching aid recognition method.
[0048] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes, and modifications made to the above embodiments based on the present invention without departing from the scope of the present invention are within the scope of the present invention.
Claims
1. A method for recognizing paper-based teaching aids based on Chinese character feature codes, characterized in that, Includes the following steps: Step S1, Feature Code Design: Select N basic recognition characters from a pre-set set of Chinese characters that have been filtered for similar-looking characters and excluded from sensitive words; Generate a basic feature code consisting of M basic identification characters for each page of the paper teaching aid, and assign a page number identification character to identify the page number; Step S2, Feature Code Embedding: Without changing the original layout of the printed teaching aid, a first combination code containing the basic feature code and the page number identifier character is added to the header area of the printed teaching aid page with a fixed font size; and a second combination code containing the basic feature code, the page number identifier character, and a check code composed of K basic identification characters is added to the footer area of the printed teaching aid page with a fixed font size. Step S3, Data Association: Establish a database in the server backend to store structured data that uniquely corresponds to the basic feature code. The structured data includes at least teaching material metadata, question area coordinate information, and the page number identifier character. Step S4, Image Recognition and Matching: Acquire an image of a printed teaching aid page containing the first and second combined codes using a terminal device; use OCR technology to recognize the first and second combined codes in the image; use the check code in the second combined code to verify the integrity of the recognition result; If the verification passes, the corresponding structured data is matched in the database based on the identified basic feature code and page number identifier character. Step S5, Region Segmentation: Based on the coordinate information of the question region in the matched structured data, the acquired image is geometrically corrected and segmented to separate the question region from the answer region.
2. The paper-based teaching aid identification method based on Chinese character codes according to claim 1, characterized in that: The preset Chinese character set contains 256 characters, the length M of the basic feature code is 12 digits, the page number identifier character uses uppercase Chinese numerals, and the length K of the check code is 4 digits.
3. The paper-based teaching aid identification method based on Chinese character codes according to claim 1, characterized in that: Step S4 also includes a similar-looking character correction step: when a similar-looking character not included in the preset Chinese character set appears in the OCR recognition result, it is automatically corrected to the standard character corresponding to the preset Chinese character set according to the preset similar-looking character mapping rules.
4. The paper-based teaching aid identification method based on Chinese character codes according to claim 1, characterized in that: In step S2, the font color of the first combination code and the second combination code is light gray or a low visual interference color similar to the background color, and the font size does not exceed 16px.
5. The paper-based teaching aid identification method based on Chinese character codes according to claim 1, characterized in that: In step S4, if the second combination code cannot be fully recognized due to image damage, obstruction, or blurring, the system will switch to using the first combination code for preliminary matching and combine it with page layout structure features for auxiliary verification; the layout structure features include at least one of title position, question type distribution pattern, and white space ratio.
6. The paper-based teaching aid identification method based on Chinese character codes according to claim 1, characterized in that: Before step S1, the paper-based teaching aid identification method also includes a test question preparation step: the teaching aid content to be processed is made into a Word format file, ensuring that the question content and answer area layout in the file are complete and without formatting errors; the Word format file is imported into a test question creation tool for test question creation, which has the functions of "teaching aid content cutting", "automatic feature code generation" and "feature code addition".
7. The paper-based teaching aid identification method based on Chinese character codes according to claim 1, characterized in that: After completing the area cutting in step S5, the following steps are also included: Step S6, Automated Grading: The cut-out answer area is intelligently analyzed, and grading results are generated according to the preset answer or scoring rules; Step S7, Data Visualization: The grading results are linked with student information and question information, and data query, statistical analysis and visualization services are provided to teachers and students through a web page.
8. A recognition device based on the paper teaching aid recognition method according to any one of claims 1-7, characterized in that, include: Teaching aid production terminal: used to execute steps S1 and S2 to generate teaching aid files with Chinese character codes; Server-side: Includes database module and processing module; The database module is used to execute step S3 and store structured data; the processing module is used to execute steps S4 and S5 to perform image recognition, verification, matching and region segmentation. User terminal: Used to perform image acquisition in step S4 and upload the image to the server.
9. An electronic device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to call and run a computer program stored in the memory to perform the steps of the paper-based teaching aid recognition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program implemented according to any one of claims 1 to 7 in the form of computer-readable instructions. When the computer program is called by a computer, it executes the steps included in the corresponding paper-based teaching aid recognition method.