A method, system, and electronic device for identifying and evaluating ancient book calculation tables.

By determining a self-recognition model based on the feature set of ancient calculation tables, and by combining image processing and recognition methods, the objective and subjective recognition information is output for comparison, which solves the problems of efficiency and accuracy in the recognition of ancient calculation tables, and realizes the optimization of the self-recognition model and the efficient capture of complex features.

CN120126162BActive Publication Date: 2025-12-02INNER MONGOLIA NORMAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510197844.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-12-02
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

Existing technologies cannot guarantee the initial compatibility between ancient book calculation tables and self-recognition models, resulting in wasted resources and low recognition efficiency. Furthermore, the recognition accuracy of self-recognition models is difficult to improve, making it difficult to capture the complex features of ancient book calculation tables.

Method used

By determining a self-recognition model based on ancient calculation tables and their corresponding feature sets, and combining image processing and recognition methods, objective recognition information is output and compared with subjective recognition information to evaluate the recognition indicators of the self-recognition model and determine the optimization direction.

Benefits of technology

It improved the recognition efficiency and accuracy of ancient book calculation tables, ensured the initial adaptability of the self-recognition model, and improved the capture rate of complex features by optimizing the direction, thereby enhancing the accuracy of text, numerical and other recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126162B_ABST
    Figure CN120126162B_ABST
Patent Text Reader

Abstract

This invention discloses a method, system, and electronic device for recognizing and evaluating ancient book calculation tables, relating to the field of ancient book calculation table recognition technology. The invention includes: S1. Staff upload ancient book calculation tables and their corresponding feature sets, and determine the self-recognition model of the ancient book calculation tables; S2. Randomly select several training samples from the ancient book calculation tables and input them into the self-recognition model, outputting objective recognition information of the ancient book calculation tables, and distributing the training samples to recognition personnel, who then output subjective recognition information of the ancient book calculation tables; S3. Evaluate the recognition indicators of the self-recognition model of the ancient book calculation tables, determine the optimization direction of the self-recognition model, and display the optimization direction of the self-recognition model. This invention avoids wasting resources on the self-recognition model and improves the recognition efficiency of ancient book calculation tables. This invention ensures the accuracy of ancient book calculation tables in text recognition, numerical recognition, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ancient book calculation table recognition technology, specifically to an ancient book calculation table recognition and evaluation method, system, and electronic device. Background Technology

[0002] Ancient mathematical tables, as crystallizations of mathematical wisdom, not only record specific calculation methods in ancient mathematics but also reflect the level, characteristics, and trends of mathematical development at that time. Studying these tables allows for a deeper understanding of the origins, evolution, and achievements of ancient mathematics, providing material evidence and theoretical support for the study of modern mathematical history. Identification and evaluation enable the systematic organization and protection of these precious cultural heritages, preventing information loss or damage due to time, environment, and other factors, and ensuring the uninterrupted transmission of mathematical knowledge. Therefore, the identification and evaluation of ancient mathematical tables are extremely necessary.

[0003] However, current technologies for recognizing ancient arithmetic tables still have some shortcomings, which can be reflected in the following aspects: Existing technologies rarely determine a self-recognition model for ancient arithmetic tables based on the ancient arithmetic tables and their corresponding feature sets. This makes it difficult to guarantee the initial adaptability between the ancient arithmetic tables and the self-recognition model, easily leading to a waste of resources in the self-recognition model and reducing the recognition efficiency of the ancient arithmetic tables. Furthermore, there is a lack of evaluation of training samples after recognizing them using the self-recognition model of the ancient arithmetic tables, and a lack of confirmation of the optimization direction of the self-recognition model. This makes it difficult to ensure the recognition accuracy of the self-recognition model of the ancient arithmetic tables and to capture the complex features of the ancient arithmetic tables. Summary of the Invention

[0004] The purpose of this invention is to provide a method, system, and electronic device for identifying and evaluating ancient book calculation tables, which solves the problems existing in the background art.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: The first aspect of the present invention provides a method for identifying and evaluating ancient book calculation tables, including: S1. Staff members upload ancient book calculation tables and their corresponding feature sets, and determine the self-identification model of the ancient book calculation tables.

[0006] The self-recognition model includes a combination of image processing methods and recognition methods.

[0007] S2. Randomly select several training samples from the ancient book calculation table and input them into the self-recognition model to output the objective recognition information of the ancient book calculation table. Then, assign several training samples of the ancient book calculation table to the recognition personnel, who will output the subjective recognition information of the ancient book calculation table.

[0008] S3. Evaluate the recognition indicators of the self-recognition model of the ancient book calculation table, determine the optimization direction of the self-recognition model of the ancient book calculation table, and display the optimization direction of the self-recognition model of the ancient book calculation table.

[0009] A second aspect of the present invention provides a system for performing the ancient book calculation table identification and evaluation method of the present invention, comprising: an ancient book calculation table uploading module, used by staff to upload ancient book calculation tables and their corresponding feature sets, and to determine the self-identification model of the ancient book calculation tables.

[0010] The self-recognition model includes a combination of image processing methods and recognition methods.

[0011] The ancient book calculation table recognition module is used to randomly select several training samples from the ancient book calculation table and input them into the self-recognition model, output the objective recognition information of the ancient book calculation table, and distribute several training samples of the ancient book calculation table to the recognition personnel, who then output the subjective recognition information of the ancient book calculation table.

[0012] The ancient book recognition and evaluation module is used to evaluate the recognition indicators of the self-recognition model of the ancient book calculation table, determine the optimization direction of the self-recognition model of the ancient book calculation table, and display the optimization direction of the self-recognition model of the ancient book calculation table.

[0013] A third aspect of the present invention provides an electronic device, comprising: a processor, a memory, and a communication bus. The memory stores a computer-readable program executable by the processor. The communication bus enables communication between the processor and the memory. When the processor executes the computer-readable program, it implements the ancient book calculation table identification and evaluation method as described in this invention.

[0014] The beneficial effects of the present invention are as follows: (1) Based on the ancient book calculation table and its corresponding feature set, the present invention determines the self-identification model of the ancient book calculation table, thereby ensuring the initial adaptability of the ancient book calculation table and the self-identification model, avoiding the waste of resources of the self-identification model, and improving the recognition efficiency of the ancient book calculation table, laying the foundation for the confirmation of the optimization direction of the self-identification model of the ancient book calculation table in the future.

[0015] (2) This invention evaluates the training samples by using the self-recognition model of ancient book calculation tables, and then confirms the optimization direction of the self-recognition model of ancient book calculation tables. To a certain extent, it improves the recognition accuracy of the self-recognition model of ancient book calculation tables, increases the capture rate of complex features of ancient book calculation tables, and ensures the accuracy of ancient book calculation tables in text recognition, numerical recognition, etc. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1This is a schematic diagram of the implementation steps of the method of the present invention.

[0018] Figure 2 This is a schematic diagram of the system structure connection of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Reference Figure 1 As shown, the first aspect of the present invention provides a method for identifying and evaluating ancient book calculation tables, including: S1. Staff members upload ancient book calculation tables and their corresponding feature sets, and determine the self-identification model of the ancient book calculation tables.

[0021] It should be noted that the feature set includes feature values ​​for complex content, complex table structure, and complex font, etc. The feature values ​​for complex content, complex table structure, and complex font are specifically values ​​between 0 and 1.

[0022] The self-recognition model includes a combination of image processing methods and recognition methods.

[0023] In a specific embodiment of the present invention, the process of determining the self-identification model of the ancient book calculation table is as follows: S100, identifying the defect parameters of the ancient book calculation table through image recognition technology.

[0024] The defect parameters are the overall characteristic parameters of the defects.

[0025] It should be noted that the overall defect characterization parameter is specifically the total defect area after normalization. For example, if the total defect area is 10 square meters and the total area of ​​the ancient book calculation table is 80 square meters, then after normalization of 0, 10, and 80, the total defect area after normalization is 0.125.

[0026] S101. Add the defect parameters of the ancient book calculation table to the feature set of the ancient book calculation table to obtain the updated feature set of the ancient book calculation table.

[0027] S102. Compare the update feature set of the ancient book calculation table with the difficulty level classification table of the ancient book calculation table stored in the web data warehouse, and obtain the difficulty level of the ancient book calculation table through matching.

[0028] It should be noted that the difficulty level classification table of the ancient book calculation table is the demand range of each element in the feature set corresponding to each difficulty level. For example, if the feature value of complex content is 0.1, the feature value of complex table structure is 0.1, the feature value of complex font is 0.1, and the overall defect representation parameter is 0.2, it falls within the demand range of each element in the feature set corresponding to the ease level of the ancient book calculation table. Therefore, the difficulty level of the ancient book calculation table is recorded as easy.

[0029] S103. Compare the difficulty level of the ancient book calculation table with the self-recognition model corresponding to each difficulty level of the ancient book calculation table stored in the web data warehouse, and obtain the self-recognition model of the ancient book calculation table through matching.

[0030] It should be noted that the self-recognition models corresponding to the various difficulty levels of the ancient book calculation tables are as follows: For example, if the difficulty level of the ancient book calculation table is easy, the image processing is denoising + binarization + tilt correction, and the recognition method is traditional OCR + rule-based table structure recognition. If the difficulty level of the ancient book calculation table is difficult, the image processing is super-resolution + contrast enhancement + table region detection, and the recognition method is deep learning OCR (such as CRNN or TrOCR) + deep learning table structure recognition (such as TableNet). The specific details are uploaded by the staff.

[0031] This invention determines the self-recognition model of ancient book calculation tables based on their corresponding feature sets, thereby ensuring the initial adaptability of ancient book calculation tables and self-recognition models, avoiding resource waste of self-recognition models, improving the recognition efficiency of ancient book calculation tables, and laying the foundation for confirming the optimization direction of self-recognition models of ancient book calculation tables in the future.

[0032] S2. Randomly select several training samples from the ancient book calculation table and input them into the self-recognition model to output the objective recognition information of the ancient book calculation table. Then, assign several training samples of the ancient book calculation table to the recognition personnel, who will output the subjective recognition information of the ancient book calculation table.

[0033] In a specific embodiment of the present invention, the objective identification information includes a text dataset, a numerical dataset, a symbolic dataset, and a calculation table structure dataset of several training samples.

[0034] The text dataset includes a header keyword set and a keyword set for each cell; the symbol dataset includes a set of operators and a set of table separators; and the calculation table structure dataset includes the number of rows, the number of columns, each merged cell, and each split cell.

[0035] The subjective recognition information includes a text dataset, a numerical dataset, a symbolic dataset, and a calculation table structure dataset of several training samples.

[0036] S3. Evaluate the recognition indicators of the self-recognition model of the ancient book calculation table, determine the optimization direction of the self-recognition model of the ancient book calculation table, and display the optimization direction of the self-recognition model of the ancient book calculation table.

[0037] In a specific embodiment of the present invention, the specific implementation steps of the self-recognition model for evaluating ancient book calculation tables and its various levels of recognition indicators are as follows: S300, based on the objective and subjective recognition information of the ancient book calculation tables, using a certain training sample as a designated sample, determine the first recognition indicator α of the designated sample of the ancient book calculation tables. _1 Second identification indicator α _2 Third identification indicator α _3 and the fourth identification indicator α _4 .

[0038] It should be noted that the second identification indicator for determining the specified sample of the ancient book calculation table is specifically determined by: using a set similarity algorithm to calculate the similarity between the numerical dataset in the objective identification information and the numerical dataset in the subjective identification information of the ancient book calculation table, and using this similarity as the second identification indicator for the ancient book calculation table.

[0039] It should also be noted that the set similarity algorithm is specifically as follows: Here, A and B refer to two sets respectively.

[0040] S301. Following this pattern, the first, second, third, and fourth recognition indices of each training sample of the ancient book calculation table are obtained. These are then averaged to obtain the average values ​​of the first, second, third, and fourth recognition indices of the ancient book calculation table, which serve as the first-level recognition index β of the self-recognition model of the ancient book calculation table. _1 Secondary identification index β _2 Level 3 identification index β _3 Level 4 Identification Indicator β _4 .

[0041] In a specific embodiment of the present invention, the first identification index α of the designated sample for determining the ancient book calculation table is... _1 The specific determination method is as follows: extract the set of header keywords and the set of keywords for each cell from the objective identification information of the specified sample of the ancient book calculation table, and extract the set of header keywords and the set of keywords for each cell from the subjective identification information of the specified sample of the ancient book calculation table.

[0042] The similarity algorithm is used to calculate the similarity between the set of header keywords in the objective identification information and the set of header keywords in the subjective identification information of a specified sample of the ancient book calculation table, and this similarity is used as the accuracy of header identification for the specified sample of the ancient book calculation table.

[0043] The similarity algorithm is used to calculate the similarity between the keyword set of each cell in the objective recognition information of a specified sample of the ancient book calculation table and the keyword set of the corresponding cell. This similarity is used as the recognition accuracy of each cell in the specified sample of the ancient book calculation table, and heterogeneous cells in the specified sample of the ancient book calculation table are filtered out.

[0044] It should be noted that the recognition accuracy is specifically a value between 0 and 1.

[0045] It should also be noted that the specific method for filtering the heterogeneous cells of the specified sample of the ancient book calculation table is as follows: the recognition accuracy of each cell of the specified sample of the ancient book calculation table is compared with the cell recognition accuracy threshold stored in the web data warehouse. If the recognition accuracy of a cell is less than the cell recognition accuracy threshold, the cell is recorded as a heterogeneous cell. The heterogeneous cells of the specified sample of the ancient book calculation table are obtained by filtering. The cell recognition accuracy threshold is specifically set by the staff of the ancient book calculation table. For example, in order to improve the text recognition accuracy of the ancient book calculation table, the cell recognition accuracy threshold is set to 0.95.

[0046] The number of heterogeneous cells M and the total number of cells M′ of a specified sample of the ancient book calculation table are summarized, and these summaries, along with the header recognition accuracy ε′ of the specified sample of the ancient book calculation table, are imported into the first recognition index model. In the process, the first identification index of a specified sample of the ancient book calculation table is output.

[0047] It should be noted that the ln() function is used in the first identification indicator model to control the value range within the range of 0-1, which facilitates subsequent analysis.

[0048] In a specific embodiment of the present invention, the third identification index α for determining the specified sample of the ancient book calculation table _3 The specific determination method is as follows: extract the set of operator symbols and the set of table delimiters from the objective identification information of the specified sample of the ancient book calculation table, and extract the set of operator symbols and the set of table delimiters from the subjective identification information of the specified sample of the ancient book calculation table.

[0049] Using a set similarity algorithm, the similarity between the set of operator symbols in the objective identification information and the set of operator symbols in the subjective identification information of a specified sample of the ancient book calculation table is calculated, which serves as the accuracy of operator symbol recognition for the specified sample of the ancient book calculation table. Similarly, the accuracy of table delimiter symbol recognition for the specified sample of the ancient book calculation table is calculated.

[0050] The accuracy η for recognizing operation symbols and the accuracy μ for recognizing table delimiters in ancient arithmetic tables are imported into the third recognition index model. In the formula, the third identification index of the specified sample of the ancient book calculation table is output. In the formula, η′ and μ′ represent the accuracy threshold for identifying operator symbols and the accuracy threshold for identifying table separator symbols stored in the web data warehouse, respectively. ∧ and ∨ are the logical symbols AND and OR, respectively.

[0051] It should be noted that the accuracy thresholds for operator recognition and table separator recognition are specifically set by the staff of the ancient book calculation table. For example, in order to improve the accuracy of operator recognition and table separator recognition in the ancient book calculation table, the accuracy thresholds for operator recognition and table separator recognition are both set to 0.95.

[0052] In a specific embodiment of the present invention, the fourth identification index α for determining the designated sample of the ancient book calculation table _4 The specific determination method is as follows: extract the number of rows, columns, merged cells and split cells of the calculation table structure dataset from the objective identification information of the specified sample of the ancient calculation table, and extract the number of rows, columns, merged cells and split cells of the calculation table structure dataset from the subjective identification information of the specified sample of the ancient calculation table, and record them as the reference number of rows, reference number of columns, reference merged cells and reference split cells of the ancient calculation table respectively.

[0053] Comparative analysis yields the row number deviation, column number deviation, merged cells for each deviation, and split cells for each deviation in a specified sample of the ancient book calculation table. The total number of merged cells and split cells for each deviation in the specified sample of the ancient book calculation table is then obtained.

[0054] It should be noted that the comparative analysis obtains the row number deviation value, column number deviation value, logical feedback value of each row and each column, each deviation merged cell, and each deviation split cell of the specified sample of the ancient book calculation table. The specific comparison process is as follows: the row number and column number of the specified sample of the ancient book calculation table are subtracted from the reference row number and reference column number, respectively, to obtain the row number deviation value and column number deviation value of the specified sample of the ancient book calculation table. Each merged cell of the ancient book calculation table is compared with each reference merged cell. If a merged cell fails to match with each reference merged cell, the merged cell is recorded as the deviation merged cell, thus obtaining each deviation merged cell. Similarly, each deviation split cell is obtained.

[0055] The number of reference merged cells and the number of reference split cells of the specified sample of the ancient book calculation table are summarized, and the fourth identification index of the specified sample of the ancient book calculation table is obtained through numerical processing.

[0056] It should be noted that the processing to obtain the fourth identification index of the specified sample of the ancient book calculation table is specifically processed as follows: the number of deviation merged cells of the specified sample of the ancient book calculation table is divided by the number of reference merged cells, and the number of deviation split cells is divided by the number of reference split cells, to obtain the deviation merged cell ratio and deviation split cell ratio of the specified sample of the ancient book calculation table, respectively.

[0057] The deviation values ​​of the number of rows, columns, merged cells, and split cells of a specified sample in the ancient book calculation table are compared with the allowed deviation values ​​of the number of rows, columns, merged cells, and split cells stored in the web data warehouse. If the deviation value of the number of rows is less than the allowed deviation value of the number of rows, the deviation value of the number of columns is less than the allowed deviation value of the number of columns, the merged cells ratio is less than the deviation merged cells ratio threshold, and the split cells ratio is less than the deviation split cells ratio threshold, then the fourth identification index of the specified sample in the ancient book calculation table is marked as 1; otherwise, it is marked as 0.

[0058] In a specific embodiment of the present invention, the method for determining the optimization direction of the self-recognition model of the ancient book calculation table is as follows: based on the first recognition index α of the ancient book calculation table. _1 Second identification index α _2 Third identification indicator α _3 and the fourth identification indicator α _4 .

[0059] If (β) _1 <β _ If ′1), then the optimization direction of the self-recognition model of the ancient book calculation table is denoted as text optimization.

[0060] If (β) _2 <β _ ′2), then the optimization direction of the self-recognition model of the ancient book calculation table is denoted as numerical optimization.

[0061] If (β) _3 If =0), then the optimization direction of the self-recognition model of the ancient book calculation table is denoted as symbol optimization.

[0062] If (β) _4 If =0), then the optimization direction of the self-recognition model of the ancient book calculation table is denoted as the calculation table structure optimization.

[0063] β _ ′1、β _ ′2 represents the convergence values ​​of the first-level identification indicators and the second-level identification indicators stored in the web data warehouse, respectively.

[0064] It should be noted that the convergence values ​​of the first and second identification indicators are specifically set by the staff of the ancient book calculation table. For example, in order to ensure the text recognition and numerical recognition effects of the self-recognition model of the ancient book calculation table, the convergence values ​​of the first and second identification indicators can both be set to 0.9.

[0065] The optimization directions for the self-recognition model of ancient book calculation tables are summarized.

[0066] This invention evaluates training samples by using a self-recognition model of ancient book calculation tables, and then confirms the optimization direction of the self-recognition model of ancient book calculation tables. To a certain extent, it improves the recognition accuracy of the self-recognition model of ancient book calculation tables, increases the capture rate of complex features of ancient book calculation tables, and ensures the accuracy of ancient book calculation tables in text recognition, numerical recognition, etc.

[0067] Reference Figure 2 As shown, a second aspect of the present invention provides a system for performing the ancient book calculation table identification and evaluation method described in the present invention, comprising:

[0068] The ancient book calculation table upload module is used by staff to upload ancient book calculation tables and their corresponding feature sets, and to determine the self-recognition model of the ancient book calculation tables.

[0069] The self-recognition model includes a combination of image processing methods and recognition methods.

[0070] The ancient book calculation table recognition module is used to randomly select several training samples from the ancient book calculation table and input them into the self-recognition model, output the objective recognition information of the ancient book calculation table, and distribute several training samples of the ancient book calculation table to the recognition personnel, who then output the subjective recognition information of the ancient book calculation table.

[0071] The ancient book recognition and evaluation module is used to evaluate the recognition indicators of the self-recognition model of the ancient book calculation table, determine the optimization direction of the self-recognition model of the ancient book calculation table, and display the optimization direction of the self-recognition model of the ancient book calculation table.

[0072] It should be noted that the present invention also includes a web data warehouse. The ancient book calculation table upload module is connected to the ancient book calculation table recognition module, the ancient book calculation table recognition module is connected to the ancient book recognition and evaluation module, and the web data warehouse is connected to both the ancient book calculation table upload module and the ancient book recognition and evaluation module.

[0073] A third aspect of the present invention provides an electronic device, comprising: a processor, a memory, and a communication bus. The memory stores a computer-readable program executable by the processor. The communication bus enables communication between the processor and the memory. When the processor executes the computer-readable program, it implements the ancient book calculation table identification and evaluation method as described in this invention.

[0074] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.

Claims

1. A method for identifying and evaluating ancient book calculation tables, characterized in that, include: S1. Staff members upload ancient book calculation tables and their corresponding feature sets, and determine the self-recognition model of the ancient book calculation tables; The self-recognition model includes a combination of image processing methods and recognition methods; The specific process for determining the self-recognition model for ancient book calculation tables is as follows: S100. Identify defect parameters in ancient calculation tables using image recognition technology; The defect parameters are overall defect characterization parameters; S101. Add the defective parameters of the ancient book calculation table to the feature set of the ancient book calculation table to obtain the updated feature set of the ancient book calculation table. S102. Compare the update feature set of the ancient book calculation table with the difficulty level classification table of the ancient book calculation table stored in the web data warehouse, and obtain the difficulty level of the ancient book calculation table through matching. S103. Compare the difficulty level of the ancient book calculation table with the self-recognition model corresponding to each difficulty level of the ancient book calculation table stored in the web data warehouse, and obtain the self-recognition model of the ancient book calculation table through matching. S2. Randomly select several training samples from the ancient book calculation table and input them into the self-recognition model to output the objective recognition information of the ancient book calculation table. Then, assign several training samples of the ancient book calculation table to the recognition personnel, who will output the subjective recognition information of the ancient book calculation table. The objective identification information includes a text dataset, a numerical dataset, a symbolic dataset, and a calculation table structure dataset of several training samples; The subjective recognition information includes a text dataset, a numerical dataset, a symbolic dataset, and a calculation table structure dataset of several training samples; S3. Evaluate the recognition index of the self-recognition model of the ancient book calculation table, determine the optimization direction of the self-recognition model of the ancient book calculation table, and display the optimization direction of the self-recognition model of the ancient book calculation table. The specific implementation steps of the various recognition indicators of the self-recognition model for evaluating ancient book calculation tables are as follows: S300. Based on the objective and subjective identification information of ancient book calculation tables, and using a certain training sample as the designated sample, determine the first identification index of the designated sample of the ancient book calculation tables. Second identification indicator Third identification indicator and the fourth identification indicator ; The first identification index for determining the specified sample of the ancient book calculation table. The specific method for determining it is as follows: Extract the header keyword set and the keyword set of each cell from the objective recognition information of a specified sample of ancient calculation tables, and extract the header keyword set and the keyword set of each cell from the subjective recognition information of a specified sample of ancient calculation tables. The similarity between the set of header keywords in the objective identification information and the set of header keywords in the subjective identification information of a specified sample of the ancient book calculation table is calculated using a set similarity algorithm. This similarity is used as the accuracy of header identification for the specified sample of the ancient book calculation table. The similarity algorithm is used to calculate the similarity between the keyword set of each cell in the objective recognition information of the specified sample of the ancient book calculation table and the keyword set of the corresponding cell. This similarity is used as the recognition accuracy of each cell in the specified sample of the ancient book calculation table. If the recognition accuracy of a cell is less than the cell recognition accuracy threshold, the cell is recorded as a heterogeneous cell, and the heterogeneous cells of the specified sample of the ancient book calculation table are filtered. The number of heterogeneous cells in a specified sample of ancient book calculation tables. and the total number of cells And compare its accuracy with the header recognition of a specified sample of ancient arithmetic tables. Import into the first identification indicator model In the process, the first identification index of a specified sample of the ancient book calculation table is output; The second identification indicator of the specified sample of the ancient book calculation table is specifically calculated by using a set similarity algorithm to calculate the similarity between the numerical dataset in the objective identification information and the numerical dataset in the subjective identification information of the ancient book calculation table, which serves as the second identification indicator of the ancient book calculation table. The third identification index for determining the specified sample of the ancient book calculation table The specific method for determining it is as follows: Extract the set of operator symbols and the set of table delimiters from the objective recognition information of a specified sample of ancient calculation tables, and extract the set of operator symbols and the set of table delimiters from the subjective recognition information of a specified sample of ancient calculation tables. The similarity algorithm is used to calculate the similarity between the set of operators in the objective identification information and the set of operators in the subjective identification information of a specified sample of the ancient book calculation table. This similarity is used as the accuracy of operator recognition for the specified sample of the ancient book calculation table. Similarly, the accuracy of table delimiter recognition for the specified sample of the ancient book calculation table is calculated. Accuracy of recognizing operation symbols in ancient calculation tables Accuracy of table delimiter recognition Import into the third identification indicator model In the formula, the third identification index of a specified sample of the ancient book calculation table is output, where , These represent the accuracy thresholds for recognizing operator symbols and table delimiters stored in the web data warehouse, respectively. , These are the logical symbols AND and OR, respectively; The fourth identification index for determining the specified sample of the ancient book calculation table. The specific method for determining it is as follows: Extract the number of rows, columns, merged cells, and split cells of the calculation table structure dataset from the objective identification information of the specified sample of the ancient calculation table, and extract the number of rows, columns, merged cells, and split cells of the calculation table structure dataset from the subjective identification information of the specified sample of the ancient calculation table, and record them as the reference number of rows, reference number of columns, reference merged cells, and reference split cells of the ancient calculation table, respectively. Comparative analysis yields the row number deviation, column number deviation, merged cells for each deviation, and split cells for each deviation in a specified sample of the ancient book calculation table. The total number of merged cells and split cells for each deviation in the specified sample of the ancient book calculation table is then obtained. The number of reference merged cells and the number of reference split cells of a specified sample of the ancient book calculation table are summarized, and the fourth identification index of the specified sample of the ancient book calculation table is obtained through numerical processing. The fourth identification index contains values ​​of 0 and 1. S301. Following this pattern, the first, second, third, and fourth recognition indicators of each training sample of the ancient book calculation table are obtained. These are then averaged to obtain the average values ​​of the first, second, third, and fourth recognition indicators of the ancient book calculation table, which serve as the first-level recognition indicators of the self-recognition model of the ancient book calculation table. Secondary identification indicators Level 3 identification indicators Level 4 identification indicators .

2. The method for identifying and evaluating ancient book calculation tables according to claim 1, characterized in that, The specific method for determining the optimization direction of the self-recognition model for ancient book calculation tables is as follows: The first identification indicator based on ancient book calculation tables Second identification indicator Third identification indicator and the fourth identification indicator ; like Then, the optimization direction of the self-recognition model of the ancient book calculation table is denoted as text optimization; like Then, the optimization direction of the self-recognition model of the ancient book calculation table is denoted as numerical optimization; like Then, the optimization direction of the self-recognition model of the ancient book calculation table is denoted as symbol optimization; like Then the optimization direction of the self-recognition model of the ancient book calculation table is denoted as the calculation table structure optimization; , These are the convergence values ​​of the primary identification indicators and the secondary identification indicators stored in the web data warehouse, respectively. The optimization directions for the self-recognition model of ancient book calculation tables are summarized.

3. A system for performing the ancient book calculation table identification and evaluation method according to any one of claims 1-2, characterized in that, include: The ancient book calculation table upload module is used by staff to upload ancient book calculation tables and their corresponding feature sets, and to determine the self-recognition model of the ancient book calculation tables; The self-recognition model includes a combination of image processing methods and recognition methods; The ancient book calculation table recognition module is used to randomly select several training samples from the ancient book calculation table and input them into the self-recognition model, output the objective recognition information of the ancient book calculation table, and distribute several training samples of the ancient book calculation table to the recognition personnel, who then output the subjective recognition information of the ancient book calculation table. The ancient book recognition and evaluation module is used to evaluate the recognition indicators of the self-recognition model of the ancient book calculation table, determine the optimization direction of the self-recognition model of the ancient book calculation table, and display the optimization direction of the self-recognition model of the ancient book calculation table.

4. An electronic device, characterized in that, include: Processor, memory, and communication bus; The memory stores a computer-readable program that can be executed by the processor; the communication bus enables communication between the processor and the memory; when the processor executes the computer-readable program, it implements the ancient book calculation table identification and evaluation method as described in any one of claims 1-2.

Citation Information

Patent Citations

  • Certificate image recognition method and device, terminal and storage medium

    CN112115748A

  • Ancient book text informatization processing method and system, electronic equipment and storage medium

    CN115410216A

  • Investment amount class table identification method based on convolutional neural network

    CN117275026A