A method for extracting tender document analysis tables
By building a table extraction model, using the BERT pre-training module and logistic regression classifier, the table information in the bidding documents is accurately identified and extracted, and the problem of inaccurate table information extraction in the existing technology is solved, and efficient and accurate table information extraction effect is achieved.
Patent Information
- Application Number
- CN202211524881.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-11-30
AI Technical Summary
It is difficult for the prior art to accurately and efficiently extract table information in bidding documents, especially when there is a large amount of interference information in the table.
A bidding document analysis table extraction method is adopted to build a table extraction model by determining key fields, including data processing modules and text classification modules. The data processing module annotates cell attribute information, and the text classification module uses the BERT pre-training module and logistic regression classifier for training, so as to accurately identify and extract relevant table information.
The accurate extraction of table information in bidding documents is achieved, the accuracy and efficiency of the extraction task is improved, and the required relevant fields can be accurately identified in the presence of interfering information.
Smart Images

Figure CN115906763B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer machine learning and image recognition, and particularly to a method for extracting tender document analysis forms. Background Art
[0002] In the current power industry, a large number of projects are carried out through bidding. The text data contained in tender procurement documents has great value for analysis and research. However, these procurement and announcement documents are often stored in the form of documents, and a large amount of labor costs are required to sort out, understand, organize, and extract information from these document-type data before they can be actually used.
[0003] With the development of artificial intelligence technology, text content extraction technology has gradually matured, and in many business requirement scenarios, machines can replace some manual labor. By training the text extraction logic through deep learning algorithms and using the model to automatically extract the required content information, the unstructured document data can be transformed into structured data that can be statistically analyzed and stored, enabling business personnel to quickly obtain the key information points and valuable data required in a large number of tender documents.
[0004] In tender documents, tables are a clearer way of expression, and a large amount of valuable information is often stored in tables. Therefore, in actual tasks, how to accurately and efficiently extract the information in the tables often determines the overall extraction effect of a document. There are multiple scenarios for table extraction, including extracting all tables in the document, extracting some tables in the document, and extracting some fields or cells in all or some tables in the document. Identifying and extracting full-scale information is similar to document content extraction, and there are some relatively mature technical supports, such as learning part of speech, text features, word order, etc. through context based on sequence labeling technology. However, for partial table extraction or extraction of partial content in a table, with the rest of the content being interference items, or when there are multiple tables in a document and only individual tables are valid information while the others are interference information, this will cause problems such as ambiguity and misalignment when extracting information. Summary of the Invention
[0005] Aiming at the above problems existing in the prior art, the technical problem to be solved by the present invention is: how to more accurately extract the table information in tender documents.
[0006] To solve the above technical problem, the present invention adopts the following technical solutions:
[0007] A method for extracting tender document analysis forms includes the following steps:
[0008] S100: Determine the key fields and select several tender documents containing the key fields; the tender documents contain tables and cell attribute information in the tables; the cell attribute information includes text data and structured data;
[0009] S200: Construct a table extraction model, which includes a data processing module and a text classification module;
[0010] The data processing module labels the cell attribute information, labels the cells with the key fields as positive sample labels, and labels the remaining cells as negative sample labels;
[0011] The text classification module includes a BERT pre-training module and a logistic regression classifier;
[0012] S300: Use several tender documents containing the key fields as the input of the data processing module, and output a positive sample set and a negative sample set with labels;
[0013] Randomly select some data from the positive sample set and the negative sample set as the training set. There are N training samples in the training set. Each training sample in the training set includes the label of the cell, text data, and structured data; the remaining data in the positive sample set and the negative sample set are used as the test set. The test samples in the test set include text data and structured data;
[0014] S400: Use the training set to train the text classification module:
[0015] S410: Let i = 1;
[0016] S420: Embed the cell label and text data in the i-th training sample into the multi-dimensional vector space of the BERT pre-training module to obtain the text vector corresponding to the i-th training sample;
[0017] S430: Use the text vector corresponding to the i-th training sample and the structured data in the i-th training sample as the input of the logistic regression classifier;
[0018] S440: Let i = i + 1. When i > N, obtain the trained text classification module and execute the next step; otherwise, return to S410; S450: Use the test set as the input of the trained text classification module, and the output is the predicted labels of all test samples;
[0019] S460: Calculate the prediction accuracy and sample recall rate of the trained text classification module according to the predicted labels of all test samples. When both the prediction accuracy and the sample recall rate exceed 70%, the finally trained table extraction model is obtained; otherwise, update the parameters of the text classification module and return to S410.
[0020] Preferably, the data processing module in S200 further includes processing of non-conventional format data, including standardizing the format of non-conventional format data, and the format standardization includes a title field, a table header, a first row, and a first column.
[0021] Non-conventional format data sometimes appears as pictures, emojis, or other byte symbols composed of punctuation marks. When such characters appear, it is necessary to perform format standardization operations, so that the content of the table can be correctly processed, and key information of cells or the entire table will not be lost due to misjudgment.
[0022] Preferably, the rules for labeling tags for the cell attribute information in S200 are as follows:
[0023] Judge the table type of the bidding document, and perform the operation of labeling tags for the cell attribute information according to the table type, specifically as follows:
[0024] If the table type is judged to be a single table, extract the cell attribute information of the cells containing the key fields in the single table as positive samples and label them, and the remaining cell attribute information as negative samples and label them;
[0025] If the table attribute is judged to be a multi-table, extract the single tables containing the key fields in the multi-table, and then execute the single table labeling tag rule.
[0026] Dividing the tables in the document first helps to improve the operation efficiency and labeling accuracy.
[0027] Preferably, the specific content of calculating the sample prediction accuracy rate and the positive sample recall rate in S460 is as follows:
[0028] S461: Calculate the sample prediction accuracy rate Accuracy, and the specific expression is as follows:
[0029]
[0030] Among them, TP represents that the true value and the predicted value of the sample label are both positive, TN represents that the true value and the predicted value of the sample label are both negative, FN represents that the true value of the sample label is positive while the predicted value of the sample label is negative, and FP represents that the true value of the sample label is negative while the predicted value of the sample label is positive;
[0031] S462: Calculate the sample recall rate Recall, and the specific expression is as follows:
[0032]
[0033] Compared with the prior art, the present invention has at least the following advantages:
[0034] 1. This technical solution processes data, tags the keyword fields to obtain positive and negative samples, and uses the BERT prediction model and logistic regression classifier to learn from the positive and negative samples containing relevant fields to ensure that the model can obtain and learn the key information of the relevant fields as much as possible. Then, the classification results of the logistic regression classifier can be used to obtain the extraction classification results of positive and negative samples, and the table corresponding to the positive samples is the final table extraction result of this model.
[0035] 2. This technical solution can obtain the table information to be extracted in advance according to the keyword fields, without traversing all documents during the extraction process, and can complete the table extraction task more quickly.
[0036] 4. Improve the accuracy of the extraction task. The extraction effect within the defined range is much better than that from the full text range.
[0037] 5. The model of the present invention can select a suitable algorithm model for replacement according to different actual situations. The extraction effect does not completely depend on a single sequence model. By optimizing the classification model to improve the extraction effect, the direction of deployable operations is more diverse. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0039] The present invention will be further described in detail below.
[0040] See Figure 1 , a method for extracting a tender document parsing table, comprising the following steps:
[0041] S100: Determine the keyword fields, select several tender documents containing the keyword fields. The keyword fields are the target information to be extracted in the tender documents, and the keyword fields are determined manually, such as equipment models, key byte profit statements, etc. The tender documents contain tables and cell attribute information in the tables. The cell attribute information includes text data and structured data. The text data is the content description in the cell, and the structured data is the cell coordinate position, quantity, list header depth, column cell depth, and row table header, etc. Among them, the identification method of the list header depth is that 0 means not using this feature, 1 means list header, and 2 means the cell corresponding to the column in the next row of the list header; the identification method of the column cell depth is that 0 means not using this feature, 1 means the previous column cell, and 2 means the previous column cell and the previous two column cells; the row table header represents the first n row cells;
[0042] S200: Construct a table extraction model. The table extraction model includes a data processing module and a text classification module;
[0043] The data processing module labels tags for the cell attribute information, labels the cells with key fields as positive sample tags, and labels the remaining cells as negative sample tags; existing methods can be used for the method of label processing.
[0044] The text classification module includes a BERT pre-training module and a logistic regression classifier; both the BERT pre-training module and the logistic regression classifier are existing technologies. In addition, when establishing a text classification model, traditional machine learning algorithms such as decision tree algorithms, logistic regression, support vector machines, etc. can also be selected, or deep learning algorithms can be selected. In this way, the most suitable algorithm can be selected for replacement according to different industry standards and data types.
[0045] The data processing module in S200 also includes the processing of non-conventional format data, including format standardization for non-conventional format data, and the format standardization includes a title field, a table header, a first row, and a first column; the non-conventional format data is other format data except for text and numbers.
[0046] The rules for labeling tags for the cell attribute information in S200 are as follows:
[0047] Perform a table type judgment on the bidding document, and perform the operation of labeling tags for the cell attribute information according to the table type, specifically as follows:
[0048] If it is determined that the table type is a single table, extract the cell attribute information containing the key field in the single table as a positive sample and label it, and label the remaining cell attribute information as a negative sample.
[0049] If it is determined that the table attribute is a multi-table, extract the single tables containing the key field in the multi-table, and then execute the single table labeling tag rule.
[0050] S300: Use several bidding documents containing key fields as the input of the data processing module, and output a positive sample set and a negative sample set with labels.
[0051] Randomly select some data from the positive sample set and the negative sample set as the training set. There are N training samples in the training set, and each training sample in the training set includes the label of the cell, text data, and structured data; the remaining data in the positive sample set and the negative sample set are used as the test set, and the test samples in the test set include text data and structured data; all positive and negative samples are randomly split, and the data is split according to the ratio of 8 to 2 for the training set and the test set.
[0052] S400: Use the training set to train the text classification module:
[0053] S410: Let i = 1;
[0054] S420: Embed the cell labels and text data in the i-th training sample into the multi-dimensional vector space of the BERT pre-training module to obtain the text vector corresponding to the i-th training sample;
[0055] S430: Use the text vector corresponding to the i-th training sample and the structured data in the i-th training sample as the input of the logistic regression classifier;
[0056] S440: Let i = i + 1. When i > N, obtain the trained text classification module and execute the next step; otherwise, return to S410;
[0057] S450: Use the test set as the input of the trained text classification module, and the output is the predicted labels of all test samples. The predicted value of this label is the prediction result of whether the sample is a positive sample or a negative sample;
[0058] S460: Calculate the prediction accuracy and sample recall rate of the trained text classification module based on the predicted labels of all test samples. When both the prediction accuracy and the sample recall rate exceed 70%, the finally trained table extraction model is obtained; otherwise, update the parameters of the text classification module and return to S410.
[0059] The specific content of calculating the sample prediction accuracy and positive sample recall rate in S460 is as follows:
[0060] S461: Calculate the sample prediction accuracy Accuracy, and the specific expression is as follows:
[0061]
[0062] Among them, TP represents that both the true value and the predicted value of the sample label are positive, TN represents that both the true value and the predicted value of the sample label are negative, FN represents that the true value of the sample label is positive while the predicted value of the sample label is negative, and FP represents that the true value of the sample label is negative while the predicted value of the sample label is positive;
[0063] S462: Calculate the sample recall rate Recall, and the specific expression is as follows:
[0064]
[0065] The objective of this invention is to accurately and quickly identify the key table information in the bidding documents of the power industry. Especially in the case of interference from other irrelevant table information, it can accurately identify the required relevant fields. The applicable scenarios can be the extraction scenarios of industry bidding documents, the extraction scenarios of industry announcement documents, etc.
[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for extracting parsing tables from tender documents, characterized in that: It includes the following steps: S100: Determine the key fields and select several tender documents containing the key fields; the tender documents contain tables and cell attribute information in the tables; the cell attribute information includes text data and structured data; S200: Construct a table extraction model, and the table extraction model includes a data processing module and a text classification module; The data processing module labels the cell attribute information, labels the cells with key fields as positive sample labels, and labels the remaining cells as negative sample labels; The data processing module also includes the processing of unconventional format data, including standardizing the format of unconventional format data, and the format standardization includes title fields, table headers, the first row, and the first column; The rules for labeling the cell attribute information are as follows: Judge the table type of the tender document and perform the operation of labeling the cell attribute information according to the table type, specifically as follows: If it is judged that the table type is a single table, extract the cell attribute information containing the key fields in the single table as a positive sample and label it, and label the remaining cell attribute information as a negative sample; If it is judged that the table attribute is a multi-table, extract the single tables containing the key fields in the multi-table, and then execute the single table labeling rules; The text classification module includes a BERT pre-training module and a logistic regression classifier; S300: Use several tender documents containing key fields as the input of the data processing module, and output a positive sample set and a negative sample set with labels; Randomly select some data from the positive sample set and the negative sample set as the training set. There are N training samples in the training set, and each training sample in the training set includes the label, text data, and structured data of the cell; the remaining data in the positive sample set and the negative sample set are used as the test set, and the test samples in the test set include text data and structured data; S400: Use the training set to train the text classification module: S410: Let i = 1; S420: Embed the cell label and text data in the i-th training sample into the multi-dimensional vector space of the BERT pre-training module to obtain the text vector corresponding to the i-th training sample; S430: Use the text vector corresponding to the i-th training sample and the structured data in the i-th training sample as the input of the logistic regression classifier; S440: Let i = i + 1. When i > N, obtain the trained text classification module and execute the next step; Otherwise, return to S410; S450: Use the test set as the input of the trained text classification module, and the output is the predicted labels of all test samples; S460: Calculate the prediction accuracy and sample recall rate of the trained text classification module according to the predicted labels of all test samples. When both the prediction accuracy and the sample recall rate exceed 70%, the finally trained table extraction model is obtained; otherwise, update the parameters of the text classification module and return to S410.
2. A method for extracting parsing tables from tender documents according to claim 1, It is characterized in that: The specific content of calculating the sample prediction accuracy and the positive sample recall rate in S460 is as follows: S461: Calculate the sample prediction accuracy Accuracy, and the specific expression is as follows: Among them, TP represents that both the true value and the predicted value of the sample label are positive, TN represents that both the true value and the predicted value of the sample label are negative, FN represents that the true value of the sample label is positive while the predicted value of the sample label is negative, and FP represents that the true value of the sample label is negative while the predicted value of the sample label is positive; S462: Calculate the sample recall rate Recall, and the specific expression is as follows:
Citation Information
Patent Citations
Vertical domain knowledge graph construction method and system
CN113177124A
PDF document cross-page table merging method and apparatus, electronic device and storage medium
WO2022105172A1
Cited By
Document image table type judgment method
CN120375402A
A method for determining the type of document image table
CN120375402B