A method, device, and storage medium for implementing table information extraction
By restoring the logical structure of the table and using machine learning algorithms, the problem of inefficient extraction of table information in the existing technology is solved, and effective classification and analysis of diversity and structural complex tables are realized, thus reducing system development costs and cycles.
Patent Information
- Application Number
- CN202111639162.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-11-04
- Filing Date
- 2021-12-29
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-12-29
AI Technical Summary
The prior art is difficult to effectively process table information of diversity and structural complexity, resulting in inefficient extraction of table information, and high cost and long cycle of system development, which cannot meet market demand.
Through the understanding and processing of the physical structure of the table, the logical structure of the table is restored and the machine learning algorithm is used to extract the table information. The specific steps include cell data annotation, cell judgment model training, table data annotation, table classification model training, new table classification and analysis to achieve effective extraction of table information.
It realizes effective classification and analysis of tables under complex types, so that table data can be used correctly for subsequent tasks, and reduces the cost and cycle of system development.
Smart Images

Figure CN114387607B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of unstructured document processing in machine learning and natural language processing (NLP), and in particular to a method and device for extracting table information based on a machine learning algorithm. Background Art
[0002] Documents are indispensable for the public disclosure or business communication of major enterprises and government departments. With the development of informatization, the storage, portability and transferability of electronic documents are much better than traditional paper documents, and thus become the current popular choice. However, the problems that come with it are the inconsistency of document formats, duplication, and difficulty in verification, which cause various corporate departments to invest too much human resources in complicated documents.
[0003] In this context, technologies such as document parsing, text content extraction, and format conversion have emerged. Among existing technologies, due to the maturity of NLP, OCR and other technologies, the text and images in documents have been well parsed. However, for tables, because they can show the relationship between data very simply and clearly, they account for a very high proportion of information in the document. Due to the diversity and structural complexity of tables, their manual processing is very time-consuming. There is currently no universal solution to problems such as how to correctly extract table information.
[0004] Due to the diversity of document types, there is no definition of writing specifications and unified standards for various types of documents, resulting in complex table styles and types. Under the existing technology, it is difficult to structure and align fields for such multi-type and multi-state data. The current more mature document parsing technology simply extracts and displays the tables in the document without excessive processing. Many applications use manual methods to process structured data.
[0005] Recently, systems that extract information from tables based on sample annotation and machine learning have been welcomed and praised by the market. However, the current system development costs are too high and the cycle is too long, which affects its ability to meet market demand. The reason is that there are many types of tables, and each type requires a separate machine learning model to be developed.
[0006] With respect to the technical problems existing in the above-mentioned prior art, no effective solutions have been proposed yet. Summary of the invention
[0007] Embodiments of the present disclosure provide a method and apparatus for implementing table information extraction based on machine learning algorithms to at least solve the technical problems existing in the prior art. The purpose of this application is to solve the problem of table information extraction under complex table types. By understanding and processing the physical structure of the table, the logical structure of the table is restored. On this basis, the machine learning model for table information extraction only needs to handle the problem of semantic consistency determination, and no longer needs to solve the problem of diverse table types.
[0008] According to one aspect of the embodiments of the present disclosure, there is provided a method for implementing table information extraction based on machine learning algorithms, including:
[0009] A cell data annotation step of obtaining a batch of cell data and obtaining classification data through machine and human-assisted annotation of whether the cell type belongs to a key;
[0010] A cell judgment model training step of training the classification data with a relevant classification model to obtain a cell judgment model;
[0011] A table data annotation step of obtaining a batch of table data and then manually annotating the type of the table data to obtain classification data;
[0012] A table classification model training step of expanding the classified data after the manual annotation to obtain a large amount of labeled data, using the cell judgment model to obtain the key distribution, and training the table classification model with the large amount of labeled data;
[0013] A new table classification step of inputting a new table into the cell judgment model to obtain the key distribution of the current table, and inputting the key distribution into the table classification model to obtain the type corresponding to the new table;
[0014] An analysis step of parsing the new table according to the type corresponding to the new table and outputting the tuple information of the current table.
[0015] According to another aspect of the embodiments of the present disclosure, there is also provided a storage medium including a stored program, wherein the method described in any one of the above is executed by a processor when the program runs.
[0016] According to another aspect of the embodiments of the present disclosure, there is also provided an apparatus for implementing table information extraction based on machine learning algorithms, including:
[0017] A cell data annotation module for obtaining a batch of cell data and obtaining classification data through machine and human-assisted annotation of the cell type;
[0018] A cell judgment model training module for training the classification data with a relevant classification model to obtain a cell judgment model;
[0019] A table data annotation module, configured to obtain a batch of table data, and then manually annotate the types of the table data to obtain classified data;
[0020] A table classification model training module, configured to expand the classified data after the manual annotation to obtain a large amount of annotated data, use a cell judgment model to obtain a key distribution, and use the large amount of annotated data to train a table classification model;
[0021] A new table classification module, configured to input a new table into the cell judgment model to obtain the key distribution of the current table, and input the key distribution into the table classification model to obtain the type corresponding to the new table;
[0022] An analysis module, configured to analyze the new table according to the type corresponding to the new table and output the tuple information of the current table.
[0023] According to another aspect of the embodiments of the present disclosure, there is also provided a device for extracting table information based on a machine learning algorithm, including:
[0024] A first processor; and
[0025] A first memory, connected to the first processor, for providing instructions for the first processor to perform the following processing steps:
[0026] A cell data annotation step, obtaining a batch of cell data, and obtaining classified data through machine and manual assistance in annotating whether the cell type belongs to a key;
[0027] A cell judgment model training step, training the classified data with a relevant classification model to obtain a cell judgment model;
[0028] A table data annotation step, obtaining a batch of table data, and then manually annotating the types of the table data to obtain classified data;
[0029] A table classification model training step, expanding the classified data after the manual annotation to obtain a large amount of annotated data, using a cell judgment model to obtain a key distribution, and using the large amount of annotated data to train a table classification model;
[0030] A new table classification step, inputting a new table into the cell judgment model to obtain the key distribution of the current table, and inputting the key distribution into the table classification model to obtain the type corresponding to the new table;
[0031] An analysis step, analyzing the new table according to the type corresponding to the new table and outputting the tuple information of the current table.
[0032] Through the technical solution of the present application, the following beneficial effects are achieved:
[0033] 1. The effective classification of tables under complex types is achieved.
[0034] 2. The parsing of different tables is realized, enabling the table data to be correctly used in subsequent tasks such as NLP. Description of the Drawings
[0035] The drawings described herein are used to provide a further understanding of the present disclosure and form a part of this application. The schematic embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure. In the drawings:
[0036] Figure 1 is a hardware structure block diagram of a computing device for implementing the method described in Embodiment 1 of the present disclosure;
[0037] Figure 2 is a schematic diagram of a system for implementing table information extraction based on machine learning algorithms according to Embodiment 1 of the present disclosure;
[0038] Figure 3 is a flowchart of a method for implementing table information extraction based on machine learning algorithms according to the first aspect of Embodiment 1 of the present disclosure;
[0039] Figure 4 is a flowchart of an algorithm for determining whether it belongs to the key determination algorithm according to Embodiment 1 of the present disclosure;
[0040] Figure 5 is a flowchart of the classification model training according to Embodiment 1 of the present disclosure;
[0041] Figure 6 is a flowchart of a table classification method based on key / value information according to Embodiment 1 of the present disclosure;
[0042] Figure 7 is a flowchart of the parsing method according to Embodiment 1 of the present disclosure;
[0043] Figure 8 is a schematic diagram of a device for implementing table information extraction based on machine learning algorithms according to the first aspect of Embodiment 2 of the present disclosure;
[0044] Figure 9 is a schematic diagram of a device for implementing table information extraction based on machine learning algorithms according to the second aspect of Embodiment 2 of the present disclosure; Detailed Embodiments
[0045] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.
[0046] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0047] Embodiment 1
[0048] According to this embodiment, a method embodiment for implementing table information extraction based on a machine learning algorithm is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0049] The method embodiment provided in this embodiment can be executed on a mobile terminal, a computer terminal, a server, or a similar computing device. Figure 1 A hardware structure block diagram of a computing device for implementing a method for implementing table information extraction based on a machine learning algorithm is shown. As Figure 1 shown, the computing device may include one or more processors (the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory for storing data, and a transmission device for communication functions. In addition to this, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computing device may further include more Figure 1more or fewer components as shown, or having a configuration different from that Figure 1 shown.
[0050] It should be noted that one or more of the above-mentioned processors and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computing device. As involved in the embodiments of the present disclosure, the data processing circuit is a processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0051] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for implementing table information extraction based on a machine learning algorithm in the embodiments of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the method for implementing table information extraction of the above application program based on a machine learning algorithm. The memory can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory can further include a memory remotely located relative to the processor, and these remote memories can be connected to the computing device through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0052] The transmission device is used to receive or send data via a network. Specific examples of the above network can include a wireless network provided by a communication provider of the computing device. In one instance, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0053] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computing device.
[0054] It should be noted here that in some alternative embodiments, the above Figure 1 shown computing device can include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1This is just an example of a specific concrete instance and is intended to show the types of components that may exist in the above computing device.
[0055] Figure 2 It is a schematic diagram of a system for implementing table information extraction based on a machine learning algorithm according to this embodiment. Refer to Figure 2 As shown, the system includes: a front-end portable electronic terminal 100 (such as a laptop), a computing device 200 for implementing a method of table information extraction based on a machine learning algorithm, and a cloud server 300. It should be noted that the computing device 200 for implementing the method of table information extraction based on a machine learning algorithm in the system can all be applicable to the above-mentioned hardware structure.
[0056] Under the above operating environment, according to the first aspect of this embodiment, a method for implementing table information extraction based on a machine learning algorithm is provided. This method is implemented by Figure 2 the computing device 200 for implementing the method of table information extraction based on a machine learning algorithm shown in
[0057] A table is to organize multiple (groups) of information with the same logical structure together in a specific pattern, having a high density of information and good visualization effects. Because of these good characteristics of the table, it has become the main data form of many important documents.
[0058] The table contains a large amount of information and can support rich business scenarios; on the other hand, compared with text, the table has relatively better structure, and it is more likely for technology to achieve sufficient accuracy.
[0059] Figure 3 shows a schematic flowchart of the method. Refer to Figure 3 As shown, the method includes:
[0060] S302: Cell data annotation step, obtain a batch of cell data, and through machine and human assistance, annotate whether the cell type belongs to a key to obtain classification data;
[0061] For example, the cell data belonging to the key screened by machine rules is manually reviewed to construct a cell classification system.
[0062] S304: Cell judgment model training step, train the classification data with relevant classification models to obtain a cell judgment model;
[0063] S306: Table data annotation step, obtain a batch of table data, and then manually annotate the type of the table data to obtain classification data.
[0064] The acquisition method of the said type is: by manually summarizing and inducing typical tables, constructing a table classification system and continuously expanding it.
[0065] S308: Steps for training the table classification model. Expand the classified data after manual annotation to obtain a large amount of labeled data. Use the cell judgment model to obtain the key distribution, and use the large amount of labeled data to train the table classification model.
[0066] The cell classification model is a model for determining whether it belongs to the key determination algorithm. For example, Figure 4 As shown, the process of determining whether it belongs to the key determination algorithm includes:
[0067] Use the classification data of the cells as training data, and the training data is labeled data belonging to key / value.
[0068] Use the training data to train the classification model to obtain a trained classification model.
[0069] After processing the table to be judged through data processing, input it into the trained classification model to obtain the key / value determination result.
[0070] The table classification model is a model based on the key determination algorithm. For example, Figure 4 As shown, the process based on the key determination algorithm includes:
[0071] Use the large amount of labeled data as training data, and the training data is labeled data belonging to the existing table types.
[0072] Use the training data to train the classification model to obtain a trained classification model.
[0073] After processing the table to be judged through data processing, input it into the trained cell classification model to obtain the key distribution; input the key distribution information into the trained table classification model to obtain the table type determination result.
[0074] In this application, the information relied on for table classification includes content and structure. Content refers to the part-of-speech of the words in each cell, whether it is an index item, whether it is an annual cycle, etc.; structure mainly refers to merged cells.
[0075] Among them, data processing includes: word segmentation, stop word removal, format processing, etc.
[0076] The function used by the key determination algorithm in this application is text classification, specifying the category to which the text belongs.
[0077] Among them, the process of using the training data to train the classification model, as Figure 5 shown, includes:
[0078] Obtain the content data of the table cells and perform text tokenization processing;
[0079] Perform data labeling and classification on the processed result;
[0080] According to the result after the data labeling and classification, set the classification model parameters and perform training;
[0081] Determine whether the classification model is reliable. If it is reliable, store the classification model; if it is not reliable, continue to adjust the classification model parameters.
[0082] S310: New table classification step. Input the new table into the cell judgment model, obtain the key distribution of the current table, and input the key distribution into the table classification model to obtain the type corresponding to the new table. As Figure 6 shown, the specific steps are as follows:
[0083] Obtain a new table or a target table. Input the new table into the cell classification model, identify the table key data, and obtain the key distribution information;
[0084] Through the key distribution information of the table, input it into the table classification model, obtain the corresponding table type and store it as the type corresponding to the new table.
[0085] S312: Parsing step. According to the type corresponding to the new table, parse the new table and output the tuple information of the current table. As Figure 7 shown, the specific steps are as follows:
[0086] Determine the table header according to the type corresponding to the new table;
[0087] Recombine the key of the new table and the extracted value result into tuple information.
[0088] In this step S312, the table is composed of multiple or multiple groups of information with the same structure. Therefore, the goal of table parsing is to output this tuple with the same structure.
[0089] At the beginning of this step S312, the object output after the table is classified by S310 is the content data of the table cells, including the physical structure and values of the table.
[0090] Thus, according to the first aspect of this embodiment, the following beneficial effects are achieved:
[0091] This application realizes the effective classification of tables under complex types. It realizes the parsing of different tables, enabling the table data to be correctly used in subsequent tasks such as NLP.
[0092] In addition, refer to Figure 1As shown, according to the third aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein when the program runs, the method described in any one of the above is executed by a processor.
[0093] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0094] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.
[0095] Embodiment 2
[0096] Figure 8 An apparatus 500 for implementing table information extraction based on a machine learning algorithm according to the first aspect of this embodiment is shown. This apparatus corresponds to the method according to the first aspect of Embodiment 1. Refer to Figure 8 As shown, the apparatus 500 includes:
[0097] A cell data annotation module 510, configured to obtain a batch of cell data, and label the cell types through machine and human assistance to obtain classification data;
[0098] A cell judgment model training module 520, configured to perform relevant classification model training on the classification data to obtain a cell judgment model;
[0099] A table data annotation module 530, configured to obtain a batch of table data, and then manually annotate the types of the table data to obtain classification data;
[0100] A table classification model training module 540, configured to expand the classification data after the manual annotation to obtain a large amount of labeled data, use the cell judgment model to obtain a key distribution, and use the large amount of labeled data to train a table classification model;
[0101] A new table classification module 550 is used to input a new table into the cell judgment model, obtain the key distribution of the current table, and input the key distribution into the table classification model to obtain the type corresponding to the new table;
[0102] An analysis module 560 is used to analyze the new table according to the type corresponding to the new table and output the tuple information of the current table.
[0103] The cell type marked with the assistance of machine and human includes: cell data belonging to keys screened by machine rules, which are manually reviewed to construct a cell classification system;
[0104] The type of the table data manually marked includes: constructing a table classification system by manually summarizing and generalizing typical tables.
[0105] Optionally, the cell classification model is a model for determining whether it belongs to the key determination algorithm. The process of the key determination algorithm includes:
[0106] Using the classification data of the cell as training data, and the training data is labeled data belonging to key / value;
[0107] Using the training data to train the classification model to obtain a trained classification model;
[0108] After processing the table to be judged through data, input it into the trained classification model to obtain the key / value determination result;
[0109] The table classification model is a model based on the key determination algorithm. The process based on the key determination algorithm includes:
[0110] Using the large amount of labeled data as training data, and the training data is labeled data belonging to existing table types;
[0111] Using the training data to train the classification model to obtain a trained classification model;
[0112] After processing the table to be judged through data, input it into the trained cell classification model to obtain the key distribution; input the key distribution information into the trained table classification model to obtain the table type determination result.
[0113] Optionally, the data processing includes: word segmentation, stop word removal, and format processing.
[0114] Optionally, the using the training data to train the classification model includes:
[0115] Obtain the content data of the table cells and perform text tokenization processing;
[0116] Perform data labeling and classification on the processed results;
[0117] According to the results after the data labeling and classification, set the classification model parameters and perform training;
[0118] Judge whether the classification model is reliable. If it is reliable, store the classification model. If it is not reliable, continue to adjust the classification model parameters and perform training.
[0119] Optionally, the new table classification module is specifically used for:
[0120] Obtain a new table, input the new table into the cell classification model, and identify the table key / value;
[0121] Judge the table type corresponding to the table through the key / value of the table and store it as the type corresponding to the new table.
[0122] Optionally, the parsing module is specifically used for:
[0123] Determine the table header according to the type corresponding to the new table;
[0124] Recombine the key of the new table and the extracted value result into tuple information and output it.
[0125] In addition, Figure 9 There is shown an apparatus 600 for extracting table information based on a machine learning algorithm according to the second aspect of the present embodiment. The apparatus 600 corresponds to the method according to the second aspect of Embodiment 1. Refer to Figure 9 As shown, the apparatus 600 includes: An apparatus for extracting table information based on a machine learning algorithm, characterized by including:
[0126] A first processor 610; and
[0127] A first memory 620, connected to the first processor, for providing instructions for the first processor to perform the following processing steps:
[0128] Cell data annotation step: Obtain a batch of cell data, and through machine and human assistance, annotate whether the cell type belongs to a key to obtain classification data;
[0129] Cell judgment model training step: Perform relevant classification model training on the classification data to obtain a cell judgment model;
[0130] Table data annotation step: Obtain a batch of table data, and then manually annotate the type of the table data to obtain classification data;
[0131] Steps for training a table classification model: expand the classified data after manual annotation to obtain a large amount of labeled data, use a cell judgment model to obtain the key distribution, and use the large amount of labeled data to train a table classification model;
[0132] Steps for classifying a new table: input the new table into the cell judgment model to obtain the key distribution of the current table, and input the key distribution into the table classification model to obtain the type corresponding to the new table;
[0133] Parsing steps: according to the type corresponding to the new table, parse the new table and output the tuple information of the current table.
[0134] Therefore, according to this embodiment, the present application adopts a device for extracting table information based on a machine learning algorithm, which realizes effective classification of tables under complex types. It realizes the parsing of different tables, enabling table data to be correctly used in subsequent tasks such as NLP.
[0135] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0136] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0137] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0138] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0139] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0140] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0141] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for realizing table information extraction based on a machine learning algorithm, characterized in that, Including: Cell data annotation step: Obtain a batch of cell data, and through machine and manual assistance, annotate whether the cell type belongs to a key to obtain classification data; Cell judgment model training step: Perform relevant classification model training on the classification data to obtain a cell judgment model; Table data annotation step: Obtain a batch of table data, and then manually annotate the type of the table data to obtain classification data; Table classification model training step: Expand the classification data after the manual annotation to obtain a large amount of labeled data, use the cell judgment model to obtain the key distribution, and use the large amount of labeled data to train the table classification model; New table classification step: Input a new table into the cell judgment model to obtain the key distribution of the current table, and input the key distribution into the table classification model to obtain the type corresponding to the new table; Parsing step: According to the type corresponding to the new table, parse the new table and output the tuple information of the current table; The cell judgment model is a model for determining whether it belongs to the key determination algorithm. The process of determining whether it belongs to the key determination algorithm includes: Using the classification data of the cell as training data, and the training data is labeled data belonging to key / value; Using the training data to train a classification model to obtain a trained classification model; After processing the table to be judged through data, input it into the trained classification model to obtain a key / value determination result; The table classification model is a model based on the key determination algorithm. The process based on the key determination algorithm includes: Using the large amount of labeled data as training data, and the training data is labeled data belonging to the existing table types; Using the training data to train a classification model to obtain a trained classification model; After processing the table to be judged through data, input it into the trained cell judgment model to obtain the key distribution; input the key distribution information into the trained table classification model to obtain a table type determination result. The parsing step specifically includes the following steps: Determine the table header according to the type corresponding to the new table; Recombine the key and the extracted value result of the new table into tuple information and output.
2. The method according to claim 1, characterized in that, The machine and manual assisted annotation of the cell type includes: Screening out the cell data belonging to the key through machine rules, and after manual review, constructing a cell classification system; The manual annotation of the type of the table data includes: Summarizing and generalizing typical tables manually to construct a table classification system.
3. The method according to claim 2, characterized in that, The data processing includes: Word segmentation, stop word removal, and format processing.
4. The method according to claim 3, characterized in that, The use of the training data to train a classification model includes: Obtain the content data of the table cell and perform text word segmentation processing; Perform data label classification on the processed result; Set classification model parameters and train according to the result after the data label classification; Judge whether the classification model is reliable. If it is reliable, store the classification model. If it is not reliable, continue to adjust the classification model parameters and train.
5. The method according to claim 4, characterized in that, The new table classification step specifically includes the following steps: Obtain a new table, input the new table into the cell judgment model, and identify the key distribution information of the table; Input the key distribution information of the table into the table classification model, judge the corresponding table type of the table and store it as the type corresponding to the new table.
6. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program runs, the method described in any one of claims 1 to 5 is executed by a processor.
7. A device for realizing table information extraction based on a machine learning algorithm, characterized in that, Comprising: A cell data annotation module, configured to obtain a batch of cell data, and obtain classification data by assisting in annotating the cell type by machine and human; A cell judgment model training module, configured to train the classification data with a relevant classification model to obtain a cell judgment model; A table data annotation module, configured to obtain a batch of table data, and then manually annotate the type of the table data to obtain classification data; A table classification model training module, configured to expand the classified data after the manual annotation to obtain a large amount of labeled data, use the cell judgment model to obtain the key distribution, and use the large amount of labeled data to train the table classification model; A new table classification module, configured to input a new table into the cell judgment model, obtain the key distribution of the current table, and input the key distribution into the table classification model to obtain the type corresponding to the new table; An analysis module, configured to analyze the new table according to the type corresponding to the new table, and output the tuple information of the current table; The cell judgment model is a model for determining whether it belongs to the key determination algorithm, and the process of determining whether it belongs to the key determination algorithm includes: Using the classification data of the cell as training data, and the training data is labeled data belonging to key / value; Training the classification model with the training data to obtain a trained classification model; After the table to be judged is processed by data, input it into the trained classification model to obtain a key / value determination result; The table classification model is a model based on the key determination algorithm, and the process based on the key determination algorithm includes: Using the large amount of labeled data as training data, and the training data is labeled data belonging to the existing table type; Training the classification model with the training data to obtain a trained classification model; After the table to be judged is processed by data, input it into the trained cell judgment model to obtain the key distribution; input the key distribution information into the trained table classification model to obtain a table type determination result; The analysis step specifically includes the following steps: Determine the table header according to the type corresponding to the new table; Recombine the key of the new table and the extracted value result into tuple information and output it.
8. An apparatus for extracting tabular information based on machine learning algorithms, characterized in that, Comprising: A first processor; And A first memory, connected to the first processor, for providing instructions for the first processor to process the following processing steps: A cell data annotation step, obtaining a batch of cell data, and assisting in annotating whether the cell type belongs to a key by machine and human to obtain classification data; A cell judgment model training step, training the classification data with a relevant classification model to obtain a cell judgment model; Steps for table data annotation: Obtain a batch of table data, and then manually annotate the types of the table data to obtain classified data; Steps for training the table classification model: Expand the classified data after manual annotation to obtain a large amount of annotated data, use the cell judgment model to obtain the key distribution, and use the large amount of annotated data to train the table classification model; Steps for new table classification: Input the new table into the cell judgment model to obtain the key distribution of the current table, and input the key distribution into the table classification model to obtain the type corresponding to the new table; Parsing steps: According to the type corresponding to the new table, parse the new table and output the tuple information of the current table; The cell judgment model is a model for determining whether it belongs to the key determination algorithm. The process of determining whether it belongs to the key determination algorithm includes: Use the classification data of the cell as the training data, and the training data is the annotated data belonging to key / value; Use the training data to train the classification model to obtain a trained classification model; After processing the table to be judged, input it into the trained classification model to obtain the key / value determination result; The table classification model is a model based on the key determination algorithm. The process based on the key determination algorithm includes: Use the large amount of annotated data as the training data, and the training data is the annotated data belonging to the existing table types; Use the training data to train the classification model to obtain a trained classification model; After processing the table to be judged, input it into the trained cell judgment model to obtain the key distribution; Input the key distribution information into the trained table classification model to obtain the table type determination result. The parsing steps specifically include the following steps: Determine the table header according to the type corresponding to the new table; Recombine the key of the new table and the extracted value result into tuple information and output.
Citation Information
Patent Citations
A Chinese table column label recovery method and system based on text classification
CN109710725A
Image recognition method and system and data processing method
CN113536856A