Table analysis method based on large model

Through the large-model-based table analysis method, unstructured tables are detected and converted, and the relationship between table data and location is established, the problem of table data not corresponding to location in the existing technology is solved, and high-precision table analysis is achieved.

CN119990084APending Publication Date: 2025-05-13SHANGHAI CONSTRUCTION FOURTH CONSTRUCTION GROUP CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510096974.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, in table recognition and analysis in unstructured documents, it is difficult to accurately ensure the relationship between the table position, resulting in errors in the analysis of the corresponding relationship between the table data and the table position.

Method used

The table analysis method based on large models is adopted to detect unstructured table areas through multi-scale algorithms, combine visual detection algorithms and optical character recognition technology to convert unstructured tables into structured tables, and establish the association relationship between each cell and its data. The general big model is fine-tuned using low-rank fine-tuning technology to obtain the table big model for analysis.

Benefits of technology

The correspondence between table location and data location is realized, the problem of serial series and other misalignment of table data and table locations is avoided, and the accuracy of table analysis is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990084A_ABST
    Figure CN119990084A_ABST
Patent Text Reader

Abstract

The invention discloses a table analysis method based on a large model. The method comprises the following steps: detecting an unstructured table region in an unstructured document through a multi-scale algorithm; converting the unstructured table into a structured table through a visual detection algorithm; identifying and extracting unstructured data in each cell in the unstructured table as structured data through an optical character identification technology, and establishing an association relationship between the unstructured table and the unstructured data in each cell in combination with a visual detection algorithm; taking the unstructured table and the data thereof, and the structured table and the data thereof as training samples, and performing fine-tuning training on the general large model by utilizing a low-rank fine-tuning technology to obtain a table large model; and analyzing the unstructured table to be analyzed and the data thereof into a structured table and the data thereof through the table large model. According to the method, the table analysis precision can be improved, and the problem that the table data and the table position do not correspond due to dislocation such as serial and serial is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of table recognition, and in particular to a table parsing method based on a large model. Background Art

[0002] At present, the recognition and understanding of tables in unstructured documents such as pictures and PDF documents are still not accurate enough. Serial or column-by-column recognition often occurs, and most current OCR recognition algorithms cannot retain the positional relationships in the table. This will cause the general large model to be unable to understand the data association in the table, resulting in errors in the parsing of the correspondence between the table data and the table position. Summary of the invention

[0003] The purpose of the present invention is to provide a table parsing method based on a large model to solve the problem of mismatch between table data and table position relationship during unstructured table recognition and parsing.

[0004] In order to solve the above technical problems, the present invention provides a table parsing method based on a large model, comprising:

[0005] Detect unstructured table regions in unstructured documents through a multi-scale algorithm;

[0006] Converting the unstructured table in the unstructured table area into a structured table by a visual detection algorithm;

[0007] The optical character recognition technology is used to identify and extract the unstructured data in each cell of the unstructured table as structured data, and the visual detection algorithm is used to establish the association relationship between the unstructured table and the unstructured data in each cell;

[0008] The unstructured table and its data, the structured table and its data of the above corresponding relationship are used as training samples, and the low-rank fine-tuning technology is used to fine-tune the general large model to obtain the table large model;

[0009] The unstructured table and its data to be parsed are parsed into structured tables and their data through the table big model.

[0010] Furthermore, the table parsing method based on the large model provided by the present invention also includes:

[0011] Detect the introduction text and variable annotation text area of ​​the unstructured table area in the unstructured document through a multi-scale algorithm;

[0012] The introduction text in the unstructured table area and the text content in the variable annotation text area are identified by combining the visual detection algorithm with the optical character recognition technology. The association relationship between the unstructured table and the unstructured data in each cell thereof is established through the logical association relationship between the introduction text in the unstructured table area and the text content in the variable annotation text area and the unstructured table and its data.

[0013] Furthermore, the present invention provides a table parsing method based on a large model. The general large model calls the table drawing software to automatically draw a structured table with a corresponding relationship according to the unstructured table with established association relationship and the unstructured data in each cell thereof, and fills the structured data therein to obtain a table large model.

[0014] Furthermore, the present invention provides a table parsing method based on a big model. The general big model is converted into codes or instructions that can be understood by the table drawing software according to the embedded instructions. The table drawing software automatically draws a structured table and fills it with structured data according to the correspondence between the unstructured table and the unstructured data in each cell therein, thereby obtaining a table big model.

[0015] Furthermore, the large model-based table parsing method provided by the present invention uses a visual detection algorithm to mark the number of rows and columns in the unstructured table in the unstructured table area, whether there are slashes in the cells, and whether the cells are merged, thereby converting the unstructured table area into a structured table.

[0016] Compared with the prior art, the present invention has the following beneficial effects:

[0017] The table parsing method based on a large model provided by the present invention detects an unstructured table area in an unstructured document through a multi-scale algorithm, accurately locates the unstructured table area, converts the unstructured table and its data into a structured table and its data through a visual detection algorithm and an optical character recognition technology, and establishes an association relationship between each cell and the data therein, realizes a correspondence between the table position and the data position, and avoids the problem of serial and serial misalignment between the table data and the table position; the table and its data before and after the conversion are used as training samples to train a general large model to obtain a table large model, thereby realizing the parsing of the unstructured table and its data into a corresponding structured table and its data through the table large model, thereby improving the table parsing accuracy and avoiding the problem that the optical character recognition technology only recognizes text data and cannot match the position of each cell of the table, resulting in the table position not corresponding to the table data. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flow chart of the table parsing method based on the large model; DETAILED DESCRIPTION

[0019] The present invention is described in detail below in conjunction with the accompanying drawings: The advantages and features of the present invention will become more apparent from the following description. It should be noted that the accompanying drawings are all in very simplified form and are not in exact proportions, and are only used to conveniently and clearly assist in explaining the purpose of the embodiments of the present invention.

[0020] Please refer to Figure 1 , an embodiment of the present invention provides a table parsing method based on a large model, which may include:

[0021] Step S1, unstructured table area detection: specifically: detecting the unstructured table area in the unstructured document by using a multi-scale algorithm.

[0022] Step S2, structured table conversion: specifically: converting the unstructured table in the unstructured table area into a structured table through a visual detection algorithm.

[0023] Step S3, table data extraction: specifically: identifying and extracting the unstructured data in each cell of the unstructured table as structured data through optical character recognition technology, and establishing the association relationship between the unstructured table and the unstructured data in each cell thereof in combination with the visual detection algorithm.

[0024] Step S4, large model training: specifically: the unstructured table and its data, the structured table and its data of the above-mentioned corresponding relationship are used as training samples, and the low-rank fine-tuning technology is used to fine-tune the training of the general large model to obtain the table large model. In order to improve the parsing accuracy of the table large model, the general large model can call the table drawing software to automatically draw the structured table of the corresponding relationship according to the unstructured table with the established association relationship and the unstructured data in each cell and fill it with structured data to obtain the trained large model, that is, the table large model. The general large model refers to the large language model. The table drawing software includes but is not limited to OFFICE software, WPS software, etc.

[0025] Step S5, table big model parsing: specifically: parsing the unstructured table and its data to be parsed into a structured table and its data through the table big model.

[0026] In order to improve the accuracy of the association between the table position and its data, the table parsing method based on a large model provided by an embodiment of the present invention may further include in step S1: detecting the introduction text and variable annotation text area of ​​the unstructured table area in the unstructured document by a multi-scale algorithm; at this time, step S3 includes: identifying the text content of the introduction text and variable annotation text area of ​​the unstructured table area by combining a visual detection algorithm with optical character recognition technology, and assisting in establishing the association relationship between the unstructured table and the unstructured data in each cell of the unstructured table through the logical association relationship between the text content of the introduction text and variable annotation text area of ​​the unstructured table area and the unstructured table and its data.

[0027] In order to further improve the parsing accuracy of the table big model, the table parsing method based on the big model provided by the embodiment of the present invention, in step S4, the general big model is converted into a code or instruction that can be understood by the table drawing software according to the embedded instructions, and the structured table is automatically drawn by the table drawing software according to the correspondence between the unstructured table and the unstructured data in each cell therein, and the structured data is filled therein to obtain the table big model. The embedded instructions are used to facilitate the automatic drawing of the structured table by the table drawing software to understand and improve the table parsing accuracy.

[0028] In order to improve the correspondence between the table position and the table data and avoid the problems of misalignment, omission, and serial sequence, the table parsing method based on a large model provided by the embodiment of the present invention uses a visual detection algorithm to mark the number of rows and columns in the unstructured table in the unstructured table area, whether there are slashes in the cells, and whether the cells are merged, thereby converting the unstructured table area into a structured table. When converting to a structured table, the structured table is converted according to the number of marked rows and columns, whether there are slashes in the cells, and the merged state of the cells, thereby improving the conversion accuracy of the structured table. Accordingly, when extracting the data in the table, it is also necessary to mark the location of the data in the table to establish an association between the table data and the table position, so as to avoid the problems of misalignment, omission, and serial sequence between the table position and the table data.

[0029] The table parsing method based on a large model provided in an embodiment of the present invention detects an unstructured table area in an unstructured document through a multi-scale algorithm, accurately locates the unstructured table area, converts the unstructured table and its data into a structured table and its data through a visual detection algorithm and an optical character recognition technology, and establishes an association relationship between each cell and the data therein, thereby realizing a correspondence between the table position and the data position, and avoiding problems such as serial misalignment between the table data and the table position; the table and its data before and after the conversion are used as training samples to train a general large model to obtain a table large model, thereby realizing the parsing of the unstructured table and its data into a corresponding structured table and its data through the table large model, thereby improving the table parsing accuracy and avoiding the problem that the optical character recognition technology only recognizes the text data and cannot match the position of each cell of the table, resulting in the table position not corresponding to the table data.

[0030] The present invention is not limited to the above-mentioned specific implementation modes. Obviously, the above-mentioned embodiments are only some embodiments of the embodiments of the present invention, but not all embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of the present invention. Those skilled in the art can make other levels of modifications and changes to the present invention. In this way, if these modifications and changes of the present invention fall within the scope of the claims of the present invention, the present invention is also intended to include these changes and changes.

Claims

1. A table parsing method based on a large model, characterized in that: include: Detect unstructured table regions in unstructured documents through a multi-scale algorithm; Converting the unstructured table in the unstructured table area into a structured table by a visual detection algorithm; The optical character recognition technology is used to identify and extract the unstructured data in each cell of the unstructured table as structured data, and the visual detection algorithm is used to establish the association relationship between the unstructured table and the unstructured data in each cell; The unstructured table and its data, the structured table and its data of the above corresponding relationship are used as training samples, and the low-rank fine-tuning technology is used to fine-tune the general large model to obtain the table large model; The unstructured table and its data to be parsed are parsed into structured tables and their data through the table big model.

2. The table parsing method based on a large model according to claim 1, characterized in that: Also includes: Detect the introduction text and variable annotation text area of ​​the unstructured table area in the unstructured document through a multi-scale algorithm; The introduction text in the unstructured table area and the text content in the variable annotation text area are identified by combining the visual detection algorithm with the optical character recognition technology. The association relationship between the unstructured table and the unstructured data in each cell thereof is established through the logical association relationship between the introduction text in the unstructured table area and the text content in the variable annotation text area and the unstructured table and its data.

3. The table parsing method based on a large model according to claim 2 is characterized in that: The general large model calls the table drawing software to automatically draw the corresponding structured table according to the unstructured table with established association relationship and the unstructured data in each cell and fill the structured data therein to obtain the table large model.

4. The table parsing method based on a large model according to claim 3 is characterized in that: The general large model is converted into codes or instructions that can be understood by the table drawing software according to the embedded instructions. According to the correspondence between the unstructured table and the unstructured data in each cell therein, the table drawing software automatically draws the structured table and fills the structured data therein to obtain the table large model.

5. The table parsing method based on a large model according to claim 1, characterized in that: The number of rows and columns in the unstructured table in the unstructured table area, whether there are slashes in the cells, and whether the cells are merged are marked through a visual detection algorithm, thereby converting the unstructured table area into a structured table.

Citation Information

Cited By

  • Multi-modal data preprocessing and fusion technology and system based on artificial intelligence

    CN120833614A