Table analysis processing method and system based on large language model

Through the tabular analysis and processing method based on large language models, tabular data in professional fields is quickly obtained, which solves the problem of time-consuming and labor-intensive acquisition of high-quality training data sets, and achieves data quality improvement and large model analysis capabilities.

CN119962499APending Publication Date: 2025-05-09AISINO CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411969239.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Obtaining high-quality training data sets is time-consuming and labor-intensive, and requires huge human, material and time costs, especially in vertical fields with complex data sources and different formats.

Method used

Provide a table analysis and processing method based on a large language model. By obtaining document content, identifying table data in the financial and taxation field, extracting fiscal and taxation keywords, analyzing table data and generating labeled data, and converting the format into markdown format text, quickly obtaining table data in the professional field as training data.

Benefits of technology

It effectively improves data quality, reduces the time and cost of obtaining high-quality training data sets, accelerates the development of industry large models, and improves the analysis and understanding of tabular data by large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005219123380000011
    Figure HDA0005219123380000011
  • Figure HDA0005219123380000021
    Figure HDA0005219123380000021
Patent Text Reader

Abstract

The invention discloses a table analysis processing method and system for a large language model. The method comprises the following steps: acquiring document content of data to be extracted, and preprocessing the document content; identifying whether table data in the finance and taxation field exists in the document content or not; if the document content has the table data of the finance and taxation field, extracting finance and taxation keywords from the table data; analyzing the table data based on the finance and tax keyword, acquiring context information related to the table data from the document content, and generating a piece of annotation data containing the table data; and performing format conversion on the annotation data, converting the annotation data into a text in a makdown format, and outputting the text. Therefore, table data related to the finance and taxation field and table context information are extracted from the document, data cleaning and format conversion are completed, and the data in the Markdown format are output in a unified mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of table analysis, and more specifically, to a table analysis processing method and system based on a large language model. Background Art

[0002] In today's digital wave, data is not only the clay that shapes the real world, but also the engine that drives innovation and technological progress. Especially in the research and development and application of large industry models, the importance of data has become increasingly prominent and has become an indispensable cornerstone.

[0003] Although the underlying capabilities of various industry big models may be similar, and the algorithm model level cannot completely determine their quality, the quality of the training data set can significantly affect the performance of the model. Scholars represented by Andrew Ng believe that "artificial intelligence should be data-centric rather than model-centric." This view emphasizes the key role of labeled high-quality data in unleashing the potential of big models. If the industry can focus more on improving data quality, the development of big models will be even faster.

[0004] For large industry models in vertical fields, meeting industry professional needs means that knowledge in specific fields must be incorporated for pre-training and alignment. This requires the model to have a certain level of professional depth and be able to understand and process data from various sources and formats, such as industry databases, professional documents, and professional websites. Since these data often come from complex sources and in different formats, and even require customized processing methods for each piece of data, obtaining high-quality training data sets is not only time-consuming and laborious, but also requires huge human, material, and time costs. Summary of the invention

[0005] According to the present invention, a table analysis and processing method and system based on a large language model are provided to solve the technical problem that data often comes from complex sources and has different formats, and even requires customized processing methods for each piece of data. Therefore, obtaining high-quality training data sets is not only time-consuming and labor-intensive, but also requires huge human, material and time costs.

[0006] According to a first aspect of the present invention, a table analysis and processing method based on a large language model is provided, comprising:

[0007] Obtaining document content of data to be extracted, preprocessing the document content, and decomposing the document content into basic units that can be processed by the model;

[0008] Identify whether there is tabular data in the field of finance and taxation in the document content;

[0009] If the document content contains table data in the field of finance and taxation, extract finance and taxation keywords from the table data;

[0010] Parsing the table data based on the finance and taxation keywords, acquiring context information related to the table data from the document content, and generating a piece of annotation data containing the table data;

[0011] The annotated data is formatted and converted into text in makdown format and outputted.

[0012] Optionally, obtaining document content of the data to be extracted, preprocessing the document content, and decomposing the document content into basic units that can be processed by the model includes:

[0013] Get the document content of the data to be extracted and store the document content in doc or pdf format;

[0014] The document content is preprocessed, i.e., word segmentation and part-of-speech tagging, to decompose the document content into basic units that can be processed by the model.

[0015] Optionally, identifying whether there is tabular data in the financial and taxation field in the document content includes:

[0016] Read the data in the document content, identify whether there is table data in the finance and taxation field in the document content, and store the table data as a list object.

[0017] Optionally, the table data is parsed based on the finance and taxation keywords, context information related to the table data is obtained from the document content, and a piece of annotation data containing the table data is generated, including:

[0018] Semantically match the financial and taxation keywords with the document content, determine the location and content of the table data in the document content, find out the context information related to the table data in the document content, and generate a piece of annotation data containing the table data.

[0019] According to another aspect of the present invention, there is also provided a table analysis and processing system based on a large language model, comprising:

[0020] A document content acquisition module is used to acquire the document content of the data to be extracted, pre-process the document content, and decompose the document content into basic units that can be processed by the model;

[0021] A table data identification module is used to identify whether there is table data in the financial and taxation field in the document content;

[0022] A module for extracting financial and tax keywords, used to extract financial and tax keywords from the table data if there is table data in the financial and tax field in the document content;

[0023] A module for generating annotated data, used for parsing the table data based on the finance and taxation keywords, obtaining context information related to the table data from the document content, and generating an annotated data containing the table data;

[0024] The format conversion output module is used to convert the format of the annotation data into text in makdown format and output it.

[0025] Optionally, obtain the document content module, including:

[0026] The document content storage submodule is used to obtain the document content of the data to be extracted and store the document content in doc or pdf format;

[0027] The document content decomposition submodule is used to pre-process the document content, namely, word segmentation and part-of-speech tagging, and decompose the document content into basic units that can be processed by the model.

[0028] Optionally, identifying a table data module includes:

[0029] The table data identification submodule is used to read the data in the document content, identify whether there is table data in the financial and taxation field in the document content, and store the table data as a list object.

[0030] Optionally, a module for generating annotation data includes:

[0031] Generate annotated data submodule, which is used to semantically match the financial and tax keywords with the document content, determine the position and content of the table data in the document content, find out the context information related to the table data in the document content, and generate an annotated data containing the table data.

[0032] According to another aspect of the present invention, there is further provided a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any of the above methods are implemented.

[0033] According to another aspect of the present invention, there is also provided an electronic device, comprising: the computer-readable storage medium; and

[0034] One or more processors are used to execute the program in the computer-readable storage medium.

[0035] Therefore, the present invention can quickly obtain tabular data in professional fields as training data according to the tabular data analysis capability requirements of the vertical field big model, and more efficiently supplement high-quality tabular analysis data in professional fields.

[0036] Retrieve table data from professional documents, determine the starting position of the table data, analyze the text in the table, use the machine learning big model to expand the search for contextual data associated with the table data, and extract these texts and tables into one piece of data; process the extracted data, use regular expressions to remove private data and spaces, convert the calculation formula format, extract the table text and convert it into markdown format text, and finally output plain text table data that is more conducive to big model recognition.

[0037] Search for the starting position of table data in the document based on specific keywords and character recognition; parse the text of the table data and identify the starting and ending positions of the table context data based on the table data content; use conversion instructions to convert the table data format and generate table data in makdown format to improve the large language model's ability to analyze and understand the table data. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:

[0039] Figure 1 A schematic diagram of a flow chart of a table analysis and processing method based on a large language model described in this embodiment;

[0040] Figure 2 This is a schematic diagram of a table analysis and processing system based on a large language model described in this embodiment. DETAILED DESCRIPTION

[0041] Now, exemplary embodiments of the present invention are described with reference to the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely and to fully convey the scope of the present invention to those skilled in the art. The terms used in the exemplary embodiments shown in the accompanying drawings are not intended to limit the present invention. In the accompanying drawings, the same units / elements are marked with the same reference numerals.

[0042] Unless otherwise specified, the terms (including technical terms) used herein have the commonly understood meanings to those skilled in the art. In addition, it is understood that the terms defined in commonly used dictionaries should be understood to have the same meanings as those in the context of the relevant fields, and should not be understood as idealized or overly formal meanings.

[0043] According to a first aspect of the present invention, a table analysis and processing method 100 based on a large language model is provided. Figure 1 As shown, the method 100 includes:

[0044] S101: Obtain document content of data to be extracted, pre-process the document content, and decompose the document content into basic units that can be processed by the model;

[0045] S102: Identify whether there is table data in the field of finance and taxation in the document content;

[0046] S103: If the document content contains table data in the field of finance and taxation, extract finance and taxation keywords from the table data;

[0047] S104: parsing the table data based on the finance and taxation keywords, acquiring context information related to the table data from the document content, and generating a piece of annotation data containing the table data;

[0048] S105: Convert the format of the annotated data into a text in makdown format and output it.

[0049] Specifically, in order to construct high-quality table analysis data for large models in the field of finance and taxation, the present invention proposes an innovative processing method for constructing table analysis pre-training data. The method accurately extracts tables and their context data from professional documents, then strictly formats and converts these data, and finally outputs high-quality table analysis pre-training data for large models.

[0050] In this way, we can not only effectively improve data quality, but also accelerate the development of industry big models, making data a powerful engine to drive innovation and technological progress. In the wave of digitalization, let us ride on the wings of data and soar into the sky of wisdom.

[0051] The present invention obtains the document from which data is to be extracted. These data generally come from professional databases, books, and test questions, and are generally stored in doc or pdf format. The document content is read, and the table data is stored as a list object. The table content is parsed, and the semantic content recognition method of the large language model can be used to determine whether the table data is financial and taxation data. The next step is to identify text paragraph data that is strongly associated with the table content in the document content, and output these paragraphs and tables as one piece of data; use the multi-layer perception context learning technology of the large model, and the data cleaning algorithm to uniformly convert the table data into text data in markdown format. Finally, these processed data can be used for large model training to improve the table learning ability of the large model in professional vertical fields.

[0052] The module for identifying tabular data in the financial and taxation fields has input conditions of text in document format and keywords of financial and taxation tabular data. The processing logic is that the first step is to call the large model interface to identify the tabular data. The second step is to identify whether it is tabular data in the financial and taxation field through the semantic understanding and analysis of the large model. The output result is the tabular data text in the financial and taxation field.

[0053] The table content parsing module has the input conditions of the recognized table and the text content of the entire document. The processing logic mainly includes: Preprocessing stage: First, preprocess the document, including word segmentation, part-of-speech tagging, etc., to decompose the document into basic units that can be processed by the model. Keyword extraction: Use the TF-IDF (term frequency-inverse document frequency) algorithm to extract keywords from the table data, such as "Company A", "Company B", "operating income", etc. Context matching: Use natural language processing (NLP) technology, such as Word2Vec or BERT model, to semantically match the extracted keywords with the document content to find out the context information related to the table data. Result output: In the end, the model outputs context information closely related to the table data.

[0054] The format conversion module takes a piece of text data containing table data as input. The processing logic includes data cleaning, table format conversion, content completion and other methods. It finally outputs text data containing tables in markdown format.

[0055] The above table analysis data processing method can help the large model quickly obtain the knowledge training data of tables in the field of finance and taxation. In the document, the location and content boundaries of the table data are quickly determined based on the keywords, and the machine learning model is used to find the context information associated with the text in the table, and a training data centered on the table data is quickly extracted; the extracted data is cleaned and formatted, and finally the text data is output in markdown format.

[0056] Thus, the starting position of the table data is searched in the document based on specific keywords and character recognition; the text of the table data is parsed, and the starting and ending positions of the table context data are identified based on the table data content; the table data is converted using conversion instructions to generate table data in the markdown format, which is used to improve the large language model's ability to analyze and understand the table data. The text classification algorithm of the large model is used to extract table analysis data that is strongly related to the financial and tax fields; the table context data can be extracted from the document based on the table's data features and prompt word engineering, such as the industry background, calculation methods, relevant regulations, etc. surrounding the table; the data is cleaned and uniformly output as text data in the markdown format, which is convenient for the large model to parse and display data.

[0057] Optionally, obtaining document content of the data to be extracted, preprocessing the document content, and decomposing the document content into basic units that can be processed by the model includes:

[0058] Get the document content of the data to be extracted and store the document content in doc or pdf format;

[0059] The document content is preprocessed, i.e., word segmentation and part-of-speech tagging, to decompose the document content into basic units that can be processed by the model.

[0060] Optionally, identifying whether there is tabular data in the financial and taxation field in the document content includes:

[0061] Read the data in the document content, identify whether there is table data in the finance and taxation field in the document content, and store the table data as a list object.

[0062] Optionally, the table data is parsed based on the finance and taxation keywords, context information related to the table data is obtained from the document content, and a piece of annotation data containing the table data is generated, including:

[0063] Semantically match the financial and taxation keywords with the document content, determine the location and content of the table data in the document content, find out the context information related to the table data in the document content, and generate a piece of annotation data containing the table data.

[0064] Therefore, the present invention can quickly obtain tabular data in professional fields as training data according to the tabular data analysis capability requirements of the vertical field big model, and more efficiently supplement high-quality tabular analysis data in professional fields.

[0065] Retrieve table data from professional documents, determine the starting position of the table data, analyze the text in the table, use the machine learning big model to expand the search for contextual data associated with the table data, and extract these texts and tables into one piece of data; process the extracted data, use regular expressions to remove private data and spaces, convert the calculation formula format, extract the table text and convert it into markdown format text, and finally output plain text table data that is more conducive to big model recognition.

[0066] Search for the starting position of table data in the document based on specific keywords and character recognition; parse the text of the table data and identify the starting and ending positions of the table context data based on the table data content; use conversion instructions to convert the table data format and generate table data in makdown format to improve the large language model's ability to analyze and understand the table data.

[0067] According to another aspect of the present invention, a table analysis and processing system 200 based on a large language model is also provided. Figure 2 As shown, the system 200 includes:

[0068] A document content acquisition module 210 is used to acquire the document content of the data to be extracted, pre-process the document content, and decompose the document content into basic units that can be processed by the model;

[0069] A table data identification module 220 is used to identify whether there is table data in the financial and taxation field in the document content;

[0070] A module 230 for extracting financial and tax keywords is used to extract financial and tax keywords from the table data if there is table data in the financial and tax field in the document content;

[0071] A tag data generating module 240 is used to parse the table data based on the finance and taxation keywords, obtain context information related to the table data from the document content, and generate a tag data containing the table data;

[0072] The format conversion output module 250 is used to convert the format of the annotation data into text in makdown format and output it.

[0073] Optionally, the module for obtaining document content 210 includes:

[0074] The document content storage submodule is used to obtain the document content of the data to be extracted and store the document content in doc or pdf format;

[0075] The document content decomposition submodule is used to pre-process the document content, namely, word segmentation and part-of-speech tagging, and decompose the document content into basic units that can be processed by the model.

[0076] Optionally, the identification table data module 220 includes:

[0077] The table data identification submodule is used to read the data in the document content, identify whether there is table data in the financial and taxation field in the document content, and store the table data as a list object.

[0078] Optionally, the module 240 for generating annotation data includes:

[0079] Generate annotated data submodule, which is used to semantically match the financial and tax keywords with the document content, determine the position and content of the table data in the document content, find out the context information related to the table data in the document content, and generate an annotated data containing the table data.

[0080] A table analysis and processing system 200 based on a large language model according to an embodiment of the present invention corresponds to a table analysis and processing method 100 based on a large language model according to another embodiment of the present invention, and will not be described in detail herein.

[0081] According to another aspect of the present invention, there is further provided a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any of the above methods are implemented.

[0082] According to another aspect of the present invention, there is also provided an electronic device, comprising: the computer-readable storage medium; and

[0083] One or more processors are used to execute the program in the computer-readable storage medium.

[0084] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.

[0085] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0086] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0088] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0089] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A table analysis and processing method based on a large language model, characterized in that: include: Obtaining document content of data to be extracted, preprocessing the document content, and decomposing the document content into basic units that can be processed by the model; Identify whether there is tabular data in the field of finance and taxation in the document content; If the document content contains table data in the field of finance and taxation, extract finance and taxation keywords from the table data; Parsing the table data based on the finance and taxation keywords, acquiring context information related to the table data from the document content, and generating a piece of annotation data containing the table data; The annotated data is formatted and converted into text in makdown format and outputted.

2. The method according to claim 1, characterized in that Obtain the document content of the data to be extracted, pre-process the document content, and decompose the document content into basic units that can be processed by the model, including: Get the document content of the data to be extracted and store the document content in doc or pdf format; The document content is preprocessed, i.e., word segmentation and part-of-speech tagging, to decompose the document content into basic units that can be processed by the model.

3. The method according to claim 1, characterized in that Identify whether the document content contains tabular data in the financial and tax fields, including: Read the data in the document content, identify whether there is table data in the finance and taxation field in the document content, and store the table data as a list object.

4. The method according to claim 1, characterized in that: The table data is parsed based on the finance and taxation keywords, context information related to the table data is obtained from the document content, and a piece of annotation data containing the table data is generated, including: Semantically match the financial and taxation keywords with the document content, determine the location and content of the table data in the document content, find out the context information related to the table data in the document content, and generate a piece of annotation data containing the table data.

5. A table analysis and processing system based on a large language model, characterized in that: include: A document content acquisition module is used to acquire the document content of the data to be extracted, pre-process the document content, and decompose the document content into basic units that can be processed by the model; A table data identification module is used to identify whether there is table data in the financial and taxation field in the document content; A module for extracting financial and tax keywords, used to extract financial and tax keywords from the table data if there is table data in the financial and tax field in the document content; A module for generating annotated data, used for parsing the table data based on the finance and taxation keywords, obtaining context information related to the table data from the document content, and generating an annotated data containing the table data; The format conversion output module is used to convert the format of the annotation data into text in makdown format and output it.

6. The system according to claim 5, characterized in that Get document content modules, including: The document content storage submodule is used to obtain the document content of the data to be extracted and store the document content in doc or pdf format; The document content decomposition submodule is used to pre-process the document content, namely, word segmentation and part-of-speech tagging, and decompose the document content into basic units that can be processed by the model.

7. The system according to claim 5, characterized in that Identify tabular data modules, including: The table data identification submodule is used to read the data in the document content, identify whether there is table data in the financial and taxation field in the document content, and store the table data as a list object.

8. The system according to claim 5, characterized in that Generate annotation data module, including: Generate annotated data submodule, which is used to semantically match the financial and tax keywords with the document content, determine the position and content of the table data in the document content, find out the context information related to the table data in the document content, and generate an annotated data containing the table data.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

10. An electronic device, characterized in that: include: The computer readable storage medium as claimed in claim 9; as well as One or more processors are used to execute the program in the computer-readable storage medium.