Intelligent table data escape and knowledge base construction method and system based on large language model
Through the intelligent tabular data escape method based on large language model, the problems of low efficiency and poor accuracy of table data extraction and knowledge base construction in the existing technology are solved, and high-precision data processing and efficient data storage and retrieval are realized.
Patent Information
- Application Number
- CN202510359742.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-17
AI Technical Summary
It is difficult for the prior art to effectively understand and convert the semantic meaning of table data, resulting in low efficiency and poor accuracy in data extraction and knowledge base construction.
The intelligent tabular data escape method based on the large language model is adopted, and the table data is identified and extracted through the table detection module, and the large language model is used for semantic understanding and data escape, and the tabular data is converted into semantic information, and merged into the knowledge base for vectorized storage.
It improves the accuracy and efficiency of data processing, realizes efficient data storage and retrieval, facilitates expansion and application, reduces manual intervention, and improves the efficiency of automated operations.
Smart Images

Figure CN120162379A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document processing, and in particular, to a method and system for intelligent tabular data transference and knowledge base construction based on a large language model. Background Art
[0002] With the development of information technology, the amount of data in documents has gradually increased. As a common form of carrying data, tables are widely used in various documents. However, traditional document processing methods usually only perform simple format extraction on tabular data, lacking intelligent semantic understanding and in-depth data mining. With the continuous progress of large language models, how to semantically understand the data in tables and convert it into a form that is convenient for retrieval and utilization has become an urgent problem to be solved.
[0003] In the prior art, although there are some tools that can extract tabular data, they usually cannot accurately understand the actual meaning of the data, and it is even more difficult to convert this data in a way suitable for storage and retrieval in a knowledge base. Summary of the Invention
[0004] The object of the present invention is to provide a method and system for intelligent tabular data transference and knowledge base construction based on a large language model, mainly aiming at the automated processing, semantic understanding of tabular data in documents and its application in a knowledge base, and solving the problems of low efficiency and poor accuracy in tabular data extraction and knowledge base construction in the prior art.
[0005] The present invention adopts the following technical solutions to solve the technical problems:
[0006] A method for intelligent tabular data transference and knowledge base construction based on a large language model includes the following steps:
[0007] Step S1, document input: Import various document formats.
[0008] Step S2, table detection and data extraction: Use a table detection module to identify the existence of a table, analyze the row and column structure of the table, locate the position of each cell, and extract the data in the table.
[0009] Step S3, semantic transference based on a large language model: Semantic understanding, the large language model can understand the data meaning of each cell according to the context and data content in the table; Data transference, after understanding the tabular data, the original data is transferred to convert it into more semantic information.
[0010] Step S4, data merging: The data after semantic transference will be merged and saved with the original data.
[0011] Step S5, Knowledge Base Construction and Data Chunking; The data after semantic transformation will be merged and saved in its original form. The merged data will serve as the input to the knowledge base, and the data will be logically chunked. Through natural language processing techniques, the transformed data will be vectorized.
[0012] Step S6, Knowledge Base Query and User Interaction: Retrieve the tabular data in the knowledge base through the large language model knowledge base system, and quickly respond to the user's query requests through the vector index of the knowledge base.
[0013] Furthermore, data cleaning is also included in Step S2: The table detection module preprocesses the extracted data, including removing redundant whitespace, abnormal data, and merged cell data, to ensure that the structure of the tabular data is clear and the format is consistent.
[0014] Furthermore, in Step S5, the vectorization methods include TF-IDF, Word2Vec, and BERT.
[0015] Furthermore, Step S6 also includes: Intelligent Retrieval. Through the vectorized data, relevant tabular information is obtained by inputting a natural language query, and matching is performed according to the semantics of the query, and relevant data is provided.
[0016] Furthermore, Step S6 also includes: Dynamic Update. As more inputs containing tabular data and semantic transformation are completed, the knowledge base is dynamically updated to ensure that the content of the knowledge base always remains up-to-date.
[0017] An intelligent tabular data transformation and knowledge base construction system based on a large language model, including:
[0018] A document input module, a table detection module, a merged data module, a large language model semantic transformation module, a knowledge base construction module, and a query and response module; among them,
[0019] The document input module is used to support the import of multiple document formats;
[0020] The table detection module is responsible for identifying tables in the document and extracting data;
[0021] The merged data module is used to merge and save the data after semantic transformation in its original form;
[0022] The large language model semantic transformation module is used to perform semantic analysis and transformation on the extracted tabular data;
[0023] The knowledge base construction module is used to merge the data after semantic analysis and transformation, and perform chunking processing on the merged data and store it in a vectorized manner;
[0024] The query and response module interacts with the knowledge base through the query interface to obtain relevant tabular data.
[0025] Advantages of the present invention:
[0026] 1. Improve the accuracy of data processing: Through the semantic understanding technology of the large language model, the tabular data can be accurately converted into information with practical significance.
[0027] 2. Efficient storage and retrieval: Merging the data as the input of the knowledge base helps to standardize the storage and structured management of the data, and supports efficient retrieval.
[0028] 3. Facilitate expansion and application: This technical solution can be widely applied to various document processing, data mining and intelligent query systems, and has strong scalability.
[0029] 4. Automated operation: Through the automated table recognition and data escape process, manual intervention is reduced and efficiency is improved. Description of the Drawings
[0030] Figure 1 It is the flowchart of the method of the present invention. Detailed Embodiments
[0031] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] Refer to Figure 1 , the present invention provides an intelligent table data escape and knowledge base construction method based on a large language model. By adopting the semantic escape technology of the large language model, the data in the table is intelligently converted, and through the operation of merging the data, it is embedded in the knowledge base to facilitate users to query and utilize, and solve the problem of insufficient accuracy in table data processing in the prior art. The method includes the following steps:
[0033] Step S1, document input: Import various document formats;
[0034] Step S2, table detection and data extraction:
[0035] First, it is necessary to automatically identify the tabular data in the document. This step is implemented by a table detection module, which can identify the tables in the document and support various common document formats, such as PDF, Word, Excel, etc.
[0036] Secondly, perform table structure analysis. The table detection module not only identifies the existence of the table, but also analyzes the row and column structure of the table, locates the position of each cell, and extracts the data in the table.
[0037] Then perform data cleaning. The table detection module preprocesses the extracted data, including removing redundant whitespace, abnormal data, merging cell data, etc., to ensure that the structure of the table data is clear and the format is consistent.
[0038] Step S3, semantic transformation based on the large language model:
[0039] The data in the table often appears in the form of numbers, text, or a mixture, lacking semantic information. To better understand the actual meaning of the data, the system performs semantic transformation on the table data by introducing a large language model.
[0040] Semantic understanding: The large language model can understand the meaning of the data in each cell according to the context and data content in the table. For example, for the "quantity" and "price" columns in the table, the system can automatically identify them as the quantity and price of the goods and convert them into corresponding standardized information.
[0041] Data transformation: After understanding the table data, the system transforms the original data, converting it into more semantic information. For example, it is transformed into structured text or natural language descriptions, so that the data in the table not only retains the original numerical values but also provides meaningful context explanations.
[0042] Step S4, merge data: The data after semantic transformation will be merged and saved with the original data.
[0043] Step S5, knowledge base construction and data chunking:
[0044] The merged data will be used as the input to the knowledge base, and the system will perform chunking on it to ensure the efficiency of data storage and retrieval.
[0045] Data chunking: Logically chunk the merged data, for example, divide it by chapters, tables, paragraphs, etc., for efficient storage and quick query in the knowledge base.
[0046] Vectorization processing: Through natural language processing (NLP) techniques, vectorize the transformed data. Vectorization is to convert data into numerical vectors so that the computer can perform storage, retrieval, and matching operations efficiently. Commonly used vectorization methods include TF-IDF, Word2Vec, BERT, etc.
[0047] Step S6, knowledge base query and user interaction:
[0048] Users can obtain tabular data in the knowledge base through the large language model knowledge base system, and the system quickly responds to users' query requests through the vector index of the knowledge base.
[0049] Intelligent retrieval: Through the vectorized data, users can obtain relevant tabular information by entering a natural language query. The system can match according to the semantics of the query and provide relevant data.
[0050] Dynamic update: With the input of more tabular data and the completion of semantic transformation, the knowledge base can be dynamically updated to ensure that the content of the knowledge base always remains up-to-date.
[0051] The present invention also provides an intelligent tabular data transformation and knowledge base construction system based on a large language model, including:
[0052] A document input module, a table detection module, a data merging module, a large language model semantic transformation module, a knowledge base construction module, and a query and response module; among them,
[0053] The document input module is used to support the import of multiple document formats;
[0054] The table detection module is responsible for identifying tables in the document and extracting data;
[0055] The data merging module is used to merge and save the data after semantic transformation with the original data;
[0056] The large language model semantic transformation module is used to perform semantic analysis and transformation on the extracted tabular data;
[0057] The knowledge base construction module is used to merge the data after semantic analysis and transformation, and perform chunking processing on the merged data and store it vectorially;
[0058] The query and response module interacts with the knowledge base through a query interface to obtain relevant tabular data.
[0059] The present invention can improve the accuracy of data processing: Through the large language model semantic understanding technology, tabular data can be accurately transformed into information with practical significance. Achieve efficient storage and retrieval: The merged data is used as the input of the knowledge base, which helps to standardize the storage and structured management of data and supports efficient retrieval. Facilitate expansion and application: This technical solution can be widely applied to various document processing, data mining, and intelligent query systems, and has strong scalability. Achieve automated operation: Through the automated table recognition and data transformation process, manual intervention is reduced and efficiency is improved.
[0060] The working principle of the present invention is as follows:
[0061] First, the document input module receives document files in various formats, and the table detection module automatically identifies the table data therein. Then, the large language model semantic transformation module is used to perform intelligent semantic transformation on the extracted table data, converting the original table data such as numbers and texts into structured information with practical significance. After that, the knowledge base construction module saves the transformed data as new merged data, and then performs chunking processing on the merged data, and uses vectorization technology to store the data in the knowledge base. Finally, the user makes a natural language query through the query and response module to obtain the required table data from the knowledge base. Through this system, the efficient extraction, semantic transformation, and knowledge base construction of table data are realized, greatly improving the data management and retrieval efficiency, and having strong scalability and practicality. This system is applicable to multiple fields such as enterprise management, scientific research data storage, and legal document analysis, and can significantly improve work efficiency and data utilization value.
[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the present invention in each embodiment.
Claims
1. A method for constructing intelligent table data translation and knowledge base based on a large language model, characterized in that: The steps include: Step S1, document input: importing multiple document formats; Step S2, table detection and data extraction: using a table detection module to identify the existence of a table, analyze the row and column structure of the table, locate the position of each cell, and extract the data in the table; Step S3, semantic interpretation based on the large language model: semantic understanding, the large language model can understand the meaning of the data in each cell according to the context and data content in the table; Data escape: after understanding the table data, the original data is escaped and converted into more semantic information; Step S4, merging data: the semantically escaped data will be merged and saved with the original data; Step S5, knowledge base construction and data segmentation: The merged data will be used as the input of the knowledge base, the data will be logically segmented, and the escaped data will be vectorized through natural language processing technology; Step S6, knowledge base query and user interaction: obtain the table data in the knowledge base through the large language model knowledge base system, and quickly respond to the user's query request through the vector index of the knowledge base.
2. According to the method of intelligent table data escape and knowledge base construction based on large language model in claim 1, it is characterized in that: Step S2 also includes data cleaning: the table detection module pre-processes the extracted data, including removing redundant blanks, abnormal data, and merging cell data to ensure that the structure of the table data is clear and the format is consistent.
3. According to claim 2, a method for constructing a smart table data escape and knowledge base based on a large language model is characterized in that: In step S5, the vectorization methods include TF-IDF, Word2Vec, and BERT.
4. According to claim 3, a method for constructing a smart table data escape and knowledge base based on a large language model is characterized in that: Step S6 also includes: intelligent retrieval, obtaining relevant table information through vectorized data by inputting natural language queries, matching according to the semantics of the queries, and providing relevant data.
5. According to the method of intelligent table data escape and knowledge base construction based on large language model in claim 4, it is characterized in that: Step S6 also includes: dynamic update: as more table data is input and semantic escaping is completed, the knowledge base is dynamically updated to ensure that the content of the knowledge base is always kept up to date.
6. An intelligent table data escape and knowledge base construction system based on a large language model, characterized in that: include: Document input module, table detection module, data merging module, large language model semantic escape module, knowledge base construction module and query and response module; among them, The document input module is used to support the import of multiple document formats; The table detection module is responsible for identifying the table in the document and extracting data; The data merging module is used to merge and save the semantically escaped data with the original data; The large language model semantic escape module is used to perform semantic analysis and escape on the extracted table data; The knowledge base construction module is used to merge the semantically analyzed and escaped data, and to process the merged data in blocks and perform vectorized storage; The query and response module interacts with the knowledge base through a query interface to obtain relevant table data.
Citation Information
Cited By
Method and system for judging disclosure state of specific form of financial report based on large model
CN122088467A