Method for optimized storage and query of table data of large model knowledge base

By performing data splicing, natural language processing, document format conversion and vectorized storage on table data, the problem of large models lacking logic in understanding tabular data is solved, and the accuracy and readability of search results are improved.

CN120336320APending Publication Date: 2025-07-18INSPUR SOFTWARE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510399054.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The large model lacks logic in understanding and retrieval generation of tabular data, resulting in poor retrieval results.

Method used

Through multi-step processing such as data splicing, natural language processing, document format conversion and vectorized storage, the logic and coherence of tabular data are enhanced, and the Rerank model optimization retrieval system is introduced.

Benefits of technology

It significantly improves the accuracy and readability of the content generated by the search table data, and improves the logic and efficiency of the search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336320A_ABST
    Figure CN120336320A_ABST
Patent Text Reader

Abstract

The invention provides a method for optimally storing and querying table data of a large model knowledge base, which belongs to the field of data processing and comprises the following steps of: performing multi-step processing including data splicing, natural language processing optimization, document format conversion and vectorization storage on the table data; the logicality and coherence of the table data during storage and retrieval are enhanced, so that the accuracy and logicality of the retrieval generated content are remarkably improved, and the retrieval result is more readable and easy to understand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly to a method for optimizing the storage and query of tabular data in a large model knowledge base. Background Art

[0002] A large model knowledge base is a huge and complex information storage and retrieval system. Its principle is to combine a pre-trained language model with an external vector database, enabling the large model to utilize additional knowledge. By connecting the relationships between entities, a large-scale knowledge network is formed to represent rich semantic relationships and achieve retrieval-augmented generation of the large model.

[0003] However, through practice, we found that although the large model has a relatively accurate understanding of natural language, when it comes to vectorizing and storing tabular data such as excel and csv, there will only be data, and the retrieved answers lack logical concepts, resulting in a poor generation effect. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a method for optimizing the storage and query of tabular data in a large model knowledge base. It conducts an optimization on the problem that the large model lacks logic in storing, retrieving, and generating tabular data, enhances the large model's understanding of tabular data, and optimizes retrieval-augmented generation.

[0005] The technical solution of the present invention is as follows:

[0006] A method for optimizing the storage of tabular data based on a large model knowledge base, which enhances the logic and coherence of tabular data during storage and retrieval through multi-step processing of tabular data, including but not limited to data splicing, natural language processing optimization, document format conversion, and vectorized storage. As a result, the accuracy and logic of the retrieved and generated content are significantly improved, making the retrieval results more readable and easier to understand.

[0007] Furthermore,

[0008] Read the uploaded data in xls, xlsx, and csv formats, and splice the data and header information in a row or column manner according to the data structure to form a data text with simple logical relationships, enabling the text to reflect the basic associations between the data.

[0009] Use a pre-trained large model to polish and generate the generated data text with simple logic. By adding prompts to guide the model to generate a text with logical concepts, coherence, and containing all the original data, the readability and logic of the data are enhanced.

[0010] Convert the data text after being processed by the large model into a PDF format file to make full use of the advantages of the PDF format, facilitate semantic understanding and further processing of the large model, and ensure the persistent preservation of the file.

[0011] During the process of converting the document into PDF format, the document content is segmented and stored and vectorized. Each segment is converted into a vector representation and finally uploaded to the vector database to optimize the retrieval efficiency and quality and ensure the accuracy and timeliness of the retrieval results.

[0012] Furthermore,

[0013] The natural language processing optimization step also includes grammar checking and correction of the generated text to ensure the grammar correctness and semantic integrity of the final text.

[0014] The data splicing step also takes into account the handling of possible missing values in the table and the accuracy of the original value verification to ensure the integrity, accuracy, and error-free nature of the generated text.

[0015] The vectorized storage step also supports the establishment of indexes for vector data to further improve the retrieval performance.

[0016] By integrating the Rerank model to enhance the performance of the retrieval system, the query results are ensured to have higher accuracy and stronger relevance.

[0017] The beneficial effects of the present invention are

[0018] This method integrates key technical means such as data splicing, natural language processing optimization, data chunking, vectorization processing, document format conversion, and adding the rerank model. With this systematic data processing strategy, the tabular data is stored in a more logical segmented manner during storage, thereby greatly improving the quality during retrieval and content generation, making the retrieval and generation results more accurate and logical. Brief Description of the Drawings

[0019] Figure 1 is a schematic diagram of the optimized storage architecture of the large model knowledge base file;

[0020] Figure 2 is a schematic diagram of the table data extraction and processing flow. Detailed Embodiments

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0022] The present invention enhances the understanding and retrieval quality of tabular data by large models through multiple steps such as data splicing, natural language processing optimization, document format conversion, and vectorized storage:

[0023] A Design for Obtaining and Splicing Tabular Data

[0024] Convert the uploaded tabular data into text with a logical structure. First, read the data from the tabular file, and then splice the data with the header information according to certain logical rules (such as by row or by column) to make it into a piece of natural language with certain patterns. This process can be described by the following formula:

[0025]

[0026] Among them, T represents the original tabular data matrix, H represents the header information, S is the spliced string, and n and m represent the number of rows and columns respectively.

[0027] B Data Processing and Transformation

[0028] First, in order to extract information from the table and convert it into natural language text, technologies optimized for the table structure are used: in the Excel environment, the cell format can be set to "text" to ensure that numerical values are not misprocessed; while in a Word document, by selecting "Convert to Text" under the "Layout" tab and specifying appropriate delimiters such as commas or tab stops, it is ensured that each row and column in the table can be correctly mapped to the corresponding text paragraphs.

[0029] After converting the tabular data into text, through rigorous prompt design:

[0030] · Clearly define the objective: Clearly inform the model of the target style, tone, and specific domain terms of the expected generated text.

[0031] · Provide context: Give sufficient background information so that the model can understand the meaning of the input data and its potential associations.

[0032] · Control the structure: Indicate the discourse structure that the model should follow, such as how the beginning, main body, and end should be organized.

[0033] · Emphasize key points: Highlight which elements are particularly important and need to be expressed more meticulously.

[0034] Invoke a large model to process the generated text to make it a logical text. This process involves using natural language processing techniques to improve the logic and coherence of the text through the generation ability of the large model. This process can be described as:

[0035] S' = LLM(Prompt + S)

[0036] Here, Prompt is the pre-input of the prompt word, S' represents the polished text, and LLM refers to the data processing function of the large model.

[0037] C PDF Conversion Design

[0038] The large model has stronger semantic understanding ability for PDF-like documents. Therefore, the text after data processing conversion is output as a PDF file for long-term preservation and subsequent processing. This step is mainly format conversion, converting the text into PDF format for the large model to read.

[0039] D Vectorized Segmentation Storage Design

[0040] Use Faiss to perform vectorized segmentation storage on the generated PDF document. Split the PDF document into multiple paragraphs or sentences, and convert each paragraph or sentence into vector form and store it in the vector database. The specific steps are as follows:

[0041] · Create an index structure: Use IndexIVFADC or a similar composite index. This is an index method based on the inverted file, which combines quantization technology and can greatly improve the search speed while ensuring a certain accuracy. It is suitable for processing large-scale data sets and balances speed and accuracy by adjusting parameters.

[0042] · Vector addition:

[0043] a. Batch insertion: To improve efficiency, vector insertion should be done in batches as much as possible rather than one by one. This can not only reduce the number of I / O operations but also make full use of the powerful parallel processing capabilities of modern CPUs / GPUs.

[0044] b. Duplicate removal: Since some paragraphs or sentences may appear repeatedly in different documents, necessary duplicate removal measures should be implemented before insertion to avoid unnecessary storage overhead.

[0045] c. Metadata association: In addition to the vectors themselves, some auxiliary information such as the original document ID, paragraph number, etc. needs to be recorded. These metadata will play an important role in future query processes and help users quickly locate the accurate original content.

[0046] E Query Optimization

[0047] When it is necessary to find the content most similar to a given query, Faiss can quickly locate relevant entries and return their location information and similarity scores. To ensure the best query experience, this solution introduces a rerank model to reorder the results obtained from the initial retrieval, so as to more accurately locate the content that the user is truly interested in. It understands the user's query intent, combines context information, and filters out the most relevant few documents from a large number of candidate documents.

[0048] The above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A method for optimizing the storage and query of tabular data in a large model knowledge base, characterized in that by performing data splicing, natural language processing optimization, document format conversion, and vectorized storage on tabular data, the logic and coherence of tabular data during storage and retrieval are enhanced.

2. The method according to claim 1, characterized in that the data in xls, xlsx, and csv formats uploaded is read, and the data and header information are spliced according to rows or columns according to the data structure to form a data text with logical relationships, so that the text can reflect the basic associations between the data.

3. The method according to claim 2, characterized in that a pre-trained large model is used to polish and generate the generated data text with logic. By adding prompt words to guide the model to generate a text with logical concepts, coherence and containing all the original data, the readability and logic of the data are enhanced.

4. The method according to claim 3, characterized in that the data text after being processed by the large model data is converted into a PDF format file.

5. The method according to claim 4, characterized in that during the process of converting the document into PDF format, the document content is stored in segments and vectorized. Each segment is converted into a vector representation and finally uploaded to the vector database.

6. The method according to any one of claims 1-5, characterized in that the natural language processing optimization further includes grammar checking and correction of the generated text to ensure the grammar correctness and semantic integrity of the final text.

7. The method according to any one of claims 1-5, characterized in that the data splicing also takes into account the handling of possible missing values in the table and the accuracy of the original value verification to ensure the integrity, accuracy, and error-free nature of the generated text.

8. The method according to claim 1, characterized in that the vectorized storage step also supports the establishment of indexes for vector data to further improve the retrieval performance.

9. The method according to claim 1, characterized in that the performance of the retrieval system is enhanced by integrating the Rerank model.

Citation Information

Cited By

  • Table data query method and device based on intelligent agent, intelligent agent and medium

    CN121705332A