RAG-based table data processing method and system
Through the RAG-based tabular data processing method, combined with vector search and keyword search, enhancement, structure and standardization are carried out, and mixed indexes and context joint indexes are established, which solves the problems of low query efficiency and poor structural adaptability in complex tabular data processing, and realizes multi-dimensional retrieval and adaptive optimization.
Patent Information
- Application Number
- CN202510489809.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-18
AI Technical Summary
When processing complex tabular data, the prior art has low query efficiency, poor structural adaptability, difficulty in cross-file association, lack of intelligent retrieval capabilities, and difficult to cope with multi-level nested and noisy data.
The RAG-based tabular data processing method is adopted to obtain user query content for vector search and keyword search, and combine dynamic template filling and data fine-tuning models to enhance, structure and standardize processing, establish mixed indexes and context joint indexes to realize multi-dimensional retrieval and adaptive optimization.
It realizes the unified feature representation of cross-modal table content, has multi-dimensional retrieval capabilities, solves the problems of semantic faults and inefficient retrieval efficiency in traditional methods, and forms a closed-loop processing system with adaptive capabilities.
Smart Images

Figure CN120354833A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of structured data processing, and particularly relates to a method and system for processing tabular data based on RAG (Retrieval-Augmented Generation). Background Art
[0002] In 2020, Meta (formerly Facebook AI) first proposed the retrieval-augmented generation framework, which combines non-parametric retrieval (such as BM25 or Dense Retrieval) with parametric generation models (such as BART). Early related models included Meta's retrieval-augmented generation (BM25 or Dense Retrieval + BART generation), Google's pre-training and fine-tuning for knowledge-intensive tasks. Advanced models in the later stage included independently encoding retrieved documents to improve generation efficiency, the model autonomously controlling the retrieval timing and content to enhance controllability. The overall evolution direction is as follows: (1) Model architecture: from single retrieval to generation pipeline to end-to-end joint optimization (such as RAG-Token and RAG-Sequence); (2) Retrieval technology: from sparse retrieval (BM25) to dense retrieval (such as DPR, ANCE); (3) Application expansion: from text Q&A to multi-modal, dialogue system and other scenarios. The existing technical solutions mainly first identify the rows and columns of the table, structure the table rows and columns, describe them as "key-value pairs", then convert the key-value pairs into describable text, and finally establish indexes according to the table cells for convenient data retrieval and generation. The main disadvantages are as follows: Complex situations such as merged cells and multi-level table headers are not considered, and the table parsing results may be incomplete. The context of the table is not considered, and the parsing of the table may be relatively one-sided, fragmenting the complete information. Only the cells of the table are indexed, which is difficult to support when dealing with complex queries. The model retrieval method is relatively single, and it is difficult to quickly retrieve the required results. For queries on numerical or statistical data, the support ability is weak, and complex query calculations cannot be supported. For the noise in the table, such as irrelevant annotations, empty cells or format errors, there is a lack of necessary processing. There is a lack of an evaluation and optimization mechanism, and it is difficult to continuously improve the retrieval accuracy and the correctness of the generation results.
[0003] Prior Art One, Application No.: CN202411046896.7 discloses a method and device for processing tabular data. The method includes: listening for the scrolling action of the scroll bar on the current page of the table list and recording the real-time scrolling distance value LR of the scroll bar, and obtaining the first control vector K1 in real time; setting the second control vector K2 corresponding to the tabular data; traversing all elements included in the tabular data, and for each element EMi, updating the cumulative height AHi = AHi-1 + Hi; forming a loading ternary number SJ = {TJ, VG, KP}, and determining the loading strategies for CLA data, PT value, and PB value based on the ternary number SJ. Although it sets a matching smooth display data according to different operation methods of the current data to avoid operation jams; however, the indexing method is single, only adjusting the data loading strategy based on the scrolling speed and cumulative height, without considering the structural sparsity or multi-level nested relationship of the table, and the efficiency is insufficient when querying complex data; the data processing ability is weak, without enhancing, semanticizing, or standardizing the data, and it is difficult to handle unstructured or cross-file tabular data association queries.
[0004] Prior Art Two, Application No.: CN202410793445.3 discloses a method, device, equipment, and storage medium for processing tabular data. The method includes: obtaining the original file; generating the file data and node data corresponding to the original file according to the original file; determining the first data corresponding to the original file according to the file data; determining the second data corresponding to the original file according to the node data; creating an online table based on the first data and the second data to obtain an online editable table file. Although it can efficiently and accurately manage tabular data, meet the requirements of modern laboratories for tabular data processing and generation, and contribute to tabular data collaboration sharing and long-term preservation; however, it lacks intelligent retrieval ability, only relying on the original file node data for storage management, without introducing a vector retrieval or dynamic index optimization mechanism, and has weak support for complex queries; it is difficult to perform cross-file association retrieval, and the solution neither involves hybrid index construction nor can optimize the association query efficiency of multiple files through context joint indexing.
[0005] Prior Art Three, Application No.: CN202410797889.4 discloses a method, apparatus, device, and storage medium for processing tabular data. In response to a trigger operation on a field control in a table page, a field interface is displayed; in the field interface, a field component and a first field are determined; based on the field component and the first field, a second field is displayed in the table page, and second data is generated in the data column corresponding to the second field. The second data is data generated using the field component based on the first data in the data column corresponding to the first field. Although the field type of the second field is determined based on the field component, improving the user's usage efficiency; however, the automation processing ability is limited. Although data can be dynamically generated according to user operations, it does not include data enhancement or structured processing methods and cannot adapt to complex tables in different formats; the query mechanism is missing, no index optimization strategy is established for tabular data, and only manual operations are relied on to generate fields, which does not support efficient multi-dimensional retrieval.
[0006] Currently, Prior Art One, Prior Art Two, and Prior Art Three have problems of low query efficiency, poor structural adaptability, and difficulty in cross-file association in the processing of complex tabular data. Therefore, the present invention provides a method and system for processing tabular data based on RAG. Summary of the Invention
[0007] To solve the above technical problems, the present invention provides a method for processing tabular data based on RAG, including the following steps:
[0008] Obtain the user's query content, convert it into a key-value combination, and perform vector retrieval and keyword retrieval; according to the results of the key-value combination, vector retrieval, and keyword retrieval, perform dynamic template filling, and instruct the data fine-tuning model to guide the generation of instructions.
[0009] Optionally, perform enhancement processing, structured processing, semantic processing, and standardization processing on the content data in the table to obtain the processed content data.
[0010] Optionally, the process of obtaining the processed content data includes the following steps:
[0011] Create a repository for storing parsing programs;
[0012] When a file with structural characteristics is obtained, issue an instruction to trigger the parsing program, and confirm the electronic file format and parsing method in sequence according to the recognition process of the parsing program.
[0013] Optionally, when the electronic file format is confirmed, perform parsing according to the corresponding interpretation method; using the electronic file format as the starting node and the parsing method as the end node, form a table parsing path in the file.
[0014] Optionally, the electronic file format is a file extension, and the files include Word, Excel, and PDF.
[0015] Optionally, after forming the table parsing path in the file, obtain the metadata, table headers, data for each row, and digital date content data in the table, and perform enhancement processing, structuring processing, semantic processing, and standardization processing according to different content data.
[0016] Optionally, the parsing program includes the electronic file format and the corresponding parsing method, as well as a file with structural characteristics that triggers the parsing program.
[0017] Optionally, according to the test cases, in combination with the query latency and accuracy, adjust the hybrid index and the context joint index to rows, columns, or sub-tables.
[0018] Optionally, the process of adjusting the hybrid index and the context joint index to rows, columns, or sub-tables includes the following steps:
[0019] Based on the combined feedback results of vector retrieval and keyword retrieval, perform quantitative analysis on the blank area distribution of the sparse table; extract the topological features of the sub-table according to the processed header hierarchy relationship; when the user's query key-value combination involves multi-level associations, automatically disassemble the original single-level index into a tree-like sub-table index network, and use the generated lightweight calculation results as the verification data set. By comparing the latency-accuracy surface of the test cases, implement three-stage optimization.
[0020] A RAG-based table data processing system provided by the present invention includes:
[0021] A parsing and conversion module, responsible for performing enhancement processing, structuring processing, semantic processing, and standardization processing on the content data in the table to obtain the processed content data;
[0022] A retrieval and adaptation module, responsible for receiving the processed content data, establishing a hybrid index and a context joint index; obtaining the user's query content, converting it into a key-value combination, and performing vector retrieval and keyword retrieval;
[0023] A generation and integration module, responsible for performing dynamic template filling according to the results of the key-value combination, vector retrieval, and keyword retrieval, and instructing the data fine-tuning model to guide the generation of instructions.
[0024] The present invention realizes an end-to-end intelligent table processing pipeline through the deep coupling of multi-stage technical features. By means of file format adaptive parsing path selection and multiple processing of content data (metadata enhancement, flattening / nesting structuring, semantic conversion, and standardization), a unified feature representation of cross-modal table content is established. In combination with a hybrid index architecture (fine-grained cell index and graph relationship storage) and a context joint index mechanism, a unified index space with multi-dimensional retrieval capabilities is constructed. Based on the collaborative optimization of dynamic query conversion (key-value combination / vector mapping / keyword matching) and generation components (template engine / lightweight calculation / fine-tuning model), an accurate mapping from query intent to structured results is achieved. Relying on the feedback adjustment mechanism of latency-accuracy, the index granularity is continuously optimized, and finally a closed-loop processing system with adaptive capabilities is formed, overcoming the key technical bottlenecks such as semantic discontinuity, low retrieval efficiency, and insufficient generation usability existing in traditional methods when processing multi-source heterogeneous table data.
[0025] Other features and advantages of the present invention will be described in the following specification, and part of them will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification and the drawings.
[0026] The technical solutions of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings
[0027] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0028] Figure 1 It is a flowchart of the method for processing table data based on RAG in Embodiment 1 of the present invention;
[0029] Figure 2 It is a schematic diagram of the method for processing table data based on RAG in Embodiment 1 of the present invention;
[0030] Figure 3 It is a process diagram of obtaining processed content data in Embodiment 2 of the present invention;
[0031] Figure 4 It is a process diagram of processing table data with hybrid index and joint retrieval in Embodiment 5 of the present invention;
[0032] Figure 5 It is a process diagram of adjusting the hybrid index and context joint index into rows, columns, or sub-tables in Embodiment 9 of the present invention;
[0033] Figure 6This is the block diagram of the RAG-based tabular data processing system in Embodiment 10 of the present invention. Detailed implementation mode
[0034] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.
[0035] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the embodiments of the present application. The singular forms "a", "the" and "said" used in the embodiments of the present application are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0036] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0037] Embodiment 1: As Figure 1 shown, the embodiment of the present invention provides a method for processing tabular data based on RAG, which includes the following steps:
[0038] S100: Obtain a file with structured characteristics and its electronic file format, and obtain the table parsing path in the file from the pre-stored parsing program according to the electronic file format; perform enhancement processing, structuring processing, semantic processing, and standardization processing on the content data in the table to obtain the processed content data; where the content data includes table metadata, table headers, each row of data, digital dates, etc.;
[0039] S200: Receive the processed content data, establish a hybrid index suitable for different table structures of sparse tables and multi-level tables and a context joint index of the content data in the table and the merged index of adjacent files; obtain the user query content, convert it into a key-value combination, and perform vector retrieval and keyword retrieval;
[0040] S300: Based on the results of key-value combination, vector retrieval, and keyword retrieval, perform dynamic template filling, and use the instruction data fine-tuning model to guide the generation of instructions while generating lightweight calculation results; according to the test cases, adjust the hybrid index and context joint index to rows, columns, or sub-tables in combination with query latency and accuracy.
[0041] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, first, a file with a structured feature and its electronic file format (file extension) are obtained, and the table parsing path in the file is obtained from the pre-stored parsing program according to the electronic file format; the content data in the table is subjected to enhancement processing, structuring processing, semantic processing, and standardization processing to obtain the processed content data; the content data includes the metadata of the table, the table header, each row of data, digital dates, etc.; secondly, the processed content data is received, and a hybrid index suitable for different table structures of sparse tables and multi-level tables and a context joint index of the content data in the table and the merged index of adjacent files are established; the user's query content is obtained, converted into a key-value combination, and vector retrieval and keyword retrieval are performed; finally, based on the results of key-value combination, vector retrieval, and keyword retrieval, dynamic template filling is performed, and the instruction data fine-tuning model is used to guide the generation of instructions while generating lightweight calculation results; according to the test cases, in combination with query latency and accuracy, adjust the hybrid index and context joint index to rows, columns, or sub-tables (for the specific principle, please refer to the appendix Figure 2 ). The above solution realizes an end-to-end intelligent table processing pipeline through the deep coupling of multi-stage technical features. Through the file format-adaptive parsing path selection and multiple processing of content data (metadata enhancement, flattening / nesting structuring, semantic conversion, and standardization), a unified feature representation of cross-modal table content is established; combined with the hybrid index architecture (fine-grained cell index and graph relationship storage) and context joint index mechanism, a unified index space with multi-dimensional retrieval capabilities is constructed; based on the collaborative optimization of dynamic query conversion (key-value combination / vector mapping / keyword matching) and generation components (template engine / lightweight calculation / fine-tuning model), the accurate mapping from query intent to structured results is realized; relying on the feedback adjustment mechanism of latency-accuracy, the index granularity is continuously optimized, and finally a closed-loop processing system with adaptive capabilities is formed, overcoming the key technical bottlenecks such as semantic discontinuity, low retrieval efficiency, and insufficient generation usability existing in traditional methods when processing multi-source heterogeneous table data.
[0042] The present invention is mainly a data processing strategy and method designed for tabular data with structural characteristics in word, excel, and pdf files in the RAG (Retrieval-Augmented Generation) project. The core idea is to retrieve relevant information from the knowledge base to assist the generation model to produce more accurate and fact-based text outputs, which can be applied to scenarios such as open-domain question answering, dialogue systems, content generation, etc., transforming tabular data from "static storage" into "dynamic knowledge assets". At the same time, it has generality and can be flexibly extended to related fields of large language models (LLM: Large Language Model).
[0043] Embodiment 2: As Figure 3 shown, on the basis of Embodiment 1, the process of obtaining the processed content data provided by the embodiment of the present invention includes the following steps:
[0044] S101: Create a repository for storing parsing programs, where the parsing programs include electronic file formats and corresponding parsing methods, as well as files with structural characteristics that trigger the parsing programs;
[0045] S102: When a file with structural characteristics is obtained, issue an instruction to trigger the parsing program, and sequentially confirm the electronic file format and parsing method according to the recognition process of the parsing program; when the electronic file format is confirmed, parse it according to the corresponding interpretation method; taking the electronic file format as the starting node and the parsing method as the ending node, form the table parsing path in the file;
[0046] S103: Obtain content data such as metadata, table headers, data in each row, and digital dates in the table, and perform enhancement processing, structuring processing, semantic processing, and standardization processing according to different content data.
[0047] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, a repository for storing parsing programs is first created. The parsing program includes electronic file formats and corresponding parsing methods, as well as files with structural characteristics that trigger the parsing program. Secondly, when a file with structural characteristics is obtained, an instruction to trigger the parsing program is issued, and the electronic file format and parsing method are confirmed in sequence according to the recognition process of the parsing program. When the electronic file format is confirmed, parsing is performed according to the corresponding interpretation method. Taking the electronic file format as the starting node and the parsing method as the end node, a table parsing path in the file is formed. Finally, content data such as metadata, table headers, data in each row, and digital dates in the table are obtained, and enhanced processing, structuring processing, semantic processing, and standardization processing are performed according to different content data. The above solution constructs a format-independent unified parsing entry through the repository preset and dynamic trigger mechanism, uses the file structure feature as the parsing routing key, and realizes the end-to-end automatic extraction from the original file to the content entity (table / metadata, etc.). The graph structure modeling of the parsing path ensures that the format recognition (starting node) and the parsing method (end node) form a deterministic mapping relationship, eliminating the redundant calculation of traditional conditional branch parsing. The intermediate representation generated based on the table parsing path drives the content extraction engine to realize the step-by-step decoupling from the physical storage format to the logical data unit (row / column / field). Integrating the four-stage processing of enhancement / structuring / semantic / standardization forms a data quality improvement pipeline, making the original unstructured data reach a machine-readable normalized state. Through the explicit construction of the parsing path and the multi-dimensional processing of content data, a complete audit chain including the original format characteristics, parsing decision logs, and data conversion history is formed. The intermediate results are passed between each step through a unified data bus to ensure data consistency and traceability during the processing process.
[0048] In this embodiment, the processing process of table parsing and its content data includes the following steps:
[0049] File parsing (parsing methods for different types of files):
[0050] Table parsing in PDF files (using pdfplumber), traverse each page of the PDF file, detect the table area in the page (based on visual features such as lines, spacing, and text alignment); extract the text content of the cells, and generate a logical table structure (table.cells) according to the table structure (row and column spacing, borders).
[0051] Merged cell processing: Identify physically merged cells (merged horizontally or vertically), and the cell range can be automatically restored through page.extract_table(merge_cells = True).
[0052] If a fault is automatically recognized (such as a borderless merging unit), page.rects can be used to calculate the overlapping area and manually reconstruct the table skeleton;
[0053] Multi-level table header processing flow (such as business reports): Detect the table header levels (such as the upper-level title spanning multiple columns and the lower-level table headers being subdivided); Reconstruct the flattened structure (such as converting "Year → Quarter" to "2023_Q1").
[0054] Excel / CSV parsing (using pandas):
[0055] Reading structured tables: For Excel files, use pd.read_excel(), and for CSV files, use pd.read_csv(). Custom delimiters (such as \t or ;) and encoding detection (such as UTF-8 / GBK) are supported;
[0056] Merged cell processing: Parse the merged areas in Excel (merged_cells attribute) and automatically fill in duplicate values (ffill()); Multi-level table header processing: Identify multi-level column indexes (MultiIndex) and reorganize them into single-level column names with hierarchical delimiters (such as "Sales.2023.Q1").
[0057] Metadata enhancement (extracting table context):
[0058] Extract structured information to enhance table retrieval and understanding, including: Basic metadata, file name, page number, and the chapter where the table is located (PDF bookmarks or chapter titles);
[0059] Table-specific metadata: Title detection, the text above the adjacent table (NLP matching such as "Table 1: " or regex / Table[\s\d]+ / i);
[0060] Source and time, footer, document properties (doc.info in PDF), or adjacent text (such as "Data source: National Bureau of Statistics").
[0061] Structured processing (table data conversion):
[0062] Simple tables (single-level table headers): Convert to key-value pair form (entity, attribute, value); Applicable scenarios: Database storage (such as RDF), inverted index in search engines (Elasticsearch);
[0063] Complex tables (nested multi-level table headers), recursively disassemble the levels to generate a nested JSON structure; Applicable scenarios: Time series data analysis, BI tools (such as Power BI hierarchical drill-down).
[0064] Semantic processing (natural language conversion):
[0065] Convert the table content into readable text to enhance semantic understanding:
[0066] Single-line description generation (for sparse tables): Template-based output "The specification of Product A is 25 ml and the sales volume is 100 bottles."; Multi-level table summary (for complex reports): Generate a summary in combination with the context, "In the first quarter of 2023, the production volume of Product A was 100 bottles; it increased to 120 bottles in the same period of 2024.";
[0067] Context enhancement technology:
[0068] Inject domain knowledge; Fine-tune the large language model (LLM) to generate coherent descriptions.
[0069] Data standardization (format unification and cleaning):
[0070] Date standardization: Parse mixed formats (such as "2023 / 01 / 01", "Jan-1-2024") → Unify to the ISO paradigm YYYY-MM-DD;
[0071] Number formatting: Unify the thousands separator (1,000 → 1000), scientific notation (1.2E3 → 1200); Unit normalization (such as "25 ml" → "25 milliliters");
[0072] Null value and noise processing: Remove meaningless placeholders (such as "-", "N / A");
[0073] Filling strategy: Mean value of adjacent rows (numerical values), most frequent item (categorical data);
[0074] Output structured paradigm: Adapt to the target system (such as SQL table, JSON-LD).
[0075] Example 3: On the basis of Example 2, the parsing program provided by the embodiment of the present invention further includes the processing process of merging cells and multi-level headers, including the following steps:
[0076] S1021: The table merges multiple adjacent cells into a large cell. By analyzing the lines or blank areas, identify the structural boundaries of the table, and judge whether there is a situation where the physical dividing line is missing; Check the coherence of the cell content. If the text of a certain cell spans multiple rows and columns, but there is no clear separation mark in the adjacent area, there may be an implicit merging relationship; Based on the positional relationship of the cells, establish a topological network, detect whether there is continuous blank or unreasonable row and column distribution, and infer potential merging areas; Combine the boundary detection and semantic analysis results to automatically merge relevant cells;
[0077] S1022: The data attributes of the multi-layer table header are distributed in a tree structure. Based on the differences in fonts and the offsets of relative spaces, the hierarchical relationship of the table header is judged. By recursively scanning the table headers at all levels, the nested title attributes are transformed into a single flat hierarchy. During the compression process, the original hierarchical relationship is retained, and a specific coding or logical indexing mechanism is adopted.
[0078] S1023: Dynamically adjust the parsing strategy according to the visual flow direction of the table content. When a cell spans multiple logical regions, comprehensively evaluate the content distribution and spatial weights, and select a partitioning scheme. Calculate the stress distribution of the internal data when the cell is split or merged, and formulate a structure association strategy.
[0079] The working principles and beneficial effects of the above technical solutions are as follows: In this embodiment, first, the table combines multiple adjacent cells into a large cell. By analyzing the lines or blank areas, the structural boundaries of the table are identified to determine whether there is a lack of physical dividing lines. Check the coherence of the cell content. If the text of a cell spans multiple rows and columns, but there is no clear separation mark in the adjacent area, there may be an implicit merging relationship. Establish a topological network based on the positional relationship of the cells, detect whether there is a continuous blank or unreasonable row and column distribution, and infer potential merging areas. Combine the boundary detection and semantic analysis results to automatically merge the relevant cells. Secondly, the data attributes of the multi-layer table header are distributed in a tree structure. Based on the differences in fonts and the offsets of relative spaces, the hierarchical relationship of the table header is judged. By recursively scanning the table headers at all levels, the nested title attributes are transformed into a single flat hierarchy. During the compression process, the original hierarchical relationship is retained, and a specific coding or logical indexing mechanism is adopted. Finally, dynamically adjust the parsing strategy according to the visual flow direction of the table content. When a cell spans multiple logical regions, comprehensively evaluate the content distribution and spatial weights, and select a partitioning scheme. Calculate the stress distribution of the internal data when the cell is split or merged, and formulate a structure association strategy. Through the collaborative analysis of multi-modal features and dynamic topology optimization, the above solutions achieve the precise parsing and dimension reconstruction of the composite structure table.
[0080] Embodiment 4: On the basis of Embodiment 2, the process of enhancing the metadata of the content data provided by the embodiment of the present invention includes the following steps:
[0081] S1031: Use a hierarchical pointer network to perform feature encoding on the table topology structure. Through discrete cosine transform, the visual layout features are transformed into a sequence of spatial vectors, and the projection relationship with the content domain is established. The positioning engine automatically identifies 16 visual parameters such as the font weighting coefficient and cell merging mode in the title area to generate a spatial fingerprint.
[0082] S1032: Develop a dynamic timestamp parser to extract the temporal information of spatial fingerprints from the following channels: the revision log watermark embedded in the document, the version hash value of the associated data stream, and the timeline alignment of the external knowledge graph; eliminate temporal conflicts through a multiple hypothesis testing algorithm, and finally generate a confidence-weighted time interval representation;
[0083] S1033: Construct a heterogeneous source recognition engine based on fuzzy string matching, including parsing the typographical features of footnote special characters, comparing the distribution probabilities of domain terms in a professional dictionary, and extracting the meta-tag patterns of document containers (such as PDF / HTML).
[0084] The working principle and beneficial effects of the above technical solutions are as follows: In this embodiment, a hierarchical pointer network is first used to perform feature encoding on the table topology structure, and the visual typographical features are transformed into a spatial vector sequence through discrete cosine transform to establish a projection relationship with the content domain; the positioning engine automatically identifies 16 visual parameters such as the font weighting coefficient and cell merging mode in the title area to generate spatial fingerprints; secondly, a dynamic timestamp parser is developed to extract the temporal information of spatial fingerprints from the following channels: the revision log watermark embedded in the document, the version hash value of the associated data stream, and the timeline alignment of the external knowledge graph; eliminate temporal conflicts through a multiple hypothesis testing algorithm, and finally generate a confidence-weighted time interval representation; finally, a heterogeneous source recognition engine based on fuzzy string matching is constructed, including parsing the typographical features of footnote special characters, comparing the distribution probabilities of domain terms in a professional dictionary, and extracting the meta-tag patterns of document containers (such as PDF / HTML). The above solution constructs a dynamically adaptive metadata enhancement architecture through systematic integration of spatial feature encoding, temporal parsing engine, and cross-domain traceability mechanism; the spatial fingerprint and the time interval representation form a spatio-temporal joint vector through tensor splicing, in which the 16-dimensional visual parameters such as the font weighting coefficient are subjected to Hadamard product operation with the confidence weights output by the timestamp parser to generate a feature kernel with discrete-continuous hybrid characteristics; the feature kernel is orthogonally processed through the domain term probability distribution matrix of the heterogeneous source recognition engine, and finally outputs a triple verification label carrying topological structure, version evolution, and source credibility. In the query stage, the system automatically activates the cross-modal attention mechanism of the feature kernel: the spatial vector sequence triggers approximate matching based on the Manhattan distance, the time interval representation starts the dynamic alignment of the sliding window, and the source feature Bloom filter performs fast pre-screening. The gradients of the three types of operation results are dynamically aggregated through a gated recurrent unit, so that the retrieval process simultaneously satisfies three constraint conditions of structural similarity, temporal continuity, and source reliability. Robustness is maintained by introducing a feature drift detection module: when it is detected that the meta-tag pattern of the document container deviates from the historical record by more than 3σ, the re-calibration process of the spatio-temporal joint vector will be automatically triggered. The re-calibration process includes: version backtracking based on the revision log watermark, updating the frequency domain features by calling discrete cosine transform, and recalculating the distribution probabilities of domain terms.
[0085] Example 5: As Figure 4 shown, based on Example 1, the process of processing tabular data for hybrid indexing and joint retrieval provided by the embodiments of the present invention includes the following steps:
[0086] S201: For a structured table with a single-level table header, perform cell-level index decomposition to map each row or key cell into an independent index unit. Each index unit records its numbers, dates, texts, and the row numbers and column numbers in the table, and maintains a mapping relationship with the overall table structure;
[0087] S202: Convert the complete table structure into a structured data format with nested relationships, and extract the data level associations therein; record the subordination relationship between the table header cells and the data cells at the index layer, and at the same time mark the dependence path of the underlying data on the high-level semantics through reverse pointers;
[0088] S203: Identify adjacent texts such as explanatory paragraphs and footnotes associated with the table, and extract key information such as term explanations and unit definitions; bind the relevant semantic information in the adjacent text to the index record in units of the cells or row data of the table to create a two-way index association.
[0089] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, first, for a structured table with a single-level table header, perform cell-level index decomposition to map each row or key cell into an independent index unit. Each index unit records its numbers, dates, texts, and the row numbers and column numbers in the table, and maintains a mapping relationship with the overall table structure; second, convert the complete table structure into a structured data format with nested relationships, and extract the data level associations therein; record the subordination relationship between the table header cells and the data cells at the index layer, and at the same time mark the dependence path of the underlying data on the high-level semantics through reverse pointers; finally, identify adjacent texts such as explanatory paragraphs and footnotes associated with the table, and extract key information such as term explanations and unit definitions; bind the relevant semantic information in the adjacent text to the index record in units of the cells or row data of the table to create a two-way index association. The process of processing tabular data for hybrid indexing and joint retrieval in the above solution essentially constructs a multi-level and strongly semantically associated dynamic index system, transforming discrete technical components (coordinate indexing, hierarchical modeling, semantic binding) into a mutually feedback enhanced loop, so that any single-point improvement (such as more refined adjacent text parsing) can synergistically act on the overall retrieval efficiency improvement within the system.
[0090] Example 6: Based on Example 5, the process of creating a two-way index association provided by the embodiments of the present invention includes the following steps:
[0091] S2031: Based on the generated hierarchical structured data, identify the cells in the table that play a semantic hub role (such as the headers spanning rows and columns, cells with statistical indicator names); through the dependency paths marked by reverse pointers, determine the high-level semantic scope to which each data unit needs to be bound (such as the monthly data columns subordinate to the "annual revenue" cell); after processing the adjacent text, complete the normalized marking of terms and units;
[0092] S2032: Automatically divide the semantic influence domain according to the table layout characteristics. When the column relationship is mainly vertical, bind the footnote interpretation information within 3 rows directly below the header to the corresponding column; for horizontally multi-page tables, capture the continuous explanatory text before and after the page break; adopt a dynamic sliding window mechanism to make the cell position coordinates inversely weighted with the physical distance of the text block;
[0093] S2033: Establish a synonymous hash between the original content of the cell and the canonical expression in the structured index; convert the unit definitions extracted from the adjacent text into constraint condition expressions and attach them to the metadata layer of the corresponding data unit; generate a semantic tree through the header subordination relationship, enabling the footnote information to drill down along the tree path to the end cell, forming a two-way addressing ability for cells, standard terms, and explanatory texts.
[0094] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, first, based on the generated hierarchical structured data, identify the cells in the table that play a semantic hub role (such as the headers spanning rows and columns, cells with statistical indicator names); through the dependency paths marked by reverse pointers, determine the high-level semantic scope to which each data unit needs to be bound (such as the monthly data columns subordinate to the "annual revenue" cell); after processing the adjacent text, complete the normalized marking of terms and units; second, automatically divide the semantic influence domain according to the table layout characteristics. When the column relationship is mainly vertical, bind the footnote interpretation information within 3 rows directly below the header to the corresponding column; for horizontally multi-page tables, capture the continuous explanatory text before and after the page break; adopt a dynamic sliding window mechanism to make the cell position coordinates inversely weighted with the physical distance of the text block; finally, establish a synonymous hash between the original content of the cell and the canonical expression in the structured index; convert the unit definitions extracted from the adjacent text into constraint condition expressions and attach them to the metadata layer of the corresponding data unit; generate a semantic tree through the header subordination relationship, enabling the footnote information to drill down along the tree path to the end cell, forming a two-way addressing ability for cells, standard terms, and explanatory texts. The above solution constructs a three-order verification chain of geometric position association, semantic category matching, and logical constraint effectiveness, enabling a reversible closed-loop mapping between the physical structure of the table data and the business context logic. Among them, the cell row and column coordinates serve as the underlying positioning benchmark, the structured hierarchical relationship undertakes the semantic relay function, and the adjacent text provides dynamic verification rules; through progressive conditional constraints, coordinated operation is achieved, and finally, the two-way traceability of the index record and the interpretation system is achieved.
[0095] Embodiment 7: On the basis of Embodiment 1, the process of performing vector retrieval and keyword retrieval provided by the embodiment of the present invention includes the following steps:
[0096] S204: Key-value combination parsing logic. In the face of a free-text query input by the user, identify the core entities and attribute requirements in the free-text query, and convert them into structured key-value pairs; automatically construct a query logic tree based on the key-value combination, and inject it into the index system to perform directional retrieval;
[0097] S205: If a fuzzy term is detected in the key-value parsing stage, activate the context joint index, infer the corresponding time interval by retrieving adjacent text, supplement the semantics to the query condition, and generate an extended key-value combination;
[0098] S206: Use the composite index structure of the hybrid index and the context joint index to perform parallel retrieval, directly call the point query mechanism of the fine-grained index, and return the matching cell or row data; through the hierarchical traversal of the nested index, locate the subtree structure that meets the conditions.
[0099] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, first, the key-value combination parsing logic is used. In the face of a free-text query input by the user, identify the core entities and attribute requirements in the free-text query, and convert them into structured key-value pairs; automatically construct a query logic tree based on the key-value combination, and inject it into the index system to perform directional retrieval; secondly, if a fuzzy term is detected in the key-value parsing stage, activate the context joint index, infer the corresponding time interval by retrieving adjacent text, supplement the semantics to the query condition, and generate an extended key-value combination; finally, use the composite index structure of the hybrid index and the context joint index to perform parallel retrieval, directly call the point query mechanism of the fine-grained index, and return the matching cell or row data; through the hierarchical traversal of the nested index, locate the subtree structure that meets the conditions. Through the coupled design of the hybrid index and the joint retrieval, the above solution realizes that whether the original table is sparse (single-level header) or complex (multi-level nesting), efficient retrieval can be completed through a unified interface; the context binding mechanism ensures that the retrieval result does not deviate from its original semantic environment (such as avoiding digital misunderstandings that are divorced from the unit explanation); the index strategy self-optimizes according to the query mode, avoiding local hotspots or resource waste caused by a fixed architecture.
[0100] Embodiment 8: On the basis of Embodiment 7, the process of generating an extended key-value combination provided by the embodiment of the present invention includes the following steps:
[0101] S2051: Detect the ambiguous expressions contained in the attribute requirement field based on the key-value pair structure in the parsed free text query, locate the corresponding logical node in the indexing system; extract the associated semantic fragments of this node in the adjacent text according to the activated context joint index to form a candidate supplementary information pool; rely on the entity type tags in the structured key-value pairs as filtering conditions to exclude unmatched context fragments;
[0102] S2052: Perform semantic quantization conversion on the identified ambiguous terms. Temporal expressions are mapped to a closed interval value range by retrieving the reference year in the context; degree expressions are converted into numerical threshold conditions in combination with the statistical descriptions of the adjacent text; utilize the nested hierarchical structure of the hybrid index, trace back along the subordination path of the table header - data to calibrate the consistency of magnitude units;
[0103] S2053: Evolve the original key-value pairs according to the triple rule, replace the direct fuzzy terms with specific intervals; attach derivative constraints to the original exact key-values; through the subtree location mechanism, inject the supplementary conditions into the associated sub-index domain to generate new key-value combinations with complete context constraints.
[0104] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, first, based on the key-value pair structure in the parsed free text query, detect the ambiguous expressions contained in the attribute requirement field, locate the corresponding logical node in the indexing system; extract the associated semantic fragments of this node in the adjacent text according to the activated context joint index to form a candidate supplementary information pool; rely on the entity type tags in the structured key-value pairs as filtering conditions to exclude unmatched context fragments; secondly, perform semantic quantization conversion on the identified ambiguous terms. Temporal expressions are mapped to a closed interval value range by retrieving the reference year in the context; degree expressions are converted into numerical threshold conditions in combination with the statistical descriptions of the adjacent text; utilize the nested hierarchical structure of the hybrid index, trace back along the subordination path of the table header - data to calibrate the consistency of magnitude units; finally, evolve the original key-value pairs according to the triple rule, replace the direct fuzzy terms with specific intervals; attach derivative constraints to the original exact key-values; through the subtree location mechanism, inject the supplementary conditions into the associated sub-index domain to generate new key-value combinations with complete context constraints. The above solution forms a chain process of fuzzy recognition, context extraction, quantization conversion, and condition fusion, enabling the extended key-value combination to not only inherit the original query intention but also integrate the context information of the table itself; among them, the structured key-value pairs serve as the control framework, the context joint index provides semantic filling materials, and the hierarchical characteristics of the hybrid index undertake the function of ensuring logical integrity.
[0105] Example 9: As Figure 5 shown, on the basis of Example 1, the process of adjusting the hybrid index and context joint index to rows, columns, or sub-tables provided by the embodiment of the present invention includes the following steps:
[0106] S301: Based on the combined feedback results of vector retrieval and keyword retrieval, perform quantitative analysis on the distribution of blank areas in the sparse table; when it is detected that the proportion of consecutive vertical blank cells exceeds the threshold, trigger the column index mode, and at this time, establish a horizontal combined index between the data types of each column and the table header semantics of adjacent files; if the horizontal blank cells are dense, switch to the row index mode, and use the semantic correlation degree between the key values of this row and the data of the upper and lower rows as the weight factor; the blank area consists of multiple blank cells;
[0107] S302: Extract the topological features of the sub-table according to the header hierarchy relationship after step processing; when the user query key-value combination involves multi-level associations, automatically decompose the original single-level index into a tree-shaped sub-table index network: the upper-level index node stores the sub-table position pointer, and the lower-level node retains the statistical features of the numerical range of specific cells;
[0108] S303: Use the generated lightweight calculation results as the validation data set, and implement three-stage optimization by comparing the latency-accuracy surface of the test cases:
[0109] First-stage optimization: For the index units involved in high-frequency queries (such as the "net profit" column of the financial statement), increase the weight of their vector retrieval dimension;
[0110] Second-stage optimization: When there are cross-file reference conflicts in the context combined index (such as differences in measurement units of the same table header in different files), trigger the standardization compensation mechanism and append unit conversion coefficient anchor points at the index layer;
[0111] Third-stage optimization: Based on the heat map distribution of the query results, perform compression and reorganization on the sub-table indexes below the response threshold, and reduce them to Boolean combination expressions of row / column indexes.
[0112] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, based on the combined feedback results of vector retrieval and keyword retrieval, a quantitative analysis is first performed on the blank area distribution of the sparse table. When it is detected that the proportion of consecutive vertical blank cells exceeds the threshold, the column index mode is triggered. At this time, a horizontal combined index is established between the data types of each column and the semantic meaning of the table headers of adjacent files. If the horizontal blank cells are dense, the row index mode is switched, and the semantic correlation degree between the key values of this row and the data of the upper and lower rows is used as the weight factor. Secondly, according to the hierarchical relationship of the table headers after step processing, the topological features of the sub-tables are extracted. When the user's query key-value combination involves multi-level associations, the original single-level index is automatically disassembled into a tree-shaped sub-table index network: the upper-level index node stores the position pointer of the sub-table, and the lower-level node retains the statistical features of the numerical range of specific cells. Finally, the generated lightweight calculation results are used as the verification data set. By comparing the latency-accuracy surface of the test cases, three-stage optimization is implemented: First-stage optimization: For the index units involved in high-frequency queries (such as the "net profit" column in the financial statement), the vector retrieval dimension weight is increased. Second-stage optimization: When there are cross-file reference conflicts in the context combined index (such as differences in measurement units of the same table header in different files), the standardization compensation mechanism is triggered, and unit conversion coefficient anchor points are added at the index layer. Third-stage optimization: Based on the heat map distribution of the query results, the sub-table indexes below the response threshold are compressed and reorganized, and they are reduced to boolean combination expressions of row / column indexes. The dynamic selection of the index granularity of the above solution directly determines the organization form of the combined index, and the real-time feedback data in turn restricts the strategy of index fusion. The entire process is enhanced progressively through quantitative indicators (such as blank cell threshold, response latency, etc.), meeting the core requirement of improving data processing efficiency through structural innovation.
[0113] Embodiment 10: As Figure 6 shown, based on Embodiments 1 - 9, the RAG-based table data processing system provided by an embodiment of the present invention includes:
[0114] A parsing and conversion module, responsible for obtaining a file with structural characteristics and its electronic file format, obtaining the table parsing path in the file from the pre-stored parsing program according to the electronic file format; performing enhancement processing, structuring processing, semantic processing, and standardization processing on the content data in the table to obtain the processed content data; where the content data includes metadata of the table, table headers, data of each row, and digital dates, etc.
[0115] A retrieval and adaptation module, responsible for receiving the processed content data, establishing a hybrid index suitable for different table structures of sparse tables and multi-level tables and a context combined index of the content data in the table and the merged index of adjacent files; obtaining the user's query content, converting it into a key-value combination, and performing vector retrieval and keyword retrieval.
[0116] The generation and integration module is responsible for performing dynamic template filling based on the results of key-value combination, vector retrieval, and keyword retrieval, guiding the generation of instructions through fine-tuning of instruction data for the model, and simultaneously generating lightweight calculation results; adjusting the hybrid index and context joint index to rows, columns, or sub-tables according to test cases, in combination with query latency and accuracy.
[0117] The working principle and beneficial effects of the above technical solution are as follows: The parsing and conversion module of this embodiment obtains a file with a structured feature and its electronic file format, and obtains the table parsing path in the file from the pre-stored parsing program according to the electronic file format; performs enhancement processing, structuring processing, semantic processing, and standardization processing on the content data in the table to obtain the processed content data; where the content data includes table metadata, table headers, data for each row, and digital dates, etc.; the retrieval and adaptation module receives the processed content data, establishes a hybrid index suitable for different table structures of sparse tables and multi-level tables, and a context joint index for merging the content data in the table with the adjacent file index; obtains the user's query content, converts it into a key-value combination, and performs vector retrieval and keyword retrieval; the generation and integration module performs dynamic template filling based on the results of key-value combination, vector retrieval, and keyword retrieval, guides the generation of instructions through fine-tuning of instruction data for the model, and simultaneously generates lightweight calculation results; adjusts the hybrid index and context joint index to rows, columns, or sub-tables according to test cases, in combination with query latency and accuracy. The above-mentioned RAG-based table data processing system constructs a complete table data processing pipeline through the collaborative operation of the parsing and conversion module, the retrieval and adaptation module, and the generation and integration module. The parsing and conversion module realizes the unified structured representation of cross-format table data, providing a standardized input for downstream processing; the retrieval and adaptation module establishes a hybrid index system based on structured data, realizing unified support for multi-modal query methods; the generation and integration module completes the interpretable conversion of query results through templated recombination of retrieval results and lightweight calculation. This embodiment establishes a mapping channel from the original table to the final result through feature transformation and data processing in each link; through the combined use of hybrid index and dynamic template, it solves the contradiction between query flexibility and result accuracy in traditional table processing methods; through the index adjustment mechanism based on performance feedback, it realizes the adaptive balance between processing quality and response efficiency of the system, and finally forms an intelligent closed-loop system covering the entire life cycle of data processing.
[0118] The application scenarios of this embodiment include: finance and business analysis, healthcare, scientific research and academic research, enterprise operation and supply chain, customer service and sales, and law and compliance, etc.;
[0119] Among them, financial and business analysis includes financial report analysis and risk monitoring; financial report analysis includes automatically extracting tabular data such as revenue, cost, and profit from enterprise financial reports; risk monitoring includes real-time retrieving loan application forms, assessing risks, and generating warning prompts.
[0120] Medical and health includes patient medical record management and scientific research data analysis; patient medical record management includes analyzing the patient's condition by combining electronic medical record forms (examination indicators, medication records); scientific research data analysis includes extracting the relationship between drug dosage and efficacy from clinical trial forms and generating statistical conclusions.
[0121] Scientific research and academic research includes experimental data collation and literature review support; experimental data collation includes analyzing experimental record forms (temperature, pressure, reaction time); literature review support includes aggregating data across paper forms and generating trend analysis of the field.
[0122] Enterprise operation and supply chain includes inventory management and supply chain optimization; inventory management includes generating replenishment suggestions based on inventory forms (SKU, quantity, shelf life); supply chain optimization includes analyzing logistics cost forms and recommending the optimal supplier combination.
[0123] Customer service and sales includes CRM data query and sales report generation; CRM data query includes quickly retrieving customer information tables and answering customer questions; sales report generation includes automatically generating a visual summary report based on sales performance tables.
[0124] Law and compliance includes contract clause retrieval and regulation comparison; contract clause retrieval includes locating clauses related to "liability for breach of contract" from contract clause forms and generating a compliance risk summary; regulation comparison includes analyzing forms in legal documents of multiple countries and outputting a list of differences.
[0125] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of equivalent technologies of the present invention, the present invention also intends to include these changes and modifications.
Claims
1. A method for processing tabular data based on RAG, characterized in that It includes the following steps: Obtain the user's query content, convert it into key-value combinations, perform vector retrieval and keyword retrieval; according to the results of the key-value combinations, vector retrieval and keyword retrieval, perform dynamic template filling, and use the instruction data fine-tuning model to guide the generation of instructions.
2. The RAG-based tabular data processing method according to claim 1, wherein Perform enhancement processing, structuring processing, semantic processing and standardization processing on the content data in the table to obtain the processed content data.
3. The method for processing tabular data based on RAG according to claim 2, wherein The process of obtaining the processed content data includes the following steps: Create a repository for storing parsing programs; When a file with structural characteristics is obtained, issue an instruction to trigger the parsing program, and confirm the electronic file format and parsing method in sequence according to the recognition process of the parsing program.
4. The RAG-based tabular data processing method according to claim 3, wherein When the electronic file format is confirmed, perform parsing according to the corresponding interpretation method; using the electronic file format as the starting node and the parsing method as the end node, form the table parsing path in the file.
5. The method for processing tabular data based on RAG according to claim 4, wherein, The electronic file format is the file extension, and the files include word, excel and pdf.
6. The method for processing tabular data based on RAG according to claim 3, wherein After forming the table parsing path in the file, obtain the metadata, table headers, each row of data and digital date content data in the table, and perform enhancement processing, structuring processing, semantic processing and standardization processing according to different content data.
7. The method for processing tabular data based on RAG according to claim 3, wherein The parsing program includes the electronic file format and the corresponding parsing method, as well as the file with structural characteristics that triggers the parsing program.
8. The RAG-based tabular data processing method according to claim 1, wherein, According to the test cases, combined with the query latency and accuracy, adjust the hybrid index and the context joint index to rows, columns or sub-tables.
9. The method for processing tabular data based on RAG according to claim 8, wherein The process of adjusting the hybrid index and the context joint index to rows, columns or sub-tables includes the following steps: Based on the combined feedback results of vector retrieval and keyword retrieval, perform quantitative analysis on the blank area distribution of the sparse table; according to the header hierarchy relationship after step processing, extract the topological features of the sub-table; When the user's query key-value combination involves multi-level associations, automatically disassemble the original single-level index into a tree-like sub-table index network, use the generated lightweight calculation results as the verification data set, and implement three-stage optimization by comparing the latency-accuracy surface of the test cases.
10. A RAG-based tabular data processing system, characterized in that, It includes: A parsing and conversion module, responsible for performing enhancement processing, structuring processing, semantic processing and standardization processing on the content data in the table to obtain the processed content data; A retrieval adaptation module, responsible for receiving the processed content data, establishing a hybrid index and a context joint index; obtaining the user's query content, converting it into key-value combinations, and performing vector retrieval and keyword retrieval; A generation and integration module, responsible for performing dynamic template filling according to the results of the key-value combinations, vector retrieval and keyword retrieval, and using the instruction data fine-tuning model to guide the generation of instructions.
Citation Information
Patent Citations
Table data processing method and device
CN118567540B
Table data processing method and device, equipment and storage medium
CN118627486A
Table data processing method and device, equipment and storage medium
CN118780253A
Class case retrieval system and method based on retrieval enhancement generation technology
CN118260391A
Retrieval method suitable for PDF and Excel coexistence in RAG scene
CN118643053A
Cited By
Large-scale unstructured data joint processing method and system
CN120892608A
Customs risk early warning method and system based on cross-modal retrieval enhancement
CN120975681A
Table query method and device
CN121579471A
Web application automatic generation and data processing method based on spreadsheet mapping
CN121579812A
Web application automatic generation and data processing method based on spreadsheet mapping
CN121579812B