Data query method and apparatus
Patent Information
- Application Number
- CN202610912017.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-29
AI Technical Summary
这类文档通常篇幅冗长、格式复杂,关键信息以自然语言形式分散呈现,难以通过传统查询手段高效获取
在本技术方案中,在接收到用户查询语句后,首先对其进行意图解析以获得查询意图,并根据该查询意图确定查询复杂度及所需的证据依赖结构,从而避免了对所有数据库进行无差别检索;进一步地,基于查询复杂度与证据依赖结构,有选择性地调用向量数据库、图数据库和倒排索引数据库中的一个或多个目标数据库执行检索,使得系统能够根据查询的实际需求精准匹配最合适的检索机制;进而,系统依据证据依赖结构对多源检索结果进行结构化推理,而非简单拼接或排序,从而生成逻辑连贯、事实可追溯的答案,实现了在保障答案准确性与可解释性的同时,降低不必要的计算开销,提升了整体查询效率;其中,在确定目标数据库时,基于证据依赖结构识别所需证据项及其逻辑关联,结合各证据项的数据特征匹配数据库类型,并依据查询复杂度筛选高效且必要的目标库;同时,根据证据间的依赖关系动态规划数据库的调用顺序与协同方式,不仅提升多源异构金融文档中关键信息的召回准确率与上下文完整性,避免全库扫描带来的计算冗余,还通过依赖驱动的执行顺序保障多跳推理的逻辑严谨性,增强系统在专业金融场景下的实用性与可靠性;此外,在构建多个数据库时,通过对富文本文档进行结构化解析,将其转化为包含类型与位置信息的布局元素,依元素类型适配提取策略生成关联文本块并确定语义层级,进而构建反映原文逻辑的文档树;基于该文档树同步构建图数据库、倒排索引数据库和向量数据库,提升了金融智能问答在准确性、可解释性与多维检索协同方面的性能。
Smart Images

Figure CN122838477A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a data query method and apparatus. Background Technology
[0002] In the modern financial system, the timeliness, accuracy, and relevance of information are crucial for risk control, investment decisions, compliance, and customer service. Financial institutions rely on a wide range of data sources, including structured transaction and account data, as well as a large number of unstructured or semi-structured rich text documents such as prospectuses, audit reports, regulatory letters, research reports, and corporate announcements. These documents are typically lengthy and complex in format, with key information presented in a scattered manner using natural language, making them difficult to retrieve efficiently using traditional search methods.
[0003] Therefore, how to achieve efficient, accurate and interpretable data querying in massive heterogeneous financial data has become the key to supporting core applications such as intelligent investment research, intelligent risk control, and compliance management. Summary of the Invention
[0004] This disclosure provides a data query method and apparatus to at least partially solve one of the technical problems in the related art. The technical solution of this disclosure is as follows:
[0005] According to a first aspect of the present disclosure, a data query method is provided, comprising: in response to receiving a user query statement, performing intent parsing on the query statement to obtain a query intent; determining query complexity and the evidence dependency structure required for the query based on the query intent; determining at least one target database to be queried from a plurality of databases based on the query complexity and the evidence dependency structure, and performing a retrieval operation on the at least one target database; wherein the plurality of databases includes a vector database for semantic retrieval, a graph database for structural navigation, and an inverted index database for full-text content retrieval; and performing an inference operation corresponding to the evidence dependency structure based on the retrieval results returned from the at least one target database to generate an answer in response to the user query statement.
[0006] According to a second aspect of the present disclosure, a data query apparatus is provided, comprising: a parsing module, configured to, in response to receiving a user query statement, perform intent parsing on the query statement to obtain a query intent; a first determining module, configured to, based on the query intent, determine query complexity and the evidence dependency structure required for the query; a second determining module, configured to, based on the query complexity and the evidence dependency structure, determine at least one target database to be queried from a plurality of databases, and perform a retrieval operation on the at least one target database; wherein the plurality of databases includes a vector database for semantic retrieval, a graph database for structural navigation, and an inverted index database for full-text content retrieval; and a generating module, configured to, based on the retrieval results returned from the at least one target database, perform a reasoning operation corresponding to the evidence dependency structure to generate an answer in response to the user query statement.
[0007] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the data query method as described in the first aspect of the present disclosure.
[0008] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform a data query method as described in the first aspect of the present disclosure.
[0009] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising: a computer program that, when executed by a processor, implements the data query method as described in the first aspect of the present disclosure.
[0010] The technical solutions provided by the embodiments of this disclosure bring at least the following beneficial effects: In this technical solution, upon receiving a user query, the system first performs intent parsing to obtain the query intent, and then determines the query complexity and required evidence dependency structure based on this intent, thereby avoiding indiscriminate searching of all databases. Furthermore, based on the query complexity and evidence dependency structure, the system selectively calls one or more target databases from vector databases, graph databases, and inverted index databases to perform the search, enabling the system to accurately match the most suitable search mechanism according to the actual needs of the query. Subsequently, the system performs structured reasoning on the multi-source search results based on the evidence dependency structure, rather than simply concatenating or sorting them, thereby generating logically coherent and factually traceable answers. This achieves the goal of reducing unnecessary computational overhead and improving overall query efficiency while ensuring the accuracy and interpretability of the answers. Specifically, when determining the target database, the system identifies the required evidence items and their logical relationships based on the evidence dependency structure, combined with various... The system matches the data features of evidence items to database types and selects efficient and necessary target databases based on query complexity. Simultaneously, it dynamically plans the database call order and collaboration method according to the dependencies between evidence items. This not only improves the recall accuracy and contextual integrity of key information in multi-source heterogeneous financial documents, avoiding computational redundancy caused by full database scanning, but also ensures the logical rigor of multi-hop reasoning through dependency-driven execution order, enhancing the system's practicality and reliability in professional financial scenarios. Furthermore, when constructing multiple databases, it performs structured parsing of rich text documents, transforming them into layout elements containing type and location information. It then generates associated text blocks and determines semantic levels based on element type-adaptive extraction strategies, thereby constructing a document tree that reflects the logic of the original text. Based on this document tree, it simultaneously constructs a graph database, an inverted index database, and a vector database, improving the performance of financial intelligent question answering in terms of accuracy, interpretability, and multi-dimensional retrieval collaboration.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0013] Figure 1 This is a flowchart illustrating the data query method shown in the first embodiment of this disclosure; Figure 2 This is a flowchart illustrating the data query method shown in the second embodiment of this disclosure; Figure 3 This is a flowchart illustrating the data query method shown in the third embodiment of this disclosure; Figure 4This is a schematic flowchart of the data query method shown in the fourth embodiment of this disclosure; Figure 5 This is a flowchart illustrating the data query method shown in the fifth embodiment of this disclosure; Figure 6 This is a flowchart illustrating the data query method shown in the sixth embodiment of this disclosure; Figure 7 This is a schematic flowchart of the data query method shown in the seventh embodiment of this disclosure; Figure 8 This is a schematic diagram illustrating the principle of the data query method shown in the embodiments of this disclosure; Figure 9 This is a schematic diagram illustrating the principle of the data query method shown in the embodiments of this disclosure; Figure 10 This is a schematic diagram of the structure of the data query device shown in the eighth embodiment of this disclosure; Figure 11 This is a schematic diagram of the structure of an electronic device shown in an exemplary embodiment of the present disclosure. Detailed Implementation
[0014] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0015] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0016] It should be noted that the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution disclosed herein are all carried out with the consent of the user, and all comply with the provisions of relevant laws and regulations, and do not violate public order and good morals.
[0017] In the context of the rapid development of large language model applications, Retrieval-Augmented Generation (RAG) systems have become a core technological support for enterprise knowledge management, intelligent question answering, and document analysis. RAGs effectively alleviate the hallucination problem and improve the timeliness of knowledge by retrieving relevant document fragments from external knowledge bases and inputting them as context into a large language model to generate answers.
[0018] However, in practical applications of rich text documents for enterprise use, enterprise documents are usually long and complex in structure, containing various elements such as multi-level headings, nested lists, tables, images, formulas and cross-chapter references. Their information expression is highly dependent on the overall document layout and hierarchical logic. Most related technologies focus on plain text processing, lacking effective multimodal content understanding mechanisms when faced with rich text documents containing multimodal elements such as tables, images, and formulas. For example, semantic connections are often not established between tables and their corresponding titles and annotations, resulting in the inability to return structurally complete and context-rich information units during the retrieval process.
[0019] Meanwhile, related technologies struggle to accurately identify multi-level heading hierarchies within documents and fail to construct clear and consistent document tree structures. This results in the segmented text chunks lacking necessary hierarchical context information, making it difficult for users to determine the location, scope, and relative importance of search results within the original document.
[0020] Furthermore, current mainstream segmentation strategies typically prioritize segmentation efficiency or fixed-length constraints, neglecting the natural division of semantic boundaries, resulting in the fragmentation of key arguments or logical chains. Consequently, search results fail to reflect the structural context of the original document, hindering users' ability to quickly locate and accurately understand information.
[0021] More importantly, most existing RAG systems use a single embedding retrieval strategy, which cannot dynamically select the most suitable retrieval mechanism based on the complexity of the user's query intent or the type of evidence required (such as fact fragments, structural paths, or full-text context). This results in unstable retrieval performance in different scenarios, making it difficult to balance accuracy, recall, and interpretability.
[0022] To address any of the above problems, this disclosure proposes a data query method and apparatus.
[0023] The data query method and apparatus of this disclosure are described below with reference to the accompanying drawings.
[0024] Figure 1 This is a flowchart illustrating the data query method shown in the first embodiment of this disclosure.
[0025] It should be noted that this embodiment of the disclosure uses a data query method configured in a data query device as an example. This data query device can be applied to an electronic device to enable the electronic device to perform data query functions.
[0026] like Figure 1 As shown, the data query method includes: Step 110: In response to receiving a user query statement, perform intent parsing on the query statement to obtain the query intent.
[0027] As one possible implementation, when a user inputs a query in natural language, a large language model or a dedicated semantic understanding module can be invoked to parse the query intent and identify its purpose. This purpose can include information need type, target object, time range, and logical relationships. For example, when a user asks, "How much did a listed company's R&D expenses increase year-on-year in 2023?", the system needs to accurately extract the specific company, the year 2023, and the required information: the year-on-year change rate of R&D expenses. It must also determine that the question belongs to a numerical calculation query. By transforming unstructured statements into structured query intents, the system can accurately capture the user's true needs.
[0028] Step 120: Determine the query complexity and the evidence dependency structure required for the query based on the query intent.
[0029] Furthermore, based on the query intent, the processing difficulty and evidence organization logic of this query are determined. Query complexity is used to distinguish between single-point fact extraction and multi-hop reasoning tasks. For example, a problem requiring only one data item is considered low complexity, while problems involving cross-year comparisons, multi-indicator aggregation, or conditional filtering are considered high complexity. Simultaneously, the system constructs an evidence dependency structure. This structure is a formalized description of the required evidence items and their interrelationships, built by the system to generate an accurate, complete, and logically consistent answer to a user query. The evidence dependency structure not only specifies which specific facts (i.e., "evidence items") are needed but also clarifies how these facts are logically related, combined, or dependent, thus guiding the subsequent retrieval and reasoning process.
[0030] For example, when a user asks "What is the percentage of bond investment balance to total investment assets listed in the investment assets section of an insurance company's 2022 annual report?", the system first needs to locate the "Investment Asset Composition" or similar section in the "2022 Annual Report" (Evidence Item 1). Then, it extracts the two values, "Bond Investment Balance" and "Total Investment Assets," from the tables or text within that section (Evidence Item 2). Finally, it calculates the percentage based on these two values and generates the answer. Here, the extraction of Evidence Item 2 depends entirely on the document scope and section location determined by Evidence Item 1. If the target section cannot be accurately identified, valid data cannot be obtained, thus forming a clear chain of evidence dependencies.
[0031] Step 130: Based on query complexity and evidence dependency structure, determine at least one target database to be queried from multiple databases, and perform a retrieval operation on at least one target database.
[0032] These databases include a vector database for semantic retrieval, a graph database for structural navigation, and an inverted index database for full-text content retrieval.
[0033] Next, based on the assessment of query complexity and evidence dependency structure, the system dynamically selects one or more of the most suitable databases for collaborative retrieval. It should be noted that graph databases only store the hierarchical structure information of documents (such as headings at each level and their parent-child relationships), vector databases only store semantic vector representations (such as the embedding vectors of chunk text), and inverted index databases are used to store the specific text content of all chunks after segmentation and their metadata. Therefore, in actual retrieval, these three types of databases usually need to be used in combination, and efficient linkage between structure, semantics, and content is achieved through a globally unique chunk_id. As an example, if the query focuses on semantic similarity, such as a user wanting to find "other deposit product documents that are similar in style to the terms and conditions of a certain bank's large-denomination certificate of deposit product", the system first calls the vector database to perform semantic vector matching, recalls several candidate chunk_ids, and then obtains the corresponding original content through the inverted index database to support subsequent answer generation.
[0034] As another example, if the query depends on the internal hierarchical structure of the document, such as locating "the rate of return on fixed-income assets under the 'Investment Return Analysis' section in an insurance company's annual report," the system first calls the graph database, using its stored title tree to quickly navigate to the section node, obtain its associated chunk_id list, and then extracts the specific text content through the inverted index database. If the query emphasizes precise keyword matching, such as searching for "financial notes paragraphs containing the word 'handling fee'," the system directly calls the inverted index database to perform a full-text search, returning the matching chunk_id and its complete context.
[0035] In addition, in complex query scenarios, the three types of databases can also work together in depth. For example, the system can first limit the scope of the "asset management business" chapter through the graph database, then filter the most semantically relevant paragraphs within that scope through the vector database, and finally verify the keyword coverage and extract the final original text through the inverted index database. Thus, this multi-database collaborative architecture, which combines functions on demand and decouples functions, not only ensures the accuracy of financial document retrieval in structural navigation and the flexibility in semantic understanding, but also ensures the completeness and readability of the final returned answers, thereby improving the practicality and reliability of complex rich text document question-answering systems.
[0036] In some embodiments, such as Figure 2 As shown, determining at least one target database may include the following steps: Step 1301: Based on the evidence dependency structure, determine the multiple evidence items that need to be retrieved to generate the answer and the dependencies between the multiple evidence items.
[0037] As one possible implementation, the evidence dependency structure is parsed to obtain the specific evidence items required to answer the user's query and the dependencies between these evidence items. Evidence items indicate the smallest unit of fact necessary to generate an answer; each evidence item corresponds to an information requirement point, such as a financial indicator value at a specific point in time, a policy description in a chapter, or the specific content of a product clause. Dependencies indicate the logical organization and execution order of the evidence items, including but not limited to parallel relationships, chain dependencies, conditional branches, or aggregation calculation relationships, to guide the system on how to combine, verify, or compute multiple evidence items to form the final answer during the retrieval and reasoning process.
[0038] For example, if a user queries "What were the net profits of a certain bank each year from 2021 to 2023?", the system needs to identify three independent but co-originating evidence items: net profit in 2021, net profit in 2022, and net profit in 2023, and understand that they are parallel aggregates with no sequential dependency. However, if the query is "The management fee income disclosed by a certain fund in its first quarterly report after completing its Series A financing," the system needs to first locate the "Completion Time of Series A Financing," then determine the "First Quarterly Report" based on that time, and finally extract the "Management Fee Income," forming a chain dependency relationship. By explicitly modeling evidence items and their dependencies, the system can accurately plan subsequent retrieval paths.
[0039] Step 1302: Based on the data characteristics of multiple evidence items, determine the candidate database type for retrieving multiple evidence items.
[0040] Furthermore, for each piece of evidence, the system analyzes its content attributes and retrieval requirements to match the most suitable database type as a candidate. For example, if a piece of evidence is a descriptive text (such as "Product Risk Warning Statement") and users are concerned with its semantic similarity, then a vector database is a candidate; if the piece of evidence needs to be extracted from a document with a clear hierarchical structure (such as "Capital Expenditure Details under Chapter 3, Section 2 of the Annual Report"), then a graph database becomes a candidate due to its support for title trees; if the piece of evidence relies on precise term matching (such as a sentence containing the phrase "Net Interest Margin"), then an inverted index database is listed as a candidate due to its efficient keyword indexing capabilities. This process ensures that each piece of evidence can be associated with the database type most likely to match its content.
[0041] Step 1303: Determine the target database type from the candidate database types based on the query complexity.
[0042] As one possible approach, decision optimization can be performed by considering the overall query complexity. For example, for low-complexity queries (such as those requiring only one piece of evidence), the system directly selects the optimal database type corresponding to that piece of evidence as the target; while for high-complexity queries (such as those involving multiple pieces of evidence, cross-chapter queries, or queries requiring joint semantic and structural judgments), the system may retain multiple candidate types and design a combination strategy based on dependencies.
[0043] For example, when a query requires locating a chapter before filtering semantically relevant content, even if both databases are candidates, the combination of the graph database and the vector database will be prioritized. It should be noted that this decision-making process is driven by a large language model or rule engine to ensure that resource investment matches the difficulty of the problem.
[0044] Step 1304: Based on the target database type and dependencies, determine at least one target database to be queried from multiple databases.
[0045] In this context, dependencies are used to determine the execution order of at least one target database retrieval.
[0046] Ultimately, the system maps the target database type to a specific database instance and arranges the retrieval execution order based on the dependencies between evidence items. For example, in a chained dependency scenario, the system first calls the graph database to obtain the list of chunk_ids for the target chapter, then uses the IDs in the chunk_id list as range constraints to pass to the vector database for semantic filtering, and finally extracts the original text through the inverted index database. In a parallel dependency scenario, multiple evidence items can be retrieved concurrently in their respective target databases. The dependencies not only determine whether phased execution is required, but also control how intermediate results are passed to the next stage.
[0047] In summary, based on the evidence dependency structure, the specific evidence items and logical relationships required to answer the query are determined. Then, the most suitable candidate database type is matched by combining the data characteristics of each evidence item. Based on the overall query complexity, the efficient and necessary target database type is selected. Finally, the calling order and collaboration method of the target database are dynamically determined according to the dependencies between evidence. This not only improves the recall accuracy and contextual integrity of key information in multi-source heterogeneous financial documents, but also effectively avoids the computational redundancy caused by full database scanning. At the same time, the dependency-driven execution order ensures the logical rigor of multi-hop reasoning, enhancing the practicality and reliability of the intelligent question answering system in professional financial scenarios.
[0048] It should be noted that before determining at least one target database to be queried from multiple databases, such as... Figure 3 As shown, multiple databases can be constructed using the following steps: Step 101: Obtain the rich text document and perform structured parsing on the rich text document to obtain multiple document layout elements.
[0049] One possible approach is to retrieve rich text documents from a data source, such as PDF versions of bank annual reports, insurance product brochures, or securities company research reports. These documents not only contain plain text but also embed various visual and logical elements such as titles, paragraphs, tables, images, headers, footers, and lists. A document parsing engine is then used to perform structured parsing of the documents, identifying each independent layout unit and labeling it with attributes such as element type, position coordinates, and font level. A set of structured document layout elements is then output according to the document's reading order. The element types may include, but are not limited to, titles, body text, and tables / images.
[0050] Step 102: Iterate through multiple document layout elements in sequence, and use a text extraction strategy that matches the element type of the currently traversed document layout element to determine the associated text block and its semantic level.
[0051] Furthermore, the system dynamically selects the most suitable text extraction strategy based on the type of document layout elements, and determines the associated text blocks and semantic levels of each document layout element. For example, for heading elements, the system combines font size, bolding status, and hierarchical indentation to infer the corresponding semantic level (such as first-level heading, second-level heading, etc.).
[0052] In some embodiments, such as Figure 4 As shown, step 102 may include the following steps: Step 1021: In response to the fact that the element type of the currently traversed document layout element is a non-text type, search and calculate the spatial distance between the current document layout element and the neighboring candidate text blocks in the preset constraint direction based on the center point of the bounding box of the currently traversed document layout element.
[0053] As one possible implementation, during the structured parsing of rich text documents, the system processes each identified document layout element. When encountering a non-text element, such as an image showing asset allocation ratios or a table containing quarterly financial data, since the image or table does not directly provide readable text, the system needs to infer its semantic level based on its spatial position on the page. Therefore, the system first extracts the bounding box of the non-text element and calculates the coordinates of its center point. Then, based on the reading order, it sets a preset constraint direction, such as performing a local search upwards or downwards in the vertical direction, and expanding to the horizontal direction if necessary. Within this defined area, the system identifies all parsed text blocks and calculates the spatial distance between the bounding box of each text block and its center point.
[0054] Step 1022: Determine the description text of the currently traversed document layout element from the candidate text blocks based on the spatial distances.
[0055] Furthermore, based on the principle of minimum distance, candidate text blocks can be filtered to determine the descriptive text most likely to serve as the title or annotation of the non-text element. For example, when a bank's annual report contains a bar chart located in the middle of the page, with a line of bold text 5 millimeters directly above it... Figure 3-2The chart displays "Net Interest Income of Each Business Line in 2023," with a short descriptive text located 5 millimeters below. The system prioritizes the text that is closer in distance and more prominently formatted as the chart's description. If multiple candidates with similar distances exist, further consideration can be given to font size, whether they contain leading words such as "chart" or "table," and whether they are in the same column to improve matching accuracy.
[0056] Step 1023: Based on the description text, determine the associated text block and its semantic level that is associated with the currently traversed document layout element.
[0057] As one possible implementation, the descriptive text can be merged with the content generated after parsing the current non-text element to form an associated text block of the document layout element; in addition, the source text block that references the descriptive text can be determined based on the descriptive text, and the semantic level to which the document layout element belongs can be determined based on the source text block.
[0058] As an example, semantic parsing is performed on the currently traversed document layout element to obtain its semantic content; based on the semantic content and descriptive text, associated text blocks are generated; based on the descriptive text, a backreference lookup is performed in the context of adjacent pages of the page containing the currently traversed document layout element to locate the source text block that references the currently traversed document layout element; based on the document structure hierarchy of the source text block, the semantic hierarchy to which the currently traversed document layout element belongs is determined.
[0059] In other words, when processing non-text document layout elements being traversed, semantic parsing is performed on these elements to extract their inherent semantic content. For example, for table elements, this process may include identifying row and column structures, headers, and data cells, and converting them into structured natural language descriptions. For chart elements, optical character recognition (OCR), legend recognition, or pre-trained multimodal models may be used to generate textual summaries of data trends or structures. Furthermore, this semantic content is compared with previously spatially determined descriptive text (such as "..."). Figure 3-2 The system integrates the description text (e.g., "changes in net profit for each quarter of 2023") to generate a related text block for a document layout element. To further pinpoint the logical hierarchy of this document layout element within the entire text, the system also performs a backreference search on the current page and its preceding and following adjacent pages based on this description text, searching for instances of phrases like "as...". Figure 3-2 When a source text block is located, the source text block is analyzed to determine its structural path in the document tree, such as under "Chapter 5 > Section 5.1", and this path is used as the semantic level to which the current non-text element belongs.
[0060] In other embodiments, such as Figure 5 As shown, step 102 may include the following steps: Step 102a: In response to the fact that the element type of the currently traversed document layout element is text type, the text content is segmented into sentences based on the reading order of the text content of the currently traversed document layout element to obtain multiple sentence units.
[0061] As one possible implementation, when traversing document layout elements, if the current element is identified as body text (i.e., continuous paragraph text), then the paragraph is segmented into sentences according to the reading order based on the grammar rules and punctuation of natural language. For example, the text "In 2023, the bank achieved a net profit of 45.6 billion yuan, a year-on-year increase of 8.2%. The net interest margin was 1.85%, a decrease of 5 basis points from the previous year." from a bank's annual report would be segmented into two independent sentence units. The segmentation process can consider terminators such as periods, question marks, and exclamation marks, and exclude interference points in abbreviations or numbers, such as the decimal point in "1.85%", to ensure that each sentence unit is relatively complete semantically and has clear boundaries.
[0062] Step 102b: Based on the sentence number, sentence content, and element type of multiple sentence units, generate input sequences for multiple sentence units, and predict the semantic level of multiple sentence units based on the multiple input sequences.
[0063] As one possible implementation, each sentence unit is combined with its sequence number in the paragraph (e.g., sentence 1, sentence 2), the original text content, and the type of the main text element to construct a structured input sequence. This input sequence is then fed into a pre-trained language model or hierarchical classifier to predict its semantic level in the logical structure of the entire text.
[0064] For example, in an insurance company's annual report, the first sentence of a paragraph, "This section mainly analyzes the credit risk exposure of investment assets," is predicted to be a level three heading; while the subsequent sentence, "As of the end of 2023, AAA-rated bonds accounted for 62%," is determined to be a level four content under that section. It should be noted that the model has learned the hierarchical patterns of a large number of financial documents during the training phase and can make joint judgments by combining contextual semantics and positional features.
[0065] This avoids the traditional method of extracting structures that relies solely on explicit headings, allowing implicit structural information to be identified and assigned a reasonable hierarchy, thus improving the ability to understand the structure of paragraphs without explicit numbering.
[0066] In some embodiments, such as Figure 6 As shown, step 102b may include the following steps: Step 102b1: In response to the fact that the length of multiple sentence units is greater than the maximum context capacity of a single inference of a large language model, based on the reading order, a sliding window with a preset window size and a preset sliding step size is used to slide and segment multiple sentence units.
[0067] As one possible implementation, if the total text length corresponding to the total number of sentence units obtained after the system completes sentence segmentation for a document layout element of the main text type exceeds the maximum context length that a large language model can handle (e.g., exceeding 8192 tokens), then it cannot input all the content into the model for semantic level prediction at once. Therefore, the system can divide the sentence sequence into several overlapping or continuous subsequences based on the original reading order, with each subsequence called a sliding window. The window size is determined by the model's maximum context capacity, while the sliding step size controls the offset between adjacent windows. It can be smaller than the window size to retain some overlapping content and ensure contextual coherence. This segmentation strategy adapts long texts to the model's input constraints without losing information.
[0068] Step 102b2: For the i-th sliding window, obtain the semantic level of the predicted sentence units in the first i-1 sliding windows, and construct the historical context based on the semantic level of the predicted sentence units in the first i-1 sliding windows.
[0069] Where i is a positive integer greater than 1.
[0070] As one possible implementation, when processing the i-th sliding window (i > 1), the system no longer only looks at the sentence in the current window, but actively backtracks the reasoning results of the previous i-1 windows.
[0071] As an example, the system collects the predicted semantic level labels of all sentence units in the previous i-1 windows and organizes them into a structured context summary, such as aggregating key turning points or the latest level state by hierarchical path, which constitutes the historical context to represent the overall structural direction of the document paragraphs before the current window. In this way, the system can inject global structural information into the subsequent reasoning process in a compressed form, avoiding hierarchical jumps or logical breaks caused by window fragmentation.
[0072] Step 102b3: Concatenate the historical context with the input sequence corresponding to the i-th sliding window to form an enhanced input.
[0073] Next, the constructed historical context is concatenated with the input sequence of the i-th sliding window itself (including the sequence number, content, and element type of each sentence within the window) to generate an enhanced input that integrates local details and global structural information. This enhanced input retains the original semantic content of the current window while carrying the previously inferred structural evolution trajectory. For example, if the historical context indicates that the preceding text is at the sub-topic level of "Section 4.2 xx", the introductory sentence in the current window is more likely to be identified as a continuation of that sub-topic rather than opening a new chapter. The length of the concatenated input remains within the model's context capacity, ensuring executability.
[0074] Step 102b4: Provide the enhanced input to the large language model to predict the semantic level of each sentence unit in the i-th sliding window.
[0075] Finally, the system feeds the enhanced input into the large language model for inference, and the model outputs the semantic hierarchy prediction result for each sentence unit in the i-th sliding window. Since the input already contains historical structural information, the model can maintain the hierarchical logic consistent with the preceding text while preserving local semantic understanding. This mechanism can effectively alleviate the "contextual forgetting" problem in long document segmentation processing and improve the coherence and accuracy of semantic hierarchy prediction.
[0076] In summary, for long text paragraphs exceeding the single context capacity of a large language model, a sliding window strategy is employed based on the reading order to orderly segment sentence units, ensuring comprehensive content coverage. Then, for the i-th window (i > 1), the predicted semantic hierarchy results from the previous i-1 windows are aggregated to construct a structured historical context, representing the document's previous logical evolution path. This historical context is then concatenated with the input sequence of the current window to form an enhanced input that integrates global structural cues and local semantic details. Finally, this enhanced input is fed into the large language model, achieving context-aware prediction of the semantic hierarchy of each sentence unit within the current window. This effectively overcomes the problems of hierarchical jumps, logical breaks, or structural misjudgments caused by context fragmentation in long text segmentation. Even with strictly limited model input length, it maintains cross-window semantic coherence and hierarchical consistency, improving the accuracy of semantic hierarchy prediction and the completeness of document tree construction, providing a structured foundation for subsequent multi-database collaborative retrieval and complex question-answering reasoning.
[0077] Step 102c: Based on the semantic level of multiple sentence units and the positional continuity of multiple sentence units in the reading order, aggregate sentence units that belong to the same semantic level and are in consecutive positions into a text fragment, and use the text fragment as an associated text block.
[0078] Furthermore, the continuity of multiple sentence units in the reading order is analyzed. For example, if several adjacent sentences are determined to be at the same semantic level, they are merged into a coherent text segment. For instance, in a text containing five sentences, if sentences 2 through 4 are predicted to belong to the same subtopic under "Section 4.2 xx Analysis" and are consecutive, these three sentences will be aggregated into a single text segment. This text segment serves as the smallest indexable unit, i.e., a related text block, for subsequent storage and retrieval. The aggregation process ensures semantic cohesion of the content, avoiding the fragmentation of the same argument into multiple isolated units.
[0079] Step 102d: Determine that the semantic level of the text fragments is the same semantic level.
[0080] Finally, the overall semantic level of the text fragment is uniformly set to the level to which all the sentence units it contains belong. For example, if a text fragment consisting of three consecutive sentences is determined to belong to "Chapter 5 > Section 5.3 > Subtopic B", then the semantic level of the fragment is officially marked as that path.
[0081] In summary, by segmenting document layout elements of the main text type into sentences according to reading order, sentence units are obtained. Then, a structured input sequence is constructed by combining sentence number, content semantics, and element type, and a pre-trained model is used to predict the semantic level of each sentence unit in the logical structure of the whole text. Furthermore, based on the consistency of semantic level and the positional continuity of reading order, adjacent and same-level sentence units are aggregated into text fragments, and these text fragments are used as associated text blocks. This avoids the limitations of traditional methods that rely solely on explicit headings for hierarchical division, and can still achieve high-precision structured organization in financial texts without numbering or formatting marks. At the same time, the generated associated text blocks have both semantic integrity and structural traceability, thereby improving the accuracy, recall, and interpretability of multi-database collaborative retrieval.
[0082] Step 103: Generate the document tree of the rich text document based on the associated text blocks and their semantic levels.
[0083] In the document tree, leaf nodes indicate associated text blocks, and branch nodes indicate headings at corresponding semantic levels.
[0084] As one possible implementation, a document tree for the rich text document is generated based on the associated text blocks and their semantic levels. The leaf nodes of the document tree indicate the atomic-level text blocks obtained after parsing, including body paragraphs, table content parsing results, and alternative text or captions for images. The branch nodes indicate the titles of the corresponding semantic levels. The parent-child relationship of the document tree strictly follows the logical nesting structure of the document. For example, if a paragraph is a child of "Section 3.2", its leaf node is attached to the "Section 3.2" branch node.
[0085] Step 104: Based on the document tree, construct a graph database for structural navigation, an inverted index database for full-text content retrieval, and a vector database for semantic retrieval.
[0086] Furthermore, based on the document tree, information of different dimensions is injected into three types of heterogeneous databases. For example, the title node and its parent-child relationship are imported into a graph database (such as Neo4j) to support structural navigation based on chapter paths; the text blocks of all leaf nodes and their metadata (such as title paths and element types) are written into an inverted index database (such as Elasticsearch) to support precise keyword matching; at the same time, a semantic vector is generated for each text block and stored in a vector database (such as Milvus) to support semantic similarity retrieval. The graph database, inverted index database, and vector database are interconnected through globally unique text block IDs.
[0087] As an example, non-leaf nodes in the document tree are stored as title nodes in a graph database. Each title node contains a title name, its associated document identifier, and its level depth, and is connected to adjacent title nodes via parent-child relationship edges. Each title node is associated with one or more associated text block identifiers, indicating the associated text blocks belonging to that title node. Each associated text block is stored as an index unit in an inverted index database. Each associated text block corresponds to an index record in the inverted index database, and the index record contains the associated text block identifier, the original text, the summary text, the title path, and the element type. Original text embedding vectors and summary embedding vectors are generated for each associated text block, and these embedding vectors are bound to the corresponding associated text block identifiers and stored in a vector database. The graph database, the inverted index database, and the vector database are linked through the associated text block identifiers.
[0088] In other words, non-leaf nodes in the document tree are stored as title nodes in the graph database. Each title node contains the title name, a unique identifier of the document it belongs to, and its hierarchical depth in the document structure. It is connected to parent and child title nodes via directed edges, forming a hierarchical graph structure reflecting the original document's chapter organization. Each title node is also associated with one or more related text block identifiers, explicitly pointing to all content units belonging to that title. Simultaneously, all related text blocks are written as basic index units to the inverted index database. Each index record contains the related text block identifier, the complete original text, the semantic summary text generated by the large language model, the complete path from the root node to the title (e.g., "Chapter 3 > Section 3.2"), and the element type (e.g., body text, table parsing text, or figure caption). Furthermore, the system generates original-text-based embedding vectors and summary-based embedding vectors for each related text block, storing these two vectors along with the corresponding related text block identifiers in the vector database to support semantic similarity retrieval at different granularities. The graph database, inverted index database, and vector database are linked across databases through the related text block identifiers.
[0089] Thus, the graph database preserves the complete logical structure of the document, supporting efficient chapter navigation and hierarchical filtering; the inverted index database ensures precise keyword matching and high recall for full-text retrieval; and the vector database empowers semantic-level fuzzy matching and intent understanding. The graph database, inverted index database, and vector database work together through shared identifiers, enabling the system to accurately locate the content range based on the structural path when responding to complex queries, and to perform multi-path fusion sorting by combining keywords and semantic vectors, thereby improving the accuracy, contextual completeness, and interpretability of professional Q&A in the financial field.
[0090] In summary, by performing structured parsing on rich text documents, the system first decomposes them into multiple document layout elements with clear type and location information. Then, based on the type of each element, it dynamically adapts the corresponding text extraction strategy to accurately generate related text blocks that are semantically consistent with the element, and determines their semantic level based on the context. Next, it constructs a document tree that fully reflects the logical organization of the original document by using all related text blocks as leaf nodes and headings at all levels as branch nodes. Finally, based on this document tree, it simultaneously constructs three types of heterogeneous databases. This not only preserves the original chapter logic and information integrity of financial documents but also provides multi-dimensional collaborative retrieval capabilities for complex queries, improving the accuracy, interpretability, and response efficiency of the intelligent question-answering system in professional financial scenarios.
[0091] Step 140: Based on the retrieval results returned from at least one target database, perform inference operations corresponding to the evidence dependency structure to generate an answer in response to the user's query.
[0092] Finally, based on the evidence-dependent structure, structured reasoning is performed on the search results returned from at least one target database, rather than directly concatenating the original text. For example, for a query like "changes in the revenue share of a certain product line over the past three years," the system first verifies whether the corresponding data blocks for the three years have been successfully retrieved. Then, it extracts the revenue value and total revenue from each block, calculates the share, and organizes it into a trend description according to the time series. If the evidence-dependent structure requires contextual traceability, the system will also integrate chapter path information from the graph database to ensure that the answer has a source basis. The entire reasoning process can be completed by the large language model under controlled prompts, strictly constraining it to generate content only based on retrieved evidence.
[0093] In some embodiments, such as Figure 7 As shown, step 140 may include the following steps: Step 1401: Based on the evidence dependency structure, map the retrieval results returned from at least one target database to the corresponding evidence items.
[0094] Next, based on the pre-constructed evidence dependency structure, these search results are precisely assigned to their corresponding logical evidence items. For example, if a user queries "What were the net fees and commissions income of a certain bank in 2021, 2022, and 2023 respectively?", the evidence dependency structure defines three independent evidence items, each corresponding to the income data for the three years. When the search returns three text blocks containing financial data from different years, the system analyzes the time tags and indicator names in each text block and maps them to the evidence items corresponding to 2021, 2022, and 2023 respectively. This mapping process ensures that the data used for subsequent reasoning is strictly aligned with the semantic structure of the original question.
[0095] Step 1402: Obtain the evidence content of at least one evidence item and the structured context information associated with the at least one evidence item from the search results returned by at least one target database.
[0096] The evidence content is used to indicate the factual data corresponding to the at least one piece of evidence.
[0097] Furthermore, not only is the core text extracted from the search results used as evidence content, i.e., the specific facts supporting the answer, such as "net interest income in 2023 was 58.23 billion yuan", but also the structured context information bound to the evidence items is obtained simultaneously. This context information includes the document name, chapter path (e.g., "financial report > Chapter 3 > Section 3.1"), element type (text, table or chart), publication time, indicator unit, and related title nodes, etc. It should be noted that the evidence content carries factual data, while the structured context explains the source location of the facts and their organizational relationship in the document.
[0098] Step 1403: Based on the dependencies defined in the evidence dependency structure, perform multi-stage aggregate reasoning according to the evidence content and structured context information to generate an answer in response to the user's query.
[0099] Furthermore, based on the dependency relationships defined in the evidence dependency structure, the evidence content and structural context information are integrated and reasoned in stages.
[0100] For example, if a user asks "the average annual growth rate of a fund's management fee income over the past three years", the system first verifies whether the three-year data is complete, then extracts the values from each piece of evidence, confirms the year is correct in the context, and then calculates the growth rate of adjacent years and takes the average. If the dependency relationship is a chain structure, such as "first determine the latest quarterly report, and then extract the risk warnings", the system will verify the consistency of the context in sequence before performing content extraction. The entire reasoning process can be completed by the large language model under controlled prompts, but it is strictly limited by the evidence and context constraints already obtained.
[0101] In summary, the system first accurately maps multi-source retrieval results to predefined logical evidence items based on the evidence dependency structure, ensuring that each factual fragment strictly corresponds to its semantics in the question. Then, it synchronously extracts evidence content and its structured contextual information from the retrieval results, preserving the accuracy of the original facts while capturing key contextual information such as source location, document level, and metadata. Finally, the system performs multi-stage aggregation reasoning on the evidence content and context according to the dependencies defined in the evidence dependency structure, achieving controllable generation from discrete data to coherent answers. This improves the accuracy, logical consistency, and interpretability of responses to complex financial queries, effectively avoiding answer bias caused by result mismatch, missing context, or unconstrained generation. Simultaneously, the joint constraints of structured context and dependencies enable the system to output high-quality answers with traceable sources and verifiable logic, thus meeting the high standards of rigor and reliability required by financial institutions for intelligent question-answering systems.
[0102] The data query method of this disclosure, upon receiving a user query, first performs intent parsing to obtain the query intent, and then dynamically determines the query complexity and required evidence dependency structure accordingly, thereby avoiding indiscriminate retrieval of all databases. Furthermore, based on the query complexity and evidence dependency structure, it selectively calls one or more target databases from vector databases, graph databases, and inverted index databases to perform the retrieval, enabling the system to accurately match the most suitable retrieval mechanism according to the actual needs of the query. Subsequently, the system performs structured reasoning on the multi-source retrieval results based on the evidence dependency structure, rather than simply concatenating or sorting them, thereby generating logically coherent and factually traceable answers. This achieves the goal of ensuring the accuracy and interpretability of the answers while reducing unnecessary computational overhead and improving overall query efficiency.
[0103] Based on any of the above embodiments, the data query method of this disclosure may further include the following steps: I. For example Figure 8 As shown, the following steps 1 to 3 are used to construct the graph database, inverted index database, and vector database; Step 1: Multimodal document parsing The core of this step is to use a visual language model (VLM, such as the LayoutLM series or a multimodal large model with layout understanding capabilities) to parse the input rich text document (such as PDF, image format document).
[0104] Specifically, the document is input page by page into the VLM (Visual Library). Based on its pixel-level understanding of the document, the VLM outputs a sequence of layout elements containing positional information. Each layout element contains at least: element type (such as Title, Header, Footer, Text, Table, Figure, Caption, Note, etc.), bounding box coordinates, and page number, and is output in the order of reading.
[0105] Step 2: Precisely Associate Layout Elements 2.1. Spatial Anchoring Based on Center Point Distance The system extracts the center coordinates (X_img, Y_img) of the bounding box of the table / image layout block. It then iterates through text-based layout blocks on the same and adjacent pages (handling cross-page cases), calculating the spatial distance between their center points (X_txt, Y_txt) and (X_img, Y_img). Finally, based on the layout type "Caption" and "Note," the system binds the table / image title / note to these elements, creating a complete table / image chunk.
[0106] 2.2. Enhanced Chart / Table Content To make images / tables searchable, a summary needs to be generated for enhancement: a multimodal large model is called to directly perform semantic understanding on the image / table to generate a natural language summary text; at the same time, a specialized table parsing model is called to accurately parse the table image into structured text in Markdown or HTML format, preserving the row and column logic.
[0107] 2.3. Chart / Table References Images or tables are often isolated from their context, and their chapter affiliation cannot be determined solely by physical distance. Based on the previously extracted and bound chart "title" text, and combined with the page number P where the chart is located, a context layout window (e.g., pages P-1 to P+1) is defined. This layout block within the defined scope is then imported into the LLM (Local Library), and the command "Find which paragraph in the above text explicitly references this chart" is executed. The LLM can recognize implicit or explicit references such as "as shown in the table below" or "see 3." Once the source text block is located, it is bidirectionally bound to the image / table block. Through this mechanism, the chart can assist in determining the document level to which its associated image / table chunk (text block) belongs during subsequent tree construction.
[0108] Step 3: Constructing a structured document tree and hierarchical segmentation 3.1. Reconstructing Text Sequences and Layout Type Injection The plain text layout blocks output by VLM are organized according to the reading order. First, a text sequence is constructed based on the layout blocks. Then, the layout blocks are segmented into sentences by punctuation, and the mapping relationship between sentences and layout blocks is recorded. Next, a standardized input sequence of "[sentence number]@[sentence content]@[layout block type]" (e.g., 15@company annual strategic plan@title) is constructed, incorporating task prompts, and fed into LLM for title and hierarchy prediction. The output format is defined as "[sentence number]@[title level]", where the title level is a number from 0 to 9, where 0 represents the body text, 1 represents a first-level title, 2 represents a second-level title (a subheading of a first-level title), and so on. Our key innovation lies in forcibly injecting the visual typography prior parsed by VLM into LLM through the type tag at the end, which improves the accuracy of hierarchy prediction.
[0109] 3.2. Sliding Window Titles and Hierarchical Reasoning LLM inputs have context length limitations. To handle very long documents, a sliding window with a length of N sentences is used. When processing the i-th window, the system does not isolate the input but dynamically constructs the "historical context." The specifically constructed input consists of two parts: Prefix: The title and hierarchy results already identified in the preceding window (e.g., {15@Company Annual Strategic Planning@Level One, 22@Part One: Market Analysis@Level Two}); Main body: The standardized sequence of the current window.
[0110] Based on this input, LLM predicts the heading level within the current window. After processing all windows, the sentence-level prediction results are mapped back to the original layout block level, achieving stable splicing of the global structure.
[0111] 3.3. Constructing the Document Tree and Chunk Splitting The document tree is constructed based on the mapping results: each title is a branch node (independent block, with chapter summaries generated by LLM), the main text / figures are attached to it as leaf nodes, and "title path" metadata (e.g., root node > Chapter 1 > Section 1.1) is generated for all nodes.
[0112] During the segmentation phase, the system uses branch nodes as absolute boundaries. When the main text content under a certain chapter exceeds the set length threshold, the system only performs secondary segmentation within the leaf nodes according to semantic boundaries, and the segmentation never cuts off the title branches. This fundamentally eliminates the phenomenon of "mixed content across chapters" in traditional segmentation.
[0113] like Figure 9 As shown, the following steps are used to query the data: Step 4: Intent-Driven Adaptive Retrieval 4.1 Retrieving Data Storage To facilitate subsequent adaptive retrieval, this system does not store all data in a single database. Instead, it splits and stores the data according to its functional attributes in a graph database, an inverted index database (such as Elasticsearch), and a vector database. The specific storage content and locations are as follows: Graph database: Storage location: Graph database (such as Neo4j).
[0114] Stored content: Only branch nodes of the document tree (i.e., header nodes at all levels) are stored.
[0115] Specific fields: Each node stores the title name, its document ID, its level depth, and parent-child hierarchical relationships between nodes. It does not store the original text of the main content chunks or long text summaries; it only retains a related field (such as a list of related_chunk_ids) that records the set of IDs for all main content / figure chunks under that title node.
[0116] Inverted index database: Storage location: Elasticsearch (ES).
[0117] Storage content: Stores the specific text content of all chunks after splitting.
[0118] Specific fields: Each chunk creates an indexed document in Elasticsearch, with fields including: a unique chunk ID (chunk_id), the full text of the chunk, the chunk summary text, the title path (e.g., "Chapter 1 > Section 1.1"), and element type (text / table / image), etc. These are used to support subsequent precise keyword retrieval.
[0119] Vector database: Storage location: Vector database (such as Milvus or Faiss).
[0120] Storage content: Stores the dense vector after Chunk text conversion.
[0121] Specific fields: Two vectors are generated for each chunk (original text embedding vector and summary embedding vector), and each vector is bound to its corresponding chunk_id for storage. This is used to support subsequent semantic similarity retrieval.
[0122] The above three elements are linked together by a globally unique identifier, chunk_id. During retrieval, the system can quickly traverse title nodes in the graph database through parent-child relationships to obtain a list of chunk_ids for the target chapter. Then, it can use these IDs to accurately extract the corresponding original text and vectors from Elasticsearch or a vector database. Alternatively, it can retrieve the corresponding chunks in Elasticsearch or a vector database using a query, and then obtain the chapter structure of the corresponding document from the graph database based on the chunk_id, thereby achieving the separation and coordination of "structure navigation" and "content reading".
[0123] 4.2 Adaptive Retrieval Planning The system encapsulates graph retrieval, Elasticsearch retrieval, and vector retrieval into independent API tools. For example... Figure 9 As shown, after receiving a user query, the system does not directly trigger a search, but instead uses the reasoning capabilities of a large language model to perform intent parsing and strategy planning.
[0124] Based on the complexity of the problem and the dependencies on evidence, the system adaptively matches different retrieval paths: For simple single-evidence queries: the system adaptively plans to call a single tool, directly and accurately hitting the target chunk (text block) through vector or ES tools.
[0125] For complex multi-hop / multi-evidence queries (taking "What were the net profits of Company XX for the four quarters of its 2025 financial report?" as an example), the system adaptively plans to use serial multi-tool calls: Thinking and Action 1 (Macro Positioning): Determine the specific document range needed, call tool_vector_summary, and lock in "XX Company's 2025 Financial Report" through the summary semantic vector.
[0126] Thinking and Action 2 (Structure Navigation): Determining that data needs to be extracted across chapters, call `tool_graph_navigate`, inputting the document ID and keywords. The graph database traverses the branch nodes in seconds, returning the paths to the corresponding chapters for the four quarters and a list of associated `chunk_id` values.
[0127] Thinking and Action 3 (Microscopic Extraction): After determining that the exact location has been obtained, call tool_es_fetch(chunk_ids=[...]) to directly and accurately pull the original data of these 4 chunks from ES.
[0128] 4.3 Search Result Judgment and Answer Generation After each round of tool calls to obtain results, the system forces the large model to perform a "sufficiency check" on the search results: verifying whether the currently recalled content contains all the key entities and data required to answer the query. If the evidence is insufficient and there are relevant tools that have not been called, the system automatically plans the next round of retrieval; if the evidence is sufficient, or if all relevant tools have been exhausted and still cannot meet the requirements, the retrieval process is dynamically terminated. Finally, high-quality recalled chunks and their "title paths" are input into the question-answering module to generate the answer. This mechanism, with on-demand token consumption, achieves complex question-answering capabilities that traditional RAGs cannot match.
[0129] Corresponding to the data query method provided in the above embodiments, this disclosure also provides a data query device. Since the data query device provided in this disclosure corresponds to the data query method provided in the above embodiments, the implementation of the data query method is also applicable to the data query device provided in this disclosure, and will not be described in detail in this disclosure.
[0130] Figure 10 This is a schematic diagram of the structure of the data query device shown in the eighth embodiment of this disclosure.
[0131] like Figure 10 As shown, the data query device 1000 includes: a parsing module 1010, a first determining module 1020, a second determining module 1030, and a generating module 1040.
[0132] The system includes a parsing module 1010, which, in response to a received user query, performs intent parsing on the query to obtain the query intent; a first determining module 1020, which determines the query complexity and the required evidence dependency structure based on the query intent; a second determining module 1030, which, based on the query complexity and evidence dependency structure, determines at least one target database to be queried from multiple databases and performs a retrieval operation on the at least one target database; wherein the multiple databases include a vector database for semantic retrieval, a graph database for structural navigation, and an inverted index database for full-text content retrieval; and a generating module 1040, which, based on the retrieval results returned from the at least one target database, performs inference operations corresponding to the evidence dependency structure to generate an answer in response to the user query.
[0133] As one possible implementation, the second determining module 1030 is used to determine, based on the evidence dependency structure, multiple evidence items to be retrieved for generating an answer and the dependencies between the multiple evidence items; determine candidate database types for retrieving the multiple evidence items based on the data characteristics of the multiple evidence items; determine the target database type from the candidate database types based on the query complexity; and determine at least one target database to be queried from the multiple databases based on the target database type and the dependencies; wherein, the dependencies are used to determine the execution order of retrieving at least one target database.
[0134] As one possible implementation, the generation module 1040 is used to map the retrieval results returned from at least one target database to corresponding evidence items according to the evidence dependency structure; to obtain the evidence content of at least one evidence item and the structured context information associated with at least one evidence item from the retrieval results returned from at least one target database; wherein the evidence content is used to indicate the factual data corresponding to at least one evidence item; and to perform multi-stage aggregation reasoning based on the dependency relationship defined in the evidence dependency structure, according to the evidence content and the structured context information, to generate an answer in response to the user's query statement.
[0135] As one possible implementation, multiple databases are constructed using the following modules: a processing module and a construction module.
[0136] The processing module is used to acquire rich text documents and perform structured parsing on them to obtain multiple document layout elements. It then iterates through these elements, employing a text extraction strategy that matches the element type of the currently traversed element to determine the associated text blocks and their semantic levels. Based on the associated text blocks and their semantic levels, it generates a document tree for the rich text document. The leaf nodes of the document tree indicate associated text blocks, and the branch nodes indicate the titles of the corresponding semantic levels. The construction module, based on the document tree, builds a graph database for structural navigation, an inverted index database for full-text content retrieval, and a vector database for semantic retrieval.
[0137] As one possible implementation, the processing module is configured to, in response to the fact that the element type of the currently traversed document layout element is a non-text type, search and calculate the spatial distance between the element and the neighboring candidate text blocks in a preset constraint direction based on the center point of the bounding box of the currently traversed document layout element; determine the descriptive text of the currently traversed document layout element from the candidate text blocks based on the spatial distances; and determine the associated text blocks and their semantic levels associated with the currently traversed document layout element based on the descriptive text.
[0138] As one possible implementation, the processing module performs semantic parsing on the currently traversed document layout element to obtain its semantic content; generates associated text blocks based on the semantic content and description text; performs a backreference lookup in the context of adjacent pages of the page containing the currently traversed document layout element based on the description text to locate the source text block that references the currently traversed document layout element; and determines the semantic level to which the currently traversed document layout element belongs based on the document structure hierarchy of the source text block.
[0139] As one possible implementation, the processing module, in response to the fact that the element type of the currently traversed document layout element is text type, segments the text content into sentences based on the reading order of the text content of the currently traversed document layout element, obtaining multiple sentence units; generates an input sequence of multiple sentence units based on the sentence number, sentence content, and element type of the multiple sentence units, and predicts the semantic level of the multiple sentence units based on the multiple input sequences; aggregates sentence units belonging to the same semantic level and with consecutive positions according to the semantic level of the multiple sentence units and the positional continuity of the multiple sentence units in the reading order into a text fragment, and treats the text fragment as an associated text block; and determines that the semantic level of the text fragment is the same semantic level.
[0140] As one possible implementation, the processing module, in response to multiple sentence units having a length greater than the maximum context capacity of a single inference by a large language model, uses a sliding window with a preset window size and a preset sliding step size to slide and segment multiple sentence units based on the reading order; for the i-th sliding window, it obtains the semantic levels of the predicted sentence units in the previous i-1 sliding windows, and constructs a historical context based on the semantic levels of the predicted sentence units in the previous i-1 sliding windows; where i is a positive integer greater than 1; the historical context is concatenated with the input sequence corresponding to the i-th sliding window to form an enhanced input; the enhanced input is provided to the large language model to predict the semantic level of each sentence unit in the i-th sliding window.
[0141] As one possible implementation, a module is constructed to store non-leaf nodes in the document tree as title nodes in a graph database. Each title node contains a title name, its associated document identifier, and its level depth, and connects adjacent title nodes via parent-child relationship edges. Each title node is associated with one or more associated text block identifiers, indicating the associated text blocks belonging to the title node. Each associated text block is stored as an index unit in an inverted index database. Each associated text block corresponds to an index record in the inverted index database, containing the associated text block identifier, the original text, the summary text, the title path, and the element type. Original text embedding vectors and summary embedding vectors are generated for each associated text block, and these embedding vectors are bound to their corresponding associated text block identifiers and stored in a vector database. The graph database, the inverted index database, and the vector database are linked through the associated text block identifiers.
[0142] The data query device of this disclosure, upon receiving a user query, first performs intent parsing to obtain the query intent, and then dynamically determines the query complexity and the required evidence dependency structure accordingly, thereby avoiding indiscriminate retrieval of all databases. Furthermore, based on the query complexity and evidence dependency structure, it selectively calls one or more target databases from vector databases, graph databases, and inverted index databases to perform the retrieval, enabling the system to accurately match the most suitable retrieval mechanism according to the actual needs of the query. Subsequently, the system performs structured reasoning on the multi-source retrieval results based on the evidence dependency structure, rather than simply concatenating or sorting them, thereby generating logically coherent and factually traceable answers. This achieves the goal of ensuring the accuracy and interpretability of the answers while reducing unnecessary computational overhead and improving overall query efficiency.
[0143] In an exemplary embodiment, an electronic device is also proposed.
[0144] The electronic devices include: processor; Memory used to store processor-executable instructions; The processor is configured to execute instructions to implement the data query method as proposed in any of the foregoing embodiments.
[0145] As an example, Figure 11 This is a schematic diagram of the structure of an electronic device 1100 shown in an exemplary embodiment of this disclosure, as follows: Figure 11 As shown, the aforementioned electronic device 1100 may further include: The memory 1110 and the processor 1120 are connected by a bus 1130, which connects different components (including the memory 1110 and the processor 1120). The memory 1110 stores a computer program, and when the processor 1120 executes the program, it implements the data query method described in the embodiments of this disclosure.
[0146] Bus 1130 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0147] Electronic device 1100 typically includes a variety of electronic device readable media. These media can be any available media that can be accessed by electronic device 1100, including volatile and non-volatile media, removable and non-removable media.
[0148] Memory 1110 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 1140 and / or cache memory 1150. Electronic device 1100 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 1160 may be used to read and write non-removable, non-volatile magnetic media (… Figure 11 Not shown; usually referred to as a "hard drive"). Although Figure 11 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 1130 via one or more data media interfaces. Memory 1110 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.
[0149] A program / utility 1180 having a set (at least one) of program modules 1170 may be stored, for example, in memory 1110. Such program modules 1170 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 1170 typically perform the functions and / or methods described in the embodiments of this disclosure.
[0150] Electronic device 1100 can also communicate with one or more external devices 1190 (e.g., keyboard, pointing device, display 1191, etc.), and with one or more devices that enable a user to interact with electronic device 1100, and / or with any device that enables electronic device 1100 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 1192. Furthermore, electronic device 1100 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1193. As shown, network adapter 1193 communicates with other modules of electronic device 1100 via bus 1130. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 1100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0151] The processor 1120 performs various functional applications and data processing by running programs stored in the memory 1110.
[0152] It should be noted that the implementation process and technical principles of the electronic device in this embodiment are explained in the foregoing description of the data query method of this disclosure embodiment, and will not be repeated here.
[0153] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, which can be executed by a processor of an electronic device to perform the data query method proposed in any of the above embodiments. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0154] In an exemplary embodiment, a computer program product is also provided, including a computer program / instructions, characterized in that the computer program / instructions, when executed by a processor, implement the data query method proposed in any of the above embodiments.
[0155] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0156] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A data query method, characterized in that, include: In response to receiving a user query, the query is parsed to obtain the query intent; Based on the query intent, determine the query complexity and the evidence dependency structure required for the query; Based on the query complexity and the evidence dependency structure, at least one target database to be queried is determined from multiple databases, and a retrieval operation is performed on the at least one target database; wherein, the multiple databases include a vector database for semantic retrieval, a graph database for structural navigation, and an inverted index database for full-text content retrieval; Based on the retrieval results returned from the at least one target database, an inference operation corresponding to the evidence dependency structure is performed to generate an answer in response to the user's query.
2. The method according to claim 1, characterized in that, The step of determining at least one target database to be queried from multiple databases based on the query complexity and the evidence dependency structure includes: Based on the evidence dependency structure, determine the multiple evidence items required to generate the answer and the dependencies between the multiple evidence items; Based on the data characteristics of the multiple evidence items, determine the candidate database type for retrieving the multiple evidence items; Based on the query complexity, determine the target database type from the candidate database types; Based on the target database type and the dependencies, at least one target database to be queried is determined from the plurality of databases; wherein the dependencies are used to determine the execution order of the retrieval of the at least one target database.
3. The method according to claim 1, characterized in that, The step of performing inference operations corresponding to the evidence dependency structure based on the retrieval results returned from the at least one target database to generate an answer in response to the user query includes: Based on the evidence dependency structure, the retrieval results returned from the at least one target database are mapped to the corresponding evidence items; From the search results returned by the at least one target database, obtain the evidence content of at least one evidence item and the structured context information associated with the at least one evidence item; wherein, the evidence content is used to indicate the factual data corresponding to the at least one evidence item; Based on the dependencies defined in the evidence dependency structure, multi-stage aggregate reasoning is performed according to the evidence content and the structured context information to generate an answer in response to the user's query.
4. The method according to claim 1, characterized in that, The multiple databases are constructed using the following steps: Obtain a rich text document and perform structured parsing on the rich text document to obtain multiple document layout elements; The document layout elements are traversed sequentially, and a text extraction strategy matching the element type of the currently traversed document layout element is used to determine the associated text block and its semantic level associated with the currently traversed document layout element. Based on the associated text block and its semantic level, a document tree of the rich text document is generated; wherein, the leaf nodes of the document tree are used to indicate the associated text block, and the branch nodes are used to indicate the title of the corresponding semantic level; Based on the document tree, a graph database for structural navigation, an inverted index database for full-text content retrieval, and a vector database for semantic retrieval are constructed.
5. The method according to claim 4, characterized in that, The text extraction strategy, which uses elements matching the type of the currently traversed document layout element, determines the associated text blocks and their semantic levels associated with the currently traversed document layout element, including: In response to the fact that the element type of the currently traversed document layout element is a non-text type, the spatial distance between the element and the neighboring candidate text blocks is searched and calculated in the preset constraint direction based on the center point of the bounding box of the currently traversed document layout element. Based on the spatial distances, the description text of the currently traversed document layout element is determined from the candidate text blocks; Based on the description text, determine the associated text block and its semantic level that is associated with the currently traversed document layout element.
6. The method according to claim 5, characterized in that, The step of determining the associated text block and its semantic level related to the currently traversed document layout element based on the description text includes: Semantic parsing is performed on the currently traversed document layout elements to obtain the semantic content of the currently traversed document layout elements; The associated text block is generated based on the semantic content and the descriptive text; Based on the description text, a backreference lookup is performed in the context of adjacent pages of the page where the currently traversed document layout element is located, in order to locate the source text block that references the currently traversed document layout element. Based on the document structure hierarchy of the source text block, determine the semantic hierarchy to which the currently traversed document layout element belongs.
7. The method according to claim 4, characterized in that, The text extraction strategy, which uses elements matching the type of the currently traversed document layout element, determines the associated text blocks and their semantic levels associated with the currently traversed document layout element, including: In response to the fact that the element type of the currently traversed document layout element is text type, the text content is segmented into sentences based on the reading order of the text content of the currently traversed document layout element to obtain multiple sentence units; Based on the sentence number, sentence content, and element type of the multiple sentence units, an input sequence of the multiple sentence units is generated, and the semantic level of the multiple sentence units is predicted based on the multiple input sequences. Based on the semantic hierarchy of the multiple sentence units and the positional continuity of the multiple sentence units in the reading order, sentence units belonging to the same semantic hierarchy and in consecutive positions are aggregated into a text segment, and the text segment is used as an associated text block; The semantic level of the text fragment is determined to be the same semantic level.
8. The method according to claim 7, characterized in that, The step of predicting the semantic level of the plurality of sentence units based on the plurality of input sequences includes: In response to the fact that the length of the multiple sentence units is greater than the maximum context capacity of a single inference of a large language model, based on the reading order, the multiple sentence units are slidally segmented using a sliding window with a preset window size and a preset sliding step size. For the i-th sliding window, obtain the semantic level of the predicted sentence units in the first i-1 sliding windows, and construct the historical context based on the semantic level of the predicted sentence units in the first i-1 sliding windows; where i is a positive integer greater than 1. The historical context is concatenated with the input sequence corresponding to the i-th sliding window to form an enhanced input; The enhanced input is provided to the large language model to predict the semantic level of each sentence unit in the i-th sliding window.
9. The method according to claim 4, characterized in that, The construction of a graph database for structural navigation, an inverted index database for full-text content retrieval, and a vector database for semantic retrieval based on the document tree includes: The non-leaf nodes in the document tree are stored as title nodes in the graph database; wherein, the title node contains a title name, a document identifier, and a level depth, and is connected to title nodes at adjacent levels through parent-child relationship edges; the title node is associated with one or more associated text block identifiers to indicate the associated text blocks belonging to the title node; Each of the associated text blocks is stored as an index unit in the inverted index database; wherein, each of the associated text blocks corresponds to an index record in the inverted index database, and the index record includes the associated text block identifier, the original text, the summary text, the title path, and the element type; For each of the associated text blocks, a text embedding vector and a summary embedding vector are generated respectively, and the embedding vectors are bound to the corresponding associated text block identifiers and stored in the vector database; The graph database, inverted index database, and vector database are linked through associated text block identifiers.
10. A data query device, characterized in that, include: The parsing module is used to respond to a received user query statement by parsing the query statement to obtain the query intent; The first determining module is used to determine the query complexity and the evidence dependency structure required for the query based on the query intent. The second determining module is used to determine at least one target database to be queried from multiple databases based on the query complexity and the evidence dependency structure, and to perform a retrieval operation on the at least one target database; wherein, the multiple databases include a vector database for semantic retrieval, a graph database for structural navigation, and an inverted index database for full-text content retrieval; The generation module is configured to perform inference operations corresponding to the evidence dependency structure based on the retrieval results returned from the at least one target database, in order to generate an answer in response to the user query.