Large and small model combined unstructured document extraction method
The Doctopus method solves the cost and quality problems in the conversion of unstructured documents to structured data through index construction and dynamic programming algorithms, combining the advantages of non-large language models and large language models, and realizes efficient and accurate attribute extraction under budget constraints.
Patent Information
- Application Number
- CN202510692663.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-29
AI Technical Summary
In the process of unstructured documents to structured data conversion, the prior art has problems of high cost and unstable extraction quality, especially when using large language models, it is difficult to optimize the extraction quality in different attribute extraction scenarios.
The Doctopus method is adopted to optimize document fragment retrieval and policy selection through index construction, verification set evaluation and dynamic programming algorithms, combining the low cost of non-large language models and the high accuracy of large language models, and optimize document fragment retrieval and strategy selection to achieve efficient and accurate attribute extraction.
Given a budget, the extraction efficiency and accuracy of unstructured documents to structured data is significantly improved, the cost of using large language models is reduced, and the efficiency and accuracy of the extraction strategy is ensured.
Smart Images

Figure CN120386864A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data retrieval, and more particularly, to an unstructured document extraction method combining large and small models. Background Art
[0002] Modern enterprises have a large amount of unstructured data, accounting for 80%-90% of all data. To gain valuable insights, many applications typically need to convert a collection of unstructured documents into structured relational tables and then run analytical SQL queries or machine learning predictions. However, the complexity and diversity of documents pose significant challenges in the process of extracting the attribute values of these tables. In Figure 1 this case, the user's goal is to extract attributes from the painter's documents, such as their personal information and artistic genre (Genre), which can be solved by different methods, such as Open Information Extraction (OpenIE), Pre-trained Language Models (PLMs), Code Generation (Codegen), and Large Language Models (LLMs). However, the performance of these methods varies significantly in different scenarios. For example, generating code through LLMs is superior to PLMs when extracting the "Lifespan" attribute because it requires logical reasoning (i.e., Lifespan equals death_date minus birth_date), which is relatively easy for code but difficult for PLM methods with limited reasoning ability. However, for the "Genre" attribute, for example, in the text: "Brull's early works were realistic, but later he focused on symbolism." The code may misidentify "Realistic" as the value. And the PLM-based method can perform well due to its complex semantic understanding ability. Therefore, none of the above methods can always outperform others. On the other hand, large language models LLMs perform well in converting unstructured documents into structured formats, but they are costly when dealing with a large number of documents containing a large amount of data.
[0003] In fact, it is not necessary to input the entire document into LLMs. Instead, only inputting the fragments related to the attributes can significantly reduce the cost without sacrificing the quality significantly. In some cases, for example, when extracting the value of "Lifespan", using a cheap method (such as Codegen) may be equivalent to the expensive LLMs. Then, a natural question is: when considering multiple possible attribute extraction strategies, how to optimize the extraction quality under the cost constraints specified by the user? To solve this problem, two challenges need to be faced: (1) Designing an effective fragment retrieval strategy is very important because missing relevant fragments will seriously affect the performance of LLMs; (2) Evaluating the extraction quality of various strategies, including different input fragment strategies based on LLMs, is crucial for selecting the appropriate strategy.
[0004] In response to this, the present invention proposes a novel method - Doctopus, which optimizes the extraction accuracy (i.e., quality) under the constraint of a cost budget (i.e., the number of tokens consumed by LLMs). Generally speaking, Doctopus first constructs an index on document fragments, estimates the accuracy of different strategies, and assigns the most suitable strategy and input fragments for extracting specific attributes of the document. Summary of the Invention
[0005] In view of this, the present invention proposes a method for extracting unstructured documents by combining large and small models. In the presence of a given budget, through index construction, quality assessment, and dynamic multi-strategy selection, an efficient method for extracting unstructured document data is realized to solve the problems existing in the above-mentioned prior art.
[0006] To achieve the above object, the present invention proposes a method for extracting unstructured documents by combining large and small models, which is characterized by including:
[0007] Obtaining document attribute values based on an attribute-enhanced retrieval method;
[0008] Estimating the accuracy rate of different extraction strategies when extracting different document attribute values based on a validation set;
[0009] Based on the dynamic programming algorithm, selecting the most suitable extraction strategy for each document attribute value within the cost budget.
[0010] Further, the process of obtaining document attribute values based on the attribute-enhanced retrieval method includes:
[0011] Dividing the document into semantically coherent text fragments and constructing a high-dimensional vector index for the text fragments;
[0012] Based on the validation set, obtaining reference sentences that can extract attributes, using the summary of the reference sentences as search keywords for enhanced query, obtaining relevant text fragments and inputting them into a large language model to obtain the document attribute values.
[0013] Further, the process of estimating the accuracy rate of different strategies when extracting different attributes based on the validation set includes:
[0014] Using the existence probability to represent the possibility that an attribute exists in the input text, and using the extraction probability to represent the possibility of accurately extracting the attribute value using a certain strategy, to construct the overall accuracy rate.
[0015] Further, the overall accuracy rate is as follows:
[0016]
[0017] Among them, represents the overall accuracy rate, p(a∈t) represents the existence probability, represents the extraction probability, represents the true value of the j-th attribute extracted from document d i whereas represents the value extracted by policy s ∈ S, and a represents the attribute value.
[0018] Furthermore, the process of obtaining the existence probability includes:
[0019] Obtaining training data based on the validation set, fine-tuning the pre-trained language model using the training data, and inputting the augmented queries and relevant text fragments related to the attribute values into the fine-tuned pre-trained language model to obtain the existence probability.
[0020] Furthermore, the process of calculating the extraction probability includes:
[0021]
[0022] where represents the true value of the j-th attribute extracted from document d i whereas represents the value extracted by policy s ∈ S, a represents the attribute value, and D v represents the validation set.
[0023] Furthermore, the process of selecting the most suitable extraction policy for each attribute within the cost budget includes:
[0024] Calculating the consumption budget of each extraction policy and adjusting the remaining budget for the subsequent extraction process. The iterative process is as follows:
[0025]
[0026] where dp[g][b] represents the maximum accuracy achievable when considering the first g attributes with a budget b ≤ B, represents the number of input tokens required to extract the attribute value v ij using policy s, B is the given total budget, represents the true value of the j-th attribute extracted from document d i whereas represents the value extracted by policy s ∈ S.
[0027] Furthermore, the process of selecting the most suitable extraction policy for each attribute within the cost budget also includes:
[0028] Taking accuracy as the budget to minimize the cost, which is specifically as follows:
[0029]
[0030] where B is the given total budget, Denote the true value of the j-th attribute extracted from document d i among them, Denote the value extracted by policy s ∈ S, Denote the number of input tokens required to extract the attribute value v using policy s ij f(v ij ) = s indicates that each v ij is extracted by a unique policy.
[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0032] The method for extracting unstructured documents by combining large and small models proposed by the present invention can accurately and efficiently extract structured data from a large number of unstructured documents, and can adjust the policy selection scheme that meets the requirements according to the budget situation. By combining the advantages of low cost of non-large language models and high accuracy of large language models, Doctopus significantly improves the efficiency of the extraction policy in dealing with large data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered as limiting the present invention. In the drawings:
[0034] Figure 1 is a schematic diagram of attribute extraction by combining large and small models according to the present invention;
[0035] Figure 2 is a schematic diagram of the processing flow framework of the Doctopus method described in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. Hereinafter, the present invention will be described in detail with reference to the drawings and in combination with the embodiments.
[0037] This embodiment proposes an unstructured document extraction method that combines large and small models - Doctopus. This method aims to perform high-quality attribute extraction from unstructured documents under the cost constraints specified by the user. First, Doctopus adopts an index-based method to efficiently identify and process relevant text fragments, thereby reducing the usage cost of large models. Then, with the help of a validation set and either human effort or large language models, this method further evaluates the quality of each extraction method when extracting various attributes. Finally, based on the cost and estimated quality, Doctopus selects the optimal extraction strategy through a dynamic programming algorithm to ensure maximizing the accuracy of extracted attributes while meeting the user's cost budget. The specific steps are as follows:
[0038] (1) Index-based attribute extraction while enhancing the information content of retrieval;
[0039] (2) Quality estimation of multiple extraction strategies;
[0040] (3) Budget-based optimization of attribute extraction, dynamically selecting the most accurate and cost-effective combination of extraction strategies.
[0041] Combined with Figure 1 、 Figure 2 , the specific implementation process of the present invention is elaborated as follows:
[0042] Problem definition: Given a set of documents D = {d i | i ∈ [1, n]} and the attributes A = {a j | j ∈ [1, m]} specified by the user, the goal of this embodiment is to generate a table T containing n tuples and m attributes, where each tuple is extracted from the document according to the attribute set A. Use v ij to represent the value of the j-th attribute extracted from the document d i .
[0043] Quality: The accuracy is used to measure the extraction quality. Formally, this embodiment uses to represent the true value of the j-th attribute extracted from the document d i . Then, the overall accuracy is expressed as follows:
[0044]
[0045] where 1{·} represents an indicator function that returns 1 if its argument is true; otherwise, it returns 0. Although the overall accuracy requires knowing the true value in advance, in real-world scenarios, this is usually not available. Therefore, this embodiment uses the expected accuracy to measure the extraction accuracy, as follows:
[0046]
[0047] Among them, represents the probability of correctly extracting a specific attribute v from the document. ij of.
[0048] Problem of this embodiment: Use the set S to represent possible extraction strategy choices. The value extracted by the strategy s ∈ S is denoted as Similarly, represents the number of input tokens required to extract v using the strategy s. Then, use a mapping function ij such that for each v , there exists a strategy s satisfying f(v ij ) = s, which indicates that each v ij is extracted by a unique strategy. For LLM-based strategies, using different texts as the input to the LLM to extract the same attribute can be regarded as different strategies. So the optimization goal of this embodiment is expressed as: ij
[0049]
[0050] where B is the given total budget. Next are the specific operation steps:
[0051] Step (1): Index-based attribute extraction. Doctopus indexes relevant text fragments to achieve efficient and effective retrieval and only inputs these fragments into the LLMs. This RAG (Retrieval-Augmented Generation) style method reduces the number of input tokens by segmenting and embedding semantically coherent text fragments while maintaining high accuracy. Considering that accurate retrieval also depends on whether the query is informative enough, Doctopus also enhances the informativeness of the query by generating attribute-related reference sentences from the sampled documents. This step mainly includes the following contents:
[0052] Step (1-1): Index construction. Doctopus first uses existing NLP tools (such as the SemanticChunker function in LangChain) to automatically divide each document into semantically coherent text fragments to ensure that each attribute is likely to be extracted from a single fragment. Subsequently, these fragments are encoded into fixed-length embedding vectors by the E5Model and organized into a high-dimensional vector index, enabling fast retrieval of relevant fragments based on the semantic similarity between the attribute and the fragment.
[0053] Step (1-2): Retrieval based on attribute enhancement. Accurate retrieval not only depends on the semantic representation of text fragments but also on whether the query (i.e., the attribute) is informative enough. However, the embedding vectors of relevant fragments are not necessarily similar to those of the attribute. To address this issue, this embodiment proposes to enhance the query. Specifically, reference sentences that can extract the attribute are identified, and the query is enhanced by using the summary of these reference sentences as search keywords, thereby achieving more accurate retrieval. This identification can be achieved by analyzing a validation set (some sampled documents) constructed by LLMs or manually.
[0054] Step (2): Quality estimation. In this part, how to estimate the accuracy of different strategies in extracting different attributes based on the validation set is discussed. As discussed in Step 1, the accuracy of the LLM-based strategy depends on two probability factors: (1) the probability of correctly identifying whether a certain attribute value exists in a given text block; (2) if it exists, the probability of correctly extracting the value.
[0055] Formally, we use the existence probability p(a∈t) to represent the possibility that the attribute a exists in the input text t (i.e., one or more text blocks). For non-LLM strategies, since they do not need to consider token consumption, the entire document can be used as input. For LLM-based strategies, since the goal of this embodiment is to save costs, usually only relevant text blocks are input into the LLM. It should be noted that using relevant text t from different parts of a single document as input can be regarded as different strategies. Then, the extraction probability is used to represent the possibility of accurately extracting the attribute value using strategy s when the input text t contains the attribute a. Therefore, in Figure 1 , the overall accuracy can be expressed as the product of them:
[0056]
[0057] It should be noted that in the formula in the problem definition part is actually equivalent to because the input text t can be regarded as part of strategy s. Next, we will discuss in detail how to calculate the existence probability and the extraction probability.
[0058] Step (2-1): Existence probability calculation. The core idea is to fine-tune a pre-trained language model, which takes an enhanced query related to attribute a and several text fragments t as input and predicts p(a∈t). First, prepare some query-text pairs from the validation set as training data. Since different fragments in the document can generate a large number of combinations, sufficient labeled training data can be obtained. Then, this embodiment uses this data to fine-tune the T5-base model. Next, model inference is performed. Given a document and an attribute a, first retrieve the top k relevant fragments according to the embedding similarity between the enhanced attribute and the fragments. In fact, the value of attribute a may be extracted from a single fragment or multiple fragments. Therefore, consider the combinations in these k fragments as different input texts t, and call the fine-tuned model to predict their existence probabilities. Then, if an LLM-based strategy is adopted, the best combination (with the most appropriate number of tokens and existence probability) will be selected as the input text subsequently.
[0059] Step (2-2): Extraction probability calculation. Use a simple statistical method to calculate the extraction probability. Apply all strategies to extract the specified attribute value on the documents in the validation set D v and then compare these extracted values with the true
[0060] values. For each attribute a and strategy s, the formula for calculating the extraction probability is:
[0061]
[0062] Step (3): Budget-aware optimization. After obtaining the estimated accuracy of each strategy, the last step is to select the most appropriate strategy for each attribute within the cost budget, which is proven to be an NP-hard problem. To solve this problem, Doctopus uses a dynamic programming algorithm that can optimally allocate strategies given the budget. Each strategy calculates the budget consumed and adjusts the remaining budget for subsequent extractions, thus ensuring the highest extraction accuracy within the current budget. Specifically, this method defines a table dp, where dp[g][b] represents the maximum accuracy that can be achieved when considering the first g attributes under the budget b≤B. The iterative process can be described as follows: In addition, for another practical scenario, that is, minimizing the cost with accuracy as the budget, it is obtained that:
[0063]
[0064] Doctopus can handle such tasks in a similar way, only need to adjust the dynamic programming table to update the minimum cost at all times.
[0065] The overall process of the method of the present invention is as Figure 2As shown, the present invention divides, embeds, and maps the documents in the original dataset into a high-dimensional vector space to construct an index, and then mines reference sentences related to attributes to enhance the information content of subsequent queries. After evaluating the extraction quality of each strategy, we use the dynamic programming algorithm to update the maximum extraction accuracy under a given budget, and finally obtain the optimal combination of extraction strategy solutions, ensuring the accuracy and effectiveness of the extraction process.
[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for extracting unstructured documents by combining large and small models, characterized in that, Including: Obtaining document attribute values through an attribute-enhanced retrieval method; Estimating the accuracy rates of different extraction strategies for extracting different document attribute values based on a validation set; Based on the dynamic programming algorithm, selecting the most suitable extraction strategy for each document attribute value within the cost budget.
2. The method for extracting unstructured documents by combining size models according to claim 1, characterized in that The process of obtaining document attribute values through the attribute-enhanced retrieval method includes: Dividing the document into semantically coherent text segments and constructing a high-dimensional vector index for the text segments; Obtaining reference sentences that can extract attributes based on the validation set, using the summary of the reference sentences as search keywords to enhance the query, obtaining relevant text segments and inputting them into a large language model to obtain the document attribute values.
3. The method for extracting unstructured documents by combining size models according to claim 1, characterized in that The process of estimating the accuracy rates of different strategies for extracting different attributes based on the validation set includes: Using the existence probability to represent the possibility of an attribute existing in the input text, and using the extraction probability to represent the possibility of accurately extracting the attribute value using a certain strategy, and constructing the overall accuracy rate.
4. The method for extracting unstructured documents by combining size models according to claim 3, characterized in that, The overall accuracy rate is as follows: Among them, represents the overall accuracy rate, and p(a∈t) represents the existence probability. represents the extraction probability. represents the j-th true value of the attribute extracted from the document d i and represents the value extracted by the strategy s∈S, where a represents the attribute value. 5. The unstructured document extraction method for combining size models according to claim 3, characterized in that The process of obtaining the existence probability includes: Obtaining training data based on the validation set, fine-tuning a pre-trained language model using the training data, and inputting the enhanced query and relevant text segments related to the attribute value into the fine-tuned pre-trained language model to obtain the existence probability.
6. The method for extracting unstructured documents by combining size models according to claim 3, characterized in that, The calculation process of the extraction probability includes: Among them, represents the true value of the j-th attribute extracted from document d i , represents the value extracted by policy s ∈ S, a represents the attribute value, and D v represents the validation set.
7. The method for extracting unstructured documents by combining size models according to claim 1, wherein The process of selecting the most suitable extraction strategy for each attribute within the cost budget includes: Calculating the consumption budget of each extraction strategy and adjusting the remaining budget for the subsequent extraction process. The iterative process is as follows: Among them, dp[g][b] represents the maximum accuracy that can be achieved when considering the first g attributes under the budget b ≤ B. Indicates extracting the attribute value v using the strategy s ij The number of input tokens required, and B is the given total budget. Indicates the j-th attribute extracted from the document d i The true value of Indicates the value extracted by the strategy s ∈ S.
8. The method for extracting unstructured documents by combining size models according to claim 1, characterized in that The process of selecting the most suitable extraction strategy for each attribute within the cost budget also includes: Taking accuracy as the budget to minimize the cost, specifically as follows: where B is the given total budget, denotes the true value of the j-th attribute extracted from document d i and denotes the value extracted by policy s ∈ S, and ij denotes the number of input tokens required to extract the attribute value v using policy s, f(v ij ) = s indicates that each v ij is extracted by a unique policy.