Method for extracting data in literature
By combining literature segmentation tools and large language models, an end-to-end material literature data extraction workflow is established, which solves the confusion and mismatch problems of different parameters associated with the same material in the prior art, and improves the efficiency of literature extraction and analysis.
Patent Information
- Application Number
- CN202411299086.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-09-18
AI Technical Summary
The prior art lacks an end-to-end material literature data extraction workflow that can quickly run modified, resulting in confusion and mismatch problems associated with different parameters to the same material, reducing the efficiency of literature extraction and analysis.
By combining literature segmentation tools and large language models, we establish an end-to-end material literature data extraction workflow. First, use the document segmentation tool to divide the document into text and charts to form a structured table. Then, data related to the material data template parameters are extracted using large language models and preset standard text prompt words to form a text data set. Finally, the text information in the graph structured table is matched with the text dataset to form the final document dataset.
The confusion and mismatch problems associated with different parameters to the same material are solved, the efficiency of literature extraction and analysis is improved, and the time requirement for users for model modification and data preprocessing is reduced.
Smart Images

Figure CN120086267A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a method for extracting data from documents. Background Art
[0002] Literature data in the field of materials is an important reference in R & D work such as material design and process optimization. However, specific numerical values, structures, curves and other data are often scattered in the literature and usually stored in unstructured formats such as pdf and xml. It is necessary to extract and organize various parameter data into structured table formats such as xlsx and csv before batch analysis and operation of the data can be carried out.
[0003] Manual reading of documents for information extraction has high accuracy and recall rates, but the cost of personnel professional knowledge training and time-consuming cost are high. Information extraction (IE) is a process of automatically extracting key information from a large number of literature resources relying on advanced technologies such as natural language processing, machine learning, and deep learning. This field has become increasingly important in recent years with the rapid development of artificial intelligence technology. Early information extraction mainly relied on methods such as rule matching and part-of-speech tagging based on predefined rules and manual annotations to identify information in specific formats or patterns, but it was difficult to handle complex and changing natural language texts.
[0004] Relying on artificial intelligence technologies such as convolutional neural networks and recurrent neural networks for information extraction has improved the model's ability to capture key information and accuracy. In recent years, large language models (LLMs) formed by the combination of neural networks and attention mechanisms have the ability to perform unsupervised learning on large-scale texts and then fine-tune for specific tasks, improving the information extraction performance. They show superiority in understanding context and semantic representation, and there are extensive work reports in the academic community.
[0005] However, there is still a lack of complete end-to-end information extraction tools driven by LLMs. The main reasons include: existing work reports focus on the accuracy of IE tools in organized text data sets and ignore the process from document format extraction to text sets; existing work reports rely on commercial non-open-source LLMs (such as GPT4) or dedicated LLMs trained for specific fields and cannot be quickly migrated based on changing domain information requirements; existing work reports lack the ability to extract the association between text and charts. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method for extracting data from documents, which can establish an end-to-end material literature data extraction workflow that can be quickly run and modified, solve the confusion and mismatch problems of different parameters associated with the same material, and improve the efficiency of literature extraction and analysis.
[0007] To solve the above problems, the present invention provides a method for extracting data from documents, including: using a document segmentation tool to segment the text and charts of the document to be extracted according to a preset segmentation method, and extracting the charts to form a text structured table including the segmented text and document information, and a chart structured table including the charts or chart storage paths, the text information of the charts and the document information; using a large language model and a preset standard text prompt to confirm whether the text structured table contains entity materials. If so, according to the parameters in the material data template, using the large language model and the preset standard text prompt to extract data related to the parameters from the text structured table to form a first text data set. The same set of data in the first text data set includes all extraction results related to the same parameter and the document information corresponding to the extraction results. The parameters of different sets of data in the first text data set are different; merging all data with different parameters corresponding to the same material name from the same document in the first text data set to form a second text data set. The same set of data in the second text data set includes all data with different parameters of the same material from the same document. The materials or source documents of different sets of data in the second text data set are different; using the large language model and a preset standard chart prompt to match the text information of the charts in the chart structured table with the data in the second text data set, and forming a final document data set according to the matching results. The same set of data in the document data set includes all text data and charts corresponding to the same material from the same document.
[0008] In one embodiment, the step of using a document segmentation tool to segment the text and charts of the document to be extracted according to a preset segmentation method includes: using the document segmentation tool to segment the document to be extracted into primary text and multiple independent charts according to the preset segmentation method; using the document segmentation tool to further segment the primary text into required text according to the preset segmentation method; the required text serves as the segmented text in the text structured table, and the independent charts serve as the charts in the chart structured table.
[0009] In one embodiment, in the step of using a document segmentation tool to segment the document to be extracted into primary text according to a preset segmentation method, the preset segmentation method corresponds to the format of the document to be extracted.
[0010] In one embodiment, in the step of using a document segmentation tool to further segment the primary text into required text according to a preset segmentation method, the preset segmentation method is one of single sentence, multiple sentences, paragraphs or chapters.
[0011] In one embodiment, the document information includes the document title.
[0012] In one embodiment, the text information of the chart includes the title of the chart and the characters included in the chart.
[0013] In one embodiment, the preset standard text prompt words include: determining whether there is entity material in the segmented text; determining whether there is a parameter in the material data template in the segmented text; determining whether there is a value of the parameter in the segmented text; determining whether there is a unit of the parameter in the segmented text.
[0014] In one embodiment, the material data template includes multiple parameters of the material to be extracted. The steps of extracting data related to the parameters from the text structured table using a large language model and preset standard text prompt words include: using the large language model and preset standard text prompt words to extract data related to each parameter from the text structured table respectively, and taking all the extraction results related to the same parameter and the source literature of the extraction results as a set of data.
[0015] In one embodiment, the extraction results include the material name, the value of the parameter, and the unit of the parameter.
[0016] In one embodiment, the preset standard chart prompt words include: determining whether the text information of the chart describes the material name in the second text dataset.
[0017] In one embodiment, after the steps of forming a text structured table including the segmented text and literature information and a chart structured table including a chart or a chart storage path, the text information of the chart, and the literature information, the following steps are further included: restoring the special format content in the text structured table and the chart structured table, and replacing the original content in the text structured table and the chart structured table.
[0018] The method for extracting data from a document provided by an embodiment of the present invention combines a document segmentation tool with a large language model, establishes an end-to-end material document data extraction workflow that can be quickly run and modified, solves the confusion and mismatch problems of different parameters associated with the same material, and furthermore, performs text extraction and chart extraction independently in steps first, and then performs matching and merging, effectively reducing the accuracy decline and video memory requirements generated by the large language model when processing long sentence sequences, and greatly reducing the time for users to modify the model and pre-process and annotate data during various document extractions, improving the efficiency of document extraction and analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for description in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 is a schematic diagram of the steps of a method for extracting data from a document provided by an embodiment of the present invention;
[0021] Figure 2 is a text structured table and a chart structured table formed according to the document to be extracted, where part (a) is the text structured table and part (b) is the chart structured table;
[0022] Figure 3 is a schematic diagram of a material data template for polishing materials;
[0023] Figure 4 is a schematic diagram of a first text data set formed by the method for extracting data from a document provided by an embodiment of the present invention;
[0024] Figure 5 is a schematic diagram of a second text data set formed by the method for extracting data from a document provided by an embodiment of the present invention;
[0025] Figure 6 is a schematic diagram of the final document data set formed by the method for extracting data from a document provided by an embodiment of the present invention. Detailed implementation manners
[0026] The following makes a detailed description of the specific implementation manners of the method for extracting data from a document provided by the present invention with reference to the drawings.
[0027] Figure 1 is a schematic diagram of the steps of a method for extracting data from a document provided by an embodiment of the present invention. Please refer to Figure 1 , the method for extracting data from a document includes the following steps:
[0028] Step S10, using a document segmentation tool to segment the text and charts of the document to be extracted according to a preset segmentation method, and extract the charts, forming a text structured table including the segmented text and document information, and a chart structured table including the charts or the storage paths of the charts, the text information of the charts, and the document information. This step is used to form a text structured table and a chart structured table for the text and charts of the document to be extracted respectively.
[0029] In some embodiments, before performing step S10, the original file of the document to be extracted is first stored in a folder at a specified path. The original file of the document to be extracted can be a file in pdf format or a document format such as word or txt.
[0030] In some embodiments, the document segmentation tool uses a combination of a text segmentation tool and a graphic optical recognition tool, and a multi-model integration tool to separately extract and organize text and charts. This type of tool includes, but is not limited to, programs, toolkits or software such as scipdfparser, grobid, openparse, llamaparse, and unstructured.
[0031] In some embodiments, the charts include pictures and tables, the document information includes the document title, the chart structured table includes the chart or the storage path of the chart, and the text information of the chart includes the title of the chart and the characters included in the chart.
[0032] In some embodiments, the steps of using a document segmentation tool to segment the text and charts of the document to be extracted according to a preset segmentation method include:
[0033] Using a document segmentation tool to segment the document to be extracted into primary text and multiple independent charts according to a preset segmentation method. In this step, the preset segmentation method corresponds to the format of the document to be extracted. The independent charts serve as the charts in the chart structured table in step S10. For example, if the format of the document to be extracted includes a title, an abstract, main text section 1, main text section 2, references, acknowledgments, citations, Figure 1 , Figure 2 , Table 1, Table 2, etc., then in this step, the document to be extracted is segmented according to the segmentation method of the title, abstract, main text section 1, main text section 2, references, acknowledgments, citations, Figure 1 , Figure 2 , Table 1, Table 2, etc. Among them, each of the text parts such as the title, abstract, main text section 1, main text section 2, references, acknowledgments, and citations is used as a primary text, Figure 1 , Figure 2 , Table 1, Table 2, etc. Each of the chart parts is used as a chart.
[0034] Use a document segmentation tool to further segment the primary text into the required text according to a preset segmentation method. In some embodiments, the preset segmentation method is one of single sentence, multiple sentences, paragraphs, or chapters. The required text serves as the segmented text in the text structure table described in step S10. In this step, the primary text is further segmented into smaller texts. When the length of a single text is less than a certain value specified in advance, this text is merged with the next segmented text, and this text serves as the segmented text in the text structure table described in step S10 to further improve the accuracy of data extraction. For example, use a document segmentation tool to segment the primary text into the required text in the way of three sentences.
[0035]
[0036] Represents the (i + 1)-th primary text obtained only through the segmentation tool in each document, T i Represents the i-th final segmented text obtained by splicing and cropping the primary text in each document, T i+1 Represents the (i + 1)-th final segmented text obtained by splicing and cropping the primary text in each document, len(T i ) represents the string length of the i-th final segmented text, l tool Represents the length threshold for merging with the next segmented text (varies according to the text segmentation tool and the number of documents, specified in advance by humans).
[0037] In some embodiments, the text structure table includes the segmented text, the paragraph where the text is located, and document information; the chart structure table includes the chart or the chart storage path, the title of the chart, the characters in the chart, and the document information.
[0038] The method for extracting data from documents provided by the embodiments of the present invention can segment the text according to the length required by the user, thereby avoiding the problems of poor user experience and insufficient flexibility caused by a fixed segmentation length, and effectively avoiding the decrease in accuracy and the video memory requirement when the large language model processes long sentence sequences in subsequent steps. At the same time, the method for extracting data from documents provided by the embodiments of the present invention uses a document segmentation tool to segment the document, reducing the time for users to perform pre-processing such as document segmentation and picture export by themselves, and greatly improving the efficiency of document extraction.
[0039] Illustrate with examples, Figure 2 is the text structure table and the chart structure table formed according to the document to be extracted, where part (a) is the text structure table and part (b) is the chart structure table. Please refer to Figure 2, the text structured table includes multiple groups of data, each group of data includes the text to be extracted (i.e., the segmented text), the paragraph where it is located, and the literature title. The chart structured table includes multiple groups of data, each group of data includes the chart storage path, the chart title, and the literature title. In Figure 2 texts and pictures from two pieces of literature are schematically listed.
[0040] Please continue to refer to Figure 1 , step S11, use a large language model and a preset standard text prompt to confirm whether there is entity material in the text structured table. If so, according to the parameters in the material data template, use the large language model and the preset standard text prompt to extract data related to the parameters from the text structured table to form a first text data set. The same group of data in the first text data set includes all extraction results related to the same parameter and the literature information corresponding to the extraction results. The parameters of different groups of data in the first text data set are different. This step is used to extract the segmented text in the text structured table and form a first text data set.
[0041] The large language model (Large Language Model, abbreviated as LLM) refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. The large language model can handle various natural language tasks, such as text classification, question answering, dialogue, etc., and is an important way to artificial intelligence. The large language model includes but is not limited to large models and large model architectures such as llama3, llama2, qwen1.5, Yi34, etc.
[0042] The standard text prompt is a series of questions in the large language model, including questions about the text. In the method for extracting data from literature provided in an embodiment of the present invention, the standard text prompt can be set in advance according to the literature type, the type of material to be extracted, and the user's needs to meet personalized requirements. In some embodiments, the preset standard text prompt includes: determining whether there is entity material in the segmented text; determining whether there is a parameter in the material data template in the segmented text; determining whether there is a value of the parameter in the segmented text; determining whether there is a unit of the parameter in the segmented text.
[0043] The material data template refers to a data list containing various parameters of a certain type or certain types of materials. For example, if the user needs to obtain data on polishing materials, the material data template is a data list of various parameters of polishing materials. For example, the material data template is a set containing parameters such as viscosity and diffusion coefficient. In some embodiments, the material data template is pre-designed by domain engineers based on empirical knowledge and covers as many potential parameters as possible.
[0044] In some embodiments, the material data template includes multiple parameters of the material to be extracted. The steps of using a large language model and a preset standard text prompt to extract data related to the parameters from the text structured table include: using the large language model and the preset standard text prompt to respectively extract data related to each parameter from the text structured table, and taking all the extraction results related to the same parameter and the source literature of the extraction results as a set of data. Multiple sets of data constitute the first text data set.
[0045] In some embodiments, the extraction results include the material name, the value of the parameter, and the unit of the parameter. In other embodiments, the preset standard text prompt can also be set according to the user's needs to obtain different extraction results and meet personalized requirements.
[0046] In some embodiments, in this step, use a large language model and a preset standard text prompt to confirm Figure 2 whether the text to be extracted in the shown text structured table contains an entity material. For example, use a large language model and a preset standard text prompt "Does the original text in the following sentence contain an entity material" to confirm Figure 2 whether the text to be extracted in the shown text structured table contains an entity material. If so, output the material names of all the entity materials in the text and take it as one of the extraction results, so as to determine the material name; after confirming Figure 2 that the text to be extracted in the shown text structured table contains an entity material, use a large language model and a preset standard text prompt "Does the original text in the following sentence contain parameter A", "Does the original text in the following sentence contain the value of parameter A", "Does the original text in the following sentence contain the unit of parameter A". If so, the large language model continues to provide commands "Use the original text in the following sentence to output parameter A", "Use the original text in the following sentence to output the value of parameter A", "Use the original text in the following sentence to output the unit of parameter A", so as to be able to extract data related to the parameter from the text structured table to form the first text data set.
[0047] For example, Figure 3 is a schematic diagram of the material data template of the polishing material, Figure 4 is a schematic diagram of the first text data set formed by the method for extracting data from a document provided in an embodiment of the present invention. Please refer to Figure 3 and Figure 4 In Figure 3The material data template for medium polishing materials only schematically gives two parameters: viscosity and diffusion coefficient. In this step, use a large language model and preset standard text prompts to confirm Figure 2 whether the text to be extracted in the text structured table shown contains solid materials. If so, output the material names of all solid materials in the text to be extracted and use them as one of the extraction results, "Does the original text in the following sentence contain viscosity?", "Does the original text in the following sentence contain the value of viscosity?", "Does the original text in the following sentence contain the unit of viscosity?". If so, the large language model continues to provide commands "Use the original text in the following sentence to output the value of viscosity" and "Use the original text in the following sentence to output the unit of viscosity". Multiple parameters are executed in the above method in sequence, so as to extract data related to the parameters from the text structured table. In Figure 4 Figure (a) shows a set of data, and figure (b) shows another set of data. The parameters corresponding to the two sets of data are different. Among them, part (a) is the data of viscosity, and part (b) is the data of diffusion coefficient.
[0048] Please continue to refer to Figure 1 , step S12, merge all the data of different parameters corresponding to the same material from the same document in the first text dataset to form a second text dataset. The same set of data in the second text dataset includes all the data of different parameters corresponding to the same material from the same document. The materials or source documents of different sets of data in the second text dataset are different. For example, please refer to Figure 5 , Figure 5 is a schematic diagram of the second text dataset formed by the method for extracting data from a document provided in an embodiment of the present invention. Merge Figure 4 all the data of different parameters corresponding to the same material from the same document in to form a second text dataset. The method for extracting data from a document provided in an embodiment of the present invention extracts each parameter separately and then merges the data. On the one hand, it reduces the information interference and confusion between different parameters, and on the other hand, it reduces the limitation on the segmentation length of the document and the consumption of video memory during program operation.
[0049] Total video memory for large language model inference = G 模型 +G KV Cache +G 中间激活值 (2)
[0050] where G 模型 is the video memory required to load the large language model and is related to the type of large language model selected. G KVCache is the video memory required for operations during the inference process of the large language model and is related to the length of the input text;
[0051] G 中间激活值The video memory required to store intermediate activation values during the inference process of the large model.
[0052] G KV Cache = B * (2 * l * h * d * (s + n) * sizeof(dtype)) (3)
[0053] Where s is the total text length of the input large language model, n is the text length of the data output by the large language model, and B, l, h, d are fixed parameters related to the large language model selection, representing batch size, the number of decoder layers, the number of attention heads per layer, and the single-head dimension respectively. sizeof(dtype) is the number of bytes for storing the data type.
[0054] s = s prompt + s 分割文本 (4)
[0055] Where s prompt is the total length of the prompt words, which is a fixed value during the operation of the large language model; s 分割文本 is the length of the segmented text. It can be seen that the embodiments of the present invention adopt the method of extracting each parameter separately to reduce s 分割文本 , thereby reducing G (KV Cache) , and finally reducing the total video memory occupied by the large language model inference and reducing the requirements for the operating device.
[0056] Step S13: Use the large language model and the preset standard chart prompt words to match the text information of the chart in the chart structured table with the data in the second text dataset, and form a final literature dataset according to the matching results. The same group of data in the literature dataset includes all text data and charts corresponding to the same material from the same literature. This step is used to combine the content of the chart structured table with the data in the second text dataset to form a final literature dataset. The literature dataset can be cleaned, labeled, and analyzed as needed. The literature dataset can be a file in excel or csv format, where the charts can be stored in the form of embedded tables or in an orderly manner in the form of folders.
[0057] Standard chart prompts are a series of questions in large language models, including questions about charts. In the method for extracting data from documents provided in an embodiment of the present invention, the standard chart prompts can be preset according to the document type, the type of material to be extracted, and the user's needs to meet personalized requirements. In some embodiments, the preset standard chart prompts include: determining whether the text information of the chart describes the material name in the second text dataset. If the text information of the chart describes the material name in the second text dataset, the chart from the same document is merged with the second text dataset to form a final document dataset.
[0058] For example, Figure 6 is a schematic diagram of the final document dataset formed by the method for extracting data from documents provided in an embodiment of the present invention. Please refer to Figure 6 , using a large language model and preset standard chart prompts to Figure 1 structure the chart in the structured table. Match the text information of the chart in Figure 5 with the material names in the second text dataset in
[0059] and form a final document dataset according to the matching results. The same set of data in the document dataset includes all the text data and charts corresponding to the same material from the same document.
[0060] The method for extracting data from documents provided in an embodiment of the present invention extracts the text and charts of the document to be extracted respectively and then performs matching and merging, effectively solving the problems of decreased accuracy and video memory requirements when the large language model processes long sentence sequences, and avoiding the decreased accuracy caused by image recognition and annotation.
[0061] The method for extracting data from documents provided in an embodiment of the present invention establishes an end-to-end material document data extraction workflow that can be quickly run and modified, solves the confusion and mismatch problems of different parameters associated with the same material, matches the picture information with the parameter information, greatly reduces the time for users to modify the model and preprocess the data annotation when extracting various documents, and improves the document extraction and analysis efficiency.
[0062] In some embodiments, when special formats such as material structural formulas, formulas, and scientific notations are extracted, they will be deformed and unable to maintain the original format. After segmentation, the extraction method of the present invention also performs program processing on special formats such as material structural formulas, formulas, and scientific notations, so that the special formats in the text presented in the text structured table can maintain the original format. Specifically, after the steps of forming a text structured table including the segmented text and document information and a chart structured table including charts or chart storage paths, the text information of the charts, and the document information, the following steps are further included: restoring the special format content in the text structured table and the chart structured table, and replacing the original content in the text structured table and the chart structured table. Among them, the restoration of the special format content in the text structured table and the chart structured table can adopt existing methods, such as relying on regular expressions and verification functions. Compare the literature extraction accuracy and recall rate indicators before and after introducing the processing tool, and modify the special format processing tool.
[0063] P 准确率 = TP / (TP + FP) (5)
[0064] R 召回率 = TP / (TP + FN) (6)
[0065] Where P 准确率 is the literature extraction accuracy, and R 召回率 is the recall rate of literature extraction. Accuracy and recall rate are two indicators for evaluating the literature extraction effect. TP is the number of parameters correctly extracted by the large language model, FP is the number of parameters extracted by the large language model but inconsistent with the original text, and FN is the number of parameters that should be extracted but not extracted by the large language model.
[0066] It should be noted that the terms "including" and "having" and their variants in the documents of the present invention are intended to cover non-exclusive inclusion. The terms "first", "second", etc. are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. Unless the context clearly indicates otherwise, it should be understood that the data used in this way can be interchanged under appropriate circumstances. The term "one or more" depends at least in part on the context and can be used to describe a feature, structure, or property in the singular sense, or can be used to describe a combination of features, structures, or features in the plural sense. The term "based on" can be understood as not necessarily intended to express a set of exclusive factors, but instead, also at least in part depending on the context, allows for the existence of other factors that are not necessarily explicitly described. Additionally, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. Furthermore, in the above description, the description of well-known components and technologies is omitted to avoid unnecessarily confusing the concepts of the present invention. In each of the above embodiments, the key point of each embodiment is to illustrate the differences from other embodiments. For the same / similar parts among the various embodiments, reference can be made to each other.
[0067] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for extracting data from a document, characterized in that: include: Use a document segmentation tool to segment the text and charts of the document to be extracted according to a preset segmentation method, and extract the charts to form a text structured table including the segmented text and document information and a chart structured table including the chart or chart storage path, the text information of the chart and the document information; Using a large language model and preset standard text prompt words to confirm whether the text structured table contains physical materials, if yes, then according to the parameters in the material data template, using the large language model and preset standard text prompt words to extract data related to the parameters from the text structured table to form a first text data set, wherein the same group of data in the first text data set includes all extraction results related to the same parameter and the literature information corresponding to the extraction results, and different groups of data in the first text data set have different parameters; Merging all data of different parameters corresponding to the same material name from the same document in the first text data set to form a second text data set, wherein the same set of data in the second text data set includes all data of different parameters of the same material from the same document, and different sets of data in the second text data set have different materials or source documents; The text information of the chart in the chart structured table is matched with the data in the second text data set using a large language model and preset standard chart prompt words, and a final document data set is formed according to the matching results, wherein the same set of data in the document data set includes all text data and charts corresponding to the same material from the same document.
2. The method for extracting data from a document according to claim 1, characterized in that: The steps of using the document segmentation tool to segment the text and charts of the document to be extracted according to the preset segmentation method include: Use the document segmentation tool to segment the document to be extracted into primary text and multiple independent charts according to the preset segmentation method; Use the document segmentation tool to segment the primary text into required texts according to the preset segmentation method; The required text is used as the segmented text in the text structured table, and the independent chart is used as the chart in the chart structured table.
3. The method for extracting data from a document according to claim 2, characterized in that: In the step of using a document segmentation tool to segment the document to be extracted into primary texts according to a preset segmentation method, the preset segmentation method corresponds to the format of the document to be extracted.
4. The method for extracting data from a document according to claim 2, characterized in that: In the step of using a document segmentation tool to segment the primary text into required texts according to a preset segmentation method, the preset segmentation method is one of single sentence, multiple sentences, paragraphs or chapters.
5. The method for extracting data from a document according to claim 1, characterized in that: The document information includes the document title.
6. The method for extracting data from a document according to claim 1, characterized in that: The text information of the chart includes the title of the chart and the characters included in the chart.
7. The method for extracting data from a document according to claim 1, characterized in that: The preset standard text prompt words include: judging whether the segmented text contains physical materials; judging whether the segmented text contains parameters in the material data template; judging whether the segmented text contains the numerical value of the parameter; judging whether the segmented text contains the unit of the parameter.
8. The method for extracting data from a document according to claim 1, characterized in that: The material data template includes multiple parameters of the material to be extracted, and the step of extracting data related to the parameters from the text structured table using a large language model and preset standard text prompt words includes: using the large language model and preset standard text prompt words to extract data related to each parameter from the text structured table respectively, and taking all extraction results related to the same parameter and the source documents of the extraction results as a group of data.
9. The method for extracting data from a document according to claim 1, characterized in that: The extraction result includes the material name, the value of the parameter and the unit of the parameter.
10. The method for extracting data from a document according to claim 1, characterized in that: The preset standard chart prompt words include: judging whether the text information of the chart describes the material name in the second text data set.
11. The method for extracting data from a document according to claim 1, characterized in that: After the steps of forming a text structured table including the segmented text and document information and a chart structured table including a chart or a chart storage path, the text information of the chart and the document information, the following steps are also included: restoring the special format content in the text structured table and the chart structured table, and replacing the original content in the text structured table and the chart structured table.
Citation Information
Patent Citations
Low-cost bulletin data extraction method driven by large language model
CN117592470A
Literature chart extraction and classification method and system, computer equipment and storage medium
CN118135582A
Multimode vertical domain knowledge question-answering method for visually impaired people based on shortest editing distance
CN118606459A
Multi-lingual natural language generation
US20240127008A1
System and methods for safe, scalable, artificial general intelligence (AGI)
WO2024182285A2