Information extraction and integration method based on RAG technology

Through the combination of RAG technology and large models, the automation of information extraction and integration is achieved, solving the problems of time-consuming, high cost and insufficient accuracy in traditional methods, and improving efficiency and accuracy.

CN120257950APending Publication Date: 2025-07-04SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510289355.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art consumes time, has high labor costs, manual operations are prone to errors and subjective judgments in the process of information extraction and fusion, resulting in low efficiency and insufficient accuracy.

Method used

RAG technology is used to standardize the split of documents in various formats, use large models to search and fill information, and automatically complete information extraction and integration through prompt word templates and RAG system.

Benefits of technology

The full process automation of information extraction and integration has been realized, which greatly improves efficiency, reduces manpower and material costs, and improves accuracy and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257950A_ABST
    Figure CN120257950A_ABST
Patent Text Reader

Abstract

The invention discloses an information extraction and integration method based on the RAG technology, and belongs to the field of large model application, and the method comprises the following steps: carrying out standardized splitting on documents in various formats to obtain a plurality of standard documents; inputting the plurality of standard documents into an information retrieval table filling pipeline to obtain a table with partial filled contents; and after carrying out cue word retrieval and filling on the information which is not filled, obtaining complete content output in combination with the table of which part of content is filled. Compared with manual filling, the judgment basis is more accurate and unified, the error rate and the omission rate are reduced, and the obtained results are convenient to integrate and be used for subsequent automatic processing work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large model applications, and particularly relates to a method for information extraction and integration based on RAG technology. Background Art

[0002] In application scenarios of a large amount of information extraction and fusion, it often requires multiple experts to spend a lot of time to consult and understand the data, and to organize and merge the content to form data in a standardized format. This process not only consumes a large amount of human and material resources, but also there will be a certain probability of making mistakes and inconsistent subjective judgments in manual operations.

[0003] RAG (Retrieval-Augmented Generation) technology is an advanced technology that combines a retrieval system and a generative model. It allows the model to refer to specific information in an external knowledge base when generating answers, thereby improving the relevance and accuracy of the answers. RAG can handle complex queries and can adapt to unstructured data environments. Since 2020, RAG technology has been evolving with the development of large language models such as GPT. The present invention applies this technology to the field of information extraction and fusion to solve many problems existing in traditional methods.

[0004] The existing technology still mainly relies on manual operations to consult a large amount of information, and extract key content for processing and integration. This process has problems and limitations such as long time consumption, high labor cost, requirement to master knowledge in related fields, and differences in subjective judgments of content, and a set of effective paradigms are needed to generate standardized outputs. For the same application scenario, traditional methods may require multiple experts to conduct hundreds of hours of data analysis and screening. Therefore, the present invention proposes a method for information extraction and integration based on RAG technology. Summary of the Invention

[0005] To solve the above technical problems, the present invention proposes a method for information extraction and integration based on RAG technology to solve the problems existing in the above-mentioned prior art.

[0006] To achieve the above object, the present invention provides a method for information extraction and integration based on RAG technology, including:

[0007] Standardly splitting documents in multiple formats to obtain a number of standard documents;

[0008] Inputting a number of the standard documents into an information retrieval table filling pipeline to obtain a table with some content filled;

[0009] After performing prompt word retrieval and filling on the unfilled information, combining the table with some content filled to obtain a complete content output.

[0010] Optionally, the documents in multiple formats include: text describing a single piece of content, text describing multiple pieces of content, and a comprehensive document of multiple pieces of content plus tabular data type.

[0011] Optionally, the process of standardizing and splitting documents in multiple formats into several standard documents includes:

[0012] Extract a single document from the documents in multiple formats, and parse the single document into paragraph text;

[0013] Retrieve and analyze the paragraph text based on the prompt template to obtain a retrieval result. When the retrieval result contains a description of the keywords in the keyword table, convert the paragraph text into a paragraph text with keyword discrimination information;

[0014] Traverse the paragraph text with keyword discrimination information. When all the retrieved keywords in the paragraph text with keyword discrimination information cover the keyword table, split the paragraph text with keyword discrimination information to obtain the split fragments of the current paragraph;

[0015] Traverse all paragraphs, and obtain several standard documents based on the split fragments of several paragraphs.

[0016] Optionally, the process of inputting several of the standard documents into the information retrieval table filling pipeline to obtain a table with some content filled in includes:

[0017] The split standard documents include documents with tables and pure text documents;

[0018] Transfer the pure text part of the standard document to the RAG system, and perform text embedding using the m3e-base model to obtain embedded text;

[0019] Perform information filling on the documents with tables and the documents without tables in the embedded text respectively to obtain a table with some content filled in.

[0020] Optionally, the process of performing information filling on the documents with tables in the embedded text includes:

[0021] Extract and fill in information according to the entries in the table;

[0022] If the title of the column where the current cell in the table appears in the header of the table to be filled, fill the current cell into the column corresponding to the current filling row;

[0023] Skip the remaining unmentioned content in the table to obtain a table with some content filled in.

[0024] Optionally, for a document without a table in the embedded text, standard description entries are filled to obtain a table with partially filled content.

[0025] The present invention also provides a computer terminal device, including:

[0026] One or more processors;

[0027] A memory, coupled to the processor, for storing one or more programs;

[0028] When the one or more programs are executed by the one or more processors, the one or more processors implement an information extraction and integration method based on the RAG technology.

[0029] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, an information extraction and integration method based on the RAG technology is implemented.

[0030] Compared with the prior art, the present invention has the following advantages and technical effects:

[0031] The present invention can be automatically generated in the whole process, and it only takes dozens of minutes to solve the information retrieval and filling work of hundreds of documents. These works require dozens of hours of reading, analysis and filling by multiple experts using traditional methods. Therefore, the present invention greatly optimizes the efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0033] Figure 1 is the overall technical route flowchart of the embodiment of the present invention;

[0034] Figure 2 is the standardized splitting flowchart of the document of the embodiment of the present invention;

[0035] Figure 3 is the flowchart of the information retrieval and table filling part of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0037] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0038] Embodiment 1

[0039] The present invention is a complete process for extracting and integrating equipment information based on the RAG technology, which is used to solve the problems of low efficiency, high cost, unstable accuracy, inconsistent descriptions, etc. in manual operations when a large amount of unstructured data needs to be extracted and integrated in equipment. With the help of the RAG technology based on large models, the present invention has established a complete process paradigm, which can efficiently and accurately process a large amount of unstructured data and perform structured sorting and integration on it.

[0040] The present invention first inputs various forms of documents, and based on certain rules, divides them into several standard documents by means of large model reasoning. Subsequently, these standard documents are processed in various ways for other format information such as tables, and the RAG technology is used for content filling. The situation of whether there are certain filling specifications for the filled content is also fully considered and processed separately.

[0041] Based on the existing RAG technology, the present invention designs a relatively complete equipment information processing process to specifically solve the difficulties in a large number of information extraction and integration tasks, and makes relevant adjustments to some application-level designs. In the implementation of the RAG technology in the present invention, some prompt words for specific fields are designed to better realize the information retrieval and integration functions of the RAG technology.

[0042] As Figure 1 shown, in this embodiment, an information extraction and integration method based on the RAG technology is provided, including the following steps: standardizing and splitting documents in various formats to obtain several standard documents; inputting the several standard documents into an information retrieval table filling pipeline to obtain a table with some content filled; after performing prompt word retrieval and filling on the unfilled information, combining the table with some content filled to obtain a complete content output.

[0043] As a specific implementation manner of this embodiment, first, various formats of documents go through a standardization and splitting process. The formats of the documents include: text containing single-content descriptions, text containing multiple-content descriptions, and comprehensive documents of multiple-content descriptions + tables and other data types. For a single document, it is parsed into paragraph texts, and each paragraph is separately subjected to large model reasoning. Specific prompt words are used in the large model reasoning part for interrogation to determine whether a certain keyword description is included in this paragraph. The prompt words used here are as follows:

[0044] prompt = f"""Please read the following passage and determine whether the passage contains a description of {keyword}.

[0045] If it contains, please answer "Yes", if it does not contain, please answer "No", and do not output anything else.

[0046] The content of the passage is as follows:\n{paragraph.text}"""

[0047] It can be seen that this prompt template allows the large model to retrieve and analyze the text of each paragraph to determine whether it contains a description of a certain keyword; and the keyword comes from the keyword list, which is a summary of the input document after analysis. For example, a document describing equipment information should contain key points such as "current status of capabilities" and "capability requirements", and integrating them gives the keyword list. After the large model makes inferences, the paragraph text is transformed into paragraph text with keyword discrimination information. At this time, starting from the first paragraph, traverse. If the keywords contained in the currently retrieved paragraphs cover all the keywords in the keyword list, it is considered that a complete description entry is obtained here, and it is split, that is, the accumulated paragraphs are stored and cleared. When all paragraphs have been traversed, several fragments containing a certain number of paragraphs are obtained for subsequent process processing. The complete process is as Figure 2 shown.

[0048] Subsequently, the obtained segmented document is input into the information retrieval table filling pipeline. At this time, the segmented document has two cases, namely a document with a table and a pure text document. First, the pure text part of the document (regardless of whether it contains table information) is passed to the RAG system, and the m3e-base model for the Chinese corpus is used for text embedding, and finally the embedded text is formed and stored in the subfolder corresponding to this document. The RAG system formed here can obtain answers through large model inferences by designing some inquiry templates. The inference large model used here is the Qwen2.5 - 32b model. After the RAG system is constructed, the following information integration work is carried out.

[0049] 1. For the detected document with a table, extract and fill in the information according to the entries in the table. The rule is that if the title of the column where a certain cell in the table appears in the header of the table to be filled, then fill this cell into the column corresponding to the current filling row. And the content not mentioned in the remaining table is skipped temporarily and waits to complete the content filling in subsequent steps;

[0050] 2. For documents without tables, first fill in the content of the standard description entries. The standard description entries here refer to: some of the information to be filled may have certain description specifications, such as rules that must be selected from certain entries. This step requires passing the standard description entries to the RAG system for traversal, and continuously comparing them in pairs during the RAG query process until the large model retrieves the most suitable answer and selects it for filling. The standard description entries mentioned here often come from some documents introducing professional terms, such as the type of ability items, etc. When filling in some content in the table, the description content in the document must be strictly followed. Taking the type of ability items as an example, design a script to extract all the ability item type names, levels, descriptions, etc. in the document, and integrate them into a json format document storing all ability item entries, so as to facilitate the large model to look up the table and reason. After obtaining the ability item entries, maintain a current best answer, with its initial value being the first item in the ability item entries. Then start traversing all the entries from the second item, extract the information of the current item each time, and deliver it to the large model for comparison and reasoning to judge which ability item describes the article more accurately. The prompt template used here is as follows:

[0051] query = f"""Please select the one you think best fits the content of this article from the following two standard ability items and their descriptions. Note that one option may be a more specific description of the other option. Please carefully judge the degree of detail of the ability item described in the document content. Just answer "Answer: A" or "Answer: B", do not output other content, and do not give explanations. \n

[0052] A. Ability item name: {current_capability['name']}. Ability item description: {current_capability['definition']}\n

[0053] B. Ability item name: {capability['name']}. Ability item description: {capability['definition']}"""

[0054] Among them, the name (ability item name) and definition (ability item description) information in current_capability (current best answer) and capability (current ability item entry) are used, and the prompt template is improved for some problems that have occurred in practice (such as poor attention to small differences between two ability items, repeated output of answers, or output of irrelevant content, etc.). After testing, the current version can complete the work smoothly. When all the ability item entries have been traversed, the current best answer is the best answer among all the ability item entries.

[0055] After performing the above tasks, a table with some of its content filled can be obtained. There may still be information in the table that has not been filled. At this time, traverse the unfilled cells in it and deliver them to the RAG system to retrieve and fill the content using certain prompts until all information has been correctly and completely filled. There are certain designs for the system prompt, Q&A prompt, and answer refinement prompt during this process, and the specific definitions are as follows:

[0056] System prompt:

[0057] SYSTEM_PROMPT = """You are an expert in military equipment information extraction tasks."""

[0058] Q&A prompt:

[0059] qa_prompt_tmpl_str = (

[0060] "The context information is as follows.\n"

[0061] "---------------------\n"

[0062] "{context_str}\n"

[0063] "---------------------\n"

[0064] "Please reason based on the context information and answer the following queries without relying on prior knowledge and speculation."

[0065] "In addition, as an expert in military equipment information extraction tasks, your answer should be as rigorous as possible and comply with the equipment information description specifications.\n"

[0066] "Query: {query_str}\n"

[0067] "Answer:" )

[0069] Answer refinement prompt:

[0070] refine_prompt_tmpl_str = (

[0071] "The original query is as follows: {query_str}"

[0072] "The existing answer is now provided for you: {existing_answer}"

[0073] "You can use the following additional context materials to improve the existing answer."

[0074] ""

[0075] ------------

[0076] "{context_msg}"

[0077] ""

[0078] ------------

[0079] "Please comprehensively consider the new context materials to optimize the original answer and obtain better query results. If you think the provided context is not helpful for the answer, please directly return the original answer."

[0080] "RefinedAnswer:" )

[0082] Through the design of these prompt words and via the prompt word templates in the llama library, a complete RAG system prompt word architecture is constructed to facilitate the model to better answer questions in this specific field. Finally, the obtained complete content is used as the structured output of the entire pipeline, and the process is as Figure 3 shown.

[0083] Compared with the traditional method, this technology can be automatically generated throughout the process after testing. It only takes dozens of minutes to complete the information retrieval and filling of hundreds of documents, which would require dozens of hours of reading, analysis, and filling by multiple experts using the traditional method. Therefore, this invention greatly optimizes the efficiency.

[0084] This invention does not require the labor cost and learning cost of experts in the relevant field and only relies on the semantic recognition ability of the large model, thus also reducing the input of labor and material costs.

[0085] Traditional methods may cause some mistakes and information omissions due to human factors. This method is based on the reasoning of the large model and automated information retrieval, and has very high accuracy and reliability.

[0086] Compared with the traditional method, due to the differences in individual understanding and expression among multiple experts, there will be certain deviations in the understanding and judgment of the text. In this invention, this part of the work is predicted by the same large model based on the context sequence, which can ensure a stronger degree of standardization in expression from a principle perspective and also facilitate subsequent integration or other work.

[0087] The above is only a preferred specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An information extraction and integration method based on RAG technology, characterized in that, It includes the following steps: Standardize and split documents in multiple formats to obtain a number of standard documents; Input a number of the standard documents into an information retrieval form filling pipeline to obtain a form with some content filled; After performing prompt word retrieval and filling on the information that has not been filled, combine the form with some content filled to obtain a complete content output.

2. The information extraction and integration method based on the RAG technology according to claim 1, wherein The documents in multiple formats include: text describing a single piece of content, text describing multiple pieces of content, and comprehensive documents of the data type of multiple pieces of content plus tables.

3. The information extraction and integration method based on the RAG technology according to claim 1, characterized in that The process of standardizing and splitting documents in multiple formats to obtain a number of standard documents includes: Extract a single document from the documents in multiple formats, and parse the single document into paragraph text; Perform retrieval and analysis on the paragraph text based on a prompt word template to obtain a retrieval result. When the retrieval result contains a description of a keyword in the keyword table, convert the paragraph text into paragraph text with keyword discrimination; Traverse the paragraph text with keyword discrimination. When all the retrieved keywords in the paragraph text with keyword discrimination cover the keyword table, split the paragraph text with keyword discrimination to obtain split segments of the current paragraph; Traverse all paragraphs, and obtain a number of standard documents based on the split segments of several paragraphs.

4. The information extraction and integration method based on the RAG technology according to claim 3, characterized in that, The process of inputting a number of the standard documents into an information retrieval form filling pipeline to obtain a form with some content filled includes: The split standard documents include documents with tables and pure text documents; Transfer the pure text part of the standard documents to the RAG system, and perform text embedding using the m3e-base model to obtain embedded text; Perform information filling on the documents with tables and the documents without tables in the embedded text respectively to obtain a form with some content filled.

5. The information extraction and integration method based on the RAG technology according to claim 4, wherein The process of performing information filling on the documents with tables in the embedded text includes: Extract and fill information according to the entries existing in the table; If the title of the column where the current cell in the table appears in the header of the table to be filled, fill the current cell into the column corresponding to the current filled row; Skip the remaining unmentioned content in the table to obtain a form with some content filled.

6. The information extraction and integration method based on the RAG technology according to claim 5, characterized in that Fill standard description entries into the documents without tables in the embedded text to obtain a form with some content filled.

7. A computer terminal device, characterized in that, It includes: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the information extraction and integration method based on the RAG technology as described in any one of claims 1-6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the information extraction and integration method based on the RAG technology as described in any one of claims 1-6.

Citation Information

Cited By

  • Environmental literature intelligent screening system and method based on large language model

    CN121436175A