Water conservancy field retrieval enhancement generation method and device based on knowledge graph and medium

By constructing a knowledge graph in the water conservancy field and combining large model technology, the problem of insufficient efficiency retrieval, noise processing and recall accuracy in the search enhancement generation method in the water conservancy field is solved, and high accuracy and comprehensive knowledge Q&A services are achieved.

CN120144743APending Publication Date: 2025-06-13SHANDONG ZHIYANG SHANGSHUI INFORMATION TECH CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510170990.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the large-scale scanned document processing scenario in the field of water conservancy, the search-enhanced generation method based on knowledge graphs has problems such as efficient and accurate retrieval difficulties, insufficient noise interference processing, and insufficient recall mechanism and generation model accuracy.

Method used

By processing data on documents in the water conservancy field, building a knowledge graph, using large models to extract and summarize and generate entity relationships, combining Leiden clustering algorithm to divide communities, and using multi-graph query and multi-path sorting fusion to recall knowledge graphs during user query, optimizing the answering process.

Benefits of technology

It has realized the construction of high-quality knowledge graphs in the water conservancy field, accurately recalled knowledge entities and relationships that are highly related to user query, improved the accuracy and comprehensiveness of the generated content, and significantly improved the intelligent response capabilities and information service level of the system in the application of the water conservancy field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144743A_ABST
    Figure CN120144743A_ABST
Patent Text Reader

Abstract

The invention provides a water conservancy field retrieval enhancement generation method and device based on a knowledge graph and a medium. The method comprises the steps of extracting entity relationships in a water conservancy field document; constructing a directed unweighted graph by utilizing the entities and the relationships; summarizing description information of each entity and relation description between the entity and other entities by using a large model to generate an entity abstract, performing multi-level semantic modeling on the entity abstract, and converging semantic information of adjacent entities as graph embedding representation of the entities; associating the entity abstract and the graph embedding representation thereof with entity nodes in the directed unweighted graph to obtain an optimized graph; dividing the entities into a plurality of communities according to the modularity among the entities in the optimization graph; carrying out community summarization on each community by utilizing the large model to obtain a community abstract; when a user query request is received, carrying out knowledge graph recall by utilizing multi-graph inquiry and multi-path sorting fusion; and according to the sorting score of the retrieved atlas information, optimizing the task cue word to guide the generation model to optimize the answer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of retrieval enhancement, and particularly to a retrieval enhancement generation method, device, and medium for the water conservancy field based on a knowledge graph. Background Art

[0002] The retrieval enhancement generation method can effectively alleviate problems such as the "hallucination" phenomenon of large models, lack of domain-specific knowledge, and outdated information by introducing an external knowledge base. At the same time, the retrieval enhancement generation method based on a knowledge graph can capture complex entity relationships and their dependencies highly relevant to the question through the graph structure data in the external knowledge base, providing clearer and more accurate context information, enabling the large model to better utilize structured knowledge to avoid redundant information interference or information omission.

[0003] However, in the scenario of processing large-scale scanned documents in the water conservancy field, the retrieval enhancement generation method based on a knowledge graph still has many problems. First, in the face of a knowledge graph database containing millions of entity relationships, the existing query mechanism is difficult to achieve efficient and accurate retrieval. Second, when the existing entity relationship extraction method based on a large model processes scanned documents, it fails to effectively remove the noise interference introduced by table and picture data in the documents. Finally, there are still problems with insufficient accuracy in the recall mechanism and generation model of the knowledge graph during the optimization process of answer generation. The existing large-scale corpus retrieval enhancement generation technology based on a knowledge graph has the following deficiencies: ignoring the interference effect of redundant information, not fully considering the common redundant information in specific domain scanned documents, such as table titles, table contents, table notes, and figure titles, figure notes, on the accuracy of entity relationship extraction tasks. Ignoring the recall accuracy of the knowledge graph: Limited by the input and computing resources of the large model, the retrieval enhancement generation method based on a knowledge graph focuses on recalling a very small number of knowledge graphs for question answering. Ignoring answer optimization: Traditional methods do not fully consider the relative importance of different text blocks or graphs in the recalled content, but treat all recalled results equally. The recalled answers should be differentially processed according to their relevance and weights to ensure that answers with higher weights occupy a more important position in the final result, thereby improving the overall accuracy of the answer. Summary of the Invention

[0004] To solve the above technical problems or at least partially solve the above technical problems, the present invention provides a retrieval enhancement generation method, device, and medium for the water conservancy field based on a knowledge graph.

[0005] In a first aspect, the present invention provides a retrieval enhancement generation method for the water conservancy field based on a knowledge graph, including:

[0006] Performing data processing on water conservancy field documents to obtain text blocks without table content and table content;

[0007] Constructing a knowledge graph based on non-tabular content and tabular content, including: for non-tabular content, using a large model to extract entity relationships; for tabular content, mapping key-value pairs of the table content based on the positional relationship between text units in the table to extract entity relationships; using the extracted entities as nodes of the graph, and the relationships between entities as directed edges to construct a directed unweighted graph; using the large model to summarize the description information of each entity and the relationship description between it and other entities to generate an entity summary, performing multi-level semantic modeling on the entity summary, and aggregating the semantic information of adjacent entities as a graph embedding representation of the entity; associating the entity summary and its graph embedding representation with the entity nodes in the directed unweighted graph, and embedding the entity summary and its corresponding graph embedding representation into the corresponding node position in the graph structure to obtain an optimized graph; using the Leiden clustering algorithm to divide the entities into multiple communities according to the modularity between the entities in the optimized graph; using the large model to perform community summarization on each community to obtain a community summary;

[0008] When receiving a user query request, the knowledge graph is recalled using multi-graph query and multi-path ranking fusion to obtain the entity relationship that best matches the query and provide corresponding community summary information; based on the ranking score of the retrieved graph information, the task prompt words are optimized to guide the generation model to give priority to high-priority graph information and supplement relevant content during the answering process to optimize the answer.

[0009] Furthermore, data processing of water conservancy documents obtains text blocks with non-table content and table content including:

[0010] Identify the text content and coordinate area of ​​the smallest text unit of documents in the water conservancy field;

[0011] Perform wireframe table structure recognition on the table data of the document, and form one or more tables according to the coordinate positions of the inseparable wireframes recognized on a single page;

[0012] Determine whether the text unit is inside the table based on the intersection-and-union ratio of the minimum text unit area and the table area; if so, delete the text unit in the table from all the text units, so as to achieve accurate screening of the content in the table; for redundant data of table titles, figure titles, table notes, and figure notes that meet the conditions, judge based on the distance relationship between them and the preceding minimum text unit to screen out redundant data; identify and eliminate irrelevant or repeated content through strong rule matching;

[0013] The obtained document with non-tabular content is split into multiple text blocks. During the block segmentation process, the title, subtitle, paragraph, table, and picture elements are used as the basis for segmentation. At the same time, the semantic structure of the document is considered for segmentation to ensure that each text block has a clear semantic structure.

[0014] Furthermore, when using a large model to extract entity relationships from non-tabular content, through multiple rounds of Q&A with the large model, entities and relationships in the knowledge graph are identified and extracted from the segmented text blocks. The process includes:

[0015] Design entity relationship extraction task prompts based on the task description and entity extraction examples. The relationship extraction task prompts control the large model to first extract the entity name entity_name, entity type entity_type, and entity description entity_description in the text block, and then extract entity pairs with obvious relationships. The entity pair includes the source entity source_entity and target entity target_entity with a relationship, and the relationship is described in detail. To ensure that no entity is missed during the entity extraction process, each text block will undergo three rounds of independent extraction operations. Subsequently, the original text block and the extracted entity relationship triples are used as inputs, and the large model is used to verify the integrity of the extraction results. If the extraction results are detected to be incomplete, entity relationship extraction continues until all entities and their relationships are fully extracted.

[0016] Furthermore, use the large model to generate entity summaries for the description information of each entity and its relationship descriptions with other entities, including: using the description information of each entity and its relationship descriptions with other entities as inputs, and relying on the summarization ability of the large model, controlling the large model to summarize and generalize each entity through summarization task prompts to generate entity summaries.

[0017] Furthermore, for each community, input the entities, relationships, and corresponding descriptions of the same community into the large model, and the large model performs community summarization under the control of community summarization task prompts to obtain community summaries;

[0018] The community summarization task prompts specify the content and structure of the community summary. Among them, the content of the community summary includes: the community name title representing the key entity of the community; the executive summary executive_summary of the overall community structure, and the executive summary executive_summary summarizes how the community entities are related to each other and the important points related to its entities; the score rating of the relevance of the text to water resource management, water conservancy projects, historical background, and professional community dynamics; a one-sentence explanation rating_explanation for configuring the rating score; the key insights findings of the community, and each insight has an internal summary insight_1_summary and an internal explanation insight_1_explanation.

[0019] Furthermore, the process of using multi-graph query and multi-path ranking fusion for knowledge graph recall includes:

[0020] Extract key entity information from the query, and use the bge-m3 model to encode the query and the entities it contains. The encoded feature vectors are then used to calculate the similarity with entities, entity summaries, and community summaries in multiple dimensions of each knowledge graph to evaluate the relevance between the query and the graph content. Return the entity relationships that best match the query from each knowledge graph and provide corresponding community summary information; multi-path ranking initially ranks the relevance between the query and multi-dimensional information in the knowledge graph through the BM25 algorithm based on statistical features. With the efficient modeling of term frequency and document length by the BM25 algorithm, calculate the relevance scores between the query and each entity, entity summary, and community summary in the graph; based on the relevance scores, further encode the query, community summary, and original text block through the bge-rerankerv2-minicpm-layerwise model, and use the deep learning model to further optimize the ranking results.

[0021] Furthermore, for the graph information with the highest priority, add the keywords "focus on" and "most relevant" when forming the prompt words, so as to ensure that the generation model places the graph information with the highest priority at the front of the answer under the control of optimizing the task prompt words and quotes more during the generation process; the graph information with low priority and supplementary content that conforms to the theme are used as supplements to the final answer to enhance the diversity and richness of the answer.

[0022] Furthermore, the optimized task prompt words include: context, question query_str, and corresponding reference content answer_str. The reference content is the graph information with the highest retrieved relevance; it is required to supplement the original answer based on the provided context content and reference content, retain every character of the reference content, and add new supplementary content based on the reference content answer_str to make the answer more in-depth and multi-dimensional; the supplementary content is closely related to the reference content and must not deviate from the themes of the reference content answer_str and the question query_str, and add more details, terms, or relevant entries.

[0023] In a second aspect, the present invention provides a water conservancy field retrieval enhanced generation device based on a knowledge graph, including: at least one processing unit, which connects the processing unit and the storage unit through a bus unit. The storage unit, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules. The processing unit realizes the above-mentioned water conservancy field retrieval enhanced generation method by running the software programs, computer-executable programs, and modules stored in the storage unit.

[0024] In a third aspect, the present invention provides a computer-readable storage medium storing a computer program which, when executed, implements the method for enhanced retrieval generation in the water conservancy field based on a knowledge graph as described above.

[0025] The above technical solutions provided by the embodiments of the present invention have the following advantages compared with the prior art:

[0026] This application processes the documents in the water conservancy field to obtain text blocks and table contents that are not in table form; constructs a knowledge graph based on the non-table contents and table contents, including: for non-table contents, uses a large model to extract entity relationships, and for table contents, performs key-value pair mapping of the table contents based on the positional relationships between the text units in the table and extracts entity relationships; takes the extracted entities as the nodes of the graph, takes the relationships between the entities as directed edges, and constructs a directed unweighted graph; uses a large model to generate entity summaries for the description information of each entity and its relationship descriptions with other entities, performs multi-level semantic modeling on the entity summaries, and aggregates the semantic information of adjacent entities as the graph embedding representation of the entity; associates the entity summaries and their graph embedding representations with the entity nodes in the directed unweighted graph, and embeds the entity summaries and their corresponding graph embedding representations into the corresponding node positions in the graph structure to obtain an optimized graph; uses the Leiden clustering algorithm to divide the entities into multiple communities according to the modularity between the entities in the optimized graph; uses a large model to generate community summaries for each community; when receiving a user query request, uses multi-graph interrogation and multi-path ranking fusion to recall the knowledge graph, obtains the entity relationship that best matches the query, and provides the corresponding community summary information; according to the sorting scores of the retrieved graph information, guides the generation model to give priority to high-priority information during the answering process by optimizing the task prompt words.

[0027] By constructing a high-quality knowledge graph in the water conservancy field and integrating the multi-dimensional information in the knowledge graph, the present invention can achieve accurate graph recall, ensure effective matching of knowledge entities and relationships highly relevant to the user query, thereby improving the accuracy and comprehensiveness of the generated content. Combining the knowledge graph with the large model technology, the present invention optimizes the generation process of the question-answering task in the water conservancy field, generates more accurate, comprehensive and practically required answers, and significantly improves the intelligent response ability and information service level of the system in the application of the water conservancy field. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0030] Figure 1 It is a flowchart of a retrieval enhancement generation method in the water conservancy field based on a knowledge graph provided by an embodiment of the present invention;

[0031] Figure 2 It is a flowchart of constructing a knowledge graph according to non-table content and table content provided by an embodiment of the present invention;

[0032] Figure 3 It is a flowchart of using multi-graph query and multi-path ranking fusion for knowledge graph recall provided by an embodiment of the present invention;

[0033] Figure 4 It is a schematic diagram of a retrieval enhancement generation device in the water conservancy field based on a knowledge graph provided by an embodiment of the present invention. Detailed implementation manners

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0035] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0036] Embodiment 1

[0037] The present invention aims to propose a retrieval-enhanced generation method in the water conservancy field based on a knowledge graph to solve the limitations of the prior art in knowledge-based question answering in the water conservancy field. By constructing a high-quality knowledge graph in the water conservancy field and integrating multi-dimensional information in the knowledge graph, the present invention can achieve accurate graph recall, ensure effective matching with knowledge entities and relationships highly relevant to user queries, thereby improving the accuracy and comprehensiveness of the generated content. Combining the knowledge graph with large model technology, the present invention optimizes the generation process of question answering tasks in the water conservancy field, generates more accurate, comprehensive and practical answers, and significantly improves the intelligent response ability and information service level of the system in the application of the water conservancy field.

[0038] As Figure 1 shown, according to the data characteristics and question answering characteristics in the water conservancy field, the retrieval-enhanced generation method in the water conservancy field based on a knowledge graph of the present invention includes:

[0039] Perform data processing on water conservancy field documents to obtain text blocks without table content and table content, and construct a knowledge graph based on the non-table content and table content. Among them, the process of performing data processing on water conservancy field documents to obtain text blocks without table content and table content includes: reading water conservancy field documents, screening redundant data, and document chunking.

[0040] Reading water conservancy field documents: Water conservancy field documents include scanned documents and electronic documents. Among them, for electronic documents, the Fitz tool is used, and for scanned documents, the RapidOCR tool is used to identify the text content and coordinate areas of the smallest text units.

[0041] Since the proportion of table content in the publicly available data in the water conservancy field is relatively large, all the extracted text units contain both the content of the non-table part and the content of the table part. The content of the non-table part is actually structurally and semantically complete, and the content of the table part is redundant data for the content of the non-table part.

[0042] Redundant data screening includes: Since the proportion of table content in the publicly available data in the water conservancy field is relatively large, therefore, the present invention performs wireframe table structure recognition on the table data of the document, and splices the coordinate positions of the inseparable wireframes identified on a single page to form one or more tables. According to the intersection and union ratio of the smallest text unit area and the table area, it is judged whether the text unit is inside the table. If so, the text units inside the table are deleted from all the text units, thereby realizing the accurate screening of the content in the table. For chart redundant data such as table titles, figure titles, table notes, and figure notes that meet the conditions, it is judged according to the distance relationship with the previous smallest text unit to screen out the redundant data. Identify and eliminate irrelevant or duplicate content through strong rule matching to improve the information density of the data.

[0043] Document chunking includes: The purpose of document chunking is to split a long document into multiple text chunks, ensuring that each text chunk has a clear semantic structure, and the extraction and processing of information can focus on key content, avoiding semantic loss or confusion caused by excessive or improper segmentation. During the chunking process of the obtained non-table content, titles, sub-titles, paragraphs, tables, and picture elements are used as the basis for segmentation, and the semantic structure of the document is considered for segmentation, so that each document chunk contains complete context information and avoids the loss of semantic information. For multi-page documents, document chunking not only deals with the continuity of the text but also pays attention to the logical relationship between pages to ensure the consistency of the chunked content in the case of cross-page.

[0044] Construct a knowledge graph based on non-table content and table content, as Figure 2 shown, including:

[0045] For non-table content, the process of constructing a knowledge graph includes using a large model for entity relationship extraction. For table content, based on the positional relationship between text units in the table, key-value pair mapping of the table content is performed for entity relationship extraction.

[0046] Among them, using a large model to perform entity relationship extraction on non-table content includes: using multiple question-and-answer sessions of the large model to identify and extract entities and relationships in the knowledge graph from the divided text chunks. In the specific implementation process, entity relationship extraction task prompt words are designed through task descriptions and entity extraction examples. First, entity names entity_name, entity types entity_type, and entity descriptions entity_description in the text chunks are extracted, and then entity pairs with obvious relationships are extracted. The entity pair contains the head entity source_entity and the tail entity target_entity with a relationship, and the relationship is described in detail. To ensure that no entity is missed during the entity extraction process by the large model, each text chunk will go through three rounds of independent extraction operations. Subsequently, the original text chunk and the extracted entity relationship triples are used as input, and the large model is used to verify the integrity of the extraction results. If it is detected that the extraction results are incomplete, entity relationship extraction continues until all entities and their relationships are fully extracted.

[0047] An exemplary entity relationship extraction task prompt word is as follows:

[0048] “-Target-

[0049] Given a relevant text document and a list of entity types, identify all entities of each type in the text document corresponding to the list of entity types and all relationships between the identified entities.

[0050] - Step - 1. Identify all entities. For each identified entity, extract the following information:

[0051] - entity_name: The name of the entity

[0052] - entity_type: The type of the entity, with the category name as general as possible to avoid being too specific - entity_description: The description of the entity, which is a comprehensive description of the entity's attributes. Format each entity as ("entity"{tuple_delimiter}<entity_name>{tuple_delimiter}<entity_type>{tuple_delimiter}<entity_description>)

[0053] {tuple_delimiter} is the tuple delimiter.

[0054] 2. From the entities identified in Step 1, identify all *obviously related* entity pairs (source_entity, target_entity).

[0055] For each pair of related entities, extract the following information:

[0056] - source_entity: The name of the source entity, as described in Step 1

[0057] - target_entity: The name of the target entity, as described in Step 1

[0058] - relationship_description: The reason why the source entity and the target entity are related to each other

[0059] - relationship_strength: A numerical score representing the strength of the relationship between the source entity and the target entity;

[0060] Format each relationship as ("relationship"{tuple_delimiter}<source_entity>{tuple_delimiter}<target_entity>{tuple_delimiter}<relationship_description>{tuple_delimiter}<relationship_strength>)

[0061] 3. Return the output in Chinese as a single list of all entities and relationships determined in Steps 1 and 2. Use **{record_delimiter}** as the list separator.

[0062] 4. After completion, feedback {completion_delimiter} to indicate completion.

[0063] #

[0064] - Example -

[0065] {example}

[0066] #

[0067] - Real Data -

[0068] #

[0069] Text: {input_text}

[0070] #

[0071] Output: ”

[0072] For the text cells inside the table, perform key-value pair mapping of the table content based on the positional relationships between the text cells to form structured triples (<head entity, relationship, tail entity>).

[0073] After completing entity relationship extraction, use the extracted entities as nodes in the graph and the relationships between entities as directed edges to construct a directed unweighted graph.

[0074] Use a large model to generate entity summaries for the description information of each entity and its relationship descriptions with other entities, including: taking the description information of each entity and its relationship descriptions with other entities as input, leveraging the summarization ability of the large model, and controlling the large model to summarize each entity through summarization task prompts to generate concise and information-rich entity summaries.

[0075] An exemplary summarization task prompt is as follows:

[0076] “You are an expert in social scientists of water resources and hydraulic engineering. You are good at analyzing social networks, identifying key participants, and understanding the dynamics of professional communities. You are good at helping people map the relationships and structures within a professional field and gain in-depth understanding of the collaboration patterns and knowledge flows in the field of 'water resources and hydraulic engineering'.

[0077] Please use your water conservancy expertise to comprehensively summarize the data provided below.

[0078] Given one or two entities and a list of descriptions, all descriptions are related to the same entity or entity pair.

[0079] Please connect all of these into **one** concise description. Ensure that the information collected from all descriptions is included.

[0080] If the provided descriptions are contradictory, resolve the contradiction and provide a coherent summary.

[0081] Ensure that it is written in the third person and includes the entity names so that we can understand the complete context. It is very important to enrich it with relevant information from the nearby text as much as possible.

[0082] If there is no possible answer, or the description is empty, then only convey the information provided in the text.

[0083] #

[0084] -Data-

[0085] Entity: {entity_name}

[0086] Description list: {description_list}

[0087] #

[0088] Output: ”

[0089] After that, use the BGE-M3 model to perform multi-level semantic modeling on the entity summary, aggregate the semantic information of adjacent entities as the graph embedding representation of the entity, and use the graph embedding representation to further characterize the potential relevance and context information between entities.

[0090] Associate the entity summary and its corresponding graph embedding representation with the entity nodes in the directed unweighted graph, and embed the entity summary and its corresponding graph embedding representation into the corresponding node positions in the graph structure to obtain an optimized graph.

[0091] Adopt the Leiden clustering algorithm to divide the entities into multiple communities according to the modularity between entities in the optimized graph. The entities within each community have strong internal connections in the graph structure, while the connections between different communities are relatively sparse. By adopting a bottom-up clustering strategy and a step-by-step multi-level optimization process, a more accurate community structure is obtained.

[0092] For each community, the entities, relationships, and corresponding descriptions within the same community are input into a large model, which performs community summarization under the control of community summarization task prompts to obtain a community summary. The community summarization task prompts specify the content and structure of the community summary. The content of the community summary includes: the community name title representing the key entities of the community; an executive summary of the overall community structure, which summarizes how the community entities are related to each other and the important points related to those entities; a score rating indicating the relevance of the text to water resource management, water conservancy projects, historical background, and professional community dynamics; a one-sentence explanation rating_explanation for configuring the rating score; key insights findings about the community, with each insight having an internal summary insight_1_summary and an internal explanation insight_1_explanation.

[0093] An exemplary community summarization task prompt is as follows:

[0094] "You are a professional sociologist of water resources and water conservancy projects. You are good at analyzing social networks, identifying key participants, and understanding the dynamics of professional communities. You are skilled at helping people map the relationships and structures within a professional field and gain in-depth understanding of the collaboration patterns and knowledge flows in the 'water resources and water conservancy projects' field.

[0095] #Objective

[0096] Write a comprehensive community report listing the entities belonging to the community, their relationships, and optional related statements. This report will be used to inform decision-makers about information related to the community and its potential impacts. The content of this report includes an overview of the community's key entities, its legal compliance, technical capabilities, reputation, and notable statements.

[0097] #Report Structure

[0098] The report should include the following sections:

[0099] -title: The community name representing its key entities - The title should be short but specific. If possible, include representative named entities in the title.

[0100] -summary: An executive summary of the overall community structure, which summarizes how the community entities are related to each other and the important points related to those entities.

[0101] -rating: A floating score between 0 and 10 indicating the relevance of the text to water resource management, water conservancy projects, historical background, and professional community dynamics, where 1 indicates negligible or irrelevant, and 10 indicates very important, profound, and influential for understanding the field and its implications.

[0102] -rating_explanation: Explain the rating score in one sentence.

[0103] -findings: List the key insights about the community (no more than 10). Each insight has an internal summary insight_1_summary and an internal explanation insight_1_explanation, with comprehensive content. Add a data record reference at the **end** of the internal explanation insight_1_explanation.

[0104] The citation format is

[0105] records: <record_source>(<record_id_list>),...,<record_source>(<record_id_list>). If there are more than 5 record_sources or record_id_lists, only show the first 5 most relevant records. If there are no relevant roles or records, use the empty list "[]".

[0106] The following is an example of the internal explanation:

[0107] "Entity A is the central entity in this community. This entity is the common bond between all other entities, indicating its importance in the community.

[0108] [records: Entities(1,2,3), Claims(2,5), Relationships(10,12)]"

[0109] Where 1, 2, 3, 2, 5, 10, 12 represent the IDs in the relevant data records.

[0110] Return the output community comprehensive report in the form of a correctly formatted JSON string as follows. Do not use any unnecessary escape sequences and do not add extra spaces in the keys. The output is a single JSON object that can be parsed by json.loads.

[0111] {

[0112] "title": "<report_title>",

[0113] "summary": "<executive_summary>",

[0114] "rating": "<threat_severity_rating>",

[0115] "rating_explanation":"<rating_explanation>"

[0116] "findings":

[0117] {

[0118] "summary":"<insight_1_summary>",

[0119] "explanation":"<insight_1_explanation>"

[0120] },

[0121] {

[0122] "summary":"<insight_2_summary>",

[0123] "explanation":"<insight_2_explanation>"

[0124] }

[0126] }

[0127] # Example input

[0128] {example}

[0129] # Real data

[0130] Answer the question using the following text. Do not fabricate anything in the answer.

[0131] Text:

[0132] {input_text}

[0133] Output:”

[0134] When receiving a user query request, use the fusion of multi-graph interrogation and multi-path ranking for knowledge graph recall, as Figure 3 shown. The process includes:

[0135] ​Extract key entity information from the query and encode the query and the entities it contains using the bge-m3 model; the encoded feature vectors are then used to calculate the similarity with entities, entity summaries, and community summaries in multiple dimensions of each knowledge graph to evaluate the relevance between the query and the graph content; return the entity relationships that best match the query from each knowledge graph and provide the corresponding community summary information. Sort the relevance between the query and the multi-dimensional information in the knowledge graph using the BM25 algorithm based on statistical features. Leveraging the efficient modeling of term frequency and document length by the BM25 algorithm, calculate the relevance scores between the query and each entity, entity summary, and community summary in the graph. On this basis, further encode the query, community summary, and original text block using the bge-rerankerv2-minicpm-layerwise model to further optimize the sorting results using a deep learning model.

[0136] According to the sorting scores of the retrieved graph information, optimize the answer by guiding the generation model to prioritize high-priority graph information and supplement relevant content during the answering process through optimized task prompts. For the highest-priority graph information, add the keywords "focus on" and "most relevant" when forming the prompts, so as to ensure that the generation model places the highest-priority graph information at the front of the answer under the control of the optimized task prompts and references it more during the generation process. The low-priority graph information is used as a supplement to the final answer to enhance the diversity and richness of the answer. The optimized task prompts include: context, question query_str, and the corresponding reference content answer_str, where the reference content is the graph information with the highest retrieved relevance; it is required to supplement the original answer based on the provided context content and reference content, retain every character of the reference content, and add new supplementary content based on the reference content answer_str to make the answer more in-depth and multi-dimensional; the supplementary content is closely related to the reference content and must not deviate from the themes of the reference content answer_str and the question query_str, and add more details, terms, or relevant entries.

[0137] An example of an optimized task prompt is as follows:

[0138] "Context:

[0139] ---------{top1_content_str}---------

[0140] You will see a question query_str and the corresponding reference content answer_str. Please supplement the original answer based on the provided context and reference content to ensure your response is more complete and detailed. Try to preserve every character of the reference content and, on this basis, reasonably add new supplementary content to make the answer more in-depth and multi-dimensional. The supplementary content should be closely related to the reference content, not deviate from the topic, but can add more details, terms, or relevant entries to help answer the question more comprehensively.

[0141] Your task is:

[0142] Focus on the reference content.

[0143] Based on the reference content, reasonably expand and supplement more information. You can use more terms, relevant backgrounds, details, or examples to enhance the quality and comprehensiveness of the answer.

[0144] Ensure that the supplementary part connects naturally with the reference content to form a smooth and coherent overall answer.

[0145] Do not add extra irrelevant information or speculative content on your own. The answer should focus on the fields and knowledge involved in the question.

[0146] Question:

[0147] {query_str}

[0148] Reference content:

[0149] {answer_str}

[0150] New answer:

[0151] (Please generate a longer and more complete answer here, supplementing and expanding the content of the reference answer while maintaining the natural fluency and consistency of the structure)”

[0152] The technology of the present invention can extract core knowledge points from massive data, automatically identify and remove redundant information, ensuring the simplicity and efficiency of the knowledge graph, and avoiding the interference of information duplication and invalid content. Based on the user's query, through the preliminary screening of multi-graph interrogation and the fusion technology of multi-path ranking, the present invention quickly identifies and recalls the entity relationships in the knowledge graph that are highly relevant to the query. Combining deep semantic analysis and context understanding, it ensures that the retrieval results are not limited to keyword matching, and can accurately capture the core meaning of the question at the semantic level, providing more accurate and relevant graph information. Utilizing the key role of the effective information in the graph information in information retrieval enhancement as the main information source, and using the remaining recalled graph information as supplementary resources, a more appropriate response that conforms to the user's query intention is generated.

[0153] In a knowledge - based Q&A project in the water conservancy field, the retrieval - enhanced generation method based on the present invention is adopted, and a corresponding knowledge graph is constructed by combining 172 books in the water conservancy field. First, through the data pre - processing method of the present invention, massive character data is screened, and a character set of tens of millions is obtained after processing. Subsequently, using the knowledge - graph construction method, millions of entities and their relationship information are successfully extracted. For the user's query problem, the system retrieves the four most relevant graph contents to the problem through the knowledge - graph recall method, and fills the template according to the answer optimization strategy, so as to generate an accurate user response. Experimental results show that compared with GraphRAG, the response time of this application in the Q&A process in the water conservancy field is shortened by 20%, and at the same time, it also shows a significant improvement in the Rouge - L evaluation index.

[0154] Embodiment 2

[0155] Refer to Figure 4 As shown, an apparatus for retrieval - enhanced generation in the water conservancy field based on a knowledge graph according to an embodiment of the present invention includes: at least one processing unit, which connects the processing unit and the storage unit through a bus unit. The communication unit communicates with the unmanned aerial vehicle. The storage unit, as a computer - readable storage medium, can be used to store software programs, computer - executable programs, and modules, such as the software program, computer - executable program, and module corresponding to a method for retrieval - enhanced generation in the water conservancy field based on a knowledge graph in an embodiment of the present invention. The processing unit realizes the above - mentioned method for retrieval - enhanced generation in the water conservancy field based on a knowledge graph by running the software programs, computer - executable programs, and modules stored in the storage unit, including:

[0156] Perform data processing on water - conservancy - field documents to obtain text blocks and table contents that are not in tabular form;

[0157] Construct a knowledge graph according to the non - tabular content and the table content, including: for the non - tabular content, use a large - model to extract entity relationships; for the table content, perform key - value pair mapping of the table content based on the positional relationship between text units in the table and extract entity relationships; take the extracted entities as nodes of the graph, take the relationships between entities as directed edges, and construct a directed acyclic graph; use a large - model to generate entity summaries for the description information of each entity and its relationship description with other entities, perform multi - level semantic modeling on the entity summaries, and converge the semantic information of adjacent entities as the graph - embedding representation of the entity; associate the entity summaries and their graph - embedding representations with the entity nodes in the directed acyclic graph, and embed the entity summaries and their corresponding graph - embedding representations into the corresponding node positions in the graph structure to obtain an optimized graph; adopt the Leiden clustering algorithm to divide the entities into multiple communities according to the modularity between entities in the optimized graph; use a large - model to perform community summary for each community to obtain a community summary;

[0158] When receiving a user query request, use the fusion of multi-graph interrogation and multi-path ranking for knowledge graph recall, obtain the entity relationship that best matches the query, and provide corresponding community summary information; according to the sorting scores of the retrieved graph information, guide the generation model to give priority to high-priority graph information and supplement relevant content during the answering process by optimizing the task prompt words to optimize the answer.

[0159] Certainly, the computer program stored in the storage unit of a water conservancy field retrieval enhancement generation device based on a knowledge graph provided by an embodiment of the present invention is not limited to the method operations described above, and can also execute related operations in a water conservancy field retrieval enhancement generation method provided by any embodiment of the present invention.

[0160] Embodiment 3

[0161] An embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed, it implements the water conservancy field retrieval enhancement generation method based on a knowledge graph, including:

[0162] Perform data processing on water conservancy field documents to obtain text blocks and table contents that are not in tabular form;

[0163] Construct a knowledge graph according to the non-tabular content and the table content, including: for the non-tabular content, use a large model to extract entity relationships; for the table content, perform key-value pair mapping of the table content based on the positional relationships between the text units in the table, and extract entity relationships; use the extracted entities as the nodes of the graph, use the relationships between the entities as directed edges to construct a directed acyclic graph; use a large model to generate entity summaries for the description information of each entity and its relationship descriptions with other entities, perform multi-level semantic modeling on the entity summaries, and converge the semantic information of adjacent entities as the graph embedding representation of the entity; associate the entity summaries and their graph embedding representations with the entity nodes in the directed acyclic graph, and embed the entity summaries and their corresponding graph embedding representations into the corresponding node positions in the graph structure to obtain an optimized graph; use the Leiden clustering algorithm to divide the entities into multiple communities according to the modularity between the entities in the optimized graph; use a large model to summarize each community to obtain community summaries;

[0164] When receiving a user query request, use the fusion of multi-graph interrogation and multi-path ranking for knowledge graph recall, obtain the entity relationship that best matches the query, and provide corresponding community summary information; according to the sorting scores of the retrieved graph information, guide the generation model to give priority to high-priority graph information and supplement relevant content during the answering process by optimizing the task prompt words to optimize the answer.

[0165] A computer-readable storage medium provided by an embodiment of the present invention stores a computer program that is not limited to the method operations described above, and can also execute relevant operations in a retrieval enhancement generation method in the water conservancy field based on a knowledge graph provided by any embodiment of the present invention.

[0166] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of structures or units can be in electrical, mechanical or other forms.

[0167] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0168] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0169] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A water conservancy field retrieval enhancement generation method based on knowledge graph, characterized in that: include: Data processing is performed on documents in the field of water conservancy to obtain text blocks and table contents of non-tabular content, and a knowledge graph is constructed based on the non-tabular content and the table content, including: for non-tabular content, entity relationship extraction is performed using a large model, and for table content, key-value pairs of table content are mapped based on the positional relationship between text units in the table to extract entity relationships; the extracted entities are used as nodes of the graph, and the relationships between entities are used as directed edges to construct a directed unweighted graph; the description information of each entity and the relationship description between it and other entities are summarized using a large model to generate an entity summary, multi-level semantic modeling is performed on the entity summary, and the semantic information of adjacent entities is aggregated as a graph embedding representation of the entity; the entity summary and its graph embedding representation are associated with the entity nodes in the directed unweighted graph, and the entity summary and its corresponding graph embedding representation are embedded in the corresponding node position in the graph structure to obtain an optimized graph; the Leiden clustering algorithm is used to divide the entities into multiple communities according to the modularity between the entities in the optimized graph; the large model is used to perform a community summary on each community to obtain a community summary; When receiving a user query request, the knowledge graph is recalled using multi-graph query and multi-path ranking fusion to obtain the entity relationship that best matches the query and provide corresponding community summary information; based on the ranking score of the retrieved graph information, the task prompt words are optimized to guide the generation model to give priority to high-priority graph information and supplement relevant content during the answering process to optimize the answer.

2. The method for enhancing retrieval in the water conservancy field based on knowledge graph according to claim 1 is characterized in that: Data processing of water conservancy documents to obtain non-table text blocks and table contents include: Identify the text content and coordinate area of ​​the smallest text unit of documents in the water conservancy field; Perform wireframe table structure recognition on the table data of the document, and form one or more tables according to the coordinate positions of the inseparable wireframes recognized on a single page; Determine whether the text unit is inside the table based on the intersection-and-union ratio of the minimum text unit area and the table area; if so, delete the text unit in the table from all the text units, so as to achieve accurate screening of the content in the table; for redundant data of table titles, figure titles, table notes, and figure notes that meet the conditions, judge based on the distance relationship between them and the preceding minimum text unit to screen out redundant data; identify and eliminate irrelevant or repeated content through strong rule matching; The obtained document with non-tabular content is split into multiple text blocks. During the block segmentation process, the title, subtitle, paragraph, table, and picture elements are used as the basis for segmentation. At the same time, the semantic structure of the document is considered for segmentation to ensure that each text block has a clear semantic structure.

3. The method for retrieval enhancement and generation in the water conservancy field based on knowledge graph according to claim 1 is characterized in that: When using the big model to extract entity relationships from non-table content, multiple questions and answers of the big model are used to identify and extract entities and relationships in the knowledge graph from the divided text blocks. The process includes: The prompt words for the entity relationship extraction task are designed based on the task description and entity extraction examples. The prompt words for the relationship extraction task control the large model to first extract the entity name entity_name, entity type entity_type and entity description entity_description in the text block, and then extract entity pairs with obvious relationships. The entity pairs include the head entity source_entity and the tail entity target_entity with a relationship, and describe the relationship in detail. To ensure that the large model does not miss any entity during the entity extraction process, each text block will undergo three independent rounds of extraction operations. Subsequently, the original text block and the extracted entity relationship triples are used as input, and the integrity of the extraction results is verified using the large model. If it is detected that the extraction result is incomplete, the entity relationship extraction continues until all entities and their relationships are fully extracted.

4. The method for retrieval enhancement and generation in the water conservancy field based on knowledge graph according to claim 1 is characterized in that: The big model is used to summarize the description information of each entity and the relationship description between it and other entities to generate an entity summary, including: taking the description information of each entity and the relationship description between it and other entities as input, using the summarization ability of the big model, and controlling the big model through summary task prompt words to summarize each entity to generate an entity summary.

5. The method for retrieval enhancement and generation in the water conservancy field based on knowledge graph according to claim 1 is characterized in that: For each community, the entities, relations and corresponding descriptions of the same community are input into the big model. The big model summarizes the community under the control of the community summary task prompt word to obtain the community summary. The community summary task prompt specifies the content and structure of the community summary, where the content of the community summary includes: the community name title representing the key entity of the community; the executive summary executive_summary of the overall structure of the community, which summarizes how the community entities are related to each other and the important points related to its entities; the score rating of the relevance of the text to water resources management, water conservancy projects, historical background, and professional community dynamics; a one-sentence explanation rating_explanation of the configuration rating score; key insights about the community findings, each of which has an internal summary insight_1_summary and an internal explanation insight_1_explanation.

6. The method for retrieval enhancement generation in the water conservancy field based on knowledge graph according to claim 1 is characterized in that: The process of using multi-graph query and multi-path ranking fusion to perform knowledge graph recall includes: Key entity information is extracted from the query, and the query and the entities it contains are encoded using the bge-m3 model. The encoded feature vector is then used to calculate the similarity with entities, entity summaries, and community summaries in multiple dimensions in each knowledge graph to evaluate the correlation between the query and the graph content, return the entity relationship that best matches the query from each knowledge graph, and provide corresponding community summary information; multi-path sorting uses the BM25 algorithm based on statistical features to preliminarily sort the correlation between the query and the multi-dimensional information in the knowledge graph, and uses the BM25 algorithm to efficiently model the term frequency and document length, and calculates the correlation score between the query and each entity, entity summary, and community summary in the graph; based on the correlation score, the query, community summary, and original text block are further encoded using the bge-rerankerv2-minicpm-layerwise model, and the sorting result is further optimized using a deep learning model.

7. The method for retrieval enhancement and generation in the field of water conservancy based on knowledge graph according to claim 1 is characterized in that: According to the ranking scores of the retrieved graph information, the generation model is guided to give priority to high-priority graph information in the answering process. For the graph information with the highest priority, the keywords "focus" or "most relevant" are added when forming the prompt words, so as to ensure that the generation model puts the graph information with the highest priority at the front of the answer under the control of the optimization task prompt words, and cites it more in the generation process; Low-priority graph information and supplementary content that is in line with the topic are used to supplement the final answer to enhance the diversity and richness of the answer.

8. The method for retrieval enhancement and generation in the water conservancy field based on knowledge graph according to claim 7 is characterized in that: The optimization task prompts include: context, question query_str and corresponding reference content answer_str. The reference content is the graph information with the greatest relevance retrieved. It is required to supplement the original answer based on the provided context content and reference content. It is required to retain every character of the reference content and add new supplementary content based on the reference content answer_str to make the answer more in-depth and multi-dimensional. The supplementary content is closely related to the reference content and cannot deviate from the subject of the reference content answer_str and the question query_str. Runxun adds more details, terms or related entries.

9. A water conservancy field retrieval enhancement generation device based on knowledge graph, characterized in that: include: At least one processing unit is connected to a storage unit via a bus unit. The storage unit is a computer-readable storage medium that can be used to store software programs, computer executable programs, and modules. The processing unit runs the software programs, computer executable programs, and modules stored in the storage unit, thereby implementing the knowledge graph-based retrieval enhancement generation method in the field of water conservancy as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, it implements the water conservancy field retrieval enhancement generation method based on the knowledge graph as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Knowledge graph recall-based agent question and answer method, device, equipment and product

    CN120578749A

  • Community analysis method and device and electronic equipment

    CN120781944A

  • Internal and external rule matching method and system facing system compliance scene

    CN121119075A

  • Enhanced retrieval generation optimization method fusing knowledge graph

    CN121350280A

  • Power grid power transformation engineering knowledge graph construction and retrieval method and system

    CN121561116A