Method and Device for Extracting Information from Subsurface Pipeline Hidden Danger Management Documents Based on Large Language Models

By constructing a thinking tree prompt template and a multi-level verification mechanism, the information extraction method of underground pipeline hidden danger management file is solved, and efficient and accurate information extraction and hidden danger identification are achieved.

CN119903025BActive Publication Date: 2025-07-22BEIJING SCI & TECH PATENT OFFICE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411993528.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-07-22
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

When handling underground pipeline hidden danger management files, the prior art has problems such as diverse data formats, dispersed information, high manual analysis costs, low efficiency and easy introduction of human errors, and large models rely on high-quality labeled data and high computing resources.

Method used

The information extraction method based on the big model is adopted, and by constructing a thinking tree prompt template and a multi-level verification mechanism, including text division, entity relationships and logic verification, information extraction is ensured in combination with pre-trained big models to ensure accuracy and efficiency.

Benefits of technology

It realizes efficient and accurate extraction of structured information from underground pipeline hidden danger management documents, improves information retrieval efficiency and hidden danger identification accuracy, and reduces manual operation and time consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119903025B_ABST
    Figure CN119903025B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for extracting file information on hidden dangers of underground pipelines based on a large model. The method includes: dividing the text in the file on hidden dangers of underground pipelines into text blocks; constructing a thinking tree prompt template; inputting the text blocks and the thinking tree prompt template into the large model, so that the large model analyzes the text blocks based on the thinking tree prompt template, realizes the extraction of entity relationship triple information of the text and outputs the result; constructing an entity relationship verification prompt template and a logical verification prompt template; inputting the output result and the entity relationship verification prompt template into the large model to verify whether each triple is accurately extracted, and calculating the verification score of the triple according to the verification result; inputting the output result and the logical verification prompt template into the large model to verify whether the extraction process is logically reasonable, and calculating the verification score of the extraction process according to the verification result; determining the reliability of the output result according to the obtained verification scores, and efficiently and accurately realizing the extraction of file information on hidden dangers of underground pipelines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and device for extracting information from underground pipeline hidden danger management documents based on a large model. Background Art

[0002] Underground pipelines are an important part of urban infrastructure. The identification, analysis, and treatment of their hidden dangers involve a large amount of complex document and data information. These hidden danger management documents usually exist in unstructured or semi-structured forms and cover relevant knowledge in the field of hidden danger management. In addition, other documents such as underground pipeline planning documents may also contain important information related to hidden danger management, such as current national standards, design contents, etc. Such documents are collectively referred to as underground pipeline hidden danger management documents. However, the diverse data formats and scattered information make manual parsing and information extraction costly, inefficient, and prone to human errors. With the rapid development of artificial intelligence and natural language processing technologies, large model-based technologies have demonstrated powerful capabilities in the field of text understanding and information extraction. Through learning on massive data, pre-trained large models possess excellent language understanding and generation capabilities, providing new ideas for the intelligent parsing of unstructured data.

[0003] Although large models have achieved certain results in the field of massive data processing in the prior art, there are still the following limitations: on the one hand, the existing large model data processing methods require a certain amount of high-quality labeled data, which consumes a large amount of human and time costs, and to a certain extent limits the wide application of this method; on the other hand, it requires a large amount of computing resources, with high computing power costs and great implementation difficulties. Summary of the Invention

[0004] In view of the above problems, the present invention is proposed to provide a method and device for extracting information from underground pipeline hidden danger management documents based on a large model that can solve the above technical problems or at least partially solve the above technical problems.

[0005] In one aspect of the present invention, there is provided a method for extracting information from underground pipeline hidden danger management documents based on a large model, the method comprising:

[0006] Obtain the text in the underground pipeline hidden danger management document, and divide the text into a plurality of text blocks;

[0007] Construct a thinking tree prompt template, which includes a task description, output format constraints, output example descriptions, and a thinking tree structure describing the task analysis process. The task description is used to guide the perspective and output format of the large model during task execution. Each child node in the thinking tree structure represents a thinking task, and the thinking tree structure is used to guide the large model to sequentially execute the corresponding thinking tasks based on the connection relationships between the child nodes in the thinking tree structure to achieve the extraction of entity relationship triple information in the text;

[0008] Input each divided text block and the thinking tree prompt template into a preset large model, so that the large model analyzes each text block based on the thinking tree prompt template, completes the extraction of entity relationship triple information according to the context, and formats and outputs the information extraction results;

[0009] Construct an entity relationship verification prompt template, which includes a problem description for verifying the given entity pair and relationship in the entity relationship triple;

[0010] Input the output result and the entity relationship verification prompt template into the large model to verify whether each entity relationship triple in the output result is accurately extracted, and calculate the first verification score for the corresponding entity relationship triple according to the verification result of each entity relationship triple;

[0011] Construct a logical verification prompt template, which includes a problem description for verifying whether there are logical conflicts in the entire extraction process;

[0012] Input the output result and the logical verification prompt template into the large model to verify whether the information extraction process of the output result is logically reasonable, and calculate the second verification score for the information extraction process according to the logical verification result of the information extraction process;

[0013] Analyze the overall reliability score of the output result according to the second verification score of the information extraction process and the first verification score of each entity relationship triple in the output result. When the overall reliability score meets the preset score threshold, use the output result as the information extraction result of the underground pipeline hidden danger management file.

[0014] Optionally, the method further includes:

[0015] During the hidden danger information retrieval process, obtain all actual attribute values of the underground pipeline to be analyzed;

[0016] Judge whether each actual attribute value matches the information extraction result. If all actual attribute values match the information extraction result, it is determined that the underground pipeline to be analyzed meets the standards of the underground pipeline hidden danger management file and has no problem hidden danger, otherwise it is determined that the underground pipeline to be analyzed has a problem hidden danger.

[0017] Optionally, the step of determining whether each actual attribute value matches the information extraction result includes:

[0018] Determine the attribute type of each actual attribute value;

[0019] For an actual attribute value with a text attribute type, compare whether the actual value is consistent with the standard information in the information extraction result. If they are consistent, it is determined that the current actual attribute value matches the information extraction result;

[0020] For an actual attribute value with a numerical attribute type, compare the actual value with the standard value in the information extraction result. For an information extraction result with a difference range requirement, determine whether the actual value is greater than or less than the standard value, and whether the difference between the actual value and the standard value is less than the difference range requirement in the information extraction result. If the actual value is greater than or less than the standard value and the difference between the actual value and the standard value is less than the difference range requirement, it is determined that the current actual attribute value matches the information extraction result.

[0021] Optionally, the construction of the thinking tree prompt template includes:

[0022] Perform a thinking tree chain decomposition on the preset information extraction task to obtain a thinking tree structure with multiple nodes;

[0023] When constructing the root node of the thinking tree, the first prompt designed is: problem description, clarifying the goal of the extraction task;

[0024] When constructing the first sub-node of the thinking tree, the second prompt designed is: Identify the key entities and entity types in the file paragraph by paragraph, ensuring that only the entities clearly mentioned in the text are recorded. When there is only one entity in a certain category, classify that entity into the "other" category;

[0025] When constructing the second sub-node of the thinking tree, the third prompt designed is: Analyze the logical associations between entities. The relationship between entities must be a clearly expressed logical association;

[0026] When constructing the third sub-node of the thinking tree, the fourth prompt designed is: Verify the logical consistency and data integrity based on the identified entities and relationships;

[0027] When constructing the leaf node of the sub-node, the fifth prompt designed is: Finally, output a JSON-formatted result that meets the requirements.

[0028] Optionally, each child node in the thought tree structure corresponds to three different task analysis branches to simulate three different underground pipeline experts. Each task analysis branch independently completes the analysis of the task at this node based on the task analysis results of the previous node, and determines the final task analysis result of this node by sharing and discussing with other analysis branches.

[0029] Optionally, the problem description in the entity relationship verification prompt template for verifying the given entity pair and relationship in the entity relationship triple includes:

[0030] Entity pair and Is it consistent with the preset domain and context; relationship Does it correctly describe the entity and The actual connection between; and / or, whether there are missing, redundant or inconsistent entity or relationship descriptions.

[0031] Optionally, the problem description in the logical verification prompt template for verifying whether there are logical conflicts in the entire extraction process includes:

[0032] Is each entity relationship triple in the output result logically self-consistent, without contradictions or repeated descriptions; and / or, is there a reasonable context semantic association between each entity relationship triple in the output result.

[0033] Optionally, the analysis of the overall reliability score of the output result based on the second verification score of the information extraction process and the first verification score of each entity relationship triple in the output result includes:

[0034] Obtain the reliability score of the corresponding entity relationship triple by multiplying the first verification score of each entity relationship triple by the second verification score of the information extraction process;

[0035] Obtain the sum of the reliability scores of each entity relationship triple in the output result to get the overall reliability score of the output result.

[0036] Another aspect of the present invention provides an information extraction device for underground pipeline hidden danger management files based on a large model, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-mentioned information extraction method for underground pipeline hidden danger management files based on a large model.

[0037] In the third aspect of the present invention, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-mentioned information extraction method for underground pipeline hidden danger management files based on a large model.

[0038] The information extraction method and device for underground pipeline hidden danger management documents based on large models provided by the embodiments of the present invention can efficiently and accurately complete the information extraction in underground pipeline hidden danger management documents by combining the natural language processing ability of pre-trained large models with prompt engineering based on the thought tree, obtain structured information, and provide high-quality structured data support for the identification, analysis, and treatment of underground pipeline hidden dangers, improving the information retrieval efficiency and the accuracy of hidden danger identification.

[0039] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Brief Description of the Drawings

[0040] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. In the drawings:

[0041] Figure 1 is a flowchart of the information extraction method for underground pipeline hidden danger management documents based on large models according to the embodiments of the present invention;

[0042] Figure 2 is a schematic implementation diagram of the information extraction method for underground pipeline hidden danger management documents based on large models according to the embodiments of the present invention. Detailed Embodiments

[0043] Hereinafter, the exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0044] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined.

[0045] An embodiment of the present invention provides a method for extracting information from underground pipeline hidden danger management documents based on a large model, which is used to efficiently and accurately extract structured information from a large number of underground pipeline hidden danger management documents to support the retrieval of underground pipeline hidden danger knowledge and improve the information retrieval efficiency and the accuracy of hidden danger identification. As Figure 1 shown, the method for extracting information from underground pipeline hidden danger management documents based on a large model proposed in the embodiment of the present invention includes the following steps:

[0046] S1. Obtain the text in the underground pipeline hidden danger management document and divide the text into several text blocks.

[0047] In this embodiment, the underground pipeline hidden danger management document is generally presented in the form of a PDF document, including hidden danger management documents and other documents such as underground pipeline planning documents, such as current national standards, design contents, etc. Specifically, for the PDF-format underground pipeline hidden danger management document to be processed, the Python PyMuPDF library can be used to extract the text D therefrom. By using the advanced functions provided by PyMuPDF for extracting, analyzing, and operating on PDF contents, it is ensured that the extracted text format meets the requirements and will not affect subsequent processing and evaluation. Then, the extracted text D is divided into blocks of length L, where L is less than the maximum context window of the large model: , to balance the trade-off between the context size and the computational efficiency. Each text block is processed separately so that the large model can focus on each part, optimizing the accuracy and efficiency of information extraction.

[0048] S2. Construct a thought tree prompt template, where the thought tree prompt template includes a task description, an output format constraint, an output example description, and a thought tree structure describing the task analysis process. The task description is used to guide the perspective and output format of the large model during task execution. Each sub-node in the thought tree structure represents a thought task, and the thought tree structure is used to guide the large model to sequentially execute the corresponding thought tasks based on the connection relationships between the various sub-nodes in the thought tree structure to achieve the extraction of entity relationship triple information of the text.

[0049] Due to the unique language characteristics of underground pipeline hidden danger management documents, these documents usually contain professional terms, complex structures, and diverse formats. To accurately extract information from underground pipeline hidden danger management documents, the present invention obtains a thought tree structure with multiple nodes by performing a thought tree chain decomposition on a preset information extraction task when constructing the thought tree prompt template. The present invention adopts components such as a thought tree, a task description, an output format constraint, and an output example description in the development of the prompt to enhance the performance of the model in extracting information from underground pipeline hidden danger management documents. Among them,

[0050] Task description

[0051] Clarify the task description to establish a specific context, thus guiding the perspective and output style of the large model in task execution.

[0052] Entity extraction: Identify the key entities and entity types in the document.

[0053] Relationship extraction: Analyze the logical associations between entities.

[0054] Output format constraint,

[0055] Ensure that the extracted information is presented in a consistent and structured manner through a detailed description of the output format.

[0056] Please output the results in the following JSON format:

[0057] Entities: List all the entities mentioned in the document and their categories.

[0058] Relationships: Describe in detail the relationships between entities, including Entity 1, the relationship, and Entity 2.

[0059] Output example description,

[0060] Add examples to the prompt and provide specific descriptions of the expected output to further improve performance and help the model better understand the task.

[0061] Example file: "4.1.2 The overall design of the water supply and reclaimed water pipeline system shall comply with the relevant provisions of the current national standards 'Code for Design of Outdoor Water Supply' GB50013 and 'Code for Design of Urban Wastewater Reclamation and Reuse Project' GB50335."

[0062] Example output: {"Entities": {"1. Current national standards": ["Code for Design of Outdoor Water Supply" GB50013, "Code for Design of Urban Wastewater Reclamation and Reuse Project" GB50335], "2. Pipeline system overall design": ["Water supply and reclaimed water pipeline system overall design"]},

[0063] "Relationships": [["Water supply and reclaimed water pipeline system overall design", "shall comply with", "Code for Design of Outdoor Water Supply" GB50013], ["Water supply and reclaimed water pipeline system overall design", "shall comply with", "Code for Design of Urban Wastewater Reclamation and Reuse Project" GB50335]]}

[0064] Thought tree rules,

[0065] Conduct task analysis in the following node order. For each child node in the thought tree structure, there are three different task analysis branches corresponding to simulate three different underground pipeline experts. Each task analysis branch that simulates an underground pipeline expert independently completes the analysis of the task of this node based on the task analysis results of the previous node, and shares and discusses with the underground pipeline experts simulated by other analysis branches. If an expert is found to have made an incorrect analysis during this process, the expert withdraws to determine the final task analysis result of this node. Provide a detailed description of the analysis processes of the three underground pipeline experts for each node task:

[0066] When constructing the root node of the thought tree, the first prompt designed is: Problem description, clearly extract the goal of the task;

[0067] When constructing the first child node 1 of the thought tree, the second prompt designed is: Identify key entities and entity types in the document paragraph by paragraph, ensuring that only entities clearly mentioned in the text are recorded. When there is only one entity in a certain category, that entity can be classified into the "other" category;

[0068] When constructing the second child node 2 of the thought tree, the third prompt designed is: Analyze the logical associations between entities. The relationship between entities must be a clearly logically associated statement;

[0069] When constructing the third child node 3 of the thought tree, the fourth prompt designed is: Verify logical consistency and data integrity based on the identified entities and relationships;

[0070] When constructing the leaf node of the child node, the fifth prompt designed is: Finally, output the JSON formatted result that meets the requirements.

[0071] Complete the analysis of all nodes and finally output the result.

[0072] In addition, the thought tree prompt template also includes separator descriptions: Use separators such as "***" and "===" to clearly separate and highlight the key parts of the prompt, helping the model distinguish different parts of the instruction, and maintaining attention and goals for each part.

[0073] S3. Input each of the divided text blocks and the thought tree prompt template into a preset large model, so that the large model analyzes each text block based on the thought tree prompt template, extracts the entity relationship triple information of the text according to the context, and formats and outputs the information extraction result.

[0074] In this embodiment, after obtaining the thought tree prompt template, input the thought tree prompt template and the text block into the large model, let it perform step-by-step analysis of the task in the thought tree manner, complete the information extraction of the text according to the context, and format and output it in json form.

[0075] S4. Construct an entity relationship verification prompt template, where the entity relationship verification prompt template includes a problem description for verifying a given entity pair and relationship in an entity relationship triple.

[0076] S5. Input the output result and the entity relationship verification prompt template into a large model to verify whether each entity relationship triple in the output result is accurately extracted, and calculate the first verification score for the corresponding entity relationship triple according to the verification result of each entity relationship triple.

[0077] In this embodiment, after obtaining the formatted output result, an entity relationship verification (Q1) process is further included. Specifically, by constructing an entity relationship verification prompt template, that is, an entity relationship verification prompt, the output result and the entity relationship verification prompt are input into a large model to verify whether each extracted entity relationship triple is accurate. The large model outputs Yes / No to evaluate whether the entity and its relationship are accurately extracted.

[0078] S6. Construct a logical verification prompt template, where the logical verification prompt template includes a problem description for verifying whether there are logical conflicts in the entire extraction process.

[0079] S7. Input the output result and the logical verification prompt template into a large model to verify whether the information extraction process of the output result is logically reasonable, and calculate the second verification score of the information extraction process according to the logical verification result of the information extraction process.

[0080] In this embodiment, a logical verification (Q2) process is further included. Specifically, by constructing a logical verification prompt template, that is, a logical verification prompt, the output result and the logical verification prompt are input into a large model to verify whether the extraction process is accurate. The large model outputs Yes / No to evaluate whether there are logical conflicts in the entire extraction process, and comprehensively verify all extracted triples to ensure that the process extraction results are consistent.

[0081] S8. Analyze the overall reliability score of the output result according to the second verification score of the information extraction process and the first verification score of each entity relationship triple in the output result. When the overall reliability score meets a preset score threshold, the output result is used as the information extraction result of the underground pipeline hidden danger management document.

[0082] In the embodiment of the present invention, in order to ensure the absolute accuracy of the information extraction result, the overall reliability score can be selected as the total score corresponding to when both entity relationship verification and logical verification pass.

[0083] The information extraction method and device for underground pipeline hidden danger management documents based on large models provided by the embodiments of the present invention can efficiently and accurately complete the information extraction in underground pipeline hidden danger management documents by combining the natural language processing ability of pre-trained large models with prompt engineering based on the tree of thought, obtain structured information, provide high-quality structured data support for the work of identifying, analyzing and treating underground pipeline hidden dangers, and improve the information retrieval efficiency and the accuracy of hidden danger identification.

[0084] In the embodiments of the present invention, the question description in the entity relationship verification prompt template for verifying the given entity pair and relationship in the entity relationship triple includes:

[0085] Entity pair and Is it consistent with the preset domain and context?

[0086] Relationship Does it correctly describe the actual connection between entities and ? And / or,

[0087] Are there any missing, redundant or inconsistent entity or relationship descriptions?

[0088] Verify whether all statements are satisfied, answer Yes / No.

[0089] Entity pair and The verification score of is expressed as , as shown in Equation (1).

[0090] (1)

[0091] where ANS_1 represents the reply of the large model to Q1.

[0092] In the embodiments of the present invention, the question description in the logical verification prompt template for verifying whether there are logical conflicts in the entire extraction process includes:

[0093] Is each entity relationship triple in the output result logically self-consistent, without contradictions or repeated descriptions?

[0094] And / or, is there a reasonable context semantic association, i.e., context consistency, between each entity relationship triple in the output result?

[0095] According to the extracted triple sequence, verify whether the entire extraction process is logically reasonable, answer Yes / No.

[0096] The verification score of the overall process consistency is expressed as V2_score(Q2), as shown in Equation (2).

[0097] (2)

[0098] Among them, ANS_2 represents the response of the large model to Q2.

[0099] Furthermore, analyzing the overall reliability score of the output result according to the second verification score of the information extraction process and the first verification score of each entity-relationship triple in the output result specifically includes: obtaining the product of the first verification score of each entity-relationship triple and the second verification score of the information extraction process to obtain the reliability score of the corresponding entity-relationship triple; obtaining the sum of the reliability scores of all entity-relationship triples in the output result to obtain the overall reliability score of the output result.

[0100] In this embodiment, the extraction result of an entity-relationship triple of the reliability score is expressed as the product of V1_score of the triple and V2_score of the whole process, as shown in formula (3). Finally, the overall reliability score is obtained by summing up the reliability scores of all extraction results, as shown in formula (4).

[0101] (3)

[0102] (4)

[0103] Where n represents the number of triples, the triples are sorted in sequence starting from 1, and k represents the serial number of the triple.

[0104] In the embodiment of the present invention, the method further includes a hidden danger information retrieval step. Specifically, during the hidden danger information retrieval process, all actual attribute values of the underground pipeline to be analyzed are obtained; it is judged whether each actual attribute value matches the information extraction result. If all actual attribute values match the information extraction result, it is determined that the underground pipeline to be analyzed meets the standard of the underground pipeline hidden danger management document and there is no problem hidden danger, otherwise it is determined that the underground pipeline to be analyzed has a problem hidden danger.

[0105] Further, determining whether each actual attribute value matches the information extraction result specifically includes: determining the attribute type of each actual attribute value; for an actual attribute value with a text attribute type, comparing the actual value with the standard information in the information extraction result. If they are consistent, it is determined that the current actual attribute value matches the information extraction result; for an actual attribute value with a numerical attribute type, comparing the actual value with the standard value in the information extraction result. For an information extraction result with a difference range requirement, it is determined whether the actual value is greater than or less than the standard value, and whether the difference between the actual value and the standard value is less than the difference range requirement in the information extraction result. If the actual value is greater than or less than the standard value and the difference between the actual value and the standard value is less than the difference range requirement, it is determined that the current actual attribute value matches the information extraction result. For an information extraction result that does not have a difference range requirement but has a greater than or less than numerical requirement, it can directly determine whether the comparison between the actual value and the standard value in the information extraction result meets the corresponding numerical requirement in the information extraction result. If it meets the requirement, it is determined that the current actual attribute value matches the information extraction result.

[0106] In this embodiment, during the hidden danger information retrieval process, the actual situation of the underground pipeline can be combined with the information extracted by the system to conduct targeted analysis on the possible problems of the underground pipeline, so as to help locate high-risk areas, shorten the time for hidden danger investigation, and improve work efficiency.

[0107] Specifically, by determining whether the actual attribute value of the underground pipeline conforms to the extracted information, the hidden dangers with problems are inspected. The formula can be described by a Boolean judgment:

[0108] Logical establishment = Condition 1 ∧ Condition 2 ∧... ∧ Condition N (5)

[0109] Among them, the judgment basis for each condition depends on the specific attribute type, and different Boolean operation methods are adopted:

[0110] Text attribute: By comparing whether the actual value is consistent with the standard information, it is determined whether the requirement is met. When the comparison between the actual value and the standard information is consistent, it is determined that the requirement is met.

[0111] Numerical attribute: By comparing the actual value with the standard value, it can be determined whether the following conditions are met: whether the actual value is greater than or less than the set standard value; whether the difference between the actual value and the standard value is less than the extracted difference range requirement.

[0112] In a specific example:

[0113] 1. Standard with difference range requirement:

[0114] For some standards, there may be an allowable range of differences. For example, if a standard requires that the difference between the actual value and the designed value of the elevation at the top of a waterway must be within ±0.5 meters, then:

[0115] If the difference between the actual value and the designed value is between -0.5 meters and +0.5 meters, it can be said that the actual elevation of this waterway matches the design requirements.

[0116] If the actual value is more than 0.5 meters higher than the designed value or more than 0.5 meters lower than the designed value, then this waterway does not meet the requirements of this difference range and cannot be considered a match.

[0117] At this time, a difference of zero is within this range, and it is determined to be a match.

[0118] 2. For those without a requirement for a difference range but need to be greater than or less than a certain standard, such as requiring that the actual value of the elevation at the top of Class I to Class V waterways should be greater than 2.0 m. When the actual value meets this requirement, it is considered a match.

[0119] At this time, if the actual value is equal to the standard value of 2.0 m instead of greater than 2.0 m, it is judged as not matching.

[0120] Figure 2 The schematic diagram of the implementation principle of the method for extracting information from underground pipeline hidden danger management files based on large models of the present invention is shown. Figure 2 The specific implementation process of the present invention through text segmentation, prompt engineering, large model information extraction, verification mechanism, and hidden danger retrieval is shown. The method for extracting information from underground pipeline hidden danger management files based on large models of the present invention has the following advantages: by combining the natural language processing capabilities of pre-trained large models, the system can efficiently and automatically process unstructured data, accurately extract information from complex underground pipeline hidden danger management files, and output it in a structured format, greatly reducing manual operations and time consumption; the Prompt engineering design based on the thought tree ensures that the model can handle the professional terms and complex structures of underground pipeline hidden danger management files and achieve highly customized processing; the present invention also uses a multi-level verification mechanism to ensure the accuracy and reliability of information extraction and ultimately improve the overall performance. In addition, this method only depends on the file content and does not use external prior knowledge, maintaining the objectivity and accuracy of the extraction results. The method of text block processing further optimizes the utilization of context and improves the extraction effect. It can provide support for retrieving underground pipeline hidden danger knowledge to significantly improve the information retrieval efficiency and the accuracy of hidden danger identification.

[0121] Taking the "Code for Urban Engineering Pipeline Comprehensive Planning" as an example, the information is extracted to illustrate the method for extracting information from underground pipeline hidden danger management documents based on large models of the present invention, so as to describe in detail the specific implementation processes such as text preprocessing, prompt engineering, large model information extraction, verification mechanism, and hidden danger retrieval.

[0122] The following is part of the text of the "Code for Urban Engineering Pipeline Comprehensive Planning":

[0123] "4.1.8 The engineering pipelines laid at the bottom of the river should be selected in stable river sections, and the pipeline elevation should be determined according to the principle of not hindering the regulation of the river course and pipeline safety, and should comply with the following provisions:

[0124] 1 When laid under the waterways of Class I to Class V, the top elevation should be below 2.0 m of the bottom elevation of the long-term planned waterway;

[0125] 2 When laid under the waterways of Class VI and Class VII, the top elevation should be below 1.0 m of the bottom elevation of the long-term planned waterway;

[0126] 3 When laid under other river courses, the top elevation should be below 0.5 m of the designed bottom elevation of the river course"

[0127] I. Text preprocessing,

[0128] Step a1: Text extraction,

[0129] For the PDF file to be processed, use the PyMuPDF library of Python to extract the text D from it.

[0130] Step a2: Text division,

[0131] Select an appropriate text block length L = 300, and divide the extracted text D into blocks of size L:

[0132] = "3.0.1 Urban engineering pipeline comprehensive planning... laid underground in ways such as."

[0133] II. Prompt engineering,

[0134] Adopt a method for constructing a prompt template based on the thought tree. The following is an example of a prompt:

[0135] "The task is to extract all relevant entities and their explicit relationships from the file and output the results in a structured manner.

[0136] Entity extraction: Identify the key entities and entity types in the file.

[0137] Relationship extraction: Analyze the logical associations between entities.

[0138] Submit the final response in JSON format. The keys should be "Entity" categories or "Relationship", and the values should be arrays of entities or relationships of that category. Divide the description of the output example with ===:

[0139] ===

[0140] Example file: "4.1.2 The overall design of the water supply and reclaimed water pipeline system shall comply with the relevant provisions of the current national standards 'Code for Design of Outdoor Water Supply' GB50013 and 'Code for Design of Urban Sewage Reclamation and Reuse Project' GB50335."

[0141] Example output: {"<Entity>":{"<1. Current national standards>":["Code for Design of Outdoor Water Supply" GB50013, "Code for Design of Urban Sewage Reclamation and Reuse Project" GB50335],"<2. Overall design of pipeline system>":["Overall design of water supply and reclaimed water pipeline system"]},

[0142] "<Relationship>": [[Overall design of water supply and reclaimed water pipeline system, shall comply with, "Code for Design of Outdoor Water Supply" GB50013], [Overall design of water supply and reclaimed water pipeline system, shall comply with, "Code for Design of Urban Sewage Reclamation and Reuse Project" GB50335]]}

[0143] ===

[0144] Analyze in the following node order. For each node task, assume there are three different underground pipeline experts. Each expert independently completes the analysis of the node task based on the processing results of the previous node, then shares and discusses with other experts. If an expert is found to have made an incorrect analysis during this process, that expert withdraws. Give a detailed description of the analysis process of the three experts for each node task. Divide the description of this part with ***:

[0145] ***

[0146] Root node: Problem description, clarifying the goal of the extraction task.

[0147] Child node 1: Identify key entities and entity types in the file paragraph by paragraph, ensuring that only entities explicitly mentioned in the text are recorded. When there is only one entity in a certain category, that entity can be classified into the "Other" category.

[0148] Child node 2: Analyze the logical associations between entities. The relationships between entities must be a clearly logically related statement.

[0149] Child node 3: Verify the logical consistency and data integrity based on the identified entities and relationships.

[0150] Leaf node: The final output is a JSON-formatted result that meets the requirements.

[0151] Complete the analysis of all nodes and output the final result.

[0152] ***”

[0153] III. Information Extraction from Large Models

[0154] Step c1: Input the Prompt and the text block into the large model, let it complete a detailed description of the analysis process in the way of a thought tree, finally generate relationship triples according to the context, and output them in JSON format.

[0155] The following is an example of the large model's answer:

[0156] "*Analysis of the root node: After receiving the task, three underground pipeline experts first clarified the goal of the extraction task, that is, to extract all relevant entities and their clear relationships from the given text, and output the result in a structured way as JSON format. The experts agreed that this required them to identify the key entities and entity types in the document, analyze the logical associations between entities, and ensure the logical consistency and data integrity of the output.

[0157] *Analysis of child node 1: Experts A, B, and C began to identify the key entities and entity types in the document paragraph by paragraph. They noticed that the text mentioned different waterway levels and river channels, as well as the elevation requirements for the laying of engineering pipelines related to these waterways and river channels.

[0158] The entities identified by Expert A include: "Class I to V waterways", "Class VI and VII waterways", "other river channels", "engineering pipelines", "bottom elevation of the long-term planned waterway".

[0159] The entities identified by Expert B are similar to those of Expert A, but further divide "engineering pipelines" into "engineering pipelines laid at the bottom of the river".

[0160] The entities identified by Expert C are the same as those of Experts A and B, but further divide "engineering pipelines" into "engineering pipelines laid at the bottom of the river", and divide "river channels" into "Class I to V river channels", "Class VI and VII river channels", and "other river channels".

[0161] After discussion, the experts agreed that the subdivision of Expert C was more accurate, so the entity classification of Expert C was adopted.

[0162] *Analysis of child node 2: Next, the experts analyzed the logical associations between entities. They noticed that the relationships between entities were based on different waterway levels and river channels, and there were different requirements for the elevation of engineering pipeline laying.

[0163] Expert A believes the relationship is: The elevation of the "engineering pipelines laid at the river bottom" needs to be determined according to the waterway level.

[0164] Expert B believes the relationship includes: The elevation requirement of the "engineering pipelines laid at the river bottom" is related to the "bottom elevation of the long-term planned waterway".

[0165] Expert C believes the relationship is: The elevation requirement of the "engineering pipelines laid at the river bottom" is related to the "waterway level" and the "river course".

[0166] After discussion, the experts agreed that Expert C's analysis is more comprehensive, so Expert C's relationship description was adopted.

[0167] *Analysis of child node 3: Based on the identified entities and relationships, the experts verified the logical consistency and data integrity. They checked whether the descriptions of the entities and relationships were consistent with the text content and whether all relevant information was covered.

[0168] Experts A and B did not find any logical contradictions during the verification process.

[0169] During the verification process, Expert C found that although the descriptions of the entities and relationships were consistent with the text content, the mention of "stable river section" was missing, which might be an important entity.

[0170] After correction, the experts agreed to add "stable river section" to the entities.

[0171] *Output of leaf node: Finally, the experts obtained the following JSON-formatted result:

[0172] {

[0173] "entities": {

[0174] "Waterway level": ["Class I - V waterways", "Class VI and VII waterways"],

[0175] "River course": ["Stable river section"],

[0176] "Engineering pipelines": ["Engineering pipelines laid at the river bottom"],

[0177] "Elevation requirement": ["Bottom elevation of the long-term planned waterway"]

[0178] },

[0179] "relationships":

[0180] ["Engineering pipelines laid at the river bottom", "should be selected in", "Stable river section"],

[0181] ["For Class I - V waterways", "the top elevation should be below", "2.0 m below the bottom elevation of the long-term planned waterway"],

[0182] ["For Class VI and Class VII waterways", "the top elevation should be below 1.0 m of the bottom elevation of the waterway in the long-term planned waterway",]

[0183] ["For other rivers", "the top elevation should be below 0.5 m of the designed bottom elevation of the river",]

[0185] }

[0186] IV. Verification mechanism

[0187] Step d1: Input the output result and the entity-relationship verification prompt into the large model to verify whether the extraction is accurate.

[0188] An example of the entity-relationship verification prompt is as follows:

[0189] "Given the entity pair and the relationship , verify whether the following statements are accurate:

[0190] Are the entities and consistent with the given domain and context?

[0191] Does the relationship correctly describe the actual connection between the entities and ?

[0192] Are there any missing, redundant, or inconsistent entity or relationship descriptions?

[0193] Verify whether all statements are satisfied, and answer Yes / No"

[0194] The final output is Yes.

[0195] Step d2: Input the output triple sequence result and the logical verification prompt into the large model to evaluate whether there are logical conflicts in the entire extraction process.

[0196] An example of the logical verification prompt is as follows:

[0197] "Given the following extracted triples , verify whether there are logical conflicts in the entire extraction process:

[0198] Are each of the triples logically self-consistent (no contradictions or duplicate descriptions)?

[0199] Is there a reasonable semantic association (context consistency) between each of the triples?

[0200] ​Verify whether the entire extraction process is logically reasonable and answer "Yes / No".

[0201] The final output is Yes.

[0202] Step d3: Obtain the overall reliability score by summing the reliability scores of all extraction results. The score range is 0 - n (n represents the number of relationships, which is 4 in this example). The final score is 4, indicating that the extracted information has high reliability.

[0203] Step d4: Check for potential problems by determining whether the text-based attributes are equal to the extracted standard values and whether the numerical deviation conditions are less than the threshold.

[0204] In the example, the judgment of text-based attribute conditions includes: The actual selected location of the engineering pipeline laid at the bottom of the river is a stable section, which meets the extracted information standard.

[0205] In the example, the judgment of numerical deviation conditions includes:

[0206] The actual value of the top elevation of Class I - V waterways is 2.5m > 2.0m, which meets the extracted information standard.

[0207] The actual value of the top elevation of Class VI and VII waterways is 1.2m > 1.0m, which meets the extracted information standard.

[0208] The actual value of the top elevation of other river channels is 1m > 0.5m, which meets the extracted information standard.

[0209] For the method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0210] Another embodiment of the present invention also provides an information extraction device for underground pipeline potential hazard management files based on a large model, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the information extraction method for underground pipeline potential hazard management files based on a large model as described in the above embodiments.

[0211] For the device embodiments, since the process of implementing the information extraction for underground pipeline potential hazard management files based on a large model is basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments, and it has corresponding technical effects.

[0212] Another embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for extracting underground pipeline hidden danger management file information based on a large model in the above embodiment are implemented.

[0213] The method and device for extracting underground pipeline hidden danger management file information based on a large model provided by the present invention can efficiently and accurately complete the information extraction in the underground pipeline hidden danger management file and obtain structured information by combining the natural language processing ability of the pre-trained large model and the prompt engineering based on the tree of thought, so as to provide high-quality structured data support for the work of identifying, analyzing and treating underground pipeline hidden dangers, and improve the information retrieval efficiency and the accuracy of hidden danger identification.

[0214] In addition, those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, any one of the claimed embodiments can be used in any combination.

[0215] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for extracting information from underground pipeline hidden danger management documents based on a large model, characterized in that, The method includes: Obtain the text in the underground pipeline hidden danger management file, and divide the text into several text blocks; Construct a thinking tree prompt template, which includes a task description, output format constraints, output example descriptions, and a thinking tree structure describing the task analysis process. The task description is used to guide the perspective and output format of the large model during task execution. Each sub-node in the thinking tree structure represents a thinking task, and the thinking tree structure is used to guide the large model to sequentially execute the corresponding thinking tasks based on the connection relationships between the sub-nodes in the thinking tree structure to achieve the extraction of entity relationship triple information of the text; Input each divided text block and the thinking tree prompt template into a preset large model, so that the large model analyzes each text block based on the thinking tree prompt template, completes the extraction of entity relationship triple information of the text according to the context, and formats and outputs the information extraction results; Construct an entity relationship verification prompt template, which includes a question description for verifying the given entity pair and relationship in the entity relationship triple; Input the output result and the entity relationship verification prompt template into the large model to verify whether each entity relationship triple in the output result is accurately extracted, and calculate the first verification score corresponding to each entity relationship triple according to the verification result of each entity relationship triple; Construct a logical verification prompt template, which includes a question description for verifying whether there are logical conflicts in the entire extraction process; Input the output result and the logical verification prompt template into the large model to verify whether the information extraction process of the output result is logically reasonable, and calculate the second verification score of the information extraction process according to the logical verification result of the information extraction process; Analyze the overall reliability score of the output result according to the second verification score of the information extraction process and the first verification score of each entity relationship triple in the output result. When the overall reliability score meets the preset score threshold, use the output result as the information extraction result of the underground pipeline hidden danger management file.

2. The method according to claim 1, characterized in that, The method further includes: During the hidden danger information retrieval process, obtain all actual attribute values of the underground pipeline to be analyzed; Judge whether each actual attribute value matches the information extraction result. If all actual attribute values match the information extraction result, it is determined that the underground pipeline to be analyzed meets the standards of the underground pipeline hidden danger management file and has no problem hidden danger, otherwise it is determined that the underground pipeline to be analyzed has a problem hidden danger.

3. The method according to claim 2, wherein The judgment of whether each actual attribute value matches the information extraction result includes: Judge the attribute type of each actual attribute value; For the actual attribute value with the attribute type of text attribute, compare the actual value with the standard information in the information extraction result. If they are consistent, it is determined that the current actual attribute value matches the information extraction result; For the actual attribute value whose attribute type is numerical attribute, compare the actual value with the standard value in the information extraction result. For the information extraction result with a requirement for the difference range, determine whether the actual value is greater than or less than the standard value, and whether the difference between the actual value and the standard value is less than the difference range requirement in the information extraction result. If the actual value is greater than or less than the standard value and the difference between the actual value and the standard value is less than the difference range requirement, it is determined that the current actual attribute value matches the information extraction result.

4. The method according to claim 1, characterized in that The constructed thinking tree prompt template includes: Performing a thinking tree chain decomposition on the preset information extraction task to obtain a thinking tree structure with multiple nodes; When constructing the root node of the thinking tree, the first prompt designed is: problem description, clarifying the goal of the extraction task; When constructing the first sub-node of the thinking tree, the second prompt designed is: identify the key entities and entity types in the document paragraph by paragraph, ensuring that only the entities explicitly mentioned in the text are recorded. When there is only one entity in a certain category, classify the entity into the "other" category; When constructing the second sub-node of the thinking tree, the third prompt designed is: analyze the logical associations between entities. The relationship between entities must be an explicitly logically associated statement; When constructing the third sub-node of the thinking tree, the fourth prompt designed is: verify the logical consistency and data integrity based on the identified entities and relationships; When constructing the leaf node of the sub-node, the fifth prompt designed is: finally output the JSON formatted result that meets the requirements.

5. The method according to claim 1, wherein In the thinking tree structure, each sub-node corresponds to three different task analysis branches to simulate three different underground pipeline experts. Each task analysis branch independently completes the analysis of the task of this node based on the task analysis result of the previous node, and determines the final task analysis result of this node by sharing and discussing with other analysis branches.

6. The method according to claim 1, characterized in that The problem description in the entity relationship verification prompt template for verifying the given entity pair and relationship in the entity relationship triple includes: Entity pair e i and e j is consistent with the preset domain and context; relationship r ij correctly describes entity e i and e j the actual connection between; and / or, whether there are missing, redundant, or inconsistent entity or relationship descriptions.

7. The method according to claim 1, wherein The problem description in the logical verification prompt template for verifying whether there are logical conflicts in the entire extraction process includes: Whether each entity relationship triple in the output result is logically self-consistent, without contradictions or duplicate descriptions; and / or, whether there is a reasonable context semantic association between each entity relationship triple in the output result.

8. The method according to claim 1, wherein The analysis of the overall reliability score of the output result based on the second verification score of the information extraction process and the first verification score of each entity relationship triple in the output result includes: Obtaining the product of the first verification score of each entity relationship triple and the second verification score of the information extraction process to obtain the reliability score of the corresponding entity relationship triple; Obtaining the sum of the reliability scores of each entity relationship triple in the output result to obtain the overall reliability score of the output result.

9. An underground pipeline hidden danger management file information extraction device based on a large model, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 8.