Knowledge-driven underground space information retrieval method, system and equipment
By constructing an underground space entity knowledge graph, semantic knowledge base and domain knowledge graph, and combining with large language models for searching, the problem of lack of semantic understanding and knowledge reasoning for underground space information management in the existing technology is solved, and efficient and accurate underground space information retrieval and decision-making support are achieved.
Patent Information
- Application Number
- CN202510694224.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The existing technology lacks semantic understanding and knowledge reasoning capabilities when managing and utilizing underground space geographic information, and is difficult to effectively manage and utilize underground pipeline knowledge. Unstructured text is stored scatteredly and lacks unified semantic correlation, resulting in inefficient cross-document knowledge retrieval.
By constructing an underground space entity knowledge graph, semantic knowledge base and domain knowledge graph, design a dual-channel search framework for semantic knowledge base and knowledge graph, combining large language models for key entity information recognition, knowledge graph traversal and semantic side search, to achieve the fusion and generation of search results.
The balance between accuracy and interpretability of search results is achieved, the depth and accuracy of knowledge inference of large language models is improved, and comprehensive data knowledge support and decision-making support are provided for underground space information retrieval and operation and maintenance.
Smart Images

Figure CN120216612A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and underground space information management, and particularly to a knowledge-driven underground space information retrieval method, system and device. Background Art
[0002] Urban underground space resources are rich, including underground pipelines, underground buildings, etc. Effectively managing and utilizing underground space geographic information is of great significance for the operation and maintenance of underground space.
[0003] There are mainly the following several underground space information management technologies in the prior art: The urban underground pipeline management system based on Geographic Information System (GIS). This solution uses GIS technology to establish an underground pipeline database, records the spatial location, attribute information, etc. of the pipelines, and conducts spatial analysis and queries. The underground space pipe network modeling technology based on the entity geometry method. This solution designs a system concept model and a data organization model of the underground space pipe network based on the entity geometry method and a three-dimensional topology integration model, uses an adjacency list to store the topological relationship, and finally stores the data in a raster-vector integrated data structure. The domain knowledge management based on knowledge graph. The main technical solution is to perform semi-automatic triple extraction through an artificially defined ontology schema, store entity relationships using a graph database, and use large language models and semantic matching for knowledge query and reasoning.
[0004] Since the operation and maintenance of underground space involve complex engineering structures, multi-source heterogeneous data and dynamic risk factors, the existing knowledge management technologies, such as the management system and modeling solution based on GIS, lack semantic understanding and knowledge reasoning capabilities. Although they can effectively express and manage spatial information, they lack semantic understanding and knowledge reasoning capabilities and are difficult to effectively manage and utilize underground pipeline knowledge. For the semantic knowledge base and knowledge graph solutions, although they can store and reason about knowledge, they lack the ability to express spatial information, are difficult to effectively manage and utilize underground space geographic information, and unstructured texts such as national standards, local specifications, accident cases, etc. are stored dispersedly, lacking unified semantic associations, and the unstructured data extraction ability is insufficient, resulting in low efficiency of cross-document knowledge retrieval. Summary of the Invention
[0005] In view of the above problems, the present invention is proposed to provide a knowledge-driven underground space information retrieval method, system and device that overcome the above problems or at least partially solve the above problems.
[0006] In one aspect of the present invention, a knowledge-driven underground space information retrieval method is provided, and the method includes: Invoke a large language model to identify and parse key entity information in the user's underground space operation and maintenance scenario instructions to understand the instruction intention of the current instruction. The key entity information includes spatio-temporal geographical objects and entity words other than the spatio-temporal geographical objects; Retrieve the spatio-temporal geographical object in the preset underground space entity knowledge graph, and traverse the characteristic attributes, association relationships, and relationships with other entity nodes of the spatio-temporal geographical object based on the knowledge graph relationships of the spatio-temporal geographical object to obtain an entity information retrieval result; The underground space entity knowledge graph is an entity model formed by semantically storing spatio-temporal geographical objects with spatio-temporal semantic information in the underground space field in the form of a knowledge graph; Retrieve relevant entities from the preset domain knowledge graph according to the key entity information, and traverse all relevant entities and association relationships along the graph relationship chain to obtain a local retrieval result on the graph side; The domain knowledge graph includes multiple communities, and each community is formed by graph clustering of entity-relationship-entity triples extracted from the preset underground space domain knowledge data and empirical knowledge data; Match relevant community summaries as context data from the community summaries of each community in the domain knowledge graph according to the key entity information, and invoke a large language model to generate a global retrieval result on the graph side based on the context data; Use the local retrieval result on the graph side and / or the global retrieval result on the graph side as vector retrieval input, and match the corresponding text unit vector representation from the preset semantic knowledge base to obtain a retrieval result on the semantic side; Fuse the entity information retrieval result, the local retrieval result on the graph side, the global retrieval result on the graph side, and the retrieval result on the semantic side to form a fused retrieval result; Input the fused retrieval result into a large language model to generate a final retrieval result.
[0007] On the other hand, the present invention also provides a knowledge-driven underground space information retrieval system, the system includes: An instruction parsing module, configured to invoke a large language model to identify and parse key entity information in the user's underground space operation and maintenance scenario instructions to understand the instruction intention of the current instruction. The key entity information includes spatio-temporal geographical objects and entity words other than the spatio-temporal geographical objects; An entity information retrieval module, configured to retrieve the spatio-temporal geographical object in the preset underground space entity knowledge graph, and traverse the characteristic attributes, association relationships, and relationships with other entity nodes of the spatio-temporal geographical object based on the knowledge graph relationships of the spatio-temporal geographical object to obtain an entity information retrieval result; The underground space entity knowledge graph is an entity model formed by semantically storing spatio-temporal geographical objects with spatio-temporal semantic information in the underground space field in the form of a knowledge graph; The local graph retrieval module is used to retrieve relevant entities from a preset domain knowledge graph according to the key entity information, and traverse all relevant entities and association relationships along the graph relationship chain to obtain the local graph retrieval result on the graph side; the domain knowledge graph includes multiple communities, and each community is formed by graph clustering of entity-relationship-entity triples extracted from preset underground space domain knowledge data and empirical knowledge data; The global graph retrieval module is used to match relevant community summaries as context data from the community summaries of each community in the domain knowledge graph according to the key entity information, and call a large language model to generate the global graph retrieval result on the graph side based on the context data; The semantic retrieval module is used to use the local graph retrieval result and / or the global graph retrieval result on the graph side as vector retrieval input, and match the corresponding text unit vector representation from a preset semantic knowledge base to obtain the semantic side retrieval result; The fusion module is used to fuse the entity information retrieval result, the local graph retrieval result on the graph side, the global graph retrieval result on the graph side, and the semantic side retrieval result to form a fusion retrieval result; The retrieval result generation unit is used to input the fusion retrieval result into a large language model to generate the final retrieval result.
[0008] On the other hand, the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor; when the computer program is executed by the processor, the steps of the knowledge-driven underground space information retrieval method described above are implemented.
[0009] On the other hand, the present invention also provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the knowledge-driven underground space information retrieval method described above are implemented.
[0010] The knowledge-driven underground space information retrieval method, system and device provided by the embodiments of the present invention combine the advantages of the underground space entity knowledge graph, the semantic knowledge base, and the domain knowledge graph. By constructing a dual-driven information model of specification guidelines and empirical knowledge, and designing a dual-channel retrieval framework for the semantic knowledge base and the knowledge graph, the balance between the accuracy and interpretability of the retrieval result is achieved, the depth and accuracy of the knowledge reasoning of the large language model are improved, and comprehensive data knowledge support and decision support are provided for underground space information retrieval and operation and maintenance.
[0011] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically describes the specific embodiments of the present invention. Brief Description of the Drawings
[0012] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered as a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to denote the same components. In the drawings: Figure 1 is a flowchart of the knowledge-driven underground space information retrieval method provided by an embodiment of the present invention; Figure 2 is a schematic diagram of the pipeline entity knowledge graph provided by an embodiment of the present invention; Figure 3 is a schematic diagram of the underground space domain knowledge graph proposed by an embodiment of the present invention; Figure 4 is a block diagram of the structure of the knowledge-driven underground space information retrieval system proposed by an embodiment of the present invention. Detailed Embodiments
[0013] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0014] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the recited features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups.
[0015] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined.
[0016] Figure 1 Schematically shows a flowchart of the knowledge-driven underground space information retrieval method according to an embodiment of the present invention. Refer to Figure 1, the knowledge-driven underground space information retrieval method according to the embodiments of the present invention specifically includes the following steps: S11. Invoke a large language model to identify and parse key entity information in the user's underground space operation and maintenance scenario instruction to understand the instruction intention of the current instruction. The key entity information includes spatio-temporal geographical objects and entity words other than the spatio-temporal geographical objects.
[0017] Specifically, when the user issues an underground space operation and maintenance scenario instruction, the current instruction intention is understood by invoking a large language model. Optionally, the instruction intention includes conventional questions, specification query questions, fault analysis questions, scenario decision questions, etc. The key entity information includes spatio-temporal geographical objects (such as pipelines, pipe points, buildings, etc.) and other relevant entities (such as time, events, attributes, etc.) in the instruction. For example, for the instruction "What are the maintenance specifications of pipeline PID-2007567", the spatio-temporal geographical object "pipeline PID-2007567" is identified, and "maintenance specifications" is an entity word other than the spatio-temporal geographical object. The extracted entities include "pipeline" and "maintenance specifications", and the instruction intention is parsed as a specification query type.
[0018] S12. Retrieve the spatio-temporal geographical object in the preset underground space entity knowledge graph, and traverse the characteristic attributes, association relationships, and relationships with other entity nodes of the spatio-temporal geographical object based on the knowledge graph relationships of the spatio-temporal geographical object to obtain an entity information retrieval result. The underground space entity knowledge graph is an entity model formed by semantically storing spatio-temporal geographical objects with spatio-temporal semantic information in the underground space field in the form of a knowledge graph. Among them, the characteristic attributes include spatial characteristics, time characteristics, and attribute characteristics.
[0019] In this embodiment, pipelines are taken as an example for illustration. Among them, the spatial characteristic is the digital absolute position of the pipeline, including longitude (X), latitude (Y), and depth (Z). These coordinates can be obtained through a GPS or GIS system and are represented using the WGS84 coordinate system.
[0020] The time characteristics include the construction timestamp and the most recent detection time. The timestamp can be represented in the ISO 8601 standard as T=YYYY-MM-DDTHH:MM:SSZ. For example, 2007-09-06T16:51:00 represents 4:51:00 pm on September 16, 2007.
[0021] The attribute characteristics include material, pipe diameter, length, and construction unit. These attributes can be stored and queried by establishing a database table or in XML / JSON format.
[0022] The association relationships include temporal association, spatial association, and semantic association. Among them, temporal association refers to establishing the temporal association relationship between spatio-temporal geographical objects. For example, the construction time, maintenance time, usage time, etc. of pipelines, and the spatio-temporal changes are expressed using the event-state mechanism. Spatial association refers to establishing the spatial association relationship between spatio-temporal geographical objects. For example, the connection relationship between pipelines and pipe points, the intersection relationship between pipelines, the distance relationship between pipelines and buildings, etc., and is described using spatial topological relationships. Semantic association refers to establishing the semantic association relationship between spatio-temporal geographical objects, defining the logical relationship between concepts through ontology. For example, the relationship between pipeline types and functions, the relationship between building types and uses, etc., including "inheritance", "belonging to", etc.
[0023] Specifically, after identifying the spatio-temporal geographical objects from the underground space operation and maintenance scenario instructions, based on the identified spatio-temporal geographical objects, retrieve the corresponding spatio-temporal geographical objects in the underground space entity knowledge graph, and traverse its four-dimensional features of spatial features, temporal features, attribute features, and association relationships, as well as the relationships with other entity nodes along the graph relationships. For example, for the pipeline entity "PID-2007567", retrieve its spatial features (longitude, latitude, depth), temporal features (construction time, maintenance time), attribute information (material, pipe diameter, length), and association relationships with other entities (such as pipe points, buildings).
[0024] S13. Retrieve relevant entities from the preset domain knowledge graph according to the key entity information, and traverse all relevant entities and association relationships along the graph relationship chain to obtain the local retrieval result on the graph side; the domain knowledge graph includes multiple communities, and each community is formed by clustering the graph with triples of entity-relationship-entity extracted from the preset underground space domain knowledge data and empirical knowledge data, and each community forms a corresponding community summary.
[0025] S14. Match relevant community summaries from the community summaries of each community in the domain knowledge graph according to the key entity information as context data, and call the large language model to generate the global retrieval result on the graph side based on the context data.
[0026] S15. Use the local retrieval result on the graph side and / or the global retrieval result on the graph side as the input for vector retrieval, and match the corresponding text unit vector representation from the preset semantic knowledge base to obtain the retrieval result on the semantic side. The semantic knowledge base includes a preset domain knowledge base and an empirical knowledge base. The semantic knowledge base is formed by the structured text data generated by semantic parsing of the pre-collected underground space vertical domain data.
[0027] In this embodiment, semantic-side retrieval means matching the text unit embeddings in the domain knowledge base and the experience knowledge base according to the graph retrieval results, and recalling relevant domain knowledge text units and experience knowledge text units.
[0028] Among them, in the semantic retrieval stage, a vector embedding model is used to generate a question vector representation of the graph-side retrieval results as the input for vector retrieval, and the semantic similarity between the vector retrieval input and the text unit vector representation matched from the semantic knowledge base is calculated through cosine similarity. The specific formula is: , where, is the graph-side retrieval result, is the text unit to be matched, and are respectively the vector representation converted from the graph-side retrieval result and the text unit vector representation converted from the text unit to be matched, and are their norms.
[0029] Furthermore, the semantic retrieval results are re-ranked, and the results are prioritized according to relevance to ensure that the most relevant results are ranked first. Optionally, the score threshold can be set to 0.5, which is used to set the similarity threshold for text unit screening, and only text segments with a score exceeding the set score are recalled. In a specific example, the weight ratio of the domain knowledge base to the experience knowledge base can be default set to 7:3, that is, the top 7 domain knowledge text units with the highest relevance to the domain knowledge base and the top 3 experience knowledge text units with the highest similarity to the experience knowledge base are recalled. Furthermore, the present invention can also dynamically adjust the knowledge weights according to the analysis of the instruction intention. For example, when a specification query type question is recognized, the weight ratio is adjusted to 9:1; for a scenario decision type question, the weight ratio is adjusted to 5:5.
[0030] S16. Integrate the entity information retrieval results, the graph-side local retrieval results, the graph-side global retrieval results, and the semantic-side retrieval results to form a fused retrieval result.
[0031] S17. Input the fused retrieval result into a large language model to generate a final retrieval result. Specifically, input the final fused retrieval result into the large language model for logical verification and evidence chain integration to generate a comprehensive, accurate, and logically strong answer. When generating the answer, the citation specification number, associated case ID, and reasoning path graph can also be attached to support the decision-making in the underground space operation and maintenance scenario.
[0032] The knowledge-driven underground space information retrieval method provided by the embodiments of the present invention combines the advantages of the underground space entity knowledge graph, the semantic knowledge base, and the domain knowledge graph. By constructing a dual-driven information model of specification guidelines and empirical knowledge, designing a dual-channel retrieval framework for the semantic knowledge base and the knowledge graph, it achieves a balance between the accuracy and interpretability of retrieval results, improves the depth and accuracy of knowledge reasoning of large language models, and provides comprehensive data knowledge support and decision-making support for underground space information retrieval and operation and maintenance.
[0033] In the embodiments of the present invention, a semantic knowledge base and a domain knowledge graph are pre-constructed, and an information retrieval is realized by adopting a semantic-graph dual-channel retrieval and recall mode. Through the combined enhancement of semantic retrieval and graph retrieval, and the dual drive of semantic rules and expert experience, a balance between the accuracy and interpretability of retrieval results is achieved. The accuracy depends on the exact matching of the knowledge graph and the generalized semantic matching of the semantic knowledge base, and the two complement each other to reduce missed detections; the graph retrieval provides a structured relationship path, and the semantic retrieval supplements the domain specifications and historical case experience basis to form an interpretable and traceable evidence chain. Among them, the graph-side retrieval and recall include the graph-side local retrieval and the graph-side global retrieval of the domain knowledge graph.
[0034] Specifically, the steps of the graph-side local retrieval are as follows: Based on the identified key entity information, use graph query statements to locate the relevant knowledge entities and text blocks in the domain knowledge graph, and then traverse all relevant entities and association relationships along the graph relationship chain to obtain the graph-side local retrieval results. Further, after obtaining the graph-side local retrieval results, sort the retrieved results according to the confidence level, and select a specified number (such as top3-10) of answers with higher confidence levels as the optimal graph-side local retrieval results, and the confidence level of the graph local retrieval The calculation formula is: , The confidence level of the graph local retrieval represents the direct association confidence level between the entity and the query, and decays based on the entity hop count in the graph. Among them, 0.8 is the decay base number, which is used to control the confidence level decline rate, and the larger the value, the slower the decay. n is the hop count, which represents the number of relationship levels that need to be crossed from the query starting point to the target entity. The confidence level decreases with the increase of the hop count, but non-zero values are retained.
[0035] Specifically, the global retrieval step on the knowledge graph side is as follows: Using the community detection algorithm, a series of community summaries from the knowledge graph are matched as context data according to the key entity information. The large language model is called to generate multiple answers based on the context data and generate an availability score for each answer, ranging from 0 to 1, where a higher value indicates stronger community reliability. The generated answers are sorted according to the availability scores, and a specified number of answers with higher availability scores are selected as the global retrieval results on the knowledge graph side.
[0036] In the embodiment of the present invention, the retrieval result fusion includes integrating the retrieval results obtained from each channel to form a complete knowledge set. The retrieval results specifically include: The retrieval result of entity information, providing relevant feature information of the entity object; The local retrieval result on the knowledge graph side, providing specific knowledge points and association relationships; The global retrieval result on the knowledge graph side, providing a domain overview and knowledge clustering; The retrieval result on the semantic side, providing a traceable document basis.
[0037] The specific implementation method includes establishing a comprehensive scoring model and a dynamic weight adjustment mechanism. The comprehensive scoring model for result ranking based on attention weights quantifies the semantic relevance, knowledge credibility, and domain adaptability of multi-source evidence on the knowledge graph side and the semantic side, and comprehensively ranks and fuses the retrieval results. Based on the dynamic weight adjustment mechanism, the channel weights are dynamically adjusted according to the instruction intention.
[0038] Specifically, the fusion of the retrieval results of entity information, the local retrieval result on the knowledge graph side, the global retrieval result on the knowledge graph side, and the retrieval result on the semantic side to form a fusion retrieval result in step S16 includes: Adopt a comprehensive scoring model for results based on attention weights to comprehensively score the local retrieval result on the knowledge graph side, the global retrieval result on the knowledge graph side, and the retrieval result on the semantic side. The result comprehensive scoring model is as follows: , where, is the comprehensive score of the retrieval result on the semantic side, is the comprehensive score of the local retrieval result on the knowledge graph side, is the comprehensive score of the global retrieval result on the knowledge graph side, is the cosine similarity of the retrieval result on the semantic side, ranging from 0 to 1, where a higher value indicates a stronger similarity between the retrieval result and the input retrieval elements; is the confidence of the local retrieval result on the knowledge graph side, ranging from 0 to 1, where a higher value indicates a stronger direct association between the entity of the local retrieval result and the query; is the availability score of the global retrieval result on the atlas side, ranging from 0 to 1. The higher the score, the stronger the relevance between the community summary and the input information. and and are the weights corresponding to the semantic retrieval channel, the local retrieval channel on the atlas side, and the global retrieval channel on the atlas side respectively. For general questions, the default values can be set to semantic 0.5 / local atlas 0.3 / global atlas 0.2. Among them, the semantic retrieval channel represents the retrieval process of obtaining the semantic retrieval result based on the semantic knowledge base, the local retrieval channel on the atlas side represents the retrieval process of obtaining the local retrieval result on the atlas side based on the domain knowledge graph, and the global retrieval channel on the atlas side represents the retrieval process of obtaining the global retrieval result on the atlas side based on the domain knowledge graph.
[0039] When splicing and fusing the entity information retrieval result with the local retrieval result on the atlas side, the global retrieval result on the atlas side, and the semantic retrieval result, the splicing order of the three retrieval results is determined according to the comprehensive scoring and sorting results of the local retrieval result on the atlas side, the global retrieval result on the atlas side, and the semantic retrieval result.
[0040] For example, for the instruction "What are the maintenance specifications of the pipeline", directly related entities / relationships A such as "pipeline - pipe point" are recalled from the local atlas retrieval with a confidence of 0.64; the pipeline network community summary B is recalled from the global atlas retrieval with a score of 0.7; the text fragment C is recalled from the semantic knowledge base rules with a similarity of 0.9, and the empirical knowledge fragment D with a similarity of 0.6. Then the comprehensive scores are: A - 0.64 * 0.3 = 0.192; B - 0.7 * 0.2 = 0.14; C - 0.9 * 0.5 = 0.45; D - 0.6 * 0.5 = 0.3, and the retrieval result sorting is: C > D > A > B.
[0041] Specifically, the present invention pre - establishes a dynamic weight adjustment mechanism, which can dynamically adjust the channel weights according to the analysis of the instruction intention. For example, for specification query - type questions, increase the weight of the semantic channel, and adjust the weights to semantic 0.7 / local atlas 0.2 / global atlas 0.1; for scenario decision - type questions, increase the weight of the atlas channel, and adjust the weights to semantic 0.4 / local atlas 0.3 / global atlas 0.3.
[0042] In the embodiments of the present invention, an underground space information model is designed that can provide comprehensive data knowledge support and decision - making support for the operation and maintenance of the underground space. The underground space information model includes an underground space entity knowledge graph (i.e., an underground space entity model), a semantic knowledge base (a domain knowledge base, an empirical knowledge base), and a domain knowledge graph.
[0043] In this embodiment, spatio-temporal geographical objects with spatio-temporal semantic information can be obtained by abstracting and modeling underground space entities in the urban underground space field, namely pipelines, buildings, etc., and endowing them with attributes such as time characteristics, space characteristics, attribute characteristics, and association relationships. All spatio-temporal geographical objects can jointly construct an underground space ontology model with spatio-temporal semantic information, and a knowledge graph is used to represent it to obtain an underground space entity knowledge graph, so as to realize entity retrieval and query based on the graph, and thus realize the efficient management and application of underground space entities. The construction steps of the underground space entity knowledge graph specifically include: 1. Abstract and model the underground space entities in the underground space field to obtain a spatio-temporal geographical object class framework.
[0044] 2. Create a class structure in the underground space ontology model according to the spatio-temporal geographical object class framework, and instantiate the class structure in the underground space ontology model into entities in the underground space ontology model. Each class structure represents a spatio-temporal geographical object model with common characteristics.
[0045] Specifically, according to the spatio-temporal geographical object class framework, define the class structure in the ontology. For example, pipeline class, pipe point class, building class, etc. Each class structure represents a spatio-temporal geographical object and contains its common characteristic attributes and relationships. Then, instantiate the specific class structure into an entity in the ontology model.
[0046] For example, instantiate the class structure of a specific pipeline into an instance of the "pipeline class" and assign it a unique identifier. For example, PID-2007567. The generation of the identifier can adopt a hash function or UUID (Universally Unique Identifier) algorithm to ensure its uniqueness and scalability. For example, the generation formula of UUID is: UUID = construction timestamp ⊕ node ID ⊕ random number.
[0047] 3. Endow each entity in the underground space ontology model with characteristic attributes and association relationships to abstract the entity into a spatio-temporal geographical object with spatio-temporal semantic information. The characteristic attributes include time characteristics, space characteristics, and attribute characteristics.
[0048] Specifically, parametric modeling technology can be used to endow the underground space entity with four-dimensional characteristics of space, time, attribute, and association relationship, and abstract it into a spatio-temporal geographical object with spatio-temporal semantic information. The spatio-temporal geographical object is the mapping object of the underground space entity in the computer.
[0049] 4. Based on the class structure, feature attributes, and association relationships in the underground space ontology model, the spatio-temporal geographic objects are transformed into a triple structure of node-edge-node and semantically stored through a knowledge graph to obtain the underground space entity knowledge graph. The nodes in the underground space entity knowledge graph represent the entities and feature attributes in the underground space ontology model, and the edges in the underground space entity knowledge graph represent the association relationships between spatio-temporal geographic objects and the attribute relationships between the entities corresponding to the spatio-temporal geographic objects themselves and their corresponding attribute features. See Figure 2 , Figure 2 which is a schematic diagram of the pipeline entity knowledge graph in the embodiment of the present invention. Further, the nodes in the underground space entity knowledge graph can also represent the entity words of other concept classes in the underground space ontology model, such as pipeline classes.
[0050] Specifically, based on the class structure, feature attributes, and relationship definitions in the underground space ontology model, the spatio-temporal geographic objects are transformed into a triple structure of node-edge-node, and semantic storage and query are realized through a graph database. The specific method is as follows: (1) Node mapping: The entity instances (such as pipeline PID-2007567) and concepts (such as pipeline classes, time features) in the ontology model are represented as nodes in the knowledge graph, and the nodes are connected by edges.
[0051] (2) Relationship mapping: The association relationships between spatio-temporal geographic objects are expressed through the edges of the knowledge graph, and types and attributes are assigned to them. Specifically: Example of representing time association: For pipeline A that undergoes spatio-temporal changes, create a node relationship of "pre-state of pipeline A - maintenance event - post-state of pipeline A", the edge type is the corresponding event, and the edge attributes can define information such as the operation unit of the maintenance time.
[0052] Example of representing space association: Describe the topological relationship through the edge type. For example, for the edge of "pipeline A - connected to - pipe point B", the edge type is connected, and the edge attributes can define information such as the interface type and connection distance; Example of representing semantic association: Express the logical relationship through the edge type. For example, for the edge of "pipeline A - belongs to - pressure pipe class", the edge type is belongs to.
[0053] In this embodiment, before constructing the semantic knowledge base and the empirical knowledge graph, data collection in the vertical field of the underground space is pre-conducted, and then the collected multi-source heterogeneous domain data is processed into structured text data that can be semantically parsed. The data in the vertical field of the underground space includes two parts: underground space domain knowledge and empirical knowledge.
[0054] The underground space domain knowledge mainly includes the following data types: (1)Regulations: Regulations related to the development, use, and management of underground space are important bases for the operation and maintenance of underground space. For example, regulations such as the "Administrative Measures for the Planning and Management of Urban Underground Space Development and Utilization".
[0055] (2)Standards and Specifications: Through the National Standard Information Public Service Platform, retrieve relevant national and local standards using keywords such as "underground space", "geotechnical engineering", and "pipelines". These standards and specifications provide technical bases for the planning, design, construction, and operation and maintenance of underground space. For example, "Requirements for Urban Underground Space Data" (GB / T 42987-2023), "Technical Specification for 3D Modeling of Urban Underground Space" (GB / T 41447-2022), "Data Specification for Urban Underground Space" (DB 3401 / T 230-2021), "Technical Standard for Geotechnical Engineering Information Model in Tianjin" (DB / T 29-292-2021), etc.
[0056] (3)Books and Documents: Collect professional books and documents in the field of underground space by accessing databases such as Wanfang Database. These materials usually contain theoretical research and practical cases in the field of underground space.
[0057] Empirical knowledge mainly includes the following data types: (1)Historical Records of Operation and Maintenance Cases: Collect accident case reports; collect underground space sensor monitoring logs, operation and maintenance records, etc.
[0058] (2)Collection of Decision-making Experience: Design an adaptive template in the field to describe specific underground space operation and maintenance scenarios, guiding experts to think about how to identify problems, retrieve specifications, match cases, calculate parameters, and generate solutions, so as to transform implicit expert experience into structured text expressions that can be retrieved by large models.
[0059] In a specific example, provide an example of collecting experience in the scenario of dealing with tunnel leakage: "In the power cabin of an urban underground utility tunnel, the temperature sensor detected an abnormal increase in the temperature of a local cable. It is necessary to further analyze the root cause and formulate a treatment plan. Please explain in turn:
Problem Identification
Specification Retrieval
Parameter Calculation
Experience Collection
[0060] Furthermore, the underground space domain knowledge data and the experience knowledge data are converted into structured text units respectively, including: For normative text data, the key identification information and chapter content structure of the text data are extracted, and the chapter content structure is segmented using a preset multi-level regular parsing template to generate regular structured text units. A hierarchical construction method is used for normative text data. The first layer extracts key identification information such as name title, release date, implementation date, standard number, and the region to which the local standard belongs; the second layer dynamically loads a multi-level regular parsing template to segment the text unit based on the boundary of "chapter-clause-sub-item" to ensure that each unit contains complete semantic constraints.
[0061] For the chart data in the standard documents, the OCR technology based on the multimodal large model is used to parse the chart data into structured text data to generate chart structured text units. For the chart data such as the "Underground Space Data Classification Table" and the "Underground Space Data Metadata Basic Content Table" in the national and local standards, the OCR technology based on the multimodal large model is used to parse them into structured text data, such as "Serial number: 1, entity set: time and space benchmark, metadata item: data collection time, description: data collection date and Beijing time in the Gregorian calendar, field type: time, constraint condition: M (mandatory)", and each picture or table is used as the boundary for segmentation.
[0062] For the term-definition pairs involved in the field of underground space, the text is segmented according to each term to generate term structured text units. For term-definition pairs extracted from national standards, local standards, and professional books, a term library is established. For example, the terms and definitions of "underground space", "geotechnical engineering", "structural health monitoring" and so on are extracted and segmented according to each term to provide a basis for subsequent knowledge retrieval and reasoning.
[0063] For historical operation and maintenance case data, perform structured extraction on the historical operation and maintenance case data, establish a case event chain description in the form of "accident type - cause - phenomenon - disposal", and associate with the urban underground space geographic coding system to obtain the spatio-temporal tags of the case. Split the combined data of the case event chain description and spatio-temporal tags of each case into words to generate case structured text units. Perform structured extraction on historical operation and maintenance case data such as historical accident reports and operation and maintenance records, establish an "accident type - cause - phenomenon - disposal" event chain, associate with the urban underground space geographic coding system, retain spatio-temporal tags such as the project location and occurrence time, and split according to each individual case. For example, for a certain tunnel water seepage accident, establish an event chain of "water seepage accident - geological structure change - top water seepage - grouting reinforcement", and record the project location and time of the accident.
[0064] For decision-making experience data, perform structured extraction on the decision-making experience data, establish a logical chain description in the form of "problem identification - specification retrieval - case matching - parameter calculation - solution generation", and split the logical chain description of each scenario into words to generate experience structured text units. The collected decision-making experience data is decomposed into a five-stage logical chain of "problem identification → specification retrieval → case matching → parameter calculation → solution generation", transformed into structured data, and split according to each individual scenario.
[0065] 2. Use a vector embedding model to represent the text units obtained from the underground space domain knowledge data as vectors to generate a domain knowledge base, and use a vector embedding model to represent the text units obtained from the experience knowledge data as vectors to generate an experience knowledge base. The domain knowledge base and the experience knowledge base constitute a semantic knowledge base. The vector embedding model can generate high-quality text vector representations and support semantic similarity calculation. The present invention uses a vector embedding model to process the text units obtained from the underground space domain knowledge data and the experience knowledge data, stores them in the knowledge base in the form of vector embeddings, and constructs a domain knowledge base and an experience knowledge base.
[0066] In this embodiment, the steps for constructing the domain knowledge graph specifically include: 1. Extract entity-relationship-entity triples from the structured text units to obtain sub-graph segments containing entity lists and relationship lists, and construct an initial knowledge graph based on each sub-graph segment.
[0067] The present invention optimizes the existing graph retrieval generation technology. First, convert the prompt template into Chinese and add underground space domain information. In a specific example, the prompt template is as follows: "Target activity You are an intelligent assistant to help human analysts analyze the statements about certain entities in the text.
[0068] Target Given a text, entity specification, and claim description that may be related to this activity, extract all entities that conform to the entity specification and all claims for these entities.
[0069] Step Extract Entities: Extract all named entities that conform to the predefined entity specification. The entity specification can be a list of entity names or a list of entity types.
[0070] Extract Claims: For each entity identified in Step 1, extract all claims related to that entity. The claims need to conform to the specified claim description, and the entity should be the subject of the claim. For each claim, extract the following information: Subject: The entity name of the claim subject, in uppercase. The subject entity is the entity that performs the action described in the claim and must be one of the named entities identified in Step 1.
[0071] Object: The entity name of the claim object, in uppercase. The object entity is the entity that is reported / processed or affected by the action described in the claim. If the object entity is unknown, use NONE.
[0072] Claim Type: The general category of the claim, in uppercase. It should be named in a repeatable way so that similar claims share the same claim type.
[0073] Claim Status: TRUE (Confirmed), FALSE (Disproved), or SUSPECTED (Pending Verification).
[0074] Claim Description: Describe in detail the reasoning process of the claim, including all relevant evidence and references.
[0075] Claim Date: The time range of the claim (start date, end date). The date format is ISO - 8601. If the claim is made on a single date rather than within a date range, set the start date and end date to the same date. If the date is unknown, return NONE.
[0076] Claim Source Text: List all original text references related to the claim.
[0077] Format each claim as: (<Subject Entity>{tuple_delimiter}<Object Entity>{tuple_delimiter}<Claim Type>{tuple_delimiter}<Claim Status>{tuple_delimiter}<Start Date>{tuple_delimiter}<End Date>{tuple_delimiter}<Claim Description>{tuple_delimiter}<Claim Source>) Return Output: Return a list of all declarations in Chinese, using {record_delimiter} as the list separator.
[0078] Completion Flag: Output {completion_delimiter} after completion.
[0079] Specific Example: Entity Specification: Underground space facilities Claim Description: Functional description related to underground space facilities Text: Trunk utility tunnel. A comprehensive utility tunnel mainly used to accommodate urban main engineering pipelines, whose main function is to provide services for urban municipal stations, can meet the normal passage of people, and has complete ancillary facilities.
[0080] Output: (Trunk utility tunnel{tuple_delimiter}Urban municipal station{tuple_delimiter}Functional description{tuple_delimiter}TRUE{tuple_delimiter}NONE{tuple_delimiter}NONE{tuple_delimiter}The trunk utility tunnel is mainly used to accommodate urban main engineering pipelines, provide services for urban municipal stations, meet the normal passage of people, and has complete ancillary facilities{tuple_delimiter}Trunk utility tunnel. A comprehensive utility tunnel mainly used to accommodate urban main engineering pipelines, whose main function is to provide services for urban municipal stations, can meet the normal passage of people, and has complete ancillary facilities.) {completion_delimiter} Actual data Generate your answer using the following input.
[0081] Entity Specification: {entity_specs} Claim Description: {claim_description} Text: {input_text} Output: ” Use the figure retrieval generation technology index to construct code, call the general large language model, extract a large number of entity-relationship-entity triples from the structured text units according to the above prompt word template, such as "Comprehensive utility tunnel - contains - Trunk utility tunnel", "Trunk utility tunnel - provides services for - Urban municipal station", etc., to obtain a sub-figure segment containing an entity list and a relationship list, and use this extracted information to construct an initial knowledge graph, as Figure 3 shown, Figure 3Shows a schematic diagram of the knowledge graph in the field of underground space in an embodiment of the present invention.
[0082] 2. Use the community detection algorithm to classify the initial knowledge graph to obtain multiple target communities, and generate community summaries for each target community to obtain the domain knowledge graph.
[0083] The present invention uses community detection technology to classify and divide the initial knowledge graph according to themes.
[0084] First, define multiple themes: (1) Ontology information: entity definitions and classification tables in regulations / standards; (2) Disease information: disease types, causes, and treatment measures in accident reports, etc. (sub-graph: accident → cause → phenomenon → treatment) (3) Operation monitoring: parameters in operation and maintenance records, maintenance treatment records, etc.; (4) Design information: design parameters in standards; (5) Project implementation: construction techniques in cases.
[0085] Among them, using the community detection algorithm to classify the initial knowledge graph to obtain multiple target communities, and generating community summaries for each target community further includes: using the community detection algorithm to cluster triples that are semantically similar in the initial knowledge graph to obtain the underlying communities, and clustering the underlying communities according to a preset multiple of underground space themes to obtain multiple target communities corresponding to different underground space themes; reading the entity, relationship, and attribute data of each clustering corresponding sub-graph, and passing them to the general large language model, using a preset structured prompt word template to guide the general large language model to identify entity relationships and restore the sub-graph logic, and generating the community summary of the target community according to the entity relationships and sub-graph logic; using the vector embedding model to convert the community summary of each target community into a vector representation for realizing the retrieval of the domain knowledge graph.
[0086] The present invention uses the community detection algorithm to cluster triples that are semantically similar, reads the entity, relationship, and attribute data of each clustering sub-graph, passes them to the general large language model, uses a structured prompt word template, and gradually guides the general large language model to identify entity relationships, restore the sub-graph logic, and generate community summaries. From the underlying communities obtained by clustering to the defined 5 theme communities, more and more comprehensive community summaries are gradually generated, and the vector embedding model is used to convert each community summary into a vector representation.
[0087] In a specific example, the prompt word template in this step is as follows: "## Role You are an expert in knowledge graph analysis and need to generate a concise and objective summary based on the entities and relationships in the sub-graph, highlighting the core theme and key information.
[0088] ## Input **[Knowledge Graph Sub-graph]** ## Skills 1. Entity Recognition - Accurately identify entity names, descriptions, and attributes, ignoring empty descriptions or meaningless entities.
[0089] - If the entity name is a document / catalog ID, directly associate its semantics (e.g., "GB / T 42987-2023§5.2" is regarded as a specification clause).
[0090] 2. Relationship Analysis - Accurately identify relationship information, including source entity name, relationship name, target entity name, and relationship description.
[0091] 3. Restore Sub-graph Logic - Restore the graph structure based on the identified entities and relationships 4. Generate Summary - Determine the theme to which the sub-graph belongs from the frequently occurring entities and their core relationships. Optional themes include "Ontology Information", "Disease Information", "Operation Monitoring", "Design Information", "Engineering Implementation".
[0092] - Summarize the information expressed in the sub-graph in accurate, professional, and concise language. The knowledge-driven underground space information retrieval system proposed by the present invention has at least the following beneficial effects: 1. Construct an urban underground space information model: Correspond underground space knowledge with geographical information, establish an entity model, and form a knowledge graph and semantic knowledge base for knowledge reasoning and semantic understanding, realizing the effective management and utilization of underground space information.
[0093] 2. Multimodal knowledge extraction method: Use text recognition technology based on multimodal large models to parse chart information into structured data, realizing multimodal data fusion in the field of underground space.
[0094] 3. Decision logic explicit framework: Design an expert experience collection and coding technical solution, extract implicit decision logic and convert it into a structured thinking chain, realizing the precipitation of implicit experience knowledge.
[0095] 4. Hybrid Retrieval Enhancement Architecture: Establish a queryable and inferable underground space knowledge graph based on the improved graph retrieval enhancement generation technology to solve the problem of slow efficiency in manually constructing the knowledge graph; design a multi-level semantic recall module to achieve hierarchical retrieval of knowledge in the underground space field through the combined retrieval enhancement of semantic retrieval and graph retrieval, improving the comprehensiveness and accuracy of knowledge retrieval. Combine the dual drives of semantic rules and expert experience to achieve a balance between the accuracy and interpretability of retrieval results and enhance the depth of knowledge reasoning.
[0096] For the method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0097] Another embodiment of the present invention also provides a knowledge-driven underground space information retrieval system, and the system includes functional modules for implementing the knowledge-driven underground space information retrieval method described in any one of the above. Figure 4 The structural block diagram of the knowledge-driven underground space information retrieval system according to another embodiment of the present invention is schematically shown. Referring to Figure 4 , the knowledge-driven underground space information retrieval system of this embodiment specifically includes an instruction parsing module 401, an entity information retrieval module 402, a local graph retrieval module 403, a global graph retrieval module 404, a semantic retrieval module 405, a fusion module 406, and a retrieval result generation unit 407, where: The instruction parsing module 401 is used to call a large language model to identify and parse key entity information in the user's underground space operation and maintenance scenario instruction to understand the instruction intention of the current instruction. The key entity information includes spatio-temporal geographical objects and entity words other than the spatio-temporal geographical objects; The entity information retrieval module 402 is used to retrieve the spatio-temporal geographical objects in a preset underground space entity knowledge graph, and traverse the characteristic attributes, association relationships, and relationships with other entity nodes of the spatio-temporal geographical objects based on the knowledge graph relationships of the spatio-temporal geographical objects to obtain an entity information retrieval result; the underground space entity knowledge graph is an entity model formed by semantically storing spatio-temporal geographical objects with spatio-temporal semantic information in the underground space field in the form of a knowledge graph; The local graph retrieval module 403 is used to retrieve relevant entities from a preset domain knowledge graph according to the key entity information, and traverse all relevant entities and association relationships along the graph relationship chain to obtain the local graph retrieval result on the graph side; the domain knowledge graph includes multiple communities, and each community is formed by clustering the triples of entity-relationship-entity extracted from the preset underground space domain knowledge data and empirical knowledge data into a graph; The global graph retrieval module 404 is used to match relevant community summaries as context data from the community summaries of each community in the domain knowledge graph according to the key entity information, and call a large language model to generate the global graph retrieval result on the graph side based on the context data; The semantic retrieval module 405 is used to use the local graph retrieval result and / or the global graph retrieval result on the graph side as the input for vector retrieval, and match the corresponding text unit vector representation from a preset semantic knowledge base to obtain the semantic side retrieval result; The fusion module 406 is used to fuse the entity information retrieval result, the local graph retrieval result on the graph side, the global graph retrieval result on the graph side, and the semantic side retrieval result to form a fusion retrieval result; The retrieval result generation unit 407 is used to input the fusion retrieval result into a large language model to generate the final retrieval result.
[0098] In the embodiment of the present invention, the local graph retrieval module 403 is further used to, after obtaining the local graph retrieval result on the graph side, sort the retrieved results according to the confidence level, and select a specified number of answers with higher confidence levels as the local graph retrieval result on the graph side, and the local graph retrieval confidence The calculation formula is: , where 0.8 is the attenuation base number, used to control the confidence level decline rate, and n is the number of hops, indicating the number of relationship levels to be crossed from the query starting point to the target entity.
[0099] In the embodiment of the present invention, the global graph retrieval module 404 is specifically used to call a large language model to generate multiple answers based on the context data and generate an availability score for each answer; sort the generated answers according to the availability score, and select a specified number of answers with higher availability scores as the global graph retrieval result on the graph side.
[0100] In the embodiment of the present invention, the fusion module 406 is specifically configured to use a result comprehensive scoring model based on attention weights to comprehensively score the local retrieval result on the graph side, the global retrieval result on the graph side, and the retrieval result on the semantic side; when splicing and fusing the entity information retrieval result with the local retrieval result on the graph side, the global retrieval result on the graph side, and the retrieval result on the semantic side, determine the splicing order of the three retrieval results according to the comprehensive scoring ranking result of the local retrieval result on the graph side, the global retrieval result on the graph side, and the retrieval result on the semantic side. The result comprehensive scoring model is as follows: , Wherein, is the comprehensive score of the retrieval result on the semantic side, is the comprehensive score of the local retrieval result on the graph side, is the comprehensive score of the global retrieval result on the graph side, is the cosine similarity of the retrieval result on the semantic side, with a range of 0-1, and the higher the value, the stronger the similarity between the retrieval result and the input retrieval element; is the confidence of the local retrieval result on the graph side, with a range of 0-1, and the higher the value, the stronger the direct association between the entity of the local retrieval result and the query; is the availability score of the global retrieval result on the graph side, with a range of 0-1, and the higher the score, the stronger the relevance between the community summary and the input information, , , are the weights corresponding to the semantic side retrieval channel, the local retrieval channel on the graph side, and the global retrieval channel on the graph side respectively. For general questions, the default values can be set to semantic 0.5 / local graph 0.3 / global graph 0.2. Among them, the semantic side retrieval channel represents the retrieval process of obtaining the retrieval result on the semantic side based on the semantic knowledge base, the local retrieval channel on the graph side represents the retrieval process of obtaining the local retrieval result on the graph side based on the domain knowledge graph, and the global retrieval channel on the graph side represents the retrieval process of obtaining the global retrieval result on the graph side based on the domain knowledge graph.
[0101] In the embodiment of the present invention, the fusion module 406 is further configured to establish a dynamic weight adjustment mechanism to dynamically adjust the channel weights according to the instruction intention.
[0102] For the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, please refer to the partial description of the method embodiment.
[0103] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0104] In addition, another embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor; when the computer program is executed by the processor, the steps of the knowledge-driven underground space information retrieval method described above are implemented.
[0105] In addition, another embodiment of the present invention further provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the knowledge-driven underground space information retrieval method described above are implemented.
[0106] Those skilled in the art can understand that although some of the embodiments herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, any one of the claimed embodiments can be used in any combination.
[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A knowledge-driven underground space information retrieval method, characterized in that The method includes: Invoking a large language model to identify and parse key entity information in the user's underground space operation and maintenance scenario instruction to understand the instruction intention of the current instruction, where the key entity information includes spatio-temporal geographical objects and entity words other than the spatio-temporal geographical objects; Retrieving the spatio-temporal geographical object in a preset underground space entity knowledge graph, and traversing the characteristic attributes, association relationships, and relationships with other entity nodes of the spatio-temporal geographical object based on the knowledge graph relationships of the spatio-temporal geographical object to obtain an entity information retrieval result; the underground space entity knowledge graph is an entity model formed by semantically storing spatio-temporal geographical objects with spatio-temporal semantic information in the underground space field in the form of a knowledge graph; Retrieving relevant entities from a preset domain knowledge graph according to the key entity information, and traversing all relevant entities and association relationships along the graph relationship chain to obtain a local retrieval result on the graph side; the domain knowledge graph includes multiple communities, and each community is formed by graph clustering of entity-relationship-entity triples extracted from preset underground space domain knowledge data and empirical knowledge data; Matching relevant community summaries as context data from the community summaries of each community in the domain knowledge graph according to the key entity information, and invoking a large language model to generate a global retrieval result on the graph side based on the context data; Taking the local retrieval result on the graph side and / or the global retrieval result on the graph side as a vector retrieval input, and matching the corresponding text unit vector representation from a preset semantic knowledge base to obtain a semantic side retrieval result; Fusing the entity information retrieval result, the local retrieval result on the graph side, the global retrieval result on the graph side, and the semantic side retrieval result to form a fused retrieval result; Inputting the fused retrieval result into a large language model to generate a final retrieval result.
2. The method according to claim 1, wherein After obtaining the local retrieval result on the graph side, the method further includes: Sort the retrieved results according to the confidence level, and select a specified number of answers with higher confidence levels as the local retrieval results on the graph side. The confidence level of local graph retrieval is calculated as follows: , Where 0.8 is the attenuation base number for controlling the confidence decline rate, and n is the number of hops, indicating the relationship levels to be crossed from the query starting point to the target entity.
3. The method according to claim 1, characterized in that, The invoking the large language model to generate a global retrieval result on the graph side based on the context data includes: Invoking a large language model to generate multiple answers based on the context data and generating an availability score for each answer; Sorting the generated answers according to the availability scores, and selecting a specified number of answers with higher availability scores as the global retrieval result on the graph side according to the sorting result.
4. The method according to claim 1, wherein Fusing the entity information retrieval result, the local retrieval result on the graph side, the global retrieval result on the graph side, and the semantic side retrieval result to form a fused retrieval result, including: Adopting a result comprehensive scoring model based on attention weights to comprehensively score the local retrieval result on the graph side, the global retrieval result on the graph side, and the semantic side retrieval result. The result comprehensive scoring model is as follows: , Among them, is the comprehensive score of the semantic-side retrieval result, is the comprehensive score of the local retrieval result on the graph side, is the comprehensive score of the global retrieval result on the graph side, is the cosine similarity of the semantic-side retrieval result; is the confidence of the local retrieval result on the graph side; is the usability score of the global retrieval result on the graph side, , , are the weights corresponding to the semantic-side retrieval channel, the local retrieval channel on the graph side, and the global retrieval channel on the graph side respectively. The semantic-side retrieval channel represents the retrieval process of obtaining the semantic-side retrieval result based on the semantic knowledge base. The local retrieval channel on the graph side represents the retrieval process of obtaining the local retrieval result on the graph side based on the domain knowledge graph. The global retrieval channel on the graph side represents the retrieval process of obtaining the global retrieval result on the graph side based on the domain knowledge graph; When splicing and fusing the entity information retrieval result with the local retrieval result on the graph side, the global retrieval result on the graph side, and the semantic side retrieval result, the splicing order of the three retrieval results is determined according to the comprehensive scoring sorting result of the local retrieval result on the graph side, the global retrieval result on the graph side, and the semantic side retrieval result.
5. The method according to claim 4, characterized in that The method further includes: Establish a dynamic weight adjustment mechanism to dynamically adjust the channel weights according to the instruction intention.
6. The method according to claim 1, wherein The construction steps of the underground space entity knowledge graph include: Abstract and model the underground space entities in the underground space field to obtain a spatio-temporal geographical object class framework; Create the class structure in the underground space ontology model according to the spatio-temporal geographical object class framework, and instantiate the class structure in the underground space ontology model into the entities in the underground space ontology model. Each class structure represents a class of spatio-temporal geographical objects with common characteristics; Assign characteristic attributes and association relationships to each entity in the underground space ontology model to abstract the entity into a spatio-temporal geographical object with spatio-temporal semantic information. The characteristic attributes include time characteristics, space characteristics, and attribute characteristics; Based on the class structure, characteristic attributes, and association relationships in the underground space ontology model, convert the spatio-temporal geographical object into a triple structure of node-edge-node and realize semantic storage through the knowledge graph to obtain the underground space entity knowledge graph. The nodes in the underground space entity knowledge graph represent the entities and characteristic attributes in the underground space ontology model, and the edges in the underground space entity knowledge graph represent the association relationships between spatio-temporal geographical objects and the attribute relationships between the entities corresponding to the spatio-temporal geographical objects themselves and the corresponding attribute characteristics.
7. The method according to claim 1, wherein The construction steps of the semantic knowledge base include: Convert the underground space domain knowledge data and empirical knowledge data into structured text units respectively; Use a vector embedding model to represent the text units obtained from the underground space domain knowledge data in vectors to generate a domain knowledge base, and use a vector embedding model to represent the text units obtained from the empirical knowledge data in vectors to generate an empirical knowledge base. The domain knowledge base and the empirical knowledge base constitute the semantic knowledge base.
8. The method according to claim 7, characterized in that, The underground space domain knowledge data includes specification text data, chart data in standard documents, and term-interpretation pairs involved in the underground space field. The empirical knowledge data includes historical operation and maintenance case data and decision-making experience data; The conversion of the underground space domain knowledge data and the empirical knowledge data into structured text units respectively includes: For the specification text data, extract the specification key identification information and the chapter content structure of the text data, and use a preset multi-level regular parsing template to segment the chapter content structure into words to generate rule-structured text units; For the chart data in standard documents, use the OCR technology based on a multi-modal large model to parse the chart data into structured text data to generate chart-structured text units; For the term-interpretation pairs involved in the underground space field, segment the words according to each term to generate term-structured text units; For the historical operation and maintenance case data, perform structured extraction on the historical operation and maintenance case data, establish a case event chain description in the form of "accident type - cause - phenomenon - disposal", and associate with the urban underground space geographical coding system to obtain the spatio-temporal tags of the case. Segment the combined data of the case event chain description and the spatio-temporal tags of each case into words to generate case-structured text units; For decision-making experience data, perform structured extraction on the decision-making experience data, establish a logical chain description in the form of "problem identification - specification retrieval - case matching - parameter calculation - solution generation", and segment the logical chain description of each scenario into words to generate experience structured text units.
9. The method according to claim 7, characterized in that, The steps for constructing the domain knowledge graph include: Extract entity-relationship-entity triples from the structured text units to obtain sub-graph segments containing entity lists and relationship lists, and construct an initial knowledge graph based on each sub-graph segment; Use a community detection algorithm to classify the initial knowledge graph to obtain multiple target communities, and generate community summaries for each target community to obtain the domain knowledge graph.
10. The method according to claim 9, wherein Using a community detection algorithm to classify the initial knowledge graph to obtain multiple target communities, and generating community summaries for each target community, including: Use the community detection algorithm to cluster triples that are semantically similar in the initial knowledge graph to obtain underlying communities, and cluster the underlying communities according to a preset number of underground space themes to obtain multiple target communities corresponding to different underground space themes; Read the entity, relationship, and attribute data of the sub-graph corresponding to each cluster, and pass it to a general large language model. Use a preset structured prompt template to guide the general large language model to identify entity relationships and restore the sub-graph logic, and generate a community summary of the target community based on the entity relationships and sub-graph logic; Use a vector embedding model to convert the community summary of each target community into a vector representation.
11. A knowledge-driven underground space information retrieval system, characterized in that, The system includes: An instruction parsing module for calling a large language model to identify and parse key entity information in the user's underground space operation and maintenance scenario instruction to understand the instruction intention of the current instruction. The key entity information includes spatio-temporal geographical objects and entity words other than the spatio-temporal geographical objects; An entity information retrieval module for retrieving the spatio-temporal geographical objects in a preset underground space entity knowledge graph, and traversing the characteristic attributes, association relationships, and relationships with other entity nodes of the spatio-temporal geographical objects based on the knowledge graph relationships of the spatio-temporal geographical objects to obtain an entity information retrieval result; the underground space entity knowledge graph is an entity model formed by semantically storing spatio-temporal geographical objects with spatio-temporal semantic information in the underground space field in the form of a knowledge graph; A graph local retrieval module for retrieving relevant entities from a preset domain knowledge graph according to the key entity information, and traversing all relevant entities and association relationships along the graph relationship chain to obtain a graph-side local retrieval result; the domain knowledge graph includes multiple communities, and each community is formed by clustering entity-relationship-entity triples extracted from preset underground space domain knowledge data and experience knowledge data; A graph global retrieval module for matching relevant community summaries from the community summaries of each community in the domain knowledge graph as context data according to the key entity information, and calling a large language model to generate a graph-side global retrieval result based on the context data; A semantic retrieval module, configured to use the local retrieval results on the graph side and / or the global retrieval results on the graph side as vector retrieval inputs, match the corresponding text unit vector representations from a preset semantic knowledge base, and obtain the semantic side retrieval results; A fusion module, configured to fuse the entity information retrieval results, the local retrieval results on the graph side, the global retrieval results on the graph side, and the semantic side retrieval results to form a fused retrieval result; A retrieval result generation unit, configured to input the fused retrieval result into a large language model to generate a final retrieval result.
12. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor; when the computer program is executed by the processor, the steps of the method according to any one of claims 1-10 are implemented.
13. A computer program product, characterized in that, A computer program is stored on the computer program product, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1-10 are implemented.
Citation Information
Patent Citations
Knowledge question-answering system based on large language model
CN119396975A
Knowledge graph searching method and system based on local semantics and global communitization
CN119577098A
Government affair processing method of big language model based on knowledge graph and electronic equipment
CN119782545A
Retrieval generation method and device based on large language model and knowledge graph
CN119848168A
Bidirectional image-text retrieval method based on multi-view joint embedding space
WO2019007041A1
Cited By
Automatic generation method and device of operation and maintenance test questions, medium and equipment
CN120448535A
Intelligent question and answer implementation method and system based on large model and semantic map
CN120723876A
An intelligent question and answer implementation method and system based on a large model and a semantic graph
CN120723876B
High-performance knowledge base system based on multi-modal mixed retrieval
CN120781933A
Smart park collaborative service system based on cloud computing and big data
CN120996458A