Geographic information knowledge graph construction method based on large language model
By adopting a staged entity recognition and relationship extraction method with a large language model and Reflxion mechanism in the construction of geographic information knowledge graph, the problem of severe dependence on manual labeled data and weak generalization ability in the existing technology is solved, and efficient and automated construction of geographic information knowledge graphs is achieved.
Patent Information
- Application Number
- CN202510656945.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing technology relies heavily on manual annotation when building geographic information knowledge graphs, which is expensive, and has weak generalization ability in low resource scenarios, making it difficult to effectively identify and extract new entities or relationships that have not been seen before.
The geographic information knowledge graph construction method based on large language models is adopted, and through data preprocessing, staged entity recognition, relationship extraction and entity alignment, combined with the Reflxion mechanism and dual-model collaborative optimization, the dependence on labeled data is reduced and the extraction effect in low-resource scenarios is improved.
It significantly reduces the dependence on manual labeled data in the construction of geographic information knowledge graphs, improves the knowledge extraction effect in low-resource scenarios, enhances the ability to identify unseen entities and relationships, and shortens the knowledge graph construction cycle.
Smart Images

Figure CN120179740A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of geographic information and artificial intelligence, and particularly relates to a method for constructing a geographic information knowledge graph based on a large language model. Background Art
[0002] As an advanced structured knowledge representation form, a Knowledge Graph (KG) is a core technology for managing and utilizing large-scale and highly complex data. It is mainly applied in general fields such as search engine optimization, intelligent question-answering systems, recommendation systems, and natural language processing. Its basic structure consists of entities, relations, and attributes. Through the basic representation unit of triples, objects, concepts, and their connections in the real world are organized into a networked knowledge system. The construction process of a knowledge graph usually covers data collection, entity recognition, relation extraction, entity alignment, knowledge storage, and updating, aiming to transform diverse unstructured or semi-structured data sources into a structured knowledge base that can be understood, computed, and reasoned by machines. Especially in fields that require precise management and analysis of significant geospatial distribution characteristics (such as urban infrastructure, natural resource management, logistics networks, etc.), a knowledge graph can effectively integrate and express complex spatial entities and their interrelationships. Although significant progress has been made in the theory and application of knowledge graph construction technologies, existing methods, especially the technical paths relying on supervised learning, still face several prominent problems and inherent drawbacks. These technologies generally highly rely on large-scale and high-quality labeled datasets. However, in the field of geographic information, the process of obtaining such data is not only extremely challenging and costly but also often requires the participation of scarce experts with in-depth domain knowledge, which makes annotation a key bottleneck. Moreover, annotation biases among different experts are likely to introduce inconsistencies in data quality, thereby negatively affecting the performance of the model and the accuracy of the knowledge graph. Further, in few-shot or zero-shot scenarios with sparse labeled data, existing model methods usually exhibit limited generalization ability and are difficult to effectively identify and extract unseen new entities or relations. In addition, when dealing with real-world data containing noise, errors, or incomplete information, the robustness of these models is often insufficient, easily resulting in incorrect knowledge extraction results.
[0003] In recent years, pre-trained models represented by large language models (LLMs) have achieved significant breakthroughs in the field of artificial intelligence. Their powerful context understanding, semantic reasoning, and text generation capabilities endow them with great potential to automatically extract semantic knowledge, identify entities, and associate relationships from unlabeled or lightly labeled data, providing a promising new technological paradigm for overcoming the bottlenecks in traditional knowledge graph construction and achieving more efficient and large-scale automated construction. However, despite the great potential, the practice of using LLMs for knowledge graph construction is still in its infancy and has obvious limitations. Current research and applications mainly focus on general domains, and the extracted knowledge tends to be common-sense or publicly available information, which limits its effectiveness in scenarios requiring in-depth domain knowledge, resulting in the lack of wide popularity of these methods. In particular, the technical frameworks and supporting toolchains related to the geographic information field also need to be matured.
[0004] For example, in the complex field of water network management, there is currently a significant lack of a systematic and universal technical route or integrated framework that can guide practitioners on how to effectively utilize the powerful capabilities of large language models to accurately and efficiently automate the construction of a professional knowledge graph that meets specific needs from diverse water network-related data sources (such as design drawings, operation logs, maintenance records, sensor data, relevant specification documents, etc.). Although significant progress has been made in knowledge graph construction technology, when dealing with specific fields such as water networks that contain complex professional knowledge and geospatial information, it faces severe challenges due to factors such as high dependence on manual annotation, high costs, and weak generalization ability in low-resource scenarios. Summary of the Invention
[0005] In view of the problems mentioned in the background art, the present invention proposes a method for constructing a geographic information knowledge graph based on a large language model. By applying the large language model (LLMs) to the construction of a geographic information knowledge graph, a systematic and universal technical route for extracting knowledge in the geographic information field of large models is proposed, providing a universal framework for making a geographic information knowledge graph, reducing the construction threshold of domain knowledge graphs, and providing a reference for the automated construction of knowledge graphs in other fields.
[0006] Technical Solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0007] A method for constructing a geographic information knowledge graph based on a large language model, comprising the following steps:
[0008] S1: Data preprocessing and knowledge graph schema definition;
[0009] S2: Stage-based entity recognition integrating the Reflxion mechanism to identify geographic information entities that conform to the predefined schema from the preprocessed text;
[0010] S3: Relationship extraction integrating the Reflxion mechanism to extract the semantic relationships between the identified entities;
[0011] S4: Perform entity alignment and attribute alignment based on semantic similarity and geographical similarity, merge the similar entities, and then store and visualize them.
[0012] Preferably, in S1, collect geospatial data, encyclopedia text data, and multi-source heterogeneous data related to water network management; perform text processing on non-text data, clean the collected text data, including removing irrelevant characters, standardizing formats, and text chunking; according to the specific water network scenario, predefined entity types include rivers, lakes, cities, locations, people, and related relationships and attribute definitions.
[0013] Preferably, the specific process of S2 is as follows:
[0014] S21: Identification of entity types;
[0015] S22: Preliminary entity recognition;
[0016] S23: Evaluation of entity recognition effect;
[0017] S24: Output of entity set.
[0018] Preferably, in S21, the process of identifying entity types is as follows:
[0019] Use a large language extraction model to conduct a preliminary delimitation of the scope of entity types for the input text chunks, judge the entity types included in the text chunks, and output a subset of relevant entity types for the text chunks.
[0020] Preferably, in S22, the specific process of preliminary entity recognition is as follows:
[0021] Extract entity instances from the subset of relevant entity types identified in S21. By constructing prompts that include text chunks and the current target entity type, instruct the large language extraction model to extract all specific instances of entities of this type and their text spans in the text.
[0022] Preferably, in S23, the specific process of evaluating the entity recognition effect is as follows:
[0023] Introduce the Reflexion mechanism for evaluation, transfer the original text, the output of the large language extraction model in S22, the target entity type definition, and the verification rules to the large language evaluation model, review the output results of the large language extraction model, identify various errors, and generate an institutionalized evaluation report or correction suggestions;
[0024] And feedback the evaluation results of the large language evaluation model to the large language extraction model for the second round of extraction of the same text block and the same entity type. At this time, the prompts will incorporate the feedback points of the large language evaluation model to guide the large language extraction model to make corrections.
[0025] Preferably, the specific process of S3 is as follows:
[0026] S31: Relationship extraction based on the large language model;
[0027] S32: Evaluation of relationship extraction effect;
[0028] S33: Output of complete triples.
[0029] Preferably, in S31, the specific content of relationship extraction based on the large language model is:
[0030] Using the verified entity list produced by S2, identify candidate entity pairs within the currently processed text block;
[0031] Using the large language extraction model as the initial relationship extractor, combining the text context, candidate entity pairs, and the complete set of predefined relationship types, determine whether there is a predefined relationship between each candidate entity pair through constructing prompt instructions for the large language extraction model, and determine the most appropriate relationship type;
[0032] Output a list of preliminarily determined relationship triples.
[0033] Preferably, in S32, the specific process of relationship extraction effect evaluation is:
[0034] Start the evaluation link in the Reflexion mechanism, and submit the list of relationship triples generated by the large language extraction model in S31 to the large language evaluation model for independent evaluation and verification;
[0035] According to the original text content, the definition and constraints of the relationship type, and the verified entity information, review each relationship triple; identify various errors, and generate evaluation results or correction suggestions.
[0036] Preferably, in S4, the specific process is:
[0037] S41: Extract all unique entity names from the obtained knowledge graph, and use geocoding technology to obtain the longitude and latitude of relevant geographical entities;
[0038] S42: Use the embedding model to convert the entity names into word vectors, calculate the cosine similarity of the two word vectors as the semantic similarity, and the specific calculation formula is:
[0039] ,
[0040] Among them, A and B respectively represent the word vectors corresponding to entities A and B, and respectively represent the norms of the word vectors A and B, i represents the current word vector dimension, and n represents the total word vector dimension;
[0041] S43: Calculate the spatial distance d of each entity, convert it into geographical similarity using the Gaussian kernel function, and measure their similarity with the distance. Specifically:
[0042] where
[0043] d represents the distance, and σ represents the adjustment parameter;
[0044] S44: Calculate the entity alignment similarity S;
[0045] where
[0046] where
[0047] isSameEntity represents whether the two entities are the same entity; α represents the weight parameter, and s represents the set threshold; when the entity alignment similarity S is greater than or equal to s, they are considered similar entities;
[0048] After merging all similar entities, use the Neo4j graph database for the storage and visualization processing of the knowledge graph.
[0049] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0050] (1) The present invention combines the context understanding and semantic reasoning capabilities of large language models (LLMs), significantly reducing the dependence on large-scale manually annotated data in the construction of knowledge graphs in the field of geographical information, and effectively reducing the participation requirements and annotation costs of domain experts.
[0051] (2) The present invention introduces large language models (LLMs) in stages and combines a dynamic feedback mechanism, utilizes the semantic understanding and reasoning capabilities of the model to learn domain patterns from limited samples, and enhances the recognition ability of unseen entities and relationships through iterative correction, significantly improving the domain knowledge extraction effect including geographical information in zero-shot or few-shot scenarios.
[0052] (3)The present invention proposes a systematic technical route and integration framework for the field of geographic information, providing practical guidance for the automated knowledge extraction and graph construction of multi-source heterogeneous data, and providing a reference for the automated construction of knowledge graphs in other fields with geospatial attributes (such as transportation networks, power systems, emergency management, etc.), with strong potential for cross-domain promotion. Through the full-process automated design and multi-model collaborative optimization, the construction cycle of the knowledge graph is shortened, and efficient and large-scale structured knowledge base generation is achieved to meet real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a flowchart of a method for constructing a geographic information knowledge graph based on a large language model;
[0054] Figure 2 is a flowchart of phased entity recognition integrating the Reflxion mechanism of the present invention;
[0055] Figure 3 is a flowchart of relation extraction integrating the Reflxion mechanism of the present invention;
[0056] Figure 4 is a flowchart of entity alignment of the present invention;
[0057] Figure 5 is a schematic diagram of the performance evaluation of the entity recognition module of the present invention;
[0058] Figure 6 is a schematic diagram of the performance evaluation of the relation extraction module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The following further clarifies the present invention in conjunction with specific embodiments. The embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0060] The method for constructing a geographic information knowledge graph based on a large language model provided in this embodiment is a knowledge graph construction method proposed in view of the problems of low accuracy, low efficiency, high demand for manual intervention, and difficulty in information integration existing in the prior art when dealing with multi-source heterogeneous unstructured geographic information data, especially for fields with complex space and professional knowledge such as water networks. Its core lies in using the dual-LLM collaborative mechanism to improve the accuracy of knowledge extraction and integrating geospatial features for entity alignment. The specific implementation steps are as follows:
[0061] S1: Data preprocessing and knowledge graph schema definition;
[0062] S11: Data preprocessing;
[0063] Collect multi-source heterogeneous data related to water network management, including text data and non-text data, convert non-text data into text, and clean the collected text data, including removing irrelevant characters, format standardization (such as date, unit), and text segmentation.
[0064] Collect geospatial data and encyclopedic text data.
[0065] S12: define the knowledge graph model;
[0066] According to the specific water network scenario, the predefined entity types include rivers (main stream, first-level tributaries, second-level tributaries), lakes, cities, related properties, places, people, and related relationships and attribute definitions.
[0067] Among them, the relationship definition of rivers includes river code, river length, and the river basin to which it belongs; the relationship definition of lakes includes lake code and water surface area; the relationship definition of cities includes prefecture-level city code and gross domestic product.
[0068] Among them, entities are defined from the perspective of spatial relationships and non-spatial relationships; spatial relationships include flow through, inflow, parent river, tributaries, related rivers, and related lakes; non-spatial relationships include alias, originate from, related people, governance / management, related culture, and related events.
[0069] S2: Phased entity recognition integrating Reflxion mechanism: Identify geographic information entities that conform to predefined patterns from preprocessed text;
[0070] Directly using a large model can easily lead to incorrect boundary demarcation of geographic information entities or confusion of hierarchical relationships due to semantic ambiguity. The correct demarcation of entity boundaries is particularly important for geographic information entities. The core goal of this step is to accurately identify geographic information entities that conform to the predefined pattern from the preprocessed text. This step uses a refined strategy that combines type scope definition and dual-model evaluation feedback to identify geographic information entities.
[0071] S21: Accurate identification of entity types;
[0072] First, for the input text block, the system uses the Large Language Extraction Model (LLM_A) to perform a preliminary definition of the entity type range. A predefined complete list of entity types (such as rivers, lakes, cities, etc.) is provided to LLM_A together with the current text block, requiring it to quickly determine the entity types that may be contained in the text block. This step is intended to narrow the scope of subsequent fine extraction, so that the model can focus more on the entity categories related to the current text, and improve processing efficiency and pertinence. The output of LLM_A is a subset of relevant entity types for the text block.
[0073] S22: preliminary entity recognition;
[0074] The system will traverse and extract and refine entity instances for the subset of relevant entity types identified in S21 above. For a specific target entity type (e.g., river), LLM_A assumes the role of performing the extraction action. By constructing a prompt that includes the text block and the current target entity type, LLM_A is instructed to extract all specific instances of entities of this type in the text and their text spans.
[0075] S23: Entity recognition effect evaluation;
[0076] The Reflexion mechanism is introduced for evaluation. The preliminary extraction results of the previous step are passed to the large language evaluation model (LLM_B) to perform independent evaluation and verification. LLM_B acts as the "evaluator". It receives the original text, the output of LLM_A, the target entity type definition, and the verification rules, strictly reviews the results of LLM_A, identifies various potential errors (such as type errors, inaccurate boundaries, entity integrity, or context consistency, etc.), and generates a structured evaluation report or correction suggestions. The evaluation results generated by LLM_B are clearly fed back to the initial extraction model LLM_A. After receiving this feedback from the "evaluator", LLM_A is instructed to perform a second round of extraction for the same text block and the same entity type. The prompt at this time will incorporate the feedback points of LLM_B to guide LLM_A to make corrections. LLM_A makes self-adjustments with the help of this external feedback obtained through the Reflexion mechanism, generates a more accurate list of entities of this type after iterative optimization, and significantly improves the accuracy of geographic information entity recognition.
[0077] S24: Output of the entity set;
[0078] After completing this "LLM_A preliminary extraction -> LLM_B evaluation -> LLM_A feedback and refinement" process based on the Reflexion mechanism for all relevant entity types, the entity recognition module will aggregate the finally refined entity lists under each type. These high-quality entity instance sets optimized through the cooperation of the two models constitute the final entity recognition result of this module.
[0079] S3: Relationship extraction integrating the Reflxion mechanism to extract the semantic relationships between the identified entities;
[0080] After obtaining high-quality entity recognition results through S2, this step is dedicated to extracting the semantic relationships existing between these identified entities and applying the Reflexion mechanism based on the feedback of the two models to improve the accuracy.
[0081] S31: Relationship extraction based on the large language model (LLM);
[0082] Using the verified entity list produced by S2, identify candidate entity pairs within the currently processed text block. Using LLM_A as the initial relation extractor, combined with the text context, candidate entity pairs, and the full set of predefined relation types, determine whether there is any predefined relation between each candidate entity pair by designing prompt instructions for LLM_A, and determine the most appropriate relation type. LLM_A outputs a list of preliminarily determined relation triples (Subject-Predicate-Object) at this step.
[0083] S32: Evaluation of relation extraction effect;
[0084] Start the evaluation session in the Reflexion mechanism and submit the list of preliminary relation triples generated by LLM_A to the large language evaluation model LLM_B for independent evaluation and verification. LLM_B plays the role of "relation evaluator". It needs to strictly review each relation triple proposed by LLM_A based on the original text content, the definition and constraints of relation types, and the verified entity information. LLM_B is responsible for identifying issues such as incorrect relation type judgments and mismatches between the subject / object and relation types, and generating detailed evaluation results or correction suggestions.
[0085] S33: Output of complete triples;
[0086] The evaluation report or correction suggestions of LLM_B are explicitly fed back to LLM_A. After receiving this feedback from the "evaluator", LLM_A conducts a second round of relation judgment and extraction for the same text block and the same candidate entity pairs (or relation scenarios). Finally, through the complete process of the relation extraction module, the system obtains a set of high-quality relation triples that have been collaboratively optimized by the two models within this text block.
[0087] S4: Entity alignment and attribute alignment based on semantic similarity and geographical similarity;
[0088] After the above steps, the system has gathered entity instances extracted and verified from all text blocks. However, due to the diversity of expression or incomplete information, entities extracted from different texts may refer to the same real-world object.
[0089] S41: Extract all unique entity names (such as entity A and entity B) from the obtained knowledge graph and obtain the longitude and latitude of relevant geographical entities using geocoding technology;
[0090] S42: Use an embedding model (such as BERT) to convert entity names into word vectors, calculate the cosine similarity of the two word vectors as the semantic similarity, and the specific calculation formula is:
[0091] ,
[0092] where A and B respectively represent the word vectors corresponding to entities A and B. 、 respectively represent the norms of the word vectors A and B, i represents the current word vector dimension, and n represents the total word vector dimension.
[0093] S43: Calculate the spatial distance d of each entity, convert it into geographical similarity using the Gaussian kernel function, and measure their similarity with the distance. Specifically:
[0094] ,
[0095] where d is the distance and σ is the adjustment parameter.
[0096] S44: Calculate the entity alignment similarity S;
[0097] ,
[0098] ,
[0099] where isSameEntity indicates whether the two entities refer to the same entity; α is the weight parameter, which can be adjusted according to actual needs; s is the set threshold.
[0100] Combine the vectorization model to convert the text data into vectors in the mathematical space. Quantify the similarity of the data by calculating the cosine similarity between the vectors. Finally, set the threshold s according to the actual situation. When the entity alignment similarity S is greater than or equal to s, the entities are considered similar.
[0101] After merging all similar entities, use the Neo4j graph database for the storage and visualization processing of the knowledge graph; for knowledge applications, including data query, reasoning analysis, knowledge Q&A, water network visualization, node partitioning, and node importance ranking.
[0102] The present invention uses three basic indicators, namely accuracy, recall rate, and F1 value, to evaluate the results of the entity extraction module and the relationship extraction module. And evaluations are carried out in the Zero-shot and Few-shot scenarios. Zero-shot means that the model has not seen any labeled samples of a specific category during the training process. Few-shot means that the model has only seen a very small number (usually several) of labeled samples of a specific category during the training process.
[0103] Precision measures how many of the results identified as positive examples by the model are truly positive examples.
[0104] Recall measures how many of all the true positive examples are successfully identified by the model.
[0105] The F1 Score is the harmonic mean of precision and recall, comprehensively measuring the accuracy and integrity of the model.
[0106] The present invention verifies the effectiveness of the method for constructing a geographic information knowledge graph based on a large language model through systematic experiments. The experimental results show that in two typical low-resource scenarios, Zero-shot and Few-shot, the large language model demonstrates remarkable performance in both named entity recognition (NER) and relation extraction (RE) tasks, preliminarily confirming the feasibility of this technical route in constructing knowledge graphs in the field of geographic information.
[0107] The experimental results of the named entity recognition module clearly show that the phased named entity recognition strategy (two-stage) proposed by the present invention is superior to the single-stage method and can more effectively identify geographic information entities. At the same time, the present invention innovatively integrates the Reflexion mechanism into the named entity recognition and relation extraction modules, significantly improving the accuracy. The introduction of the Reflexion mechanism has brought performance improvements in the experiments of both modules, especially in the extremely challenging Zero-shot scenario, where its advantages are particularly prominent.
[0108] The present invention has achieved the best performance in these two key aspects of named entity recognition and relation extraction, strongly supporting that the method for constructing a geographic information knowledge graph based on a large language model proposed by the present invention can effectively solve problems such as the serious dependence on large-scale labeled data and the weak generalization ability in low-resource scenarios of traditional methods.
[0109] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for constructing a geographic information knowledge graph based on a large language model, characterized in that: The following steps are involved: S1: Data preprocessing and knowledge graph model definition; S2: Phased entity recognition integrating Reflxion mechanism to identify geographic information entities that conform to predefined patterns from preprocessed text; S3: Relation extraction integrating Reflxion mechanism to extract semantic relations between identified entities; S4: Entity alignment and attribute alignment are performed based on semantic similarity and geographic similarity, and similar entities are merged for storage and visualization.
2. The method for constructing a geographic information knowledge graph based on a large language model according to claim 1, characterized in that: In S1, geospatial data, encyclopedia text data, and multi-source heterogeneous data related to water network management are collected; non-text data are converted into text, and the collected text data are cleaned, including removing irrelevant characters, format standardization, and text segmentation; according to the specific water network scenario, predefined entity types include rivers, lakes, cities, places, people, and related relationships and attribute definitions.
3. The method for constructing a geographic information knowledge graph based on a large language model according to claim 1, characterized in that: The specific process of S2 is: S21: Identification of entity types; S22: preliminary entity recognition; S23: Entity recognition effect evaluation; S24: Output of entity collection.
4. The method for constructing a geographic information knowledge graph based on a large language model according to claim 3, characterized in that: In S21, the entity type identification process is: The large language extraction model is used to perform preliminary entity type range definition on the input text block, determine the entity types contained in the text block, and output a relevant entity type subset of the text block.
5. The method for constructing a geographic information knowledge graph based on a large language model according to claim 4, characterized in that: In S22, the specific process of preliminary entity recognition is as follows: Entity instances are extracted for the subset of relevant entity types identified in S21. By constructing a prompt containing a text block and the current target entity type, the large language extraction model is instructed to extract all specific instances of this type of entity in the text and its text span.
6. The method for constructing a geographic information knowledge graph based on a large language model according to claim 5, characterized in that: In S23, the specific process of entity recognition effect evaluation is as follows: Introduce the Reflexion mechanism for evaluation, pass the original text, the output of the large language extraction model in S22, the target entity type definition and verification rules to the large language evaluation model, review the output results of the large language extraction model, identify various errors, and generate an institutionalized evaluation report or correction suggestions; The evaluation results of the large language evaluation model are fed back to the large language extraction model to perform a second round of extraction of the same text block and the same entity type. The prompts at this time will incorporate the feedback points of the large language evaluation model to guide the large language extraction model to make corrections.
7. The method for constructing a geographic information knowledge graph based on a large language model according to claim 1, characterized in that: The specific process of S3 is: S31: Relation extraction based on large language model; S32: Relation extraction effect evaluation; S33: Output of the complete triplet.
8. The method for constructing a geographic information knowledge graph based on a large language model according to claim 7, characterized in that: In S31, the specific content of relation extraction based on the large language model is: Using the verified entity list produced by S2, identify candidate entity pairs in the text block currently being processed; Using the big language extraction model as the initial relation extractor, combined with the text context, candidate entity pairs, and the full set of predefined relation types, the prompt instruction big language extraction model is constructed to determine whether there is a predefined relationship between each candidate entity pair and determine the most appropriate relation type; Output a list of relation triples that are initially determined.
9. The method for constructing a geographic information knowledge graph based on a large language model according to claim 8, characterized in that: In S32, the specific process of evaluating the relationship extraction effect is as follows: Start the evaluation phase in the Reflexion mechanism and submit the relation triple list generated by the large language extraction model in S31 to the large language evaluation model for independent evaluation and verification; Review each relationship triple based on the original text content, the definition and constraints of the relationship type, and the verified entity information; Identify various types of errors and generate assessment results or correction suggestions.
10. The method for constructing a geographic information knowledge graph based on a large language model according to claim 1, characterized in that: In S4, the specific process is: S41: extract all unique entity names from the acquired knowledge graph and obtain the latitude and longitude of the relevant geographic entities using geocoding technology; S42: Use the embedding model to convert the entity name into a word vector and calculate the cosine similarity of the two word vectors as the semantic similarity. The specific calculation formula is: , Among them, A and B represent the word vectors corresponding to entities A and B respectively. , Respectively represent the modulus length of word vectors A and B, i represents the current word vector dimension, and n represents the total word vector dimension; S43: Calculate the spatial distance d of each entity, convert it into geographic similarity using Gaussian kernel function, and use distance to measure their similarity, specifically: , Among them, d represents the distance, σ represents the adjustment parameter; S44: Calculate entity alignment similarity S; , , Among them, isSameEntity indicates whether two entities are the same entity; α indicates the weight parameter, and s indicates the set threshold; when the entity alignment similarity S is greater than or equal to s, it is considered to be a similar entity; After merging all similar entities, the Neo4j graph database is used to store and visualize the knowledge graph.
Citation Information
Patent Citations
Geographic information-oriented knowledge graph construction method
CN117453928A
Hydropower station equipment knowledge graph construction method based on large language model and related system
CN119783792A
Knowledge graph construction method based on fine-tuning large language model
CN119808917A
Tax field-oriented knowledge map construction method and system
WO2021196520A1
Cited By
Low-carbon community evaluation index library dynamic construction method and system
CN120448595A
Knowledge graph automatic construction system and method based on LangGraph workflow
CN120781950A
Rail transit maintenance knowledge graph construction method, terminal equipment and medium
CN121257671A