Automatic Construction Method of Multi-Source Geographic Information Knowledge Graph Based on Deep Learning

By designing HGIC-CNN and GKVLD modules, the problem of multi-source geographic information cannot be efficiently integrated, the automatic construction of geographical knowledge graphs and rich information are realized, and the efficiency and accuracy of geographic information processing are improved.

CN115129894BActive Publication Date: 2025-07-08NORTHEAST FORESTRY UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210790862.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2025-07-08
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

The existing geographic knowledge graph construction methods cannot efficiently integrate multi-source heterogeneous geographic information, cannot automatically adapt to different data sources, and rely on manual design to lead to inefficiency.

Method used

Using a deep learning-based method, the HGIC-CNN model and GKVLD module are designed to realize the unified storage structure and automatic learning of multi-source geographic information. Through the HGIC-CNN fusion of multi-source data, the GKVLD module automatically completes the embedded attribute key-value pairs of geographic information nodes.

Benefits of technology

It realizes efficient integration and accurate alignment of multi-source geographic information, improves the information richness and accuracy of the geographical knowledge map, reduces labor costs, and improves the automation level of geographic information processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115129894B_ABST
    Figure CN115129894B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic construction method for a multi-source geographic information knowledge graph based on deep learning, including: designing the storage structures of the geographic knowledge graph, geographic entities, and geographic corpora; constructing the data sources of the geographic knowledge graph through web encyclopedias and geographic databases. Extract the information from the web encyclopedia and / or geographic database and store it in the geographic entity structure. If the entity structures of the web encyclopedia and the geographic database are both described, but the information recorded content is not completely consistent, it is considered that the information in the geographic database is more accurate, and the information in the geographic database is used as the standard to align and form geographic entity nodes. The model predicts whether the geographic nodes and candidate geographic entities represent the same real-world entity and provides a confidence score for the prediction. Finally, the node selects the geographic entity with the highest confidence from the positive classification candidate pairs to establish the correct identity link. The advantages of the present invention are: improving the accuracy of information and saving labor costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic information technology, and particularly relates to a method for automatically constructing a multi-source geographic information knowledge graph based on deep learning. Background Art

[0002] Many major events such as natural disasters and epidemics worldwide have spatial specificity and involve geospatial problems. Effectively using geospatial information can help local governments, countries, or residential areas and other institutions handle more complex problems, such as effectively coping with natural disasters like wildfires and floods [1] and reducing the harm they bring, completing population prediction [2] and other work to help restrict human development, understanding the spread dynamics of epidemic diseases [3] to control and treat epidemics, and discovering spatio-temporal knowledge to enrich human understanding of the geographical world.

[0003] With the continuous development of scientific knowledge and communication technology, the quantity of geospatial information has also increased explosively, and the storage forms of geospatial information data from different sources are also different. People cannot efficiently find the information they need from the massive multi-source geospatial knowledge data. The emergence and development of Knowledge Graph (KG) technology effectively solve this problem. A knowledge graph is an organized collection that connects different entities with possible relationships between entities, which can realize the visualization of knowledge and reveal the hidden relationships between different fields. A Geographic Knowledge Graph (GKG) is a knowledge graph specifically for describing geospatial information data. It first extracts geographic entities from geospatial information and then aligns different geographic entities through relationships. Using a geographic knowledge graph can easily complete various knowledge-based questions without spending a lot of time searching for information, processing data, and solving problems.

[0004] However, the current work on constructing knowledge graphs for geospatial information is limited in quantity and still has great room for improvement in terms of performance. First, a large amount of geographic entities are contained in the massive data, and the existing models have insufficient ability to extract geographic entities and often cannot capture accurate geographic coordinates. Second, widely collecting geospatial information data from different sources can enhance the information volume of the geographic knowledge graph, but the existing geographic knowledge graphs cannot align and fuse multi-source data. Finally, the spatial relationships between the geographic entities constructed by the existing geographic knowledge graphs are not rich enough and will be restricted when solving practical problems.

[0005] A knowledge graph is a method of visualizing scientific knowledge, with a focus on using different signals to reflect the relationships between different entities. A knowledge graph constructed using specialized geographic information data can expand the functions of traditional maps [4]. Chen et al. proposed a geographic information map, using a graphical method to represent the idea of geographic information, aiming to show the spatio-temporal distribution principle of geographic objects [5]. However, this method is not much different from traditional maps and does not have the ability of automatic learning. Specialized researchers are needed to extract information from this geographic information map. In the past decade, researchers have mostly used semantic link networks (SLNs) to represent knowledge. It consists of vertices representing entities and edges representing semantic relationships between entities [6], which laid a good foundation for the formation and development of knowledge graphs later. Traditional knowledge graph representation methods rely on user-defined heuristic methods to extract features encoding the structural information of the graph, mainly divided into methods based on clustering functions and methods based on matrix decomposition strategies. Among the methods based on clustering functions, typical ones include: Bhagat proposed using an iterative traditional clustering classification algorithm and a method of randomly walking to spread labels to complete the classification of nodes in the graph [7]. Vishwanathan [8] used a method of optimizing the kernel matrix to better capture the relationships between entities and key-value pairs in the knowledge graph. Representative methods based on matrix decomposition strategies include, Drinea [9] proposed a sparse matrix decomposition method called CUR, but this method ignores the sparsity of the graph, consumes too much memory and computational cost, and has little application value. To address this drawback, Sun et al.

[10] proposed a compact matrix decomposition method called CMD, but it has insufficient ability to discover the link relationships of internal attributes of geographic entities. The work LinkedGoData

[11] and Yago2Geo

[12] also focus on improving the link discovery ability of geographic nodes in the semantic Web, but both of these works rely on manually defined schema mappings. Given the large scale and openness of the OpenStreetMap (OSM) schema, maintaining such mappings seems infeasible and unsustainable.

[0006] No matter how researchers optimize their algorithms, these methods need to be manually designed, cannot automatically adapt to tasks during the learning process, and cannot automatically make adjustments when dealing with different data. Traditional methods are not only extremely vulnerable to limitations, but also a time-consuming, laborious, and expensive process to design these functions.

[0007] Neural networks can approximate functions of arbitrary shapes by changing the network structure and have strong expressive power, so they can achieve better results from high-dimensional datasets than traditional methods. In this field, Michael

[13] et al. tried to apply deep neural models in the field of graphs, but this method still has problems with poor performance. Yan et al.

[14] proposed a latent representation learning method based on enhancing the spatial environment, which enhances the network's ability to learn the relationships between entities by enhancing the learning embedding of the context. Yan et al.

[15] used the spatial sequence pattern of geographic information as a Bayesian prior and combined it with the state-of-the-art convolutional neural network model to improve the network's ability to classify different geographic information entities. Mai, Yan, Janowicz, and Zhu

[16] incorporated geographic weights into the latent representation learning process to provide better knowledge graph embeddings for geographic graph question answering tasks. Manning

[17] et al. used the standard BM25 text retrieval model to predict links. Daiber

[16] et al. used the DBpedia SPOTLIGHT

[18] model to determine links. Sherif

[20] et al. integrated the Wombat algorithm into the LIMES framework, learned link specifications, and rated the similarity of two entities to improve link accuracy.

[0008] Although neural network models can automatically capture the relationships between different entities and achieve good results, which greatly liberates human labor. However, the existence forms of geographic information are diverse, and most of the existing neural network models are unable to extract geographic entities and spatial relationships from multi-source heterogeneous information. In addition, the neural network does not learn sufficiently about the relationships between the internal attributes of geographic entities, and the correct rate of attribute key-value pair matching is not high. Therefore, we propose a more scalable neural network-based method to learn low-dimensional latent knowledge graph representations and can process geographic data from different sources with different structures, solve the errors that occur when data from different sources are aligned, and enhance the problem-solving ability of the geographic knowledge graph. This method can also automatically mine the potential link relationships between internal attributes and ensure the correctness of link connections.

[0009] References

[0010] [1] Tomaszewski B, Szarzynski J, Schwartz D I. Serious Games for DisasterRisk Reduction Spatial Thinking.

[0011] [2]Benomar T B,Fuling B,Shalgam A M.Application of GIS for PopulationAnalysis Case Study of Zwarah,Libya[J].Journal of Applied Sciences,2006,6(3).

[0012] [3]Mwila I,Phiri J.Tuberculosis Prevention Model in DevelopingCountries based on Geospatial,Cloud and Web Technologies[J].InternationalJournal of Advanced Computer Science and Applications,2020,11(1).

[0013] [4]Dodge M,Kitchin R.Mapping cyberspace[M].Routledge,2003.

[0014] [5]Chen C.Mapping Scientific Frontiers:The Quest for KnowledgeVisualization[J].Journal of Documentation,2003,59(3).

[0015] [6]Zhuge H.Semantic linking through spaces for cyber-physical-sociointelligence:A methodology[J].Artificial Intelligence,2011,175(5-6):988-1019.

[0016] [7]Bhagat S,Cormode G,Muthukrishnan S.Node Classification in SocialNetworks[J].Computer Science,2011,16(3):115-148.

[0017] [8]Vishwanathan S, Schraudolph N N, Kondor R, et al. Graph Kernels[J]. Journal of Machine Learning Research, 2010, 11.

[0018] [9]Drineas, Petros, Kannan, et al. FAST MONTE CARLO ALGORITHMS FORMATRICES III: COMPUTING A COMPRESSED APPROXIMATE MATRIXDECOMPOSITION.[J]. SIAMJournal on Computing, 2006.

[0019]

[10] Sun J, Xie Y, Zhang H, et al. Less is more: Compact matrixdecomposition for large sparse graphs[C] / / Proceedings of the 2007 SIAMInternational Conference on Data Mining. Society for Industrial and AppliedMathematics, 2007: 366 - 377.

[0020]

[11] Stadler C, Lehmann J, Konrad et al. LinkedGeoData: A core fora web of spatial open data[J]. Semantic Web, 2012, 3(4): 333 - 354.

[0021]

[12] NikolaosKaralis, GeorgiosMandilaras, ManolisKoubarakis. Extendingthe YAGO2 Knowledge Graph with Precise Geospatial Knowledge[M]. 2019.

[0022]

[13] Michael, M, Bronstein, et al. Geometric Deep Learning: Going beyond Euclidean data[J]. IEEE Signal Processing Magazine, 2017, 34(4): 18 - 42.

[0023]

[14] Yan, B., Janowicz, K., Mai, G., & Gao, S. (2017). From ITDL to Place2Vec: Reasoning about place type similarity and relatedness by learning embeddings from augmented spatial contexts. In Proceedings of the 25th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems. New York, NY: ACM.

[0024]

[15] Yan, B., Janowicz, K., Mai, G., & Zhu, R. (2018). xNET+SC: Classifying places based on images by incorporating spatial contexts. In Proceedings of the 10th International Conference on Geographic Information Science(pp. 17:1–17:15). Melbourne, Australia.

[0025]

[16] Gengchen Mai, Bo Yan, Krzysztof Janowicz, et al. Relaxing Unanswerable Geographic Questions Using A Spatially Explicit Knowledge Graph Embedding Model[J]. Springer, Cham, 2019.

[0026]

[17] Manning C D, Raghavan P, Hinrich Schütze. Introduction to information retrieval[M]. Posts & Telecom Press, 2010.

[0027]

[18] Daiber J, Jakob M, Hokamp C, et al. Improving efficiency and accuracy in multilingual entity extraction[C] / / International Conference on Semantic Systems. ACM, 2013: 121.

[0028]

[19] Sherif M A, Ngonga A C, Lehmann J. WOMBAT – A Generalization Approach for Automatic Link Discovery[C] / / 14th Extended Semantic Web Conference (ESWC 2017). Springer, Cham, 2017.

[0029]

[20] Kimmel R, Amir A, Bruckstein A M. Finding shortest paths on surfaces, presented in curves and surfaces, Chamonix, France, June 1993. 1994. Summary of the Invention

[0030] Aiming at the defects of the prior art, the present invention provides an automatic construction method for a multi-source geographic information knowledge graph based on deep learning.

[0031] In order to achieve the above invention purpose, the technical solutions adopted by the present invention are as follows:

[0032] An automatic construction method for a multi-source geographic information knowledge graph based on deep learning, comprising the following steps:

[0033] Step 1, designing the storage structures of a geographic knowledge graph, a geographic entity, and a geographic corpus.

[0034] Step 2, constructing the data sources of the geographic knowledge graph through an online encyclopedia and a geographic database: Using the data of the online encyclopedia to supplement the geographic database can improve the functions of the geographic knowledge graph.

[0035] Extract the information from the online encyclopedia and / or the geographic database and store it in the geographic entity structure.

[0036] If the entity structures of the online encyclopedia and the geographic database are both described, but the information record contents are not completely consistent, then the information in the geographic database is considered to be more accurate, and the information in the geographic database is used as the standard to align and form geographic entity nodes. For each geographic entity node n (n ∈ C),

[0037] A set of candidate geographic entities is generated from the set of geographic entities Egeo contained in the geographic knowledge graph.

[0038] Calculate the similarity between the aligned geographic node n and the candidate geographic entity e, where n ∈ C and e ∈ E0.

[0039] Step 3, the model predicts whether the geographic node n and the candidate geographic entity e represent the same real-world entity and provides a confidence score for the prediction. The scoring rule is shown in the following formula. Finally, select the geographic entity with the highest confidence in the positive classification candidate pairs of node n to establish the correct identity link.

[0040] E′ = {e ∈ E geo ∣ distance(n, e) ≤ th block}

[0041] where distance(n, e) is a function that calculates the geographic distance between node n and geographic entity e.

[0042] Preferably, in step 1, the storage structure of the geographic knowledge graph is stored in a graph structure, and the storage structure of the geographic knowledge graph is represented as: Let E be a set of entities, R be a set of labeled directed edges, and L be a set of literals. A complete knowledge graph KG is a triple, represented as where the entities in E represent real-world entities, and the directed edges R represent real-world entity relationships and entity attributes.

[0043] Preferably, in step 1, the storage structure of the geographic entity is represented as: Let E be the sum of geographic entities. A geographic entity e (e ∈ E) will have more geographic significance only after being modified by a relationship r (r ∈ R).

[0044] Preferably, in step one, the storage structure of the geographical corpus is represented as follows: Let the geographical corpus be C, and the geographical entity n ∈ C. n can be represented as a triple n = <i, l, T>, where i is the number of the geographical entity in the corpus, which is the key basis for distinguishing it from other geographical entities. Since each entity has longitude and latitude, the l attribute is used to record the spatial location information of the geographical entity. T represents a set of multiple paired key-value pairs, which is used to describe the meaning of the geographical entity in the real space.

[0045] Further, in step two, to extract the information in the online encyclopedia, after opening the encyclopedia entry by the entity name, the attribute names and values in the information box are extracted. For the attribute values existing in the extracted entity, object characteristics are constructed according to the attribute names.

[0046] When extracting the information entities of the geographical database, the spatial database includes a list of tables, and each table contains many rows, that is, geographical graphic features. The correspondence between the entity name and the characteristic field is used to form pairs for describing the relationship.

[0047] Further, in step two, for the attribute relationships that do not exist in the geographical database, the information in the online encyclopedia is used for supplementation.

[0048] Further, in step three, the model predicts whether this pair of data represents the same real-world entity, and a GKVLD module is proposed to infer the new potential representation of the nodes in the geographical corpus. For node n,

[0049] The input layer of the GKVLD module converts the identifier n.i of each node n into a one-hot vector through the embedding technology.

[0050] The output layer of the GKVLD module uses softmax to map the potential representation to the encoded keys and values. The mathematical description form of the optimization objective is as follows:

[0051]

[0052] Among them, logp(k∣n.i) and logp(v∣n.i) represent the probabilities of the key item k and the value item v of the node matching.

[0053] Compared with the prior art, the advantages of the present invention are as follows:

[0054] 1) Aiming at the problem that the multi-source geographical information storage structures are different and cannot be fused into the same data set, this embodiment designs a unified storage structure of geographical knowledge nodes and a storage structure of geographical corpus.

[0055] 2) Aiming at the disadvantages of the existing geographical knowledge graph with less stored information and inaccurate geographical data, the present invention realizes the fusion function of multi-source heterogeneous geographical information. After aligning the widely sourced data, it can effectively supplement the existing geographical knowledge graph and improve the accuracy of information.

[0056] 3) Aiming at the disadvantages of the existing model, which relies on manual input of information, resulting in low efficiency and high cost, the invention designs a convolutional network model HGIC-CNN, which can automatically learn and capture the model of geographical entities without relying on manual input. This is conducive to constructing a larger dataset and can use this geographical knowledge graph to better guide human production activities. The invention also designs a supervised semantic classification module GKVLD, which automatically completes the embedding of attribute key-value pairs inside geographical information nodes, saving labor costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is a flowchart of HGIC-CNN in an embodiment of the present invention;

[0058] Figure 2 is a schematic structural diagram of the GKVLD module in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following provides further detailed descriptions of the present invention with reference to the drawings and by way of examples.

[0060] This embodiment is implemented through the existing deep learning framework TensorFlow and corresponding data processing operations, deployed on the GPU platform of Tesla V100, and the framework proposed in this article is implemented in the Linux operating system using the Python programming language. The specific implementation principle of the solution is as follows:

[0061] In the first step, this embodiment first designs the storage structures of the knowledge graph, geographical entities, and geographical corpus, unifying the standards for key steps such as the extraction, integration, storage, transmission, and application of subsequent geographical knowledge. This is also the physical basis that must be solved to achieve the multi-source heterogeneous data fusion function.

[0062] The storage structure of the knowledge graph can be represented as follows. Suppose E is a set of entities, R is a set of labeled directed edges, and L is a set of literals. A complete knowledge graph KG is a triple and can be represented as where the entities in E represent real-world entities, and the directed edges R represent real-world entity relationships and entity attributes.

[0063] It is necessary to store multiple attributes of geographical entities in the knowledge graph, and its storage structure is represented as follows. Suppose E is the sum of geographical entities. A geographical entity e (e ∈ E) will have more geographical significance only after being modified by a relationship r (r ∈ R).

[0064] A large number of geographical entities containing various relationships gather together to form a geographical information corpus, and the storage structure of the geographical corpus can be represented as follows. Let the geographical corpus be C, and the geographical entity n ∈ C. n can be represented as a triple n = <i, l, T>, where i is the number of the geographical entity in the corpus and is the key basis for distinguishing it from other geographical entities. Since each entity has longitude and latitude, the l attribute is used to record the spatial location information of the geographical entity. T represents a set of multiple paired key-value pairs, which is used to describe the meaning of the geographical entity in the real space.

[0065] In the second step, the HGIC-CNN (Heterogeneous geographic information classification convolution neural network) method is designed based on the storage structure proposed in the first step, as Figure 1 shown. The data sources used to construct the geographical knowledge graph are two completely different data sources: online encyclopedias or geographical databases. Online encyclopedias cover a wide range of fields, and entity features are described using tags, images, and infoboxes on each web page. Geographical databases are spatial information carriers that meet the needs of production units and the general public. The information in them is mostly stored in the form of vector data, and it is convenient to represent the spatial distribution and spatial relationships of geographical entities in the Cartesian coordinate system. Although geographical databases contain a large number of concepts, they do not contain geographical entities or common sense knowledge. Using the data from online encyclopedias to supplement these geographical databases can improve the functions of the geographical knowledge graph.

[0066] For the knowledge in the encyclopedia, after opening the encyclopedia entry by the entity name, the "attribute-value" pairs (including the attribute name and value) in the infobox are extracted. For the attribute values existing in the extracted entities, object features will be constructed according to the attribute names. When extracting entities using geographical databases, the spatial database includes a list of tables, and each table contains many rows (i.e., geographical graphic features). The corresponding relationship between the entity name and the feature field is used to form pairs to describe the relationship. If an entity is described by the data from both sources but the information recorded is not completely consistent, it is generally considered that the information in the geographical database is more accurate, and the information in the geographical database is used as the standard. For the attribute relationships that do not exist in the geographical database, the information in the encyclopedia can be used for supplementation. For each aligned geographical entity node n (n ∈ C), a set of candidate geographical entities E o Generated from the set of geographical entities \(E_{geo}\) contained in the knowledge graph. This operation realizes the fusion of multi-source heterogeneous data, effectively improving the richness and accuracy of the geographical information knowledge graph data. During the feature extraction stage, the features of both candidate entities and geographical information nodes are learned, and the similarity between the aligned geographical node \(n\) and candidate geographical entity \(e\) (\(n\in C, e\in E_0\)) is calculated. We use a more stringent similarity calculation requirement instead of the majority voting method to help ensure the correctness of the aligned entities. The model predicts whether this pair of data represents the same real-world entity and provides a confidence score for the prediction, and the scoring rule is shown in the following formula. Finally, the geographical entity with the highest confidence in the positive classification candidate pairs of node \(n\) is selected to establish the correct identity link.

[0067] \(E'=\{e\in E geo | distance(n, e)\leq th block \}\)

[0068] where \(distance(n, e)\) is a function that calculates the geographical distance between node \(n\) and geographical entity \(e\). Here, the geographical distance is measured as the geodisc distance

[20] .

[0069] During the process of the HGIC-CNN model fusing multi-source heterogeneous geographical data, the attributes of geographical entities become richer. There are many phenomena where the attributes of the geographical entities after classification by the HGIC-CNN network have unmatched values. Manually matching corresponding values for these tens of thousands of attributes is a very time-consuming and laborious task. This embodiment proposes a GKVLD module to infer the new latent representation of nodes in the geographical corpus, and effectively improves the link matching accuracy rate compared with existing models. For node \(n\) (\(n = \langle i, l, T\rangle\)), the function of this network is to encode a set of key-value pairs \(T\) describing the node into an embedding representation. The specific structure of this neural network model is as Figure 2 shown. The input layer converts the identifier \(n.i\) of each node \(n\) into a one-hot vector through the embedding technology. Embedding generates a similar representation for nodes with similar attributes, regardless of the real geographical location of the geographical entity vector. In the one-hot encoding structure, each identifier \(n.i\) corresponds to a dimension of the input layer. The corresponding entry of the vector representation is set to 1, while other entries are set to 0. The projection layer calculates the latent representation of the node. The number of neurons in this layer corresponds to the dimension in the projection, that is, the embedding size. Among them, the key-value pairs \(\langle k, v\rangle\) (\(\langle k, v\rangle\in n.T\)) used to describe the attributes of geographical entities are also mapped to two different vector spaces by the one-hot encoding method. The output layer uses softmax to map the latent representation to the encoded keys and values, and the mathematical description form of the optimization objective of this model is as follows:

[0070]

[0071] Among them, logp(k∣n.i) and logp(v∣n.i) represent the probabilities of the key item k and the value item v of the node matching.

[0072] Finally, in this embodiment, two heterogeneous geographical databases, Wikipedia and OpenStreetMap, are selected for fusion, and compared with the existing state-of-the-art algorithms. The comparison baseline models are the BM25

[17] , SPOTLIGHT

[18] , LGD

[11] , YAGO2GEO

[12] , and LIMES / Wombat

[19] methods. To evaluate the performance of the attribute relationship discovery method, we evaluate the model through the following several metrics.

[0073] Precision: Calculate the proportion of correctly linked geographical nodes among all the nodes assigned by the current method.

[0074] Recall: Among all the nodes with links in the ground truth sample dataset, calculate the proportion of geographical nodes correctly linked by this method.

[0075] F1: The harmonic mean of Recall and Precision. In this work, we consider the F1 score to be the most relevant metric because it reflects both recall and precision.

[0076] This embodiment uses a 10-fold cross-validation method to randomly sample links from the link datasets generated by different models to obtain a batch of results. Taking a batch of results as a unit, a classification model is trained on the generated link data center. After 10 batches, the macro-average value of each metric is calculated. Experiments prove that the algorithm proposed in this embodiment outperforms the existing models in all metrics. Three locations of different magnitudes are selected for the experiment: France (FR), China (ZG), and Asia (YZ). The specific experimental data are shown in Table 1.

[0077] Table 1: Results of the control experiment (%)

[0078]

[0079]

[0080] According to the results in the table, it can be calculated that compared with various existing methods, except that the LGD algorithm achieves the best result in the precision metric due to the small number of attribute key-value pair links, the algorithm proposed in this embodiment is only 9.94% lower than this algorithm. In the recall and F1 metrics, the results of the algorithm proposed in this embodiment are significantly better than other algorithms.

[0081] Those of ordinary skill in the art will realize that the embodiments described herein are provided to assist the reader in understanding the implementation methods of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.

Claims

1. An automatic construction method for a multi-source geographic information knowledge graph based on deep learning, characterized in that It includes the following steps: Step 1, design the storage structures of the geographical knowledge graph, geographical entities, and geographical corpus; Step 2, construct the data sources of the geographical knowledge graph through the online encyclopedia and geographical database: Using the data of the online encyclopedia to supplement the geographical database can improve the functions of the geographical knowledge graph; Extract the information from the online encyclopedia and / or geographical database and store it in the geographical entity structure; If the entity structures of the web encyclopedia and the geographical database are described simultaneously, but the information record contents are not exactly the same, the information in the geographical database is considered to be more accurate, and the information in the geographical database is used as the standard for alignment to form geographical entity nodes; for each geographical entity node n (n ∈ C), a group of candidate geographical entities is generated from the set of geographical entities Egeo included in the geographical knowledge graph; Calculate the similarity between the aligned geographical node n and the candidate geographical entity e, where n ∈ C and e ∈ E0; Step 3, the model predicts whether the geographical node n and the candidate geographical entity e represent the same real-world entity and provides a confidence score for the prediction. The scoring rule is shown in the following formula; Finally, select the geographical entity with the highest confidence in the positive classification candidate pairs for node n to establish the correct identity link; E′ = {e ∈ E geo | distance(n, e) ≤ th block} Among them, distance(n, e) is a function to calculate the geographical distance between node n and geographical entity e; The model predicts whether this pair of data represents the same real-world entity and proposes a GKVLD module to infer the new latent representation of the nodes in the geographical corpus. For node n, the input layer of the GKVLD module converts the identifier n.i of each node n into a one-hot vector through the embedding technology; The output layer of the GKVLD module uses softmax to map the latent representation to the encoded keys and values. The mathematical description form of the optimization objective is shown as follows: Among them, logp(k∣n.i) and logp(v∣n.i) represent the probabilities of the key item k and value item v of the node matching; 2. The automatic construction method of a multi-source geographic information knowledge graph based on deep learning according to claim 1, wherein: In Step 1, the storage structure of the geographical knowledge graph is stored in a graph structure, and the storage structure of the geographical knowledge graph is represented as follows: Let E be a set of entities, R be a set of directed edges with labels, and L be a set of literals; a complete knowledge graph KG is a triple, represented as KG = <E ∪ L, R>, where the entities in E represent the entities in the real world, and the directed edges represent the entity relationships and entity attributes in the real world.

3. The automatic construction method of a multi-source geographic information knowledge graph based on deep learning according to claim 1, characterized in that: In Step 1, the storage structure of geographical entities is represented as: Let E be the sum of geographical entities. A geographical entity e (e ∈ E) will have more geographical significance only after being modified by a relationship r (r ∈ R).

4. A method for automatically constructing a multi-source geographic information knowledge graph based on deep learning according to claim 1, characterized in that: In Step 1, the storage structure of the geographical corpus is represented as: Let the geographical corpus be C. For a geographical entity n ∈ C, n can be represented as a triple n = <i, l, T>, where i is the number of this geographical entity in the corpus and is the key basis for distinguishing it from other geographical entities; Since each entity has longitude and latitude, the l attribute is used to record the spatial location information of the geographical entity; T represents a set of multiple paired key-value pairs, and this item is used to describe the meaning of this geographical entity in the real space.

5. The automatic construction method of a multi-source geographic information knowledge graph based on deep learning according to claim 1, characterized in that: In Step 2, to extract the information from the online encyclopedia, after opening the encyclopedia entry by the entity name, extract the attribute names and values in the information box; For the attribute values existing in the extracted entities, construct object characteristics according to the attribute names; When extracting the information entities of the geographical database, the spatial database includes a list of tables, and each table contains many rows, that is, geographical graphic features. The corresponding relationship between the entity name and the feature field is used to form pairs to describe the relationship; 6. The automatic construction method of a multi-source geographic information knowledge graph based on deep learning according to claim 1, characterized in that: In Step 2, for the attribute relationships that do not exist in the geographical database, use the information in the online encyclopedia to supplement them.

Citation Information

Patent Citations

  • Hybrid knowledge graph construction method fusing geographic knowledge

    CN113139065A

  • Geographical knowledge graph

    US20190179917A1