Geographic information text intelligent extraction method based on knowledge perception
By constructing a geographic knowledge ontology library and semantic relationship network, using multi-level neural networks to identify geographic entities and semantic relationships, and building a semantic tree for knowledge perception and semantic understanding, the accuracy and efficiency problems of geographic information extraction in traditional methods are solved, and efficient and accurate information extraction and integration are achieved.
Patent Information
- Application Number
- CN202511265267.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Traditional geographic information extraction methods have difficulty in accurately identifying and extracting complex semantic relationships in geographic information texts, resulting in information omissions or incorrect extraction.
Build a geographic knowledge ontology library and semantic relationship network, use multi-level neural networks to identify geographic entities and semantic relationships, build a semantic tree for knowledge perception and semantic understanding, and combine semantic mapping and reasoning to extract information.
It improves the accuracy and efficiency of geographic information extraction, enhances semantic understanding capabilities, supports flexible and diverse geographic information extraction needs, and promotes the integration and application of geographic information.
Smart Images

Figure CN120745633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and more specifically, to a method for intelligent extraction of geographic information text based on knowledge perception. Background Art
[0002] In the field of geographic information processing, accurately and efficiently extracting geographic information from text is of vital importance. With the widespread application of geographic information technology in many fields, such as geographical research, urban planning, environmental monitoring, etc., the demand for geographic information is increasing, and higher requirements are placed on its accuracy and completeness.
[0003] However, traditional geographic information extraction methods face many challenges. On the one hand, the complexity and diversity of geographic information texts make information extraction difficult. Geographic texts contain a large number of different types of geographic entities, and the semantic relationships between them are complex and diverse. These complex relationships make it difficult to accurately identify and extract geographic information. Traditional methods often fail to fully and accurately grasp these relationships, and are prone to information omission or incorrect extraction.
[0004] In view of this, the present invention proposes a knowledge-based intelligent extraction method for geographic information text to solve the above problems. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned objectives, the present invention provides the following technical solutions:
[0006] A method for intelligent extraction of geographic information text based on knowledge perception, comprising:
[0007] Step 1: Build a geographic knowledge ontology library and build a corresponding semantic relationship network based on it;
[0008] Step 2: Obtain the textual geographic entities and textual semantic relationships corresponding to the input geographic information text;
[0009] Step 3: Based on the obtained text geographic entities and text semantic relationships, and in combination with the semantic relationship network, a corresponding initial dendrogram is constructed; based on the geographic knowledge ontology library, knowledge perception and semantic understanding are performed on the corresponding initial dendrogram to obtain a corresponding semantic tree;
[0010] Step 4: Extract information based on the obtained semantic tree to obtain the required geographic information.
[0011] Furthermore, the process of building a geographic knowledge ontology database includes:
[0012] Obtain existing geographic entity concepts, and generalize and abstract them into basic geographic entities, wherein the basic geographic entities include natural geographic entities, artificial geographic entities, and combined geographic entities; the basic geographic entities are composed of several different geographic entity categories;
[0013] Setting a data collection node, collecting data on geographic information related to corresponding basic geographic entities based on different data sources, and obtaining corresponding basic geographic data;
[0014] Perform data preprocessing on the obtained basic geographic data to obtain corresponding basic text data;
[0015] After preprocessing is completed, the geographic entities in the corresponding basic text data are identified based on the pre-built geographic entity recognition model, and the semantic relationships between the corresponding geographic entities are identified;
[0016] At the same time, based on the basic text data, the attribute information of the corresponding geographic entity is extracted to obtain the corresponding entity attributes;
[0017] Based on the obtained geographic entities, semantic relationships and entity attributes, corresponding knowledge triples are constructed and stored in the corresponding geographic entity categories, and based on them, the corresponding geographic knowledge ontology library is constructed.
[0018] Furthermore, the construction process of the geographic entity recognition model includes:
[0019] The backbone network of the geographic entity recognition model is a multi-layer neural network. The basic framework of the multi-layer neural network is an input layer, a double layer , subject-object extraction layer, fully connected layer and output layer;
[0020] The input layer is used to encode the input text sentence to obtain the encoding vector corresponding to each word in the corresponding text sentence;
[0021] The double layer Including the first Layer and second layer; the first The layer extracts semantic features from the input encoding vector by introducing a control gate mechanism to obtain the corresponding word semantic feature vector, and assigns weights to each word in the corresponding text sentence based on the obtained word semantic feature vector to obtain the corresponding weighted sentence; wherein the control gate mechanism includes a forget gate, an enhancement gate, and an output gate;
[0022] The second The layer is used to extract the word semantic feature vector corresponding to each word in the received weighted sentence, and perform feature fusion to obtain the corresponding sentence semantic feature vector;
[0023] The subject-object extraction layer is used to receive the second The word semantic feature vector output by the layer is used to obtain the main word position of the geographic entity in the corresponding text sentence based on a pre-set prediction classifier; based on the main word position of the geographic entity, the word representation vector corresponding to the corresponding main word is obtained;
[0024] Based on the subject word position of the geographic entity and the acquisition process of the word representation vector, the object word position and the corresponding word representation vector of the geographic entity in the corresponding text sentence are obtained, and the subject word and the object word of the corresponding geographic entity are marked in the corresponding text sentence;
[0025] The fully connected layer is used to receive the obtained word representation vector and sentence semantic feature vector, and merge them to obtain the corresponding entity vector;
[0026] The output layer is used to map the obtained entity vector into a category probability of a certain semantic category; the obtained category probability is compared with the corresponding category threshold; if the category probability is greater than the category threshold, the corresponding semantic category is used as the semantic relationship between the corresponding physical entities; if the category probability is not greater than the category threshold, no other operation is performed;
[0027] Obtain several sets of manually annotated geographic text sentences and construct corresponding training sample sets based on them; define As the optimizer continuously optimizes the parameters of the geographic entity recognition model during the training process, the training sample set is input into the geographic entity recognition model in batches, and the value of the corresponding loss function is recorded. When the value of the loss function of consecutive L2 batches no longer decreases or changes, the parameters of the geographic entity recognition model at this time are saved, and the training of the geographic entity recognition model is completed. L2 is a constant.
[0028] Furthermore, the formula for obtaining the word semantic feature vector is: Where, and Respectively represent the first and The word semantic feature vector corresponding to the word; represents the weight matrix of the output gate in the control gate mechanism; Indicates the first The encoding vector corresponding to each word; Represents the bias term of the output gate; is a natural number, I represents the total number of words in the text sentence; Represents vector concatenation; represents the pre-selected activation function;
[0029] Indicates the current time Next First The network status data corresponding to the layer, Where, and Both represent the weight matrix of the enhancement gate; and Both represent bias terms; the enhancement gate is used to Update the network status data of the layer; Represents the hyperbolic tangent function operation;
[0030] Represents the forget gate output corresponding to the i-th word in the text sentence, Where, and Represent the weight matrix and bias term of the forget gate respectively;
[0031] The formula for assigning weights is: ; Indicates the first The word weight corresponding to each word; Q, K, V represent the semantic feature vectors of the corresponding words respectively. The weight matrix obtained by linear transformation; represents a matrix permutation; Represents the weight matrix The matrix dimensions of .
[0032] Furthermore, the formula for obtaining the keyword position is: Where, The subject words representing geographic entities in the corresponding text sentences The probability of being at position S1, S1 represents the position identifier of the subject word of the geographic entity, that is, the range of the start and end of the subject word; Respectively represent the corresponding The probability of a word being the starting position of the main word of a geographic entity; Where, represents the training weight; represents the bias term; Indicates the weighted sentence The word semantic feature vector corresponding to each word; ;in, By the corresponding probability Decision, if appropriate If the probability is greater than the pre-set threshold, , otherwise, ; Represents the trainable parameters of the model; and They represent the start position mark and end position mark of the subject word respectively; n represents the position index of the subject word between the start position and the end position of the subject word; Indicates the total number of words in the text sentence; Represents the product operation;
[0033] Class probability Where, Indicates that the entity vector ST belongs to the first The probability of semantic categories, and Represent exponential operation and activation function operation respectively; All semantic categories representing predefined semantic relations;
[0034] Defining the loss function for the geographic entity recognition model Where, and Represents the weight ratio, represents the position probability of the object word of the geographic entity; Ld represents the first training samples, Indicates the Semantic category labels of training samples; Indicates the first The potential triples corresponding to the training samples are .
[0035] Furthermore, the process of constructing the semantic relationship network includes:
[0036] Obtain the geographic entities stored in the corresponding geographic knowledge ontology library, perform entity disambiguation on them, and mark the disambiguated geographic entities as conceptual entities; obtain different conceptual entities and their corresponding semantic relationships; and convert the corresponding conceptual entities into corresponding graph nodes, and the semantic relationships into corresponding directed edges, and mark the corresponding directions and relationship types;
[0037] Based on a pre-selected directed graph construction algorithm, and based on the graph nodes constructed by the corresponding concept entities and the edges corresponding to the semantic relationships, a semantic relationship network between the corresponding concept entities is constructed.
[0038] Furthermore, the process of obtaining the textual geographic entities and textual semantic relationships corresponding to the inputted geographic information text includes:
[0039] The input geographic information text is preprocessed and input into the geographic entity recognition model to obtain the geographic entities and semantic relationships in the corresponding geographic information text, and mark them as text geographic entities and text semantic relationships respectively.
[0040] Furthermore, the process of constructing a corresponding initial dendrogram based on the obtained textual geographic entities and textual semantic relationships and in combination with the semantic relationship network includes:
[0041] Obtain the extracted text geographic entities and text semantic relationships, and map them to the constructed semantic relationship network; and obtain the graph nodes and edges corresponding to the corresponding text geographic entities and text semantic relationships;
[0042] The graph nodes corresponding to the corresponding textual geographic entities are used as cluster centers, and other graph nodes in the corresponding semantic relationship network are clustered to obtain the graph nodes in the same cluster set as the corresponding textual geographic entities. The conceptual entities and semantic relationships of the corresponding edges corresponding to the corresponding graph nodes are obtained and marked as extended geographic entities and extended semantic relationships.
[0043] Then, all the obtained text semantic relations and extended semantic relations are obtained and structurally classified to obtain upstream and downstream relations or association relations between different text geographic entities and extended geographic entities, and a corresponding relationship determination matrix is constructed based on the relations;
[0044] Constructing an enhanced tree structure framework, wherein the enhanced tree structure framework includes a trunk hierarchical structure and horizontal semantic links;
[0045] Select the textual geographic entity with the highest level of upstream and downstream relationships in the semantic relationship network and use it as the root node in the backbone hierarchical structure; and based on the upstream and downstream relationships in the relationship matrix, add other textual geographic entities or extended geographic entities layer by layer to the corresponding backbone hierarchical structure;
[0046] After the addition is completed, the association relationship is stored as a property of the tree node in the enhanced tree structure framework based on the horizontal semantic link to obtain the initial tree structure;
[0047] Input the obtained initial dendrogram into the constructed local geographic knowledge base for semantic mapping and semantic reasoning, and obtain the corresponding semantic tree based on it;
[0048] After the semantic mapping and semantic reasoning are completed, the entity attributes stored in the corresponding address knowledge local library and the corresponding semantic reasoning results are stored in the tree nodes in the corresponding semantic tree.
[0049] Furthermore, the semantic mapping refers to mapping the node entities corresponding to the tree nodes in the initial tree and the edge semantic relationships corresponding to the edges between the tree nodes to the geographic knowledge triples in the geographic knowledge ontology library, wherein the node entities refer to textual geographic entities or extended geographic entities; the edge semantic relationships refer to textual semantic relationships or extended semantic relationships; and the semantic reasoning refers to performing semantic-level reasoning on other tree nodes adjacent to the corresponding tree nodes based on the entity attributes in the knowledge triples to which they are mapped.
[0050] Furthermore, the process of extracting information based on the obtained semantic tree to obtain the required geographic information includes:
[0051] The root node of the corresponding semantic tree is obtained as the starting node, and the corresponding semantic tree is traversed based on the preset extraction rules. The tree nodes and edges corresponding to the tree nodes in the corresponding semantic tree that meet the preset extraction rules are marked, and the entity attributes and semantic reasoning results stored in the corresponding tree nodes and the edge semantic relationships of the tree nodes are read to obtain the corresponding geographic information; the obtained geographic information is deduplicated and summarized to obtain the corresponding integrated information, and whether the semantic relationships between the geographic entities involved in the extracted integrated information meet the requirements are examined. If not, the corresponding integrated information is sorted and adjusted according to the local geographic knowledge library; after the adjustment is completed, the corresponding integrated information is converted into the required format file and visualized output.
[0052] The technical effects and advantages of the knowledge-aware intelligent extraction method of geographic information text in this invention are as follows:
[0053] 1. By building a geographic knowledge ontology and semantic relationship network, combining the geographic entities and semantic features of the input text to build a semantic tree, and using the ontology for knowledge perception and semantic understanding, the semantic relationship between geographic entities can be accurately grasped. For example, semantic mapping can be used to clarify the precise semantics of geographic entities in the text, and semantic reasoning can be used to explore deep semantics, thereby improving the accuracy of geographic information extraction.
[0054] 2. It integrates artificial intelligence and geographic information system technologies to realize the intelligent extraction and processing of geographic information texts; and has significant beneficial effects in improving the accuracy and efficiency of geographic information extraction, enhancing the semantic understanding ability of geographic information, supporting flexible and diverse geographic information extraction needs, promoting the integration and application of geographic information, and improving the intelligent level of geographic information processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 A schematic diagram of a method for intelligent extraction of geographic information text based on knowledge perception according to the present invention;
[0056] Figure 2 This is a schematic diagram of a geographic information text intelligent extraction system based on knowledge perception of the present invention. DETAILED DESCRIPTION
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0058] Example 1
[0059] See also Figure 1 As shown, the present embodiment provides a method for intelligent extraction of geographic information text based on knowledge perception, including:
[0060] Step 1: Build a geographic knowledge ontology library and build a corresponding semantic relationship network based on it;
[0061] Step 2: Obtain the textual geographic entities and textual semantic relationships corresponding to the input geographic information text;
[0062] Step 3: Based on the obtained text geographic entities and text semantic relationships, and in combination with the semantic relationship network, a corresponding initial dendrogram is constructed; based on the geographic knowledge ontology library, knowledge perception and semantic understanding are performed on the corresponding initial dendrogram to obtain a corresponding semantic tree;
[0063] Step 4: Extract information based on the obtained semantic tree to obtain the required geographic information;
[0064] It should be further explained that, in the specific implementation process, the process of building a geographic knowledge ontology database includes:
[0065] Based on a pre-built corpus, existing geographic entity concepts are obtained and the corresponding geographic entity concepts are summarized and abstracted into corresponding basic geographic entities. The basic geographic entities include natural geographic entities, artificial geographic entities, and combined geographic entities. The basic geographic entities are composed of several different geographic entity categories. Natural geographic entities refer to geographical phenomena or objects existing in nature, such as mountains and rivers. Artificial geographic entities refer to buildings or other structures produced by human activities, such as roads and bridges. Combined geographic entities generally refer to regional divisions with certain functions or significance, such as administrative divisions and nature reserves.
[0066] Setting up data collection nodes, which collect geographic information related to corresponding basic geographic entities based on different data sources to obtain corresponding basic geographic data; the data sources include social media, academic research papers and works, geographic data institutions, etc.; the basic geographic data includes vector layers provided by map services, remote sensing image data, government-published statistical data, and relevant geographic academic papers, etc.;
[0067] Performing data cleaning on the obtained basic geographic data, said data cleaning includes removing special characters, numbers, HTML tags and other non-text characters in the text, as well as processing abbreviations, abbreviations and non-standard writing methods in the text;
[0068] Then, the basic geographic data after data cleaning is preprocessed to obtain corresponding basic text data, wherein the data preprocessing includes steps such as word segmentation, removal of stop words and part of speech restoration;
[0069] After preprocessing is completed, the subject words and object words of the geographic entities in the corresponding basic text data are identified based on the pre-built geographic entity recognition model, and the semantic relationship between the subject words and object words of the corresponding geographic entities is obtained for identification, thereby obtaining the corresponding geographic entities and the semantic relationship between geographic entities; the semantic relationship includes topological relationship, subordinate relationship and temporal relationship; topological relationship includes spatial relationship such as adjacency, association, inclusion and connectivity between geographic entities; subordinate relationship refers to the logical structural relationship between geographic entities; temporal relationship includes the generation time, change time and extinction time of geographic entities;
[0070] Then, based on the geographic encyclopedia, the attribute information of the corresponding geographic entity is obtained, and based on the attribute information, the attribute extraction of the corresponding basic text data is performed, and the attribute information of the corresponding geographic entity is supplemented and updated to obtain the corresponding entity attributes; the entity attributes include basic attributes and thematic attributes. The basic attributes refer to common data attributes, including classification code, entity code, name, other names, geometric form, etc.; the thematic attributes refer to unique attributes. Taking rivers as an example, the thematic attributes include the starting point, end point, length, river level, river function, etc.
[0071] Based on the obtained geographic entities, semantic relationships and entity attributes, corresponding knowledge triples are constructed and stored under corresponding geographic entity categories, and based on them, corresponding geographic knowledge ontology libraries are constructed;
[0072] It should be further explained that, in the specific implementation process, the construction process of the geographic entity recognition model includes:
[0073] The backbone network of the geographic entity recognition model is a multi-layer neural network. The basic framework of the multi-layer neural network is an input layer, a double layer , subject-object extraction layer, fully connected layer and output layer;
[0074] The input layer is used to encode the input text sentence to obtain the encoding vector corresponding to each word in the corresponding text sentence;
[0075] The double layer Including the first Layer and second layer; the first The layer extracts semantic features from the input encoding vector by introducing a control gate mechanism to obtain the corresponding word semantic feature vector, and assigns weights to each word in the corresponding text sentence based on the obtained word semantic feature vector to obtain the corresponding weighted sentence; wherein the control gate mechanism includes a forget gate, an enhancement gate, and an output gate;
[0076] Among them, the formula for obtaining the word semantic feature vector is: Where, and Respectively represent the first The word semantic feature vector corresponding to each word; Indicates the first The encoding vector corresponding to each word; and Represent the weight matrix and bias term of the output gate respectively; the output gate is used to control the corresponding first The output content of the layer; Represents vector concatenation; represents the pre-selected activation function;
[0077] Indicates the current time Next First The network status data corresponding to the layer is used to support the perception of long-distance context by the semantic features of the word, and serve the positioning of geographic entities (subject / object words) and relationship identification; among them, Where, and Both represent the weight matrix of the enhancement gate; and Both represent bias terms; the enhancement gate is used to Update the network status data of the layer; Represents the hyperbolic tangent function operation;
[0078] Indicates the first The forget gate output corresponding to the word, Where, and Represent the weight matrix and bias term of the forget gate respectively;
[0079] The formula for assigning weights is: ; Indicates the first The word weight corresponding to the word; 、 、 Respectively represent the semantic feature vectors of the corresponding words The weight matrix obtained by linear transformation; represents a matrix permutation; Represents the weight matrix The matrix dimensions of
[0080] The second BiLSTM layer is used to extract the word semantic feature vector corresponding to each word in the received weighted sentence, and perform feature fusion to obtain the corresponding sentence semantic feature vector;
[0081] The subject-object extraction layer is used to receive the word semantic feature vector output by the second BiLSTM layer and obtain the subject word position of the geographic entity in the corresponding text sentence based on a pre-set prediction classifier. The formula for obtaining the subject word position is: Where, Represents the probability of the subject word of the geographic entity at position S1 in the corresponding text sentence x, where S1 represents the position identifier of the subject word of the geographic entity, that is, the range where the subject word starts and ends; represents the probability that the corresponding i-th word is the starting position of the main word of the geographic entity; Where, represents the training weight; express deviation; Represents the word semantic feature vector corresponding to the i-th word in the weighted sentence, Indicates the total number of words in the text sentence; represents an exponential variable and ;in, By the corresponding probability Decision, if appropriate If the probability is greater than the pre-set threshold, , otherwise, ; Represents the trainable parameters of the model; and They represent the start position mark and end position mark of the subject word respectively; n represents the position index of the subject word between the start position and the end position of the subject word; Indicates the total number of words in the text sentence; Represents the product operation;
[0082] Acquire a word representation vector corresponding to the corresponding subject word based on the subject word position of the geographic entity, wherein the word representation vector is composed of at least one word semantic feature vector;
[0083] Furthermore, based on the acquisition process of the subject word position and the corresponding word representation vector of the geographic entity, the object word position and the corresponding word representation vector of the geographic entity in the corresponding text sentence are obtained, and the subject word and the object word of the corresponding geographic entity are marked in the corresponding text sentence;
[0084] The fully connected layer is used to receive the obtained word representation vector and sentence semantic vector, and merge them to obtain the corresponding entity vector ST;
[0085] The output layer is used to map the obtained entity vector to the category probability of a certain semantic category ;
[0086] Where, Indicates that the entity vector ST belongs to the first The category probability of semantic categories, and Represent exponential operation and activation function operation respectively; All semantic categories representing predefined semantic relations;
[0087] The obtained category probability is compared with the preset category threshold. If the category probability is greater than the category threshold, the corresponding semantic category is used as the semantic relationship between the corresponding geographic entities. If the category probability is not greater than the category threshold, no other operations are performed.
[0088] Defining the loss function for the geographic entity recognition model Where, and Represents the weight ratio, The location probability of the object word represented as a geographic entity is obtained in a similar process to the location probability of the subject word, and will not be elaborated in this application. Indicates the first training samples, Indicates the Semantic category labels of training samples; Indicates the first The potential triples corresponding to the training samples are ;
[0089] Obtain several sets of manually annotated geographic text sentences and construct corresponding training sample sets based on them. Define AdaGrad as the optimizer to continuously optimize the parameters of the geographic entity recognition model during the training process. Input the training sample sets into the geographic entity recognition model in batches and record the corresponding loss function values. When the loss function values of consecutive L2 batches no longer decrease or change, save the parameters of the geographic entity recognition model at this time, indicating that the training of the geographic entity recognition model is complete. L2 is a constant.
[0090] It should be further explained that, in the specific implementation process, the construction process of the semantic relationship network includes:
[0091] Obtain the geographic entities stored in the corresponding geographic knowledge ontology library, perform entity disambiguation on them, and mark the disambiguated geographic entities as conceptual entities. Entity disambiguation refers to distinguishing and marking geographic entities with the same name but different locations, and merging duplicate knowledge triples corresponding to geographic entities with different names but the same location.
[0092] Then, different conceptual entities and their corresponding semantic relationships are obtained; the corresponding conceptual entities are converted into corresponding graph nodes, and the semantic relationships are converted into corresponding directed edges, and the corresponding directions and relationship types are marked. The relationship types include affiliation, proximity, connectivity, etc. For example, if city A belongs to province B, then there is a directed edge in the graph from node "city A" to node "province B", and the relationship type of the edge is marked as "belongs to";
[0093] Furthermore, based on a pre-selected directed graph construction algorithm, and based on the graph nodes constructed by the corresponding conceptual entities and the edges corresponding to the semantic relationships, a semantic relationship network between the corresponding conceptual entities is constructed; the role of the directed graph construction algorithm is to organize and connect the constructed graph nodes and edges; to ensure the accuracy and completeness of the relationships, such as correctly reflecting the inclusion hierarchical relationships between different geographical areas, the interaction relationships between geographical elements, etc. in the graph.
[0094] It should be further explained that, in a specific implementation process, the process of obtaining the textual geographic entities and textual semantic relationships corresponding to the input geographic information text includes:
[0095] The input geographic information text is preprocessed and input into the geographic entity recognition model to obtain the geographic entities and semantic relationships in the corresponding geographic information text, and mark them as text geographic entities and text semantic relationships respectively.
[0096] It should be further explained that, in a specific implementation process, a corresponding initial dendrogram is constructed based on the obtained textual geographic entities and textual semantic relationships and in combination with a semantic relationship network; knowledge perception and semantic understanding are performed on the corresponding initial dendrogram based on the geographic knowledge ontology library to obtain a corresponding semantic tree, which includes the following process:
[0097] Obtain the extracted text geographic entities and text semantic relationships, and map them to the constructed semantic relationship network; and obtain the graph nodes and edges corresponding to the corresponding text geographic entities and text semantic relationships;
[0098] The graph nodes corresponding to the corresponding textual geographic entities are used as cluster centers, and other graph nodes in the corresponding semantic relationship network are clustered to obtain the graph nodes in the same cluster set as the corresponding textual geographic entities. The conceptual entities and semantic relationships of the corresponding edges corresponding to the corresponding graph nodes are obtained and marked as extended geographic entities and extended semantic relationships.
[0099] The process of clustering operation includes:
[0100] Obtain the graph node corresponding to the text geographic entity in the corresponding semantic relationship network and set it as the cluster center;
[0101] Then, based on the pre-set cluster radius, other graph nodes connected to the cluster center in the semantic relationship network are traversed. The cluster radius refers to the maximum number of hops between the cluster center node in the semantic relationship network, which is usually set to 2-3 hops, that is, other graph nodes that are no more than 2-3 edges away from the cluster center are considered;
[0102] Then, based on the traversal results, the knowledge triples corresponding to each graph node within the cluster radius are obtained, and based on the semantic similarity between the corresponding graph node and the cluster center, the graph nodes with a semantic similarity greater than a preset threshold are divided into the cluster set where the cluster center is located. The cluster set retains the node attributes and edge types in the original semantic relationship network, and adds a cluster identification attribute to mark the cluster category to which different nodes belong. For example, the cluster set formed with "Yangtze River" as the cluster center contains graph nodes corresponding to geographical entities such as cities and rivers related to the Yangtze River (such as "Three Gorges of the Yangtze River", "Yangtze River Basin", "Yangtze River Estuary" and other related geographical entities), as well as semantic relationships such as "flowing through", "merging into", and "located in" between them. The semantic similarity threshold is usually set to a value between 0.6 and 0.8, which can be adjusted according to the needs of the application scenario.
[0103] Among them, the mathematical calculation formula of semantic similarity is: Where, Represents the semantic similarity between the cluster center zx and the qth graph node within the cluster radius; Represents the normalized inverse of the shortest path distance between the cluster center zx and the qth graph node within the cluster radius; Represents the node attribute similarity, which is obtained by comparing the overlap between the attributes of the corresponding nodes; ;and and represents the weight coefficient;
[0104] Furthermore, based on the text semantic relationship and the extended semantic relationship, upstream and downstream relationships between different text geographic entities, between different extended geographic entities, and between text geographic entities and extended geographic entities are obtained; for example, a "contains" relationship can be reflected as an upstream node containing a downstream node (such as a province containing a city), and a "flows through" relationship can be reflected by connecting a node representing a river with a node representing a city or region it flows through at the same level to indicate their relationship in terms of geographical location and mutual connection, thereby constructing a complete tree structure;
[0105] Then, all the obtained textual semantic relations and extended semantic relations are obtained and structurally classified to obtain upstream and downstream relations or association relations between different textual geographic entities and extended geographic entities. For example, hierarchical relations (such as "includes", "belongs to", "composed of", etc.) are mapped to specific upstream and downstream relations, and interactive relations (such as "flows through", "adjacent", "connected", etc.) are mapped to association relations with specific labels.
[0106] A relationship determination matrix is constructed based on the obtained upstream and downstream relationships and association relationships. The relationship determination matrix is used to store and clarify the upstream and downstream relationships or association relationships between different textual geographic entities and extended geographic entities. The matrix element values of the relationship determination matrix are used to represent the type and direction of the relationship between the entities. The upstream and downstream relationships are recorded using directional values (e.g., +1 for top-to-bottom, -1 for bottom-to-top), and the same-level relationships are recorded using specific identifiers (e.g., 2) and additional semantic type tags.
[0107] Furthermore, an enhanced tree structure framework is constructed, which includes two parts: a trunk hierarchical structure and horizontal semantic links. The trunk hierarchical structure strictly follows the hierarchical characteristics of the tree and is constructed based on upstream and downstream relationships. The horizontal semantic links are stored in the relevant nodes in the form of metadata, which does not affect the basic structure of the tree, but is also considered during semantic processing. For example, the "flow through" relationship is represented in the data structure as follows: the river node and the city node each maintain their original positions in the tree, while the horizontal semantic links are added to the node attributes to record their mutual association and the semantic type of "flow through".
[0108] Select the textual geographic entity with the highest hierarchical relationship in the semantic relationship network and use it as the root node of the backbone hierarchical structure. Based on the relationship, determine the upstream and downstream relationships in the matrix and add other textual geographic entities or expanded geographic entities to the corresponding backbone hierarchical structure layer by layer. During the addition process, ensure that each geographic entity appears only once in the tree.
[0109] After the addition is completed, the interactive relationships (such as "flow through", "adjacent", etc.) are stored as attributes of tree nodes in the enhanced tree structure framework based on horizontal semantic links to obtain the initial dendrogram. Each attribute contains the relationship type and direction information. For example, the attributes of the node "Yangtze River" include [Type: "flow through", Target: "Wuhan City", Direction: "one-way"], and the node "Wuhan City" also contains the corresponding attributes pointing to "Yangtze River".
[0110] It should be noted that one embodiment of the present invention further includes: when multiple semantic relationships exist between two geographic entities, the most representative relationship is selected as the basis for constructing the tree structure according to a preset semantic relationship priority order; for example, if a city is both "located" within a mountain range (geographic containment relationship) and "belongs to" a province (administrative division relationship), the "belongs to" relationship is preferentially used to determine its position in the tree; the semantic relationship priority order, from high to low, is administrative division relationship, geographic containment relationship, spatial positioning relationship, and functional interaction relationship;
[0111] Then, the obtained initial dendrogram tree is input into the constructed geographic knowledge ontology library for semantic mapping and semantic reasoning to obtain the corresponding semantic tree;
[0112] The semantic mapping refers to mapping the node entities corresponding to the tree nodes in the initial dendrogram and the edge semantic relationships corresponding to the edges between the tree nodes to the geographic entities and semantic relationships in the geographic knowledge ontology library based on a pre-selected tree algorithm. For example, an extracted textual geographic entity is mapped to the geographic entity category of "river", and its relationship with the surrounding geographic entities (such as flowing through a certain city) can also be mapped to the semantic relationship in the geographic knowledge ontology library, thereby clarifying its accurate semantics.
[0113] Semantic reasoning refers to performing semantic reasoning on adjacent tree nodes based on the entity attributes within the knowledge triples corresponding to the mapped geographic entities, and obtaining corresponding semantic reasoning results. For example, if a lake is known to be upstream of a river, the knowledge system constructed by the geographic knowledge ontology library can infer that the water resources in the area where the lake is located may have an impact on the water volume and ecology of the downstream river, thereby expanding the understanding of the meaning of geographic entities and relationships.
[0114] After the semantic mapping and semantic reasoning are completed, the entity attributes stored in the corresponding address knowledge local library and the corresponding semantic reasoning results are stored in the tree nodes in the corresponding semantic tree.
[0115] It should be further explained that, in the specific implementation process, the process of extracting information based on the obtained semantic tree and obtaining the required geographic information includes:
[0116] The root node of the corresponding semantic tree is obtained as the starting node, and the corresponding semantic tree is traversed based on the preset extraction rules, and the tree nodes and edges corresponding to the tree nodes in the corresponding semantic tree that meet the preset extraction rules are marked. Then, the entity attributes and semantic reasoning results stored in the corresponding tree nodes and the edge semantic relationship between the tree nodes and the edges are read to obtain the corresponding geographic information; wherein the preset extraction rule refers to a pre-set condition or standard for selectively extracting specific geographic information from the semantic tree; for example, the preset extraction rule can be to extract all lakes in a river basin. During the traversal process, when a river node and a lake node with a specific semantic relationship with it (such as a "contains lake" relationship, which can be defined when the semantic tree is constructed) are encountered, if the relationship meets the restrictions in the rule, the geographic information of the lake node (such as the lake name, area, location, etc.) is extracted; wherein the river node and the lake node are tree nodes corresponding to the corresponding geographic entity; the river and the lake are both one of the corresponding geographic entity categories; wherein the predicted extraction rule is determined according to actual needs.
[0117] Then, the obtained geographic information is deduplicated and aggregated to obtain the corresponding integrated information. The semantic relationships between the geographic entities involved in the extracted integrated information are examined to see if they are clear and reasonable. If not, the corresponding integrated information is sorted and adjusted based on the local geographic knowledge base. After the adjustment is completed, the corresponding integrated information is converted into the required format file and visualized.
[0118] The present invention aims to achieve efficient and accurate extraction and integration of geographic information by constructing a geographic knowledge ontology library and semantic relationship network, and using multi-level neural network to construct a geographic entity recognition model, thereby providing more powerful technical support for geographic information processing.
[0119] Example 2
[0120] See also Figure 2 As shown, for parts not described in detail in this embodiment, please refer to the description of Example 1. A geographic information text intelligent extraction system based on knowledge perception is provided; it includes:
[0121] Data construction module, used to build geographic knowledge ontology library and build corresponding semantic relationship network based on it;
[0122] A feature recognition module is used to obtain the textual geographic entities and textual semantic relationships corresponding to the input geographic information text;
[0123] The information expansion module constructs a corresponding initial dendrogram based on the obtained text geographic entities and text semantic relationships and combines the semantic relationship network; performs knowledge perception and semantic understanding on the corresponding initial dendrogram based on the geographic knowledge ontology library to obtain a corresponding semantic tree;
[0124] The information extraction module extracts information based on the obtained semantic tree to obtain the required geographic information;
[0125] The modules are connected via wired and / or wireless means to achieve data transmission between modules.
[0126] Example 3
[0127] This embodiment discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the operation mode of the above-mentioned method for intelligent extraction of geographic information text based on knowledge perception is implemented.
[0128] Since the electronic device introduced in this embodiment is an electronic device used to implement a method for intelligent extraction of geographic information text based on knowledge perception in the embodiment of this application, based on the method for intelligent extraction of geographic information text based on knowledge perception introduced in the embodiment of this application, technical personnel in this field can understand the specific implementation of the electronic device of this embodiment and its various variations, so how the electronic device implements the method in the embodiment of this application will not be introduced in detail here. As long as technical personnel in this field implement the electronic device used in the method for intelligent extraction of geographic information text based on knowledge perception in the embodiment of this application, it falls within the scope of protection to be protected by this application.
[0129] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.
[0130] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for users of ordinary skill in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A method for intelligent extraction of geographic information text based on knowledge perception, characterized in that: include: Step 1: Build a geographic knowledge ontology library and build a corresponding semantic relationship network based on it; Step 2: Obtain the textual geographic entities and textual semantic relationships corresponding to the input geographic information text; Step 3: Based on the obtained textual geographic entities and textual semantic relations, and combined with the semantic relationship network, a corresponding initial dendrogram is constructed; Performing knowledge perception and semantic understanding on the corresponding initial dendrogram tree based on the geographic knowledge ontology library to obtain a corresponding semantic tree; Step 4: Extract information based on the obtained semantic tree to obtain the required geographic information.
2. The method for intelligent extraction of geographic information text based on knowledge perception according to claim 1 is characterized in that: The process of building a geographic knowledge ontology database includes: Obtain existing geographic entity concepts, and generalize and abstract them into basic geographic entities, wherein the basic geographic entities include natural geographic entities, artificial geographic entities, and combined geographic entities; the basic geographic entities are composed of several different geographic entity categories; Setting a data collection node, wherein the data collection node collects geographic information related to the corresponding basic geographic entity based on different data sources to obtain corresponding basic geographic data; Performing data preprocessing on the obtained basic geographic data to obtain corresponding basic text data, wherein the data preprocessing includes word segmentation, stop word removal and part of speech restoration; After preprocessing is completed, the geographic entities in the corresponding basic text data are identified based on the pre-built geographic entity recognition model, and the semantic relationships between the corresponding geographic entities are identified; At the same time, based on the basic text data, the attribute information of the corresponding geographic entity is extracted to obtain the corresponding entity attributes; Based on the obtained geographic entities, semantic relationships and entity attributes, corresponding knowledge triples are constructed and stored in the corresponding geographic entity categories, and based on them, the corresponding geographic knowledge ontology library is constructed.
3. The method for intelligent extraction of geographic information text based on knowledge perception according to claim 2 is characterized in that: The process of building a geographic entity recognition model includes: The backbone network of the geographic entity recognition model is a multi-layer neural network. The basic framework of the multi-layer neural network is an input layer, a double layer , subject-object extraction layer, fully connected layer and output layer; The input layer is used to encode the input text sentence to obtain the encoding vector corresponding to each word in the corresponding text sentence; The double layer Including the first Layer and second layer; the first The layer extracts semantic features from the input encoding vector by introducing a control gate mechanism to obtain the corresponding word semantic feature vector, and assigns weights to each word in the corresponding text sentence based on the obtained word semantic feature vector to obtain the corresponding weighted sentence; wherein the control gate mechanism includes a forget gate, an enhancement gate, and an output gate; The second The layer is used to extract the word semantic feature vector corresponding to each word in the received weighted sentence, and perform feature fusion to obtain the corresponding sentence semantic feature vector; The subject-object extraction layer is used to receive the second The word semantic feature vector output by the layer is used to obtain the main word position of the geographic entity in the corresponding text sentence based on a pre-set prediction classifier; based on the main word position of the geographic entity, the word representation vector corresponding to the corresponding main word is obtained; Based on the acquisition process of the subject word position and the corresponding word representation vector of the geographic entity, the object word position and the corresponding word representation vector of the geographic entity in the corresponding text sentence are obtained, and the subject word and the object word of the corresponding geographic entity are marked in the corresponding text sentence; The fully connected layer is used to receive the obtained word representation vector and sentence semantic feature vector, and merge them to obtain the corresponding entity vector; The output layer is used to map the obtained entity vector into a category probability of a certain semantic category; the obtained category probability is compared with the corresponding category threshold; if the category probability is greater than the category threshold, the corresponding semantic category is used as the semantic relationship between the corresponding physical entities; if the category probability is not greater than the category threshold, no other operation is performed; Obtain several sets of manually annotated geographic text sentences and construct corresponding training sample sets based on them; define As the optimizer continuously optimizes the parameters of the geographic entity recognition model during the training process, the training sample set is input into the geographic entity recognition model in batches, and the value of the corresponding loss function is recorded. When the value of the batch loss function no longer decreases or changes, the parameters of the geographic entity recognition model at this time are saved, and the training of the geographic entity recognition model is completed. is a constant.
4. The method for intelligent extraction of geographic information text based on knowledge perception according to claim 3 is characterized in that: The formula for obtaining word semantic feature vector is: Where, and Respectively represent the first and The word semantic feature vector corresponding to each word; represents the weight matrix of the output gate in the control gate mechanism; Indicates the first The encoding vector corresponding to each word; Represents the bias term of the output gate; Represents vector concatenation; represents the pre-selected activation function; Indicates the current time Next First The network status data corresponding to the layer, Where, and Both represent the weight matrix of the enhancement gate; and Both represent bias terms; the enhancement gate is used to Update the network status data of the layer; Represents the hyperbolic tangent function operation; Indicates the first The forget gate output corresponding to the word, Where, and Represent the weight matrix and bias term of the forget gate respectively; The formula for assigning weights is: ; Indicates the first The word weight corresponding to the word; Respectively represent the semantic feature vectors of the corresponding words The weight matrix obtained by linear transformation; represents a matrix permutation; Represents the weight matrix The matrix dimensions of .
5. The method for intelligent extraction of geographic information text based on knowledge perception according to claim 3 is characterized in that: The formula for obtaining the keyword position is: Where, The subject words representing geographic entities in the corresponding text sentences Position within The probability, The location identifier of the subject word that represents the geographic entity, that is, the range where the subject word starts and ends; Respectively represent the corresponding The probability of a word being the starting position of the main word of a geographic entity; Where, represents the training weight; represents the bias term; Indicates the weighted sentence The word semantic feature vector corresponding to the word, Indicates the total number of words in the text sentence; ;in, By the corresponding probability Decision, if appropriate If the probability is greater than the pre-set threshold, , otherwise, ; Represents the trainable parameters of the model; and Respectively represent the start position marker and end position marker of the main word; Indicates the position index of the subject word between the start position and the end position of the subject word; Indicates the total number of words in the text sentence; Represents the product operation; Class probability Where, Represents entity vector Belongs to the semantic relationship The probability of semantic categories, and Represent exponential operation and activation function operation respectively; All semantic categories representing predefined semantic relations; Defining the loss function for the geographic entity recognition model Where, and Represents the weight ratio, Positional probabilities of object words represented as geographic entities; Indicates the first training samples, Indicates the Semantic category labels of training samples; Indicates the first The potential triples corresponding to the training samples are .
6. The method for intelligent extraction of geographic information text based on knowledge perception according to claim 3 is characterized in that: The process of constructing a semantic relationship network includes: Obtain the geographic entities stored in the corresponding geographic knowledge ontology library, perform entity disambiguation on them, and mark the disambiguated geographic entities as conceptual entities; obtain different conceptual entities and their corresponding semantic relationships; and convert the corresponding conceptual entities into corresponding graph nodes, and the semantic relationships into corresponding directed edges, and mark the corresponding directions and relationship types; Based on a pre-selected directed graph construction algorithm, and based on the graph nodes constructed by the corresponding concept entities and the edges corresponding to the semantic relationships, a semantic relationship network between the corresponding concept entities is constructed.
7. The method for intelligent extraction of geographic information text based on knowledge perception according to claim 6 is characterized in that: The process of obtaining the textual geographic entities and textual semantic relationships corresponding to the input geographic information text includes: The input geographic information text is preprocessed and input into the geographic entity recognition model to obtain the geographic entities and semantic relationships in the corresponding geographic information text, and mark them as text geographic entities and text semantic relationships respectively.
8. The method for intelligent extraction of geographic information text based on knowledge perception according to claim 7 is characterized in that: The process of constructing the corresponding initial dendrogram based on the obtained text geographic entities and text semantic relationships and combining the semantic relationship network includes: Obtain the extracted text geographic entities and text semantic relationships, and map them to the constructed semantic relationship network; and obtain the graph nodes and edges corresponding to the corresponding text geographic entities and text semantic relationships; The graph nodes corresponding to the corresponding textual geographic entities are respectively used as cluster centers, and clustered with other graph nodes in the corresponding semantic relationship network to obtain the graph nodes in the same cluster set as the corresponding textual geographic entities. The conceptual entities and semantic relationships of the corresponding edges corresponding to the corresponding graph nodes are obtained and marked as extended geographic entities and extended semantic relationships. Obtain all obtained text semantic relationships and extended semantic relationships, perform structural classification on them, obtain upstream and downstream relationships or association relationships between different text geographic entities and extended geographic entities, and construct a corresponding relationship determination matrix based on them; Constructing an enhanced tree structure framework, wherein the enhanced tree structure framework includes a trunk hierarchical structure and horizontal semantic links; Select the textual geographic entity with the highest level of upstream and downstream relationships in the semantic relationship network and use it as the root node in the backbone hierarchical structure; and based on the upstream and downstream relationships in the relationship matrix, add other textual geographic entities or extended geographic entities layer by layer to the corresponding backbone hierarchical structure; After the addition is completed, the association relationship is stored as a property of the tree node in the enhanced tree structure framework based on the horizontal semantic link to obtain the initial dendrogram tree; Input the obtained initial dendrogram into the constructed local geographic knowledge base for semantic mapping and semantic reasoning, and obtain the corresponding semantic tree based on it; After the semantic mapping and semantic reasoning are completed, the entity attributes stored in the corresponding address knowledge local library and the corresponding semantic reasoning results are stored in the tree nodes in the corresponding semantic tree.
9. The method for intelligent extraction of geographic information text based on knowledge perception according to claim 8 is characterized in that: The semantic mapping refers to mapping the node entities corresponding to the tree nodes in the initial dendrogram and the edge semantic relationships corresponding to the edges between the tree nodes to the geographical knowledge triples in the geographical knowledge ontology library, wherein the node entities refer to textual geographical entities or extended geographical entities; The edge semantic relationship refers to a text semantic relationship or an extended semantic relationship; the semantic reasoning refers to performing semantic-level reasoning on other tree nodes adjacent to the corresponding tree node based on the entity attributes in the knowledge triples mapped thereto.
10. The method for intelligent extraction of geographic information text based on knowledge perception according to claim 8, characterized in that: The process of extracting information based on the obtained semantic tree and obtaining the required geographic information includes: The root node of the corresponding semantic tree is obtained as the starting node, and the corresponding semantic tree is traversed based on the preset extraction rules. The tree nodes and edges corresponding to the tree nodes in the corresponding semantic tree that meet the preset extraction rules are marked, and the entity attributes and semantic reasoning results stored in the corresponding tree nodes and the edge semantic relationships of the tree nodes are read to obtain the corresponding geographic information; the obtained geographic information is deduplicated and summarized to obtain the corresponding integrated information, and whether the semantic relationships between the geographic entities involved in the extracted integrated information meet the requirements are examined. If not, the corresponding integrated information is sorted and adjusted according to the local geographic knowledge library; after the adjustment is completed, the corresponding integrated information is converted into the required format file and visualized output.
Citation Information
Patent Citations
Geographic knowledge acquisition method
CN112256888A
Social media address information extraction method and system fusing embedded semantics
CN118585645A
Knowledge semantic tree construction method based on text data
CN119917942A