A method and apparatus for determining spatial fuzzy location based on place name spatiotemporal derivation network
By constructing a spatiotemporal derivation network of place names and utilizing text, spatial, and temporal feature recognition methods, the problems of insufficient semantic expression and information retrieval of place names are solved, enabling accurate reasoning of spatially ambiguous locations and efficient retrieval of geographic information.
Patent Information
- Application Number
- CN202511086256.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-05
AI Technical Summary
Existing technologies fail to fully utilize the spatiotemporal derivation relationships of place names to enhance the connections between them, resulting in insufficient semantic expression and information retrieval capabilities of place names.
By constructing a spatiotemporal derivation relationship network of place names, and using text, spatial and temporal feature recognition methods, the derivation relationships between place names are identified, and spatial fuzzy location information, including orientation and distance, is determined based on a knowledge graph database.
It realizes spatial fuzzy location reasoning based on the spatiotemporal derivation relationship network of place names, which improves the accuracy and efficiency of place name semantic expression and information retrieval.
Smart Images

Figure CN120578786B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data information processing, and in particular to a method and apparatus for determining spatial fuzzy location based on a place name spatiotemporal derivation relationship network. Background Technology
[0002] Place names are the names of natural or human geographical entities located in a specific spatial location. Simply put, place names originate from the conceptualization and naming of geographical elements, entities, or places. Place names play a crucial role in Geographic Information Systems (GIS), serving not only as core references for spatial data location but also carrying rich semantic information. Place names are a representative and unique type of geographic data in GIS, providing an intuitive way to identify and access specific geographical locations, thereby enhancing the retrieval, analysis, and visualization capabilities of GIS data.
[0003] The spatiotemporal derivation relationship of place names refers to the process of creating new place names through derivation based on the attributes of existing place names during the naming of newly discovered place names. This relationship of derivation and being derived between existing and new place names constitutes the "spatial-temporal derivation relationship of place names." It not only reflects the textual similarity between two place names but also associates them with the location of geographical entities, indicating their spatial proximity. This plays a crucial role in semantic expression and geographic information retrieval. Although the spatiotemporal derivation relationship of place names has potential applications in the expression and retrieval of geographic information, existing research has been insufficient in exploring this relationship, failing to fully utilize it to enhance the correlation between place names and improve the semantic expression and information retrieval of place names. Summary of the Invention
[0004] The purpose of this application is to provide a method and apparatus for determining spatial fuzzy locations based on a place name spatiotemporal derivation relationship network, which can make full use of place name spatiotemporal derivation relationships to achieve inference and determination of spatial fuzzy locations.
[0005] To achieve the above objectives, this application provides the following solution:
[0006] Firstly, this application provides a spatial fuzzy location determination method based on a place-name spatiotemporal derivation relationship network, including:
[0007] Obtain information on the features to be inferred;
[0008] Based on the place name spatiotemporal derivation relationship network, a spatial similarity set is determined by semantic similarity according to the information of the place features to be inferred; the place name spatiotemporal derivation relationship network is determined by the place name spatiotemporal derivation relationship identification method, and is used to represent the knowledge graph database of place name spatiotemporal derivation relationships; the place name spatiotemporal derivation relationship is determined by the process of naming new place names based on the spatiotemporal relationship of the place features referred to by known place names;
[0009] Using the aforementioned place name spatiotemporal derivation relationship, spatial fuzzy location information is determined based on the spatial similarity set; the spatial fuzzy location information includes: orientation and distance; the spatial fuzzy location information is used to provide a basis for retrieving geographic information.
[0010] In one embodiment, the method for identifying the spatiotemporal derivation relationship of place names specifically includes: a text feature identification method, a spatial feature identification method, and a temporal feature identification method.
[0011] In one embodiment, the text feature recognition method employs a sequence comparison approach to calculate text similarity; the expression corresponding to the text similarity is:
[0012] ;
[0013] in, Text similarity; The longest common subsequence of the two place names; It is the original place name; It is a derived place name; This is a function for calculating the sequence length.
[0014] In one embodiment, the spatial feature recognition method uses a spatial proximity classification model to identify spatial features and obtain recognition results; the recognition results include spatial proximity relationships and non-spatial proximity relationships.
[0015] The method for determining the spatial proximity classification model specifically includes:
[0016] The topological relationships are determined based on the nine-intersection model with dimension extension;
[0017] Feature filtering is performed based on the topological relationships to obtain filtered features; the filtered features include: spatial distance between land features and the categories of primary land features and derived land features;
[0018] A dataset is constructed based on the selected features and the corresponding recognition results;
[0019] The dataset is divided into a training set and a test set;
[0020] The K-fold cross-validation method is used to train and optimize the data metrics parameters of the classification model based on the training set, resulting in a trained classification model. The data metrics parameters include precision, recall, and F1 score. The classification model uses a CART decision tree.
[0021] The trained classification model is tested based on the test set, and its generalization performance is evaluated based on the confusion matrix to obtain the tested classification model.
[0022] The tested classification model will be used as the spatial proximity classification model.
[0023] In one embodiment, the time feature recognition method converts time into timestamps, compares the sizes of any two timestamps, and determines the time feature recognition result based on the comparison result; the time feature recognition result is used to determine whether the original place name was generated earlier than the derived place name.
[0024] In one embodiment, the spatial fuzzy location information is determined based on the spatial similarity set using the spatiotemporal derivation relationship of place names, specifically including:
[0025] Based on the directional relationships between features in the spatial similarity set, the location of the feature information to be inferred is determined;
[0026] Based on the spatial constraint distance of native features, the distance between the feature information to be inferred and the native features is determined.
[0027] In one embodiment, the place name spatiotemporal derivation relationship network is stored in a place name database.
[0028] In one embodiment, it further includes:
[0029] The spatially ambiguous location information is verified to determine its accuracy.
[0030] In one embodiment, a success rate function is used to determine accuracy; the mathematical expression of the success rate function is:
[0031] ;
[0032] in, The number of correctly positioned items; The number of units whose locations are determined by spatiotemporal derivation relationships; For success rate.
[0033] Secondly, this application provides a spatial fuzzy location determination device based on a place name spatiotemporal derivation relationship network, comprising:
[0034] The information acquisition module is used to acquire information about the features to be inferred.
[0035] The spatial similarity set determination module is used to determine the spatial similarity set based on the place name spatiotemporal derivation relationship network and semantic similarity according to the information of the place features to be inferred; the place name spatiotemporal derivation relationship network is determined by the place name spatiotemporal derivation relationship identification method and is used to represent the knowledge graph database of place name spatiotemporal derivation relationships; the place name spatiotemporal derivation relationship is determined by the process of naming new place names based on the spatiotemporal relationship of the place features referred to by known place names;
[0036] The spatial fuzzy location information determination module is used to determine spatial fuzzy location information based on the spatial similarity set using the spatiotemporal derivation relationship of the place names; the spatial fuzzy location information includes: orientation and distance; the spatial fuzzy location information is used to provide a basis for retrieving geographic information.
[0037] According to the specific embodiments provided in this application, the following technical effects are disclosed:
[0038] This application provides a method and apparatus for determining fuzzy spatial locations based on a place name spatiotemporal derivation relationship network. It utilizes semantic similarity to obtain a set of spatially similar features to be inferred from the place name spatiotemporal derivation relationship network. The place name spatiotemporal derivation relationship network is determined using a place name spatiotemporal derivation relationship identification method, serving as a knowledge graph database representing place name spatiotemporal derivation relationships. Place name spatiotemporal derivation relationships are determined through the process of naming new place names based on the spatiotemporal relationships of known place names representing features. Then, spatial fuzzy location inference is performed using the spatiotemporal derivation relationships to determine fuzzy spatial location information, providing a basis for geographic information retrieval. Therefore, this application can fully utilize place name spatiotemporal derivation relationships to achieve the inference and determination of fuzzy spatial locations. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 A flowchart of a spatial fuzzy location determination method based on a place name spatiotemporal derivation network;
[0041] Figure 2 Flowchart for identifying the spatiotemporal derivation relationship of place names;
[0042] Figure 3 A schematic diagram illustrating the process of constructing a knowledge graph of spatiotemporal derivation relationships of place names;
[0043] Figure 4 This is a schematic diagram of a spatiotemporal derivation relationship network;
[0044] Figure 5 This is a schematic diagram of spatial fuzzy position reasoning;
[0045] Figure 6 This is a diagram illustrating the reasoning process for place names.
[0046] Figure 7 A structural diagram of a spatial fuzzy location determination device based on a place name spatiotemporal derivation relationship network. Detailed Implementation
[0047] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0048] The core objective of this application is to address the insufficient application of spatiotemporal derivation relationships of place names in the existing field of geographic information science. Therefore, a specific method for identifying and representing spatiotemporal derivation relationships of place names is proposed. Based on the construction of a knowledge graph of spatiotemporal derivation relationships of place names, this network is used to conduct research on spatially fuzzy location reasoning.
[0049] The significance of this application lies in the fact that researching the spatiotemporal derivation relationship of place names not only enriches the theoretical framework of this relationship but also provides new perspectives and tools for practical applications in the field of geographic information science. By incorporating the spatiotemporal derivation relationship of place names into the reasoning system, it can provide support for fields such as semantic reasoning of place names.
[0050] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0051] In an exemplary embodiment, Figure 1 As shown, a spatial fuzzy location determination method based on a place-name spatiotemporal derivation relationship network is provided, including:
[0052] Step 100: Obtain information on the features to be inferred.
[0053] Step 200: Based on the place name spatiotemporal derivation relationship network, determine the spatial similarity set according to the information of the place features to be inferred using semantic similarity. The place name spatiotemporal derivation relationship network is determined by the method of identifying place name spatiotemporal derivation relationships, and is used as a knowledge graph database to represent place name spatiotemporal derivation relationships; place name spatiotemporal derivation relationships are determined by the process of naming new place names based on the spatiotemporal relationships of the place features referred to by known place names.
[0054] The spatiotemporal derivation network of place names is stored in the place name database.
[0055] In one embodiment, the method for identifying the spatiotemporal derivation relationship of place names specifically includes: a text feature identification method, a spatial feature identification method, and a temporal feature identification method.
[0056] The text feature recognition method uses sequence comparison to calculate text similarity; the expression for text similarity is:
[0057] .
[0058] in, Text similarity; The longest common subsequence of the two place names; It is the original place name; It is a derived place name; This is a function for calculating the sequence length.
[0059] The spatial feature recognition method uses a spatial proximity relationship classification model to identify spatial features and obtain recognition results; the recognition results include spatial proximity relationships and non-spatial proximity relationships.
[0060] The method for determining the spatial proximity classification model specifically includes:
[0061] The topological relationships are determined based on the nine-intersection model with dimension extension.
[0062] Feature filtering is performed based on topological relationships to obtain filtered features; the filtered features include: spatial distance between land features and the categories of native and derived land features.
[0063] A dataset is constructed based on the selected features and the corresponding recognition results; the dataset is then divided into a training set and a test set.
[0064] The K-fold cross-validation method was used to train and optimize the data metrics parameters of the classification model based on the training set, resulting in the trained classification model. The data metrics parameters included precision, recall, and F1 score. The classification model used a CART decision tree.
[0065] The trained classification model is tested using the test set, and its generalization performance is evaluated based on the confusion matrix to obtain the tested classification model; the tested classification model is then used as the spatial proximity classification model.
[0066] The time feature recognition method converts time into timestamps, compares the size of any two timestamps, and determines the time feature recognition result based on the comparison result. The time feature recognition result is used to determine whether the original place name was generated earlier than the derived place name.
[0067] Step 300: Using the spatiotemporal derivation relationship of place names, determine the spatial fuzzy location information based on the spatial similarity set. The spatial fuzzy location information includes: orientation and distance; the spatial fuzzy location information is used to provide a basis for retrieving geographic information.
[0068] In one embodiment, spatial fuzzy location information is determined based on a set of spatial similarities using the spatiotemporal derivation relationship of place names. Specifically, this includes:
[0069] Based on the directional relationships between features in a spatially similar set, the orientation of the feature information to be inferred is determined; based on the spatial constraint distance of the original features, the distance between the feature information to be inferred and the original features is determined.
[0070] As an optional implementation method, the spatial fuzzy location determination method based on the spatiotemporal derivation relationship network of place names mentioned in this application further includes: verifying the spatial fuzzy location information to determine its accuracy.
[0071] Accuracy is determined using a success rate function; the mathematical expression for the success rate function is:
[0072] .
[0073] in, The number of correctly positioned items; The number of units whose locations are determined by spatiotemporal derivation relationships; For success rate.
[0074] This application retrieves place name data from a place name database that has a spatiotemporal derivation relationship with the place feature to be inferred, and uses semantic similarity to obtain the spatial similarity set of the place feature to be inferred from the spatiotemporal derivation relationship network. Spatial fuzzy location inference (determination) is then performed using the spatiotemporal derivation relationship. This method can effectively improve the utilization rate of unregistered place names and provides an innovative solution for place name data management.
[0075] The research focus of this application is to use the semantic and spatial relationships contained in the constructed place name spatiotemporal derivation relationship network to perform spatial fuzzy location reasoning for missing geographical entities in the place name database, and to study and discuss the influencing factors that need to be considered in the reasoning process.
[0076] This application method includes the following three steps:
[0077] The first step is to obtain place name data from the place name database that have a spatiotemporal derivation relationship with the place features to be inferred.
[0078] By using methods for identifying the spatiotemporal derivation relationships of place names, including text feature recognition methods, spatial feature recognition methods, and temporal feature recognition methods, place name data with spatiotemporal derivation relationships are determined and entered into the place name database.
[0079] Specifically, based on the definition of place name spatiotemporal derivation relationship, the characteristics and identification methods of place name spatiotemporal derivation relationship in terms of text, space and time were determined.
[0080] Spatiotemporal derivation of place names: The process of naming a new place name based on its spatiotemporal relationship with the features referred to by an existing place name is called "place name derivation". The existing place name is called the "original place name", and the newly formed place name is called the "derived place name". The relationship of derivation and being derivation between the original place name and the derived place name is called the "spatiotemporal derivation of place names".
[0081] In terms of text feature recognition, text similarity is used to identify whether place names are similar; in terms of spatial feature recognition, topological relationships and decision trees are used to identify spatial proximity relationships between features; and in terms of temporal feature recognition, timestamps are used to identify the order in which two place names were generated.
[0082] Methods for identifying the spatiotemporal derivation relationships of place names can more accurately understand the semantic connections and spatial proximity between place names, and also enrich the semantic information of place names, so that they not only reflect geographical locations, but also reflect the historical, cultural and geographical connections between place names.
[0083] The method for identifying the spatiotemporal derivation relationship of place names first uses text similarity to determine whether two place names are similar in naming; then it calculates the topological relationship between geographical entities, and uses a trained decision tree model to determine the spatial proximity relationship of geographical features with disjoint topological relationships; finally, it uses timestamps to determine the chronological order of the generation of the original place name and the derived place name. Figure 2 A flowchart for identifying the spatiotemporal derivation relationship of place names.
[0084] 1. Text feature recognition method.
[0085] Based on the textual features of the spatiotemporal derivation relationship of place names, it can be seen that the original place names and derived place names are similar in name. Therefore, this application uses text similarity to calculate the similarity between two place names and performs text feature recognition based on the place name similarity score.
[0086] Text similarity, as an indicator of the semantic or semantic closeness between two texts, indicates a greater semantic difference and lower semantic similarity between them, as a smaller value indicates greater semantic difference; conversely, a larger value indicates higher semantic similarity. Given that place names are expressed as strings, this application employs a sequence comparison-based method to calculate the similarity between two place names. This method calculates the similarity by identifying the longest common subsequence (LCS) of the two strings. This method is fast, considers the sequence length, and standardizes the results, allowing for direct comparison.
[0087] .
[0088] The similarity score (result) is standardized to a range of 0 to 1. A score of 0 indicates that the two strings have no common subsequences, meaning the two place names have a common place name relationship. A score of 1 indicates that the two strings are completely identical, meaning the two place names have a complete spatiotemporal derivation relationship. Other scores indicate that the two strings only have partial similarities, meaning the two place names have a partial spatiotemporal derivation relationship.
[0089] 2. Spatial feature recognition methods.
[0090] Spatial feature identification primarily determines whether the features referred to by place names with derived and derived relationships have spatial proximity. First, the topological relationships between geographic entities are calculated; then, features with disjoint topological relationships are determined to have spatial proximity through spatial constraint distance. This application views this process as a classification process of whether spatial proximity exists, using feature category and distance as feature factors and employing a decision tree for classification, thereby identifying spatial features.
[0091] Topological relationships: The nine-intersection model based on dimension extension proposed by Clementini E et al. is adopted to analyze geographic entities. a and geographical entities b A relation matrix is constructed using the intersection dimensions (DIM) of the interior (I), boundary (B), and exterior (E) regions to extract the topological relationships of geographic entities. .
[0092] .
[0093] I , B , E These represent the interior, boundary, and exterior of a geographic entity, respectively. DIM Indicates dimension.
[0094] Spatial Proximity Classification Model: For place names that are topologically disjoint, a final step of determining spatial proximity is needed based on the spatial constraint distance between the two features. The threshold for this spatial constraint distance is related to the feature category. Therefore, determining whether a spatial proximity relationship exists between two place names based on the categories of native and derived features and the spatial distance between them can be viewed as a text classification process, that is, classifying the relationship between features into spatial proximity and non-spatial proximity relationships based on the two features of feature category and distance.
[0095] Text classification typically involves considerations such as the training sample set used to train the classifier, feature selection, classification algorithm, and performance evaluation metrics for the classification system. These will be discussed in detail below.
[0096] (a) Feature screening.
[0097] Feature selection, also known as feature subset selection (FSS), refers to selecting the features that have the greatest impact on the classification of a study from the original features. This method increases the generality and robustness of the model, while also strengthening the relationship between features and feature sets, thereby improving the model's accuracy and interpretability. Based on the spatial characteristics of place name spatiotemporal derivation relationships and the above analysis, the features used in this application are the spatial distance between place features and the categories of primary and derived place features.
[0098] Spatial Distance: To accurately quantify the spatial relationships between geographic entities, this application uses numerical values to represent the quantitative distance between entities. Considering that geographic entities referred to by place names have three feature types—points, lines, and polygons—the distance calculation first extracts the center points of polygon and line features, and then calculates the distance between the center point and the point feature. The formula for the distance between geographic entities is:
[0099] .
[0100] in, and These represent the spatial coordinates of the geographical entities referred to by the two place names. It represents the distance between two entities.
[0101] Category: Determine the category to which a place name belongs based on its attribute information.
[0102] (ii) Classification algorithm.
[0103] Text classification algorithms mainly include association rule-based classification algorithms, logistic regression, support vector machines (SVM), Bayesian classification, decision trees, and neural network algorithms.
[0104] Classification algorithms based on association rules are called association classification. CBA (Classification Based on Association Rule) algorithms construct a classifier in two steps. The first step uses standard association rule mining algorithms to discover relevant association rules, i.e., Classification Association Rules (CARs), where the right-hand side represents the class attribute value. The second step selects high-priority rules from the discovered CARs to cover the training set. If multiple association rules have the same left-hand side but different right-hand sides, the rule with the highest confidence is selected as the possible rule. The CBA algorithm primarily constructs the classifier by discovering association rules in the training set. The classic algorithm for discovering association rules is the Apriori algorithm. Its advantage is high classification accuracy, but its main disadvantages are: when the size of the latent frequent two-itemsets is large, the algorithm becomes limited by hardware memory, resulting in excessive system I / O load and low efficiency; secondly, the time cost is too high.
[0105] Logistic regression is a classification algorithm primarily used to solve binary classification problems. The output variable of logistic regression has only two possible values, representing either or both classes. In classification tasks, logistic regression uses the logistic function (also known as the sigmoid function) to construct a prediction function. This function maps a linear combination of inputs to a probability value between 0 and 1, representing the probability that a sample belongs to a certain class. In binary classification problems, the goal is to predict probabilities as closely as possible to the actual observations. The loss function measures the difference between the model's predictions and the actual observations. Its form can be derived from the log-likelihood function. Gradient descent is used to iteratively adjust the model parameters, continuously reducing the loss function to find the parameter configuration that optimizes model performance. However, logistic regression struggles to handle imbalanced data and has relatively low accuracy.
[0106] Support Vector Machines (SVMs) represent a new development in statistical learning theory. Unlike traditional statistics, SVMs are not based on the traditional principle of empirical risk minimization, but rather on the principle of structural risk minimization, evolving into a novel structured learning method. It effectively solves the problem of constructing high-dimensional models with a limited number of samples, and the constructed models exhibit excellent predictive performance. As a powerful machine learning algorithm, SVMs possess efficient nonlinear classification capabilities, the ability to handle high-dimensional data, and strong interpretability and robustness. However, SVMs are slow when processing large-scale datasets, are sensitive to parameter selection, cannot directly handle multi-class problems, and are relatively sensitive to missing data.
[0107] Bayesian classification primarily utilizes Bayes' theorem to predict the probability that a sample of an unknown class belongs to each of the various classes, selecting the class with the highest probability as the final class for that sample. In many situations, the basic Bayesian classification algorithm is comparable to decision tree and neural network classification algorithms. This algorithm can be applied to large databases and is simple, accurate, and fast. Bayesian classifiers assume that all features are independent, meaning there are no relationships between features. In practical applications, this assumption does not always hold true; dependencies between features can affect classification accuracy, and accuracy decreases when this assumption is not met. The accuracy of Bayesian classifiers depends on the training data, especially when the feature dimensionality is high or the data noise is significant, requiring sufficient training samples to obtain reliable probability estimates.
[0108] Decision Trees (DT) are a supervised learning algorithm that builds a tree-like model through a top-down recursive construction process. The goal of this model is to learn decision rules from training data to predict the label value of a target variable. The CART decision tree classification method is convenient, easy to understand, and efficient, making it one of the mainstream classification methods. This method selects features at each node of the decision tree using the Gini coefficient method, building the tree recursively. Starting from the root node, the Gini coefficient values of all features are calculated, and the minimum value is selected as the feature of that node. The Gini coefficient of that feature is then calculated again, and the minimum value is selected as the splitting node for that feature. One of the advantages of decision trees is their ease of interpretation and visualization, as each node corresponds to a simple decision made for a particular feature. The classification method of decision tree models is concise, direct, easy to understand and interpret, and as a very fast learning and prediction algorithm, it can provide high efficiency in text classification.
[0109] The advantages of neural network methods lie in their ability to approximate any function with arbitrary precision, and classification itself is the process of determining the classification discriminant function. Neural network methods are nonlinear models, which allows them to adapt to various complex data relationships in the real world. Neural networks possess strong learning capabilities; through learning, they can automatically adjust their internal weights, gradually adapting to different tasks and environments. This enables neural networks to exhibit good generalization ability when facing new tasks. However, neural networks are often considered "black box" models, and their outputs are often difficult to interpret. This can lead to doubts about the results when dealing with neural networks. Furthermore, neural networks typically require a large number of parameters for training. These parameters include weights, biases, etc., and their number can be enormous, increasing the complexity and training difficulty of the neural network.
[0110] Compared to other classification models such as neural networks or Bayesian classification, decision trees have a simple and easy-to-understand classification principle, readily generating understandable rules, making them easier to comprehend and accept. Furthermore, decision tree classification generally does not require manual parameter setting, is fast to build, and can handle both continuous and discrete values. In addition, the decision tree method uses information principles to analyze the information content of attributes in a large number of samples, calculating the information content of each attribute, which can clearly show which features are more important for classification. Based on the characteristics of the research data, this application selects CART decision trees (Classification and Regression Trees) as the spatial proximity classifier.
[0111] (III) Model training and evaluation.
[0112] The constructed dataset is divided into training and test sets. A decision tree model is then trained using the training set data, with parameters set to control the tree depth and prevent overfitting. The trained model is then used to predict the spatial proximity relationships between ground features. For trained decision tree classification models, precision, recall, and F1 score are typically used as data metrics to evaluate the model's prediction accuracy. However, in cases of imbalanced data or significantly different error costs, simple prediction accuracy may not meet all requirements. Therefore, this application combines a confusion matrix-based generalization performance evaluation technique to comprehensively assess the model's learning ability.
[0113] Accuracy Recall rate The formula for calculating the F1 score is as follows:
[0114] .
[0115] In this system, TP stands for True Positive, representing the number of positive samples predicted as true; FP stands for False Positive, representing the number of negative samples predicted as true; and FN stands for False Negative, representing the number of positive samples predicted as false. In terms of meaning, precision measures the accuracy of positive predictions; recall measures how many of the actual positive samples were correctly predicted by the model; and the F1 score considers both precision and recall, providing a comprehensive evaluation of the model's performance in the form of a harmonic mean. A higher F1 score indicates better classification performance.
[0116] A confusion matrix is a tool used to evaluate the performance of a classification model. The diagonal elements of the confusion matrix represent the number of correctly predicted values and the off-diagonal elements represent the number of incorrectly predicted values. Therefore, comparing the predicted values with the true values in the confusion matrix can reveal the distribution of incorrect predictions across different categories. To reduce the risk of underfitting or overfitting, the decision tree model needs to be generalized. K-fold cross-validation can uniformly fit the data distribution; that is, it performs multiple splits between the training and test sets, using the resulting training and test sets for training and testing the model. This establishes a balance between the training and test sets, helping to assess the consistency level of different randomly split datasets and improve the accuracy of the proposed model. Therefore, this application uses K-fold cross-validation for model tuning. Furthermore, this application also uses an independent test set to classify land features with spatial proximity relationships to evaluate the model's generalization ability.
[0117] For an optimized decision tree model, the land cover category and spatial distance are used as feature parameters to determine whether there is a spatial proximity relationship between two land covers.
[0118] 3. Time feature recognition.
[0119] Convert the time into a timestamp, and then directly compare the size of the two timestamps to determine whether the original place name was created earlier than the derived place name.
[0120] The second step is to use semantic similarity to obtain the spatial similarity set of the objects to be inferred from the spatiotemporal derived relation network.
[0121] The spatiotemporal derivation relation network is a knowledge graph database that uses semantic similarity to calculate the similarity between the place name to be inferred and the place names in the spatiotemporal derivation relation network, and selects the place names with higher similarity scores to form a similarity set. Figure 3 The process of constructing a knowledge graph of spatiotemporal derivation relationships for place names. Data layers are constructed based on open geographic databases and the Geographic Information System platform (GeoNames) and the open geographic database (Open Street Map, OSM).
[0122] The similarity formula is the same as the expression for text similarity described above.
[0123] The third step is to use spatiotemporal derivation relationships to predict the location (determining spatially fuzzy location information).
[0124] By examining the directional relationships between spatial similarity sets, the approximate location of the object to be judged is inferred by reasoning the directional relationships between objects in the similarity sets. At the same time, the distance is determined based on the spatial constraint distance of the original object to judge, which is the farthest distance between the object to be judged and the original object. The prediction result provides the approximate location and distance.
[0125] The fourth step is to verify the proposed spatial fuzzy location reasoning method and use the success rate to evaluate the accuracy of the spatial location reasoning results proposed in this application using spatiotemporal derivation relationship networks.
[0126] The mathematical expression for the success rate function is:
[0127] ;
[0128] in, The number of correctly positioned items; The number of units whose locations are determined by spatiotemporal derivation relationships; For success rate.
[0129] For example: Figure 4 As shown, the place name "University Plot 5 of a Certain Location" does not exist in the place name database, so the geographical location of the geographical entity referred to by this place name cannot be obtained. Using the spatial fuzzy location reasoning method of this application, we first search for place name data with a spatiotemporal derivation relationship with this place name in the place name database. First, we identify the place names with the derivation relationship and mark them in the database. (Through keyword search, we search for place name data with a spatiotemporal derivation relationship with this place name) Then, we obtain the spatial similarity set of the geographical feature: University of a Certain Location, University Plot 1 of a Certain Location, University Plot 2 of a Certain Location, University Plot 3 of a Certain Location, and University Plot 4 of a Certain Location; we determine the directional relationship between the various names in the spatial similarity set; then, according to the definition of the derivation relationship, it can be clearly determined that University Plot 5 of a Certain Location is located within the constraint range of University of a Certain Location, and based on the westward linear distribution of University Plot 1, University Plot 2, University Plot 3, and University Plot 4, it is determined that University Plot 5 of a Certain Location is located due west of University Plot 4.
[0130] like Figure 5 As shown, for spatial location reasoning, this application adopts a method based on existing derivation relationships in the database to infer the geographical location of unregistered place names in the database. Location is a fundamental and important feature among the many spatial characteristic dimensions. Cognitive spatial location can answer the "Where are we?" question among the six major geographical questions and provide spatial references for answers to other geographical questions. In urban space, the main forms of location information include: coordinates, postal codes, telephone numbers, IP addresses, place names, and addresses. Among these, except for coordinates, other location data need to be spatially located before they can be converted into urban space. Among these data types that need to be converted, place name data is a relatively standardized data in natural language form with a wide range of applications, and is widely included in data from citizens' lives, government management, and enterprises. Figure 5 The derivation relationship lookup in the middle uses the following method: Figure 4 The spatiotemporal derivation network in the middle.
[0131] like Figure 6 As shown, a map is used to verify the feasibility of location reasoning. To perform location reasoning for a given place name (place name 1), first, place name data with a spatiotemporal derivation relationship with this place name is searched in the place name database; then, the spatial similarity set of the feature is obtained; finally, based on the definition of derivation relationship, it can be clearly determined that place name 1 is within the constraint range of the place name "Hongmoulu". Searching the location of place name 1 on the map reveals that it is approximately 1760 meters away from "Hongmoulu".
[0132] Evaluation of spatial fuzzy location reasoning methods.
[0133] The proposed spatial fuzzy location reasoning method was validated, and the success rate was used to evaluate the accuracy of the spatial location reasoning results using spatiotemporal derivation relation networks. Specific results are shown in Table 1.
[0134] Table 1 Detailed Results Information Table
[0135]
[0136] Because spatiotemporal derivation relationships are category-dependent, this analysis randomly selects 300 place name data points from a specific category for location inference, setting a location threshold of 2 kilometers. If the inferred location is within 2 kilometers, the inference is considered correct; otherwise, it is considered incorrect. The 300 place name data points contain 32 pairs of derived place name pairs, with 25 correctly inferred. Therefore, the success rate of spatial location inference using place name spatiotemporal derivation relationships is 78%. The limitations of spatially fuzzy location inference lie in the need for further optimization of the location threshold setting and the inference of location orientation.
[0137] The spatiotemporal derivation relationship of place names, as a special type of place name relationship, can simultaneously represent the semantic and spatiotemporal connection between two place names. Therefore, identifying the derivation relationship between place names can not only enrich the semantic expression of place names but also enable more accurate retrieval of geographic information. The success rate of spatial location reasoning using the spatiotemporal derivation relationship of place names is 78%. The results indicate that the method proposed in this application has a high success rate in spatial location reasoning and can provide a new perspective and tool for practical applications in the field of geographic information science.
[0138] In an exemplary embodiment, Figure 7 As shown, a spatial fuzzy location determination device based on a place name spatiotemporal derivation relationship network is provided, comprising:
[0139] The information acquisition module is used to acquire information about the features to be inferred.
[0140] The spatial similarity set determination module is used to determine the spatial similarity set based on the place name spatiotemporal derivation relationship network and semantic similarity according to the information of the place features to be inferred. The place name spatiotemporal derivation relationship network is determined by the place name spatiotemporal derivation relationship identification method and is used to represent the knowledge graph database of place name spatiotemporal derivation relationships. The place name spatiotemporal derivation relationship is determined by the process of naming new place names based on the spatiotemporal relationship of the place features referred to by known place names.
[0141] The spatial fuzzy location information determination module is used to determine spatial fuzzy location information based on the spatiotemporal derivation relationship of place names and a set of spatial similarities. The spatial fuzzy location information includes orientation and distance. The spatial fuzzy location information is used to provide a basis for retrieving geographic information.
[0142] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0143] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for determining spatial fuzzy location based on a place-name spatiotemporal derivation relationship network, characterized in that, include: Obtain information on the features to be inferred; Based on the spatiotemporal derivation relationship network of place names, the spatial similarity set is determined by semantic similarity according to the information of the place features to be inferred. The place name spatiotemporal derivation relationship network is determined by the place name spatiotemporal derivation relationship identification method, and is used to represent the place name spatiotemporal derivation relationship knowledge graph database; the place name spatiotemporal derivation relationship is determined by the process of naming new place names based on the spatiotemporal relationship of the place features referred to by known place names; Using the aforementioned place name spatiotemporal derivation relationship, spatial fuzzy location information is determined based on the aforementioned spatial similarity set; The spatial fuzzy location information includes: orientation and distance; the spatial fuzzy location information is used to provide a basis for retrieving geographic information. Methods for identifying the spatiotemporal derivation relationship of place names include: text feature recognition methods, spatial feature recognition methods, and temporal feature recognition methods. The text feature recognition method employs sequence comparison to calculate text similarity; the expression corresponding to the text similarity is: ; in, For text similarity; The longest common subsequence of the two place names; It is the original place name; It is a derived place name; This is a function for calculating the sequence length. The spatial feature recognition method uses a spatial proximity classification model to identify spatial features and obtain recognition results; the recognition results include spatial proximity relationships and non-spatial proximity relationships. The method for determining the spatial proximity classification model specifically includes: The topological relationships are determined based on the nine-intersection model with dimension extension; Feature filtering is performed based on the topological relationships to obtain filtered features; the filtered features include: spatial distance between land features and the categories of primary land features and derived land features; A dataset is constructed based on the selected features and the corresponding recognition results; The dataset is divided into a training set and a test set; The K-fold cross-validation method is used to train and optimize the data metrics parameters of the classification model based on the training set, resulting in a trained classification model. The data metrics parameters include precision, recall, and F1 score. The classification model uses a CART decision tree. The trained classification model is tested based on the test set, and its generalization performance is evaluated based on the confusion matrix to obtain the tested classification model. The tested classification model will be used as the spatial proximity classification model. The time feature recognition method converts time into timestamps, compares the size of any two timestamps, and determines the time feature recognition result based on the comparison result; the time feature recognition result is used to determine whether the original place name was generated earlier than the derived place name.
2. The spatial fuzzy location determination method based on place name spatiotemporal derivation relationship network according to claim 1, characterized in that, Using the aforementioned place name spatiotemporal derivation relationship, spatial fuzzy location information is determined based on the aforementioned spatial similarity set, specifically including: Based on the directional relationships between features in the spatial similarity set, the location of the feature information to be inferred is determined; Based on the spatial constraint distance of native features, the distance between the feature information to be inferred and the native features is determined.
3. The spatial fuzzy location determination method based on place name spatiotemporal derivation relationship network according to claim 1, characterized in that, The spatiotemporal derivation relationship network of place names is stored in the place name database.
4. The spatial fuzzy location determination method based on place name spatiotemporal derivation relationship network according to claim 1, characterized in that, Also includes: The spatially ambiguous location information is verified to determine its accuracy.
5. The spatial fuzzy location determination method based on place name spatiotemporal derivation relationship network according to claim 4, characterized in that, Accuracy is determined using a success rate function; the mathematical expression of the success rate function is: ; in, The number of correctly positioned items; The number of units whose locations are determined by spatiotemporal derivation relationships; For success rate.
6. A spatial fuzzy location determination device based on a place-name spatiotemporal derivation relationship network, characterized in that, include: The information acquisition module is used to acquire information about the features to be inferred. The spatial similarity set determination module is used to determine the spatial similarity set based on the place name spatiotemporal derivation relationship network and semantic similarity according to the place feature information to be inferred. The place name spatiotemporal derivation relationship network is determined by the place name spatiotemporal derivation relationship identification method, and is used to represent the place name spatiotemporal derivation relationship knowledge graph database; the place name spatiotemporal derivation relationship is determined by the process of naming new place names based on the spatiotemporal relationship of the place features referred to by known place names; The spatial fuzzy location information determination module is used to determine spatial fuzzy location information based on the spatial similarity set by employing the spatiotemporal derivation relationship of the place names. The spatial fuzzy location information includes: orientation and distance; the spatial fuzzy location information is used to provide a basis for retrieving geographic information. Methods for identifying the spatiotemporal derivation relationship of place names include: text feature recognition methods, spatial feature recognition methods, and temporal feature recognition methods. The text feature recognition method employs sequence comparison to calculate text similarity; the expression corresponding to the text similarity is: ; in, For text similarity; The longest common subsequence of the two place names; It is the original place name; It is a derived place name; This is a function for calculating the sequence length. The spatial feature recognition method uses a spatial proximity classification model to identify spatial features and obtain recognition results; the recognition results include spatial proximity relationships and non-spatial proximity relationships. The method for determining the spatial proximity classification model specifically includes: The topological relationships are determined based on the nine-intersection model with dimension extension; Feature filtering is performed based on the topological relationships to obtain filtered features; the filtered features include: spatial distance between land features and the categories of primary land features and derived land features; A dataset is constructed based on the selected features and the corresponding recognition results; The dataset is divided into a training set and a test set; The K-fold cross-validation method is used to train and optimize the data metrics parameters of the classification model based on the training set, resulting in a trained classification model. The data metrics parameters include precision, recall, and F1 score. The classification model uses a CART decision tree. The trained classification model is tested based on the test set, and its generalization performance is evaluated based on the confusion matrix to obtain the tested classification model. The tested classification model will be used as the spatial proximity classification model. The time feature recognition method converts time into timestamps, compares the size of any two timestamps, and determines the time feature recognition result based on the comparison result; the time feature recognition result is used to determine whether the original place name was generated earlier than the derived place name.
Citation Information
Patent Citations
Geographic name space-time derivation relation network construction method and system based on semantic driving
CN119312811A