A method and system for constructing geographic entities and generating relationships

By using the ANN algorithm to filter and comprehensively score to determine standard layer categories, and combining spatial topology and semantic analysis to establish a geographic relationship network, the problem of semantic inconsistency in multi-source heterogeneous data is solved, achieving efficient data consistency and management efficiency improvement.

CN120780790BActive Publication Date: 2025-12-02WUHAN FENGLING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511292228.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-12-02
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

The semantic inconsistency of multi-source heterogeneous geographic information data makes automated extraction and fusion difficult, affecting information sharing and decision-making efficiency. Existing methods are time-consuming and error-prone.

Method used

Layer names are filtered using the ANN algorithm, and a comprehensive score is obtained by combining vector space similarity, text surface similarity, and contextual feature similarity to determine the standard layer category. A geographic relationship network is then established based on spatial topology analysis and semantic association analysis to achieve attribute standardization and entity binding.

Benefits of technology

It improves data consistency and quality, supports complex spatial analysis, enhances data interoperability, and improves the efficiency of geographic information management and decision support capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780790B_ABST
    Figure CN120780790B_ABST
Patent Text Reader

Abstract

This invention relates to a method and system for constructing geographic entities and generating relationships, specifically including: identifying target vector layers from multiple sources, and extracting layer names, attribute field names, and their corresponding attribute values ​​from the target vector layers; performing multi-level filtering and comprehensive scoring based on layer names to determine the standard layer category for optimal alignment of layer names; mapping the attribute field names of the target vector layers to target standard fields using the standard layer category as a reference, and converting the attribute values ​​corresponding to the attribute field names into standard codes to achieve attribute standardization; acquiring the geometric data of the target vector layers, generating target vector entities based on the geometric data, and binding the corresponding standard attributes to the target vector entities; establishing topological relationships, hierarchical relationships, and cross-source entity referencing relationships between entities based on spatial topology analysis and semantic association analysis, thereby forming a structured geographic relationship network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of 3D reconstruction and reality modeling technology, and in particular to a method and system for constructing geographic entities and generating relationships. Background Technology

[0002] In practical geographic information integration, mapping vector layers from multiple departments or sources often suffer from inconsistencies in naming and attribute coding. For example, different units may use different names to represent the same type of feature (e.g., "water body" may be named "river," "ditch," or "water area"). This multi-source heterogeneity leads to semantic inconsistencies, which in turn affects the automated extraction and fusion of geographic entities. Currently, the main solution is to manually match layer names and attributes, which is not only time-consuming but also prone to errors. In emergency situations requiring rapid integration of spatial data from multiple departments, semantic incompatibility between different data sources can severely restrict information sharing and decision-making efficiency. Therefore, there is an urgent need for an intelligent method based on semantic understanding to automatically unify the mapping of heterogeneous layers, thereby reducing manual intervention and enabling automatic entity extraction and relationship construction across data sources. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method and system for constructing geographic entities and generating relationships, addressing the shortcomings of the prior art.

[0004] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0005] Firstly, this application discloses a method for constructing geographic entities and generating relationships, which includes the following steps:

[0006] S1. Determine the target vector layer from multiple sources, and extract the layer name, attribute field name and its corresponding attribute value from the target vector layer;

[0007] S2. Based on the layer name, perform multi-level filtering and comprehensive scoring to determine the standard layer category with the best alignment of the layer name;

[0008] S3. Based on the standard layer category, map the attribute field names of the target vector layer to the target standard fields, and convert the attribute values ​​corresponding to the attribute field names into standard codes to achieve attribute standardization.

[0009] S4. Obtain the geometric data of the target vector layer, generate a target vector entity based on the geometric data, and bind the corresponding standard attributes to the target vector entity;

[0010] S5. Based on spatial topology analysis and semantic association analysis, establish topological relationships, hierarchical relationships, and cross-source entity reference relationships between entities, thereby forming a structured geographic relationship network.

[0011] Furthermore, in step S2, the step of performing multi-level filtering and comprehensive scoring based on the layer name to determine the standard layer category for optimal alignment of the layer name includes:

[0012] S21. Based on the ANN algorithm, retrieve the Top-K candidate set that is most similar to the layer name from the standard category vector library;

[0013] S22. Obtain the similarity score of each candidate in the Top-K candidate set, and filter out candidates with scores lower than a preset threshold to obtain the first candidate set after similarity threshold filtering;

[0014] S23. Obtain the display characteristics of the layer name, and filter out candidates that do not conform to the display characteristics from the first candidate set according to the display characteristics to obtain the second candidate set after display characteristic filtering;

[0015] S24. For each candidate in the second candidate set, a weighted sum is performed based on vector space similarity, text surface similarity, and context feature similarity to obtain a comprehensive score for each candidate.

[0016] S25. The candidate with the highest overall score is selected as the standard layer category for best alignment of the layer name.

[0017] Furthermore, in step S24, the weights of vector space similarity, text surface similarity, and contextual features are dynamically adjusted based on the spatial context that covers the constraint types and scale adaptation information of adjacent layers, the business rule context that covers domain knowledge injection and administrative division features, and the temporal context that covers data version awareness, in order to balance the priority of semantic similarity, text surface similarity, and domain knowledge.

[0018] Furthermore, in step S3, the mapping of the attribute field names of the target vector layer to target standard fields based on the standard layer category, and the conversion of the attribute values ​​corresponding to the attribute field names into standard codes to achieve attribute standardization, includes:

[0019] S31. Based on the standard layer category, determine the pre-configured field correspondence, and map the attribute field names of the target vector layer to the target standard fields based on the field correspondence.

[0020] S32. Based on the standard layer category, determine the corresponding attribute value lookup table, and convert the attribute values ​​corresponding to the attribute field names into standard codes based on the attribute value lookup table to achieve attribute standardization.

[0021] Furthermore, in step S4, when it is determined that a single feature covered in the geometric data has multiple different attribute characteristics, the feature is split into multiple independent target vector entities according to the attribute differences, so as to achieve a one-to-one mapping between attributes and geometric space and avoid analytical ambiguity caused by data aliasing.

[0022] Furthermore, in step S4, when it is determined based on geometric data that there are multiple small features that are spatially continuous and have consistent attributes, the attribute consistency of these features is identified and verified, and the spatial adjacency relationship of these features is verified to confirm that they are spatially continuous and adjacent. Based on the determination that the dual conditions of attributes and space are met, the multiple target vector entities generated accordingly are geometrically merged according to the spatial adjacency relationship to obtain a merged single entity.

[0023] Furthermore, in step S4, based on the generated target vector entity, by combining spatial and attribute features, a redundant element detection algorithm is used, combined with spatial analysis and attribute similarity comparison, to identify spatially overlapping, adjacent, or attribute-similar elements, and redundant elements are filtered out according to preset rules to ensure the simplicity and accuracy of the entity.

[0024] Furthermore, in step S5, if it is determined based on spatial topology analysis technology that the first road entity and the second road entity intersect in space, then a spatial intersection topology relationship is established between the two road entities, and when constructing the geographic relationship network, a road intersection node is generated between the two road entities according to the spatial intersection topology relationship; if it is determined based on spatial topology analysis technology that the first pipeline entity and the second pipeline entity are connected in space, then a spatial connection topology relationship is established between the two pipeline entities, and when constructing the geographic relationship network, the two pipeline entities are connected according to the spatial connection topology relationship; if it is determined based on spatial topology analysis technology that a building entity and its ancillary facility entity are spatially adjacent, then a spatial proximity topology relationship is established between the two building entities, and when constructing the geographic relationship network, the ancillary association between the two building entities is established according to the spatial proximity topology relationship.

[0025] Furthermore, in step S5, if semantic association analysis determines that entity A belongs to a higher-level entity B, then a hierarchical relationship between entity A and entity B is established, and when constructing the geographic relationship network, the hierarchical relationship between the two entities is established according to this hierarchical relationship; if semantic association analysis determines that entities C and D from different sources are the same geographic entity, then a cross-source entity reference relationship is established between entity C and entity D, and when constructing the geographic relationship network, the reference relationship between the two entities is established according to this reference relationship.

[0026] Secondly, this application discloses a geographic entity construction and relationship generation system, the system comprising a layer information extraction module, a layer name standardization module, a layer attribute standardization module, a vector entity generation module, and a geographic relationship network construction module, wherein:

[0027] The layer information extraction module is used to determine target vector layers from multiple sources and extract the layer name, attribute field name and its corresponding attribute value from the target vector layers.

[0028] The layer name standardization module is used to perform multi-level filtering and comprehensive scoring based on the layer name to determine the standard layer category with the best alignment of the layer name.

[0029] The layer attribute standardization module is used to map the attribute field names of the target vector layer to the target standard fields based on the standard layer category, and to convert the attribute values ​​corresponding to the attribute field names into standard codes to achieve attribute standardization.

[0030] The vector entity generation module is used to acquire the geometric data of the target vector layer, generate a target vector entity based on the geometric data, and bind the corresponding standard attributes to the target vector entity;

[0031] The geographic relationship network construction module is used to establish topological relationships, hierarchical relationships, and cross-source entity reference relationships between entities based on spatial topology analysis and semantic association analysis, thereby forming a structured geographic relationship network.

[0032] The beneficial technical effects of this invention are as follows: it ensures the accuracy of layer names through multi-layer screening, achieves data consistency through attribute standardization, improves data quality by combining geometric data and standard attributes, and finally constructs a structured geographic relationship network, thereby supporting complex spatial analysis, enhancing data interoperability, and significantly improving the efficiency and decision support capabilities of geographic information management. Attached Figure Description

[0033] Figure 1 A flowchart illustrating a method for constructing geographic entities and generating relationships provided by this invention.

[0034] Figure 2 This is a schematic diagram of the structure of a geographic entity construction and relationship generation system provided by the present invention. Detailed Implementation

[0035] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0036] like Figure 1 As shown, this application provides a method for constructing geographic entities and generating relationships, which includes the following steps:

[0037] S1, determine target vector layers from multiple sources, and extract the layer name, attribute field name and its corresponding attribute value from the target vector layers.

[0038] Specifically, this application reads in initial vector layers from multiple sources and performs preprocessing such as format conversion (e.g., Shapefile to GeoJSON, DWG to DXF), coordinate system transformation, topology checking, and data cleaning to obtain the target vector layer. Then, based on preset extraction criteria, it extracts the corresponding layer name, attribute field name, and their corresponding attribute values ​​from the target vector layer, using these as input for subsequent layer semantic matching, attribute mapping, and standardization processes.

[0039] S2, based on the layer name, perform multi-level filtering and comprehensive scoring to determine the standard layer category with the best alignment of the layer name.

[0040] Specifically, this application uses the extracted layer names and preset filtering rules and scoring models to perform multi-level filtering and comprehensive scoring on the target vector layer. Specifically, the multi-level filtering includes a first-level filtering based on a similarity threshold to remove candidates with too low similarity; and a second-level filtering based on display features to remove candidates that do not conform to the display features. The comprehensive scoring considers a weighted sum based on vector space similarity, text surface similarity, and contextual feature weights. Finally, the candidate with the highest comprehensive score is determined as the standard layer category that best aligns with the target vector layer.

[0041] S3 uses the standard layer category as a reference to map the attribute field names of the target vector layer to the target standard fields, and converts the attribute values ​​corresponding to the attribute field names into standard codes to achieve attribute standardization.

[0042] Specifically, attribute field name mapping is a crucial step in data standardization. Since attribute field names may differ from data from different sources, mapping them to a target standard field can unify the field names. For example, different attribute field names such as "road name," "street name," and "highway name" in the source data can be uniformly mapped to the standard field name "road name."

[0043] Specifically, converting attribute values ​​into standard codes helps achieve data standardization and efficient processing. Standard codes typically use concise forms such as numbers and letters to represent specific meanings, reducing data storage space and improving data processing efficiency. For example, the textual descriptions of "road grade" attributes such as "Level 1" and "Level 2" can be converted into standard codes "01" and "02".

[0044] S4, acquire the geometric data of the target vector layer, generate a target vector entity based on the geometric data, and bind the corresponding standard attributes to the target vector entity.

[0045] Specifically, this application acquires the geometric data of the target vector layer and copies or transforms the corresponding geometric shape to the target coordinate system (such as the WGS84 geographic coordinate system) based on the geometric data. Then, it creates the entity object corresponding to the geometric shape (i.e., generates the target vector entity) in a specified spatial data storage environment (such as an ArcGIS geodatabase, a PostGIS spatial database, or an industry-specific GIS platform) and binds the associated standard attributes to the target vector entity. For example, it generates a "road entity" in an ArcGIS geodatabase based on geometric data matching the standard layer category of "road," where the assigned standard attributes include a unified road grade, name, and other encodings.

[0046] S5 establishes topological relationships, hierarchical relationships, and cross-source entity reference relationships based on spatial topology analysis and semantic association analysis, thereby forming a structured geographic relationship network.

[0047] Specifically, this application utilizes spatial topology analysis techniques, such as buffer analysis, overlay analysis, and network analysis, to accurately identify spatial relationships between geographic entities, including adjacency, intersection, overlap, and containment. Through semantic association analysis, and employing techniques such as natural language processing, it delves into the semantic relationships between entities, such as functional association, category affiliation, and logical dependency. Finally, based on the topological relationships, hierarchical relationships, and cross-source entity referencing relationships between entities, a structured geographic relationship network is formed using data structures such as graph databases or network models.

[0048] As can be seen from the above, the geographic entity construction and relationship generation method disclosed in this application ensures the accuracy of layer names through multi-level screening, achieves data consistency through attribute standardization, improves data quality by combining geometric data and standard attributes, and finally constructs a structured geographic relationship network, thereby supporting complex spatial analysis, enhancing data interoperability, and significantly improving the efficiency and decision support capabilities of geographic information management.

[0049] In one embodiment, step S2, which involves performing multi-level filtering and comprehensive scoring based on the layer name to determine the standard layer category for optimal alignment of the layer name, includes:

[0050] S21. Based on the ANN algorithm, retrieve the Top-K candidate set that is most similar to the layer name from the standard category vector library.

[0051] Specifically, the ANN algorithm is an approximate nearest neighbor search algorithm. During the retrieval process, it calculates the similarity score between the layer name and each pre-encoded dense vector in the standard category vector library, and selects the K results with the highest scores as the Top-K candidate set after sorting them in descending order of similarity score.

[0052] In one embodiment, considering the need to balance system performance and user experience, this application sets K to 10 to ensure retrieval efficiency while providing enough high-quality candidates for downstream tasks to process, avoiding waste of computing resources or difficulty in user selection due to an excessively large result set.

[0053] S22, obtain the similarity score of each candidate in the Top-K candidate set, and filter out candidates with scores lower than a preset threshold to obtain the first candidate set after similarity threshold filtering.

[0054] Specifically, since the ANN algorithm calculates the similarity score of each candidate in the Top-K candidate set after execution, in order to improve the quality of the candidate set and balance retrieval efficiency and result accuracy, this application sets a preset threshold of 0.5 and compares the similarity score of each candidate with the preset threshold item by item. During the comparison process, if it is determined that the similarity score of the corresponding candidate is lower than the preset threshold, the candidate is deleted from the Top-K candidate set to remove candidates with excessively low similarity.

[0055] S23, obtain the display characteristics of the layer name, and filter out candidates that do not conform to the display characteristics from the first candidate set according to the display characteristics, to obtain the second candidate set after display characteristic filtering.

[0056] Specifically, the display features include generic features, such as "road", "river", "boundary" and other generic names. This application uses the display features to match predefined generic name dictionaries and combine them with string pattern recognition to filter out candidates that do not conform to the display features from the first candidate set, thereby narrowing the search range of standard layer categories.

[0057] S24. For each candidate in the second candidate set, a weighted sum is performed based on vector space similarity, text surface similarity, and context feature similarity to obtain a comprehensive score for each candidate.

[0058] Specifically, this application represents the target text and candidate options as vectors, and measures the similarity between the two vectors based on cosine similarity (although other methods can be chosen to measure vector similarity in different embodiments, this application does not limit this comparison). For text surface similarity, this application specifically measures it through edit distance, where edit distance refers to the minimum number of single-character editing operations (insertion, deletion, or replacement) required to convert one string into another. Contextual features can be the current map scale, adjacent layer types, etc., and contextual feature similarity can be measured through feature matching degree, which is not limited in this application. Furthermore, if the weight coefficients of vector space similarity, text surface similarity, and contextual feature similarity are set to w1, w2, and w3 respectively, and w1+w2+w3=1, then the formula for calculating the comprehensive score of each candidate option is as follows:

[0059] Score = w1·sim_{cosine} + w2·sim_{edit} + w3·C_{context};

[0060] Where sim_{cosine} represents vector space similarity, sim_{edit} represents text surface similarity, and C_{context} represents context feature similarity.

[0061] S25, the candidate with the highest overall score is selected as the standard layer category for best alignment of the layer name.

[0062] Specifically, after obtaining the overall score of all candidates, this application will sort all candidates from highest to lowest (or lowest to highest) according to the overall score. Then, based on the sorting results, the candidate with the highest overall score will be used as the standard layer category for best alignment of the layer name.

[0063] In one embodiment, if multiple candidate options have the same overall score and are all the highest, then the specific values ​​of these candidate options in vector space similarity, text surface similarity, and contextual feature similarity are further compared, and the candidate option with the better performance is selected as the standard layer category for optimal alignment according to a preset priority order. Specifically, this application sets a corresponding priority order among vector space similarity, text surface similarity, and contextual feature similarity based on actual needs or domain knowledge. For example, since vector space similarity can more directly reflect the degree of matching of layers in geometric space, which is crucial for layer alignment tasks with high requirements for spatial positional relationships, vector space similarity is set to have the highest priority, followed by contextual feature similarity, and finally text surface similarity. During the comparison, the values ​​of vector space similarity are compared first, and the candidate option with the highest value is selected; if the values ​​of vector space similarity are the same, the values ​​of contextual feature similarity are further compared; if the values ​​of contextual feature similarity are also the same, the values ​​of text surface similarity are finally compared.

[0064] In the above embodiments, the series of steps from retrieving similar candidates based on the ANN algorithm to comprehensive multi-dimensional filtering, weighted summation of similarity, and determination of the best alignment category can significantly improve the matching accuracy. The accurate retrieval and multi-dimensional filtering ensure the accuracy of candidates, while the efficient retrieval and filtering by the ANN algorithm improves the matching efficiency while reducing invalid calculations.

[0065] In one embodiment, in step S24, the weights of vector space similarity, text surface similarity, and contextual features are dynamically adjusted based on the spatial context covering adjacent layer constraint types and scale adaptation information, the business rule context covering domain knowledge injection and administrative division features, and the temporal context covering data version awareness, in order to balance the priority of semantic similarity, text surface similarity, and domain knowledge.

[0066] In one embodiment, S3, the step of mapping the attribute field names of the target vector layer to target standard fields based on the standard layer category, and converting the attribute values ​​corresponding to the attribute field names into standard codes to achieve attribute standardization, includes:

[0067] S31. Based on the standard layer category, determine the pre-configured field correspondence, and map the attribute field names of the target vector layer to the target standard fields based on the field correspondence.

[0068] Specifically, this application first determines the pre-configured field correspondence based on the standard layer category (this step is performed on the premise that a clear mapping relationship (such as a mapping table, configuration file, or database record) has been established in advance between the standard layer category and the pre-configured field correspondence, which clarifies the correspondence between the standard layer category and the target standard field). Subsequently, based on the preset mapping rules, the attribute field names of the target vector layer are mapped to the target standard fields.

[0069] S32. Based on the standard layer category, determine the corresponding attribute value lookup table, and convert the attribute values ​​corresponding to the attribute field names into standard codes based on the attribute value lookup table to achieve attribute standardization.

[0070] Specifically, this application clarifies the standard code mapping rules by constructing an attribute value lookup table and establishes an association mechanism between standard layer categories and the attribute value lookup table. Through this association mechanism, the automatic matching of attribute values ​​and standard codes is ultimately achieved.

[0071] In the above embodiments, by mapping fields and standardizing attribute values, data quality and consistency can be significantly improved, data processing can be accelerated, cross-system data sharing and reuse can be promoted, data management costs can be reduced, and data quality monitoring capabilities can be strengthened, laying the foundation for the standardized management and efficient utilization of geographic information data.

[0072] In one embodiment, in step S4, when it is determined that a single feature covered in the geometric data has multiple different attribute features, the feature is split into multiple independent target vector entities according to the attribute differences, so as to achieve a one-to-one mapping between attributes and geometric space and avoid analytical ambiguity caused by data aliasing.

[0073] In one embodiment, in step S4, when it is determined based on geometric data that there are multiple small features that are spatially continuous and have consistent attributes, the consistency of the attributes of these features is identified and verified, and the spatial adjacency relationship of these features is verified to confirm that they are spatially continuous and adjacent. Based on the determination that the dual conditions of attributes and space are met, the multiple target vector entities generated accordingly are geometrically merged according to the spatial adjacency relationship to obtain a merged single entity.

[0074] Specifically, this application verifies the consistency of attributes of multiple elements through an attribute field comparison algorithm to ensure that the matching degree of key attribute fields exceeds a preset threshold; secondly, it uses spatial topology analysis technology to verify the spatial adjacency relationship of elements and confirm that the geometric connection or small gap of the element boundary meets the spatial continuity requirements; finally, it uses a geometric merging algorithm to merge elements that meet the conditions into a single entity.

[0075] In one embodiment, in step S4, based on the generated target vector entity, by combining spatial and attribute features, a redundant element detection algorithm is used, combined with spatial analysis and attribute similarity comparison, to identify spatially overlapping, adjacent, or attribute-similar elements, and redundant elements are filtered out according to preset rules to ensure the simplicity and accuracy of the entity.

[0076] In one embodiment, in step S5, if it is determined that the first road entity and the second road entity intersect in space based on spatial topology analysis technology, a spatial intersection topology relationship is established between the two road entities, and a road intersection node is generated between the two road entities according to the spatial intersection topology relationship when constructing the geographic relationship network; if it is determined that the first pipeline entity and the second pipeline entity are connected in space based on spatial topology analysis technology, a spatial connection topology relationship is established between the two pipeline entities, and the two pipeline entities are connected according to the spatial connection topology relationship when constructing the geographic relationship network; if it is determined that a building entity and its ancillary facility entity are spatially adjacent based on spatial topology analysis technology, a spatial proximity topology relationship is established between the two building entities, and the ancillary association between the two building entities is established according to the spatial proximity topology relationship when constructing the geographic relationship network.

[0077] Specifically, for detecting spatially intersecting topological relationships, a hierarchical bounding box method is used (to accelerate the geometric intersection detection process; when generating road intersection nodes, a straight line is formed between two points, and a certain width is extended to both sides to generate the road surface; when multiple roads intersect, the intersection point is automatically calculated). Secondly, for verifying spatially connected topological relationships, an endpoint matching algorithm is used to verify whether the endpoints of pipeline entities are connected, and a tolerance threshold is set to allow for the existence of small gaps. Simultaneously, DFS or BFS algorithms are used to verify the connectivity of the pipeline network, ensuring that the connected pipeline entities remain connected in graph theory. Thirdly, when quantifying spatially adjacent topological relationships, Euclidean distance is used to calculate the spatial proximity between buildings and their ancillary facilities, while ensuring that the ancillary facilities do not overlap with the main building, and this is verified through spatial intersection detection. Finally, after establishing the ancillary associations, a spatial constraint check is performed to ensure that the ancillary facilities are within the orthographic projection range of the building, to meet the needs of practical applications.

[0078] In one embodiment, in step S5, if it is determined based on semantic association analysis that entity A belongs to a higher-level entity B, then a hierarchical relationship between entity A and entity B is established, and when constructing the geographic relationship network, the hierarchical relationship between the two entities is established according to this hierarchical relationship; if it is determined based on semantic association analysis that entity C and entity D from different sources are the same geographic entity, then a cross-source entity reference relationship is established between entity C and entity D, and when constructing the geographic relationship network, the reference relationship between the two entities is established according to this reference relationship.

[0079] Specifically, when establishing hierarchical relationships, entities are mapped to corresponding ontology concepts to support the establishment of multi-level hierarchical relationships (such as building → community → administrative region) and to verify the transitivity of hierarchical relationships, ensuring logical consistency. At the same time, a semantic reasoning engine is used to verify the hierarchical relationships. Secondly, when establishing reference relationships between entities from different sources, a similarity threshold is set (such as the similarity of classification fields must reach 100%, and the difference of numerical fields must be less than 5%) to accurately identify the same geographic entity and record differences in data sources (such as data version, collection time, and accuracy information) for subsequent data traceability and verification.

[0080] Please refer to Figure 2 This application discloses a geographic entity construction and relationship generation system, which includes a layer information extraction module, a layer name standardization module, a layer attribute standardization module, a vector entity generation module, and a geographic relationship network construction module, wherein:

[0081] The layer information extraction module is used to identify target vector layers from multiple sources and extract the layer name, attribute field name and its corresponding attribute value from the target vector layers.

[0082] The layer name standardization module is used to perform multi-level filtering and comprehensive scoring based on the layer name to determine the standard layer category with the best alignment of the layer name.

[0083] The layer attribute standardization module is used to map the attribute field names of the target vector layer to the target standard fields based on the standard layer category, and to convert the attribute values ​​corresponding to the attribute field names into standard codes, so as to achieve attribute standardization.

[0084] The vector entity generation module is used to acquire the geometric data of the target vector layer, generate a target vector entity based on the geometric data, and bind the corresponding standard attributes to the target vector entity.

[0085] The geographic relationship network construction module is used to establish topological relationships, hierarchical relationships, and cross-source entity reference relationships between entities based on spatial topology analysis and semantic association analysis, thereby forming a structured geographic relationship network.

[0086] In one embodiment, the above modules are also used to implement the geographic entity construction and relationship generation method as described in any of the foregoing method embodiments, and this application does not limit this.

[0087] As can be seen from the above, the geographic entity construction and relationship generation system disclosed in this application ensures the accuracy of layer names through multi-level screening, achieves data consistency through attribute standardization, improves data quality by combining geometric data and standard attributes, and finally constructs a structured geographic relationship network, thereby supporting complex spatial analysis, enhancing data interoperability, and significantly improving the efficiency and decision support capabilities of geographic information management.

[0088] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for constructing geographic entities and generating relationships, characterized in that, The method includes the following steps: S1. Determine the target vector layer from multiple sources, and extract the layer name, attribute field name and its corresponding attribute value from the target vector layer; S2. Based on the layer name, perform multi-level filtering and comprehensive scoring to determine the standard layer category with the best alignment of the layer name; S3. Based on the standard layer category, map the attribute field names of the target vector layer to the target standard fields, and convert the attribute values ​​corresponding to the attribute field names into standard codes to achieve attribute standardization. S4. Obtain the geometric data of the target vector layer, generate a target vector entity based on the geometric data, and bind the corresponding standard attributes to the target vector entity; S5. Based on spatial topology analysis, establish topological relationships between entities, analyze hierarchical relationships and cross-source entity reference relationships based on semantic association, and form a structured geographic relationship network. In step S2, the step of performing multi-level filtering and comprehensive scoring based on the layer name to determine the standard layer category with the best alignment of the layer name includes: S21. Based on the ANN algorithm, retrieve the Top-K candidate set that is most similar to the layer name from the standard category vector library; S22. Obtain the similarity score of each candidate in the Top-K candidate set, and filter out candidates with scores lower than a preset threshold to obtain the first candidate set after similarity threshold filtering; S23. Obtain the display characteristics of the layer name, and filter out candidates that do not conform to the display characteristics from the first candidate set according to the display characteristics to obtain the second candidate set after display characteristic filtering; S24. For each candidate in the second candidate set, a weighted sum is performed based on vector space similarity, text surface similarity, and contextual feature similarity to obtain a comprehensive score for each candidate; wherein, contextual features include the current map scale and adjacent layer types; S25. The candidate with the highest overall score is selected as the standard layer category for best alignment of the layer name.

2. The method according to claim 1, characterized in that, In step S24, the weights of vector space similarity, text surface similarity, and contextual features are dynamically adjusted based on the spatial context that covers the constraint types and scale adaptation information of adjacent layers, the business rule context that covers domain knowledge injection and administrative division features, and the temporal context that covers data version awareness, in order to balance the priority of semantic similarity, text surface similarity, and domain knowledge.

3. The method according to claim 1, characterized in that, In step S3, the process of mapping the attribute field names of the target vector layer to target standard fields based on the standard layer category, and converting the attribute values ​​corresponding to the attribute field names into standard codes to achieve attribute standardization, includes: S31. Based on the standard layer category, determine the pre-configured field correspondence, and map the attribute field names of the target vector layer to the target standard fields based on the field correspondence. S32. Based on the standard layer category, determine the corresponding attribute value lookup table, and convert the attribute values ​​corresponding to the attribute field names into standard codes based on the attribute value lookup table to achieve attribute standardization.

4. The method according to claim 1, characterized in that, In step S4, when it is determined that a single feature covered in the geometric data has multiple different attribute characteristics, the feature is split into multiple independent target vector entities according to the attribute differences, so as to achieve a one-to-one mapping between attributes and geometric space and avoid analytical ambiguity caused by data aliasing.

5. The method according to claim 1, characterized in that, In step S4, when it is determined from geometric data that there are multiple small features that are spatially continuous and have consistent attributes, the attribute consistency of these features is identified and verified, and the spatial adjacency relationship of these features is verified to confirm that they are spatially continuous and adjacent. Based on the determination that the dual conditions of attributes and space are met, the multiple target vector entities generated accordingly are geometrically merged according to the spatial adjacency relationship to obtain a merged single entity.

6. The method according to claim 1, characterized in that, In step S4, based on the generated target vector entity, by combining spatial and attribute features, a redundant element detection algorithm is used, combined with spatial analysis and attribute similarity comparison, to identify spatially overlapping, adjacent, or attribute-similar elements, and redundant elements are filtered out according to preset rules to ensure the conciseness and accuracy of the entity.

7. The method according to claim 1, characterized in that, In step S5, if it is determined that the first road entity and the second road entity intersect in space based on spatial topology analysis technology, then a spatial intersection topology relationship between the two road entities is established, and when constructing the geographic relationship network, a road intersection node is generated between the two road entities according to the spatial intersection topology relationship. If the first pipeline entity and the second pipeline entity are determined to be spatially connected based on spatial topology analysis technology, then a spatial connection topology relationship between the two pipeline entities is established, and the two pipeline entities are connected according to this spatial connection topology relationship when constructing the geographic relationship network. If spatial topology analysis technology determines that a building entity and its ancillary facilities are spatially adjacent, then a spatial proximity topology relationship is established between the two building entities, and when constructing a geographic relationship network, the ancillary associations between the two building entities are established according to this spatial proximity topology relationship.

8. The method according to claim 1, characterized in that, In step S5, if it is determined based on semantic association analysis that entity A belongs to entity B at a higher level, then a hierarchical relationship between entity A and entity B is established, and when constructing the geographic relationship network, the hierarchical relationship between the two entities is established according to this hierarchical relationship. If semantic association analysis determines that entities C and D from different sources are the same geographic entity, then a cross-source entity reference relationship is established between entities C and D, and when constructing the geographic relationship network, the reference association between the two entities is established according to this reference relationship.

9. A geographic entity construction and relationship generation system, characterized in that, The system includes a layer information extraction module, a layer name standardization module, a layer attribute standardization module, a vector entity generation module, and a geographic relationship network construction module, wherein: The layer information extraction module is used to determine target vector layers from multiple sources and extract the layer name, attribute field name and its corresponding attribute value from the target vector layers. The layer name standardization module is used to perform multi-level filtering and comprehensive scoring based on the layer name to determine the standard layer category with the best alignment of the layer name. The layer attribute standardization module is used to map the attribute field names of the target vector layer to the target standard fields based on the standard layer category, and to convert the attribute values ​​corresponding to the attribute field names into standard codes to achieve attribute standardization. The vector entity generation module is used to acquire the geometric data of the target vector layer, generate a target vector entity based on the geometric data, and bind the corresponding standard attributes to the target vector entity; The geographic relationship network construction module is used to establish topological relationships, hierarchical relationships, and cross-source entity reference relationships between entities based on spatial topology analysis and semantic association analysis, thereby forming a structured geographic relationship network. The specific implementation of the layer name standardization module, which performs multi-level filtering and comprehensive scoring based on the layer name to determine the standard layer category with the best alignment of the layer name, is as follows: Based on the ANN algorithm, a Top-K candidate set most similar to the layer name is retrieved from a standard category vector library; Obtain the similarity score of each candidate in the Top-K candidate set, and filter out candidates with scores lower than a preset threshold to obtain the first candidate set after similarity threshold filtering; Obtain the display characteristics of the layer name, and filter out candidates that do not conform to the display characteristics from the first candidate set according to the display characteristics to obtain the second candidate set after display characteristic filtering; For each candidate in the second candidate set, a weighted sum is performed based on vector space similarity, text surface similarity, and contextual feature similarity to obtain a comprehensive score for each candidate; wherein, contextual features include the current map scale and adjacent layer types; The candidate with the highest overall score will be used as the standard layer category for best alignment of the layer name.

Citation Information

Patent Citations

  • Geographic entity generation method and system based on multi-source heterogeneous data

    CN116719898A

  • Method and system for constructing geographic entity based on DLG topographic map

    CN119917598A