A multi-source geographic vector data matching method based on topological correlation

CN121807827BActive Publication Date: 2026-09-08CHINA INST OF WATER RESOURCES & HYDROPOWER RES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511891889.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-09-08
Estimated Expiration
2045-12-16

AI Technical Summary

Technical Problem

特别是在灾害应急实战场景中,决策部门需要整合规划部门的城市道路网、国土部门的地籍图、应急部门的实时调查数据以及无人机航测的灾情矢量图,这些数据由于采集时间不同、坐标基准不一、表达精度各异,导致同一建筑物、道路在不同数据源中的位置、形状存在显著偏差,严重影响灾情研判准确性和时效性

Benefits of technology

[0062] A quality scoring mechanism based on local topological features is proposed to identify low-quality data regions and guide subsequent matching strategies. By pre-evaluating topological quality, weak links in the data are identified before matching, preventing the propagation of topological errors from low-quality data, thus overcoming the blind matching methods of existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121807827B_ABST
    Figure CN121807827B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-source geographic vector data matching methods based on topological correlation, including collecting multi-source basic geographic vector data, forming pre-processing dataset;Multi-index data quality evaluation based on local topological features is constructed;The extraction and classification of spatial topological relation edge are carried out;The target graph weighted adjacency matrix is constructed;The neighborhood topological consistency measure index is calculated, and the neighborhood topological consistency quantization and adaptive weight fusion are fused;With confidence matrix as input, establish global optimization objective function, with topological retention rate as hard constraint condition, solve the maximum weight bipartite graph matching problem with topological constraint, output global optimal matching scheme;Topological deviation measure based on topological invariant difference is constructed, determine whether there is significant topological conflict in matching scheme, repair according to conflict situation, hierarchical.This application method improves the reliability and logical self-consistency of multi-source vector data fusion, solves the topological contradiction key problem of prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical field of this invention is intelligent analysis of geographic data, specifically involving a multi-source geographic vector data matching method based on topological correlation. Background Technology

[0002] With the rapid development of geographic information system (GIS) technology and the diversification of multi-source data acquisition methods, the same geographic area often contains multi-source geographic vector data from different departments, different periods, and different scales. How to achieve the fusion and updating of heterogeneous data has become a key problem that urgently needs to be solved in the field of geographic information processing. Especially in disaster emergency response scenarios, decision-making departments need to integrate urban road networks from planning departments, cadastral maps from land departments, real-time survey data from emergency departments, and disaster vector maps from UAV aerial surveys. Due to differences in acquisition time, coordinate benchmarks, and expression precision, the location and shape of the same building or road in different data sources vary significantly, seriously affecting the accuracy and timeliness of disaster assessment.

[0003] Traditional vector data matching methods mainly rely on geometric feature similarity calculation, such as feature registration based on indicators such as distance, shape, and area. However, these methods have shortcomings: First, pure geometric matching ignores the inherent spatial topological relationships between geographic features. The graph model uses a uniform 0-1 adjacency matrix to treat all topological relationships equally, which makes it impossible to reflect the differences in importance of different spatial relationships during the optimization process. Factors such as road connectivity, adjacency relationships of land parcels, and river confluence structures can lead to matching results that, while similar in local geometry, exhibit contradictions in the overall topology. For instance, previously connected rescue roads may become broken after matching, or adjacent disaster-stricken land parcels may overlap after matching, directly impacting the reliability of emergency decision-making. Secondly, existing methods often employ greedy strategies for pairwise matching, resulting in overly strong topological constraints on isolated elements and overly loose geometric constraints on central nodes due to fixed weights. This lack of a global constraint mechanism makes them prone to getting trapped in local optima and failing to guarantee topological consistency in the matching results. Thirdly, when multi-source data exhibit significant differences in geometric precision, element density, and representation granularity, matching methods based on geometric thresholds have poor robustness, easily leading to mismatches or missed matches, and failing to meet the timeliness requirements in emergency scenarios. Summary of the Invention

[0004] To address the aforementioned problems, this invention proposes a multi-source geographic vector data matching method based on topological correlation. This method considers the topological relationships inherent in the structural correlation of geospatial vector data fusion, achieving global consistency of the topological structure during the matching process. It avoids the topological fragmentation problem caused by pure geometric matching, improves the reliability and logical self-consistency of multi-source vector data fusion, and solves the key problem of topological contradictions in existing technologies.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] A multi-source geographic vector data matching method based on topological correlation, the method comprising:

[0007] S1: Collect basic geographic vector data from multiple sources and perform preprocessing, including topological cleaning and normalization, to form a preprocessed dataset.

[0008] S2, perform multi-indicator data quality assessment on the preprocessed dataset based on local topological features to identify low-quality data areas, mark data below a predetermined threshold, and establish a tree space index for datasets that meet the quality standards.

[0009] S3 maps the point, line, and polygon geographic features in the preprocessed dataset and the target graph dataset to graph nodes respectively. Based on the topological relationship of the preprocessed data, it constructs a mixed edge set containing directed and undirected edges to extract and classify spatial topological relationship edges.

[0010] S4. Assign differentiated weights based on the type of topological relationship to quantify the degree of constraint of different spatial relationships on the matching process and construct a weighted adjacency matrix of the target graph.

[0011] S5. For the elements in the source graph dataset and the candidate elements in the target graph dataset, construct a comprehensive similarity model with geometric-topological co-constraints, quantify the dynamic neighborhood constraints in the matching process into a computable metric, calculate the neighborhood topological consistency metric, and integrate adaptive weights.

[0012] S6. A global optimization objective function is established with the confidence matrix as input. The topology preservation rate is used as a hard constraint. An improved Hungarian algorithm is used to solve the maximum weight bipartite graph matching problem with topological constraints and output the globally optimal matching scheme.

[0013] S7. Construct a quantitative evaluation index based on the difference of topological invariants, topological deviation measurement, compare the topological deviation measurement value with the set topological deviation threshold, determine whether there is a significant topological conflict in the matching scheme, and execute the corresponding hierarchical repair strategy according to the conflict situation.

[0014] Furthermore, the multi-indicator data quality assessment based on local topological features described in S2 is:

[0015] For each feature in the preprocessed dataset Calculate their topology quality scores respectively. :

[0016]

[0017] in, Indicates the first in the preprocessed dataset A geographic element can be a point element, a line element, or a polygon element; For connectivity quality indicators of linear elements such as roads and pipelines, This is a geometric normalization index for all geographic vector features. The weighting is set as follows for the feature density consistency index based on neighborhood statistics: , , ,

[0018] when When an element is identified as low-quality, its confidence weight will be reduced during subsequent matching, or it will be subject to priority manual review. Elements with a value ≥0.6 are defined as qualified data. Elements with a value of ≥0.8 are defined as high-quality data.

[0019] Furthermore, the construction of the weighted adjacency matrix of the target graph described in S4 is as follows:

[0020] For the source image , build Weighted adjacency matrix Its element definition is as follows:

[0021] The elements of the classification-weighted adjacency matrix are calculated as follows:

[0022]

[0023] In the formula, and For node indexing, , This represents the total number of nodes in the source graph. For nodes in the source graph, the weight parameter is set as follows: weight of line feature connectivity. As the strongest constraint, the adjacency weight of surface elements. This is a secondary constraint, with weights for point-to-surface inclusion relationships. This is a weak constraint. This represents the set of edges connecting all line features in the source graph. This represents the set of adjacency edges for all face features in the source graph. This represents the set of all points and faces in the source graph that contain relational edges.

[0024] Furthermore, the comprehensive similarity model for constructing geometric-topological co-constraints described in S5 includes:

[0025] First, for the nodes in the source graph dataset Candidate nodes in the target graph dataset Calculate the geometric similarity of three indicators: location distance metric, shape similarity metric, and scale ratio;

[0026] Then, the three indicators mentioned above are weighted and fused together to obtain the comprehensive geometric similarity:

[0027]

[0028] In the formula, For location distance measurement, For shape similarity measurement, The scale ratio is used; the weight parameters are set as follows, depending on the feature type: Point features: focus on location, Line elements: Emphasizing shape. Surface elements: balancing shape and scale. .

[0029] Furthermore, the calculation of the neighborhood topological consistency metric described in S5 specifically involves:

[0030] (1) Define the neighborhood set for nodes in the source graph. its neighborhood set ( ) is defined in the topological association graph In and source graph nodes There exists a set of all nodes connected by an edge:

[0031]

[0032] In the formula, For the source graph and the source graph nodes Neighboring nodes that have topological relationships and For target graph nodes The neighborhood set is ( ),

[0033]

[0034] in, For nodes in the target graph Neighboring nodes that have topological relationships and ;

[0035] (2) Matching mapping function: During the matching process, the determined matching mapping function is maintained. Map the source graph nodes to the target graph nodes. For nodes that have not yet been matched, Return the empty set;

[0036] (3) Calculate topological similarity

[0037] The neighborhood topology consistency is defined as the ratio of the intersection of the neighborhoods of nodes in the source graph to the union of the neighborhoods of nodes in the target graph under the matching mapping:

[0038]

[0039] in: Source graph node With target graph nodes The neighborhood topological consistency metric between them, with a range of values. ;molecular The denominator represents the number of nodes in the source graph's neighborhood that have successfully matched the target graph's neighborhood, reflecting the degree of topology preservation; This represents the size of the union of two neighborhoods, used as a normalization factor; This means that the neighborhood of the target graph node is back mapped back to the source graph node space through an existing matching mapping.

[0040] Furthermore, the calculation of the neighborhood topological consistency metric described in S5 also includes:

[0041] Correcting edge weighting by introducing a weighted adjacency matrix The differentiated weights are calculated for neighboring nodes, and the corrected topological similarity is:

[0042] .

[0043] Furthermore, the calculation of the neighborhood topological consistency metric and the integration of adaptive weights, as described in S5, involves dynamically adjusting the weight ratio of geometry and topology based on the centrality of nodes in the topological graph, and then calculating the overall similarity.

[0044]

[0045] in, For adaptive weighting coefficients, the range of values ​​is... ;

[0046] Based on node degree centrality (number of neighboring nodes), the adaptive weight is calculated as follows:

[0047]

[0048] in, For nodes The degree (number of neighboring nodes). The threshold for the number of neighbors, To adjust the parameters.

[0049] Furthermore, the S6 objective function optimization method includes the following steps:

[0050] First, define the decision variables, define... 3D decision matrix Its elements :

[0051]

[0052] when Time indicates that the source graph node Matched target graph nodes ,when The time indicates a mismatch;

[0053] Secondly, a maximum weight bipartite graph matching model is established, with the objective function being:

[0054]

[0055] in Represents source graph nodes With target graph nodes The confidence level of the match between them This represents the decision variable, taking values ​​of 0 or 1. , Indicates the number of features in the target graph dataset and the source graph dataset;

[0056] Finally, a hard constraint is set for the topology preservation rate: the proportion of topological edges in the source graph that are preserved in the target graph after matching is no less than a threshold. For each edge in the source graph ,like and Matched respectively and This requires that the target graph also contains corresponding edges. The optimized objective function formula is:

[0057]

[0058] Among them, indicator function The value is 1 if a corresponding edge exists in the target graph, and 0 otherwise. This represents the total number of edges in the source graph; The topology retention threshold is set to 0.85.

[0059] Furthermore, the topology deviation threshold mentioned in S7 When topological deviation is measured If a significant topological conflict is detected in the matching scheme, a subsequent hierarchical repair process will be triggered.

[0060] Furthermore, the hierarchical repair described in S7 is as follows: Level 1 repair targets connectivity disruptions, traversing low-confidence matching pairs ( First, remove matches that cause connectivity disruption and replace them with suboptimal solutions. Second, for adjacency errors, search for alternative features that satisfy adjacency constraints within a 50-meter buffer. Third, for systematic errors, reduce the topology preservation rate threshold to 0.75 and re-execute global optimization.

[0061] The beneficial effects of this invention are:

[0062] A quality scoring mechanism based on local topological features is proposed to identify low-quality data regions and guide subsequent matching strategies. By pre-evaluating topological quality, weak links in the data are identified before matching, preventing the propagation of topological errors from low-quality data, thus overcoming the blind matching methods of existing technologies.

[0063] To address the problem that existing vector matching methods use a uniform 0-1 adjacency matrix for their graph models, treating all topological relationships equally and failing to reflect the differences in importance of different spatial relationships during optimization, this invention introduces a differentiated weighting mechanism for the first time. It sets hierarchical weights based on the degree of dependence of emergency decisions on connectivity and adjacency, enabling the matching algorithm to intelligently select and protect high-priority relationships when topological conflicts occur, significantly improving the reliability and logical rationality of the matching results in practical applications.

[0064] For the first time, the dynamic neighborhood constraints in the matching process are quantified into a computable metric, changing the traditional "post-matching verification" to "pre-matching screening". This overcomes the problems of excessively strong topological constraints on isolated elements and excessively loose geometric constraints on central nodes caused by the use of fixed weights in existing technologies, significantly improving matching accuracy while ensuring topological consistency.

[0065] By incorporating topology preservation rate as a hard constraint into the global optimization framework, the "local optimum trap" is avoided, ensuring structural integrity: the road network remains connected after matching, and no dead ends occur. Based on a quantitative evaluation index of topological invariant differences, the degree to which the matching scheme preserves the source map's topology is comprehensively measured. In a matching experiment with 500 sets of real multi-source geographic vector data, compared to traditional geometric matching methods (78% accuracy, 72% topology consistency), the matching accuracy of this invention is improved to over 92%, and the topology consistency preservation rate reaches 90%, significantly improving the reliability and logical consistency of multi-source vector data fusion. This provides a high-precision data matching method for applications such as real-time integration of multi-source geographic data in disaster emergency response, comparison and updating of historical and current data in urban planning, and cross-departmental data fusion in land surveys. Attached Figure Description

[0066] Figure 1 This is a flowchart illustrating the steps of the multi-source geographic vector data matching method based on topological correlation of the present invention.

[0067] Figure 2This is a schematic diagram of the multi-source geographic vector data matching results based on topological correlation in this invention. Detailed Implementation

[0068] The following is in conjunction with the appendix Figure 1-2 The present invention will be further described in detail with reference to specific implementation methods.

[0069] A multi-source geographic vector data matching method based on topological correlation is as follows:

[0070] 1. Obtain multi-source basic geographic vector data, which comes from the following sources:

[0071] Vector data from urban planning departments: Collect basic geographic vector data provided by urban planning departments, including: urban road network data, building outline data, and landmark data. The data format is Shapefile or GeoJSON, and the coordinate system is CGCS2000 or a local independent coordinate system.

[0072] Vector data from the Ministry of Land and Resources: This includes basic vector data collected from the Ministry of Land and Resources for surveys and planning, such as land use status maps and boundary data of land features. The coordinate systems are mostly WGS84 or local coordinate systems.

[0073] Real-time survey data from emergency management departments: Collects real-time survey vector data from emergency management departments during disaster events, including: distribution of affected buildings, temporary rescue channels, emergency shelters, and the scope of disaster impact. Data collection methods include on-site collection via mobile terminals and GPS trajectory recording.

[0074] UAV aerial survey vector data: Data obtained from aerial images acquired by UAVs equipped with high-resolution cameras and then vectorized manually or automatically, including: the outlines of damaged buildings after disasters, the locations of road blockages, and areas of ground deformation.

[0075] 2. The acquired multi-source basic geographic vector data is preprocessed as follows:

[0076] (1) Coordinate system unification and geometric correction

[0077] Due to the different acquisition units, multi-source vector data has the problem of inconsistent coordinate references. The seven-parameter Bursa model is used for coordinate transformation. The seven parameters are solved by the least squares method of common control points, the transformation residuals are calculated, and outliers with residuals greater than 1 meter are removed.

[0078] (2) Topology cleaning and normalization

[0079] Topological cleaning of vector data is performed to eliminate topological errors generated during data acquisition and processing. This involves the following three steps:

[0080] The first step is to set the tolerance threshold. Meters, using a node capture method to select distances less than 1 meter. The endpoints of line segments are merged into shared nodes to ensure the connectivity of the road network;

[0081] The second step is to identify and delete suspended line segments with a length of less than 2 meters to avoid pseudo-nodes interfering with topology analysis.

[0082] The third step is to eliminate fragmented polygons with an area of ​​less than 1 square meter, fill in the voids in the face (area < 0.5 square meters), and fix self-intersection and overlap issues.

[0083] Furthermore, attribute consistency checks are performed on the data after topology cleanup, and the feature type coding and attribute field naming conventions are standardized to form a preprocessed dataset.

[0084] (3) Data quality assessment based on topological features

[0085] This invention proposes a quality scoring mechanism based on local topological features to identify low-quality data regions and guide subsequent matching strategies. Compared with the blind matching of existing technologies, this invention identifies weak links in the data before matching through topological quality pre-assessment, thus avoiding the propagation of topological errors by low-quality data.

[0086] For each feature in the preprocessed dataset For example, for roads, houses, etc., calculate their topology quality scores respectively:

[0087]

[0088] The three quality indicators are as follows:

[0089] The following formula is used to calculate the connectivity quality index for linear elements such as roads and pipelines:

[0090]

[0091] To share the number of endpoints of nodes with other line segments, This represents the total number of endpoints of the line segment, which is usually 2.

[0092] The geometric normality index for all geographic vector features is calculated using the following formula:

[0093]

[0094] It represents the standard deviation of the side lengths of a line segment or polygon (a measure of the dispersion of side lengths). The average side length is used to measure the regularity of a geometric shape.

[0095] The formula for calculating the feature density consistency index based on neighborhood statistics is as follows:

[0096]

[0097] For The element density (numbers / km²) within a buffer zone with a center radius of 100 meters. For global average density, Weight settings for the smoothing factor: , , ,when If the element is low quality, it will be marked as such and its confidence weight will be reduced during subsequent matching, or it will be subject to manual review first.

[0098] (4) Spatial index construction

[0099] Based on the calculated topology quality score, Elements with a value ≥0.6 are defined as qualified data. Elements with a value ≥0.8 are defined as high-quality data, and an R-tree spatial index is created for datasets S that meet the quality standards.

[0100] To improve the efficiency of proximity queries and topology analysis, the geospatial space is divided into a hierarchical structure of minimum bounding rectangles (MBRs), reducing the query time complexity from O(n) to O(logn).

[0101] 3. Map the point, line, and polygon geographic features in the preprocessed dataset and the target graph dataset to graph nodes respectively. Based on the topological relationships of the preprocessed data, construct a mixed edge set containing directed and undirected edges, and extract and classify the spatial topological relationship edges.

[0102] 3.1 Node Mapping and Graph Structure Initialization

[0103] Before constructing the spatial topology graph, two core datasets need to be defined:

[0104] Preprocessing datasets A benchmark dataset constructed through coordinate system unification and geometric correction, topological cleaning and normalization, quality assessment, and spatial indexing, containing... Each geographic feature serves as a reference benchmark for matching and fusion. This dataset typically originates from high-quality data sources such as existing map databases and authoritative surveying results.

[0105] Target Graph Dataset : New data sources to be matched and merged, including Each geographic feature needs to be compared with the preprocessed dataset. Perform topological relationship matching. This dataset typically originates from new surveying data, remote sensing image extraction results, crowdsourced map data, updated and supplementary data, etc.

[0106] Preprocess the dataset In Each geographic feature (including three geometric types: points, lines, and polygons) is mapped to a set of graph nodes. New data sources to be matched and integrated, such as new mapping data, crowdsourced data, and updated data, while also including the target map dataset. In Each element is mapped to a set of nodes. Each node retains the original feature's geometric coordinates, type identifier, and attribute information. For datasets with mixed feature types, a feature type index table is created. .

[0107] 3.2. Extraction and Classification of Spatial Topological Relationship Edges

[0108] Based on the topological relationships of the preprocessed data, a hybrid edge set of three types of directed and undirected edges is constructed to explicitly express the spatial relationships between elements:

[0109] (1) Connecting edge set : This refers to the endpoint connectivity between line elements. If line segments and Shared endpoint (node ​​capture tolerance) If the inside is an edge, then an edge is established. This edge set characterizes the connectivity structure of linear features such as road networks and river networks.

[0110] (2) Adjacent edge set This refers to the boundary adjacency relationships between polygonal features. and Shared common boundary, i.e., length greater than a threshold If the distance is 1 meter, then establish an edge. This edge set represents the adjacency relationship between areal features such as plots and buildings.

[0111] (3) Includes edge set This refers to the spatial containment relationship between point features and polygon features. If a point... Located in polygon Within the internal region, directed edges are established. This edge set describes the hierarchical relationship between landmarks and their respective regions. The final source graph is constructed from this edge set. and target map ,in It is a mixed edge set.

[0112] 4. Assign differentiated weights based on the type of topological relationship to quantify the degree of constraint of different spatial relationships on the matching process, and construct a weighted adjacency matrix for the target graph. This includes the following steps:

[0113] For the source image , build Weighted adjacency matrix Its element definition is as follows:

[0114] The nodes of each graph in the classification-weighted adjacency matrix are calculated as follows:

[0115]

[0116] In the formula, and For node indexing, , This represents the total number of nodes in the source graph. For nodes in the source graph, the weight parameter is set as follows: weight of line feature connectivity. As the strongest constraint, the adjacency weight of surface elements. This is a secondary constraint, with weights for point-to-surface inclusion relationships. This is a weak constraint. This represents the set of edges connecting all line features in the source graph. This represents the set of adjacency edges for all face features in the source graph. This represents the set of all points and faces in the source graph that contain relational edges.

[0117] For the target graph, the same method is used to construct its weighted adjacency matrix. .

[0118] 5. For graph nodes in the source graph dataset and candidate graph nodes in the target graph dataset, construct a comprehensive similarity model with geometric-topological co-constraints, quantify the dynamic neighborhood constraints in the matching process into a computable metric, calculate the neighborhood topology consistency metric, and integrate adaptive weights to achieve topology awareness in the matching decision stage.

[0119] 5.1 Constructing a comprehensive similarity model with geometric-topological co-constraints for graph nodes in the source graph dataset Candidate graph nodes in the target graph dataset Calculate geometric similarity It includes the following three indicators:

[0120] (1) Location distance measurement

[0121] Spatial offset is measured using Hausdorff distance and then normalized.

[0122]

[0123] in The distance attenuation parameter is set to 5 meters based on the data accuracy.

[0124] (2) Shape similarity measurement

[0125] Using the correlation coefficient of the rotation function To measure shape consistency, the boundaries of polygonal or line features are expressed as a function of cumulative rotation angle and arc length:

[0126] Traverse clockwise along the boundary and record the data at the normalized arc length position. Cumulative turning angle at the location For source graph nodes and target graph nodes Calculate their rotation angle functions respectively. and Then, the Pearson correlation coefficient is used to measure the similarity between the two function curves:

[0127]

[0128] in and These are the mean values ​​of the two rotation angle functions, respectively.

[0129] Value range: (Normalize after taking the absolute value). When When the two elements are completely identical in shape; when The time indicates that the shape is unrelated.

[0130] (3) Scale ratio

[0131] Calculate the area (surface feature) or length (line feature) ratio of the features:

[0132]

[0133] in, For source graph nodes The area (surface element) or length (line element). For target graph nodes The area (surface element) or length (line element).

[0134] The three indicators are weighted and fused together to calculate the comprehensive geometric similarity:

[0135]

[0136] Weights are dynamically adjusted based on feature type: point features are weighted more heavily based on location. Line elements emphasize shape weight The weights of shape and scale for surface elements are respectively... , .

[0137] 5.2. Neighborhood Topological Consistency Measurement

[0138] This invention proposes a neighborhood topology consistency metric, which quantifies the strength of topological constraints by comparing the degree of preservation of the neighborhood structure of nodes before and after matching. It quantifies the dynamic neighborhood constraints in the matching process into a computable metric, shifting the traditional "post-matching verification" to "pre-matching screening." The method includes the following steps:

[0139] (1) Definition of neighborhood set

[0140] For nodes in the source graph its neighborhood set Defined in the topological association graph In and source graph nodes There exists a set of all nodes connected by an edge:

[0141]

[0142] In the formula, For the source graph and the source graph nodes Neighboring nodes that have topological relationships and For nodes in the target graph The neighborhood set is ( ),

[0143]

[0144] in, For nodes in the target graph Neighboring nodes that have topological relationships and .

[0145] The neighborhood set reflects the local association patterns of the nodes in the overall topology of the graph.

[0146] (2) Matching mapping function: During the matching process, the determined matching mapping function is maintained. Map the source graph nodes to the target graph nodes. For nodes that have not yet been matched, Return to the empty set.

[0147] (3) Calculate topological similarity

[0148] The neighborhood topology consistency is defined as the ratio of the intersection of the neighborhoods of nodes in the source graph to the union of the neighborhoods of nodes in the target graph under the matching mapping:

[0149]

[0150] in: For source graph nodes With target graph nodes The neighborhood topological consistency metric between the two values ​​ranges from [0,1]; numerator The denominator represents the number of nodes in the source graph's neighborhood that have successfully matched the target graph's neighborhood, reflecting the degree of topology preservation; This represents the size of the union of two neighborhoods, used as a normalization factor; This means that the neighborhood of the target graph node is back mapped back to the source graph node space through an existing matching mapping.

[0151] when When the neighborhoods of the two nodes are completely identical, the topology can be perfectly preserved after matching; when If the neighborhood is completely non-overlapping, matching would disrupt the topological relationship. We employ Jaccard similarity to measure neighborhood topological consistency and dynamically maintain matching constraints through an iteratively updated mapping function, achieving a synergistic optimization of neighborhood structure preservation and geometric similarity.

[0152] (4) Correct the edge weighting and introduce a weighted adjacency matrix. The differentiated weights are calculated for neighboring nodes, and the corrected topological similarity is:

[0153] .

[0154] This weighted approach ensures the connection relationship. It will receive higher priority in the matching decision.

[0155] Furthermore, the neighborhood topology consistency quantization and adaptive weight fusion described in S5 involves dynamically adjusting the weight ratio of geometry and topology based on the centrality of nodes in the topology graph, and calculating the comprehensive similarity.

[0156]

[0157] in, These are adaptive weighting coefficients, with values ​​ranging from [0,1].

[0158] Based on node degree centrality (number of neighboring nodes), the adaptive weight is calculated as follows:

[0159]

[0160] in, For nodes The degree (number of neighboring nodes). The threshold for the number of neighbors, To adjust the parameters.

[0161] 5.3. Comprehensive Similarity Based on Adaptive Weight Fusion

[0162] Traditional methods, which use fixed weights to fuse geometric and topological similarity, cannot adapt to the differences in topological complexity among different graph nodes. This invention proposes a comprehensive similarity model with geometric-topological co-constraints, which achieves topological awareness in the matching decision stage through neighborhood topological consistency quantification and adaptive weight fusion.

[0163] Adaptive weighting dynamically adjusts the weight ratio of geometry and topology based on the centrality of nodes in the topological graph, and calculates the overall similarity:

[0164]

[0165] in, These are adaptive weighting coefficients, with values ​​ranging from [0,1].

[0166] Based on node degree centrality (number of neighboring nodes), the adaptive weight is calculated as follows:

[0167]

[0168] in, For nodes The degree (number of neighboring nodes) reflects its topological complexity; The threshold for the number of neighbors is used; when the number of neighbors exceeds 3, the node is considered to have strong topological constraints. To adjust the parameters and control the degree of weight transformation.

[0169] For isolated elements , Overall similarity mainly relies on geometric features because these elements lack topological constraints, and accurate matching of geometric positions should be prioritized.

[0170] For the central node , The overall similarity mainly depends on topological consistency, because these elements are key nodes in the topological structure (such as road intersections and regional centers). Their mismatch will lead to large-scale topological destruction, so topological constraints must be satisfied first.

[0171] For nodes with medium connectivity The weights smoothly transition between geometry and topology, achieving a balance constraint between the two.

[0172] This adaptive mechanism overcomes the problems of excessively strong topological constraints on isolated elements and excessively loose geometric constraints on central nodes caused by the use of fixed weights in existing technologies, significantly improving matching accuracy while ensuring topological consistency.

[0173] 6. Establish a global optimization objective function with the confidence matrix as input, take the topology preservation rate as a hard constraint, and use the improved Hungarian algorithm to solve the maximum weighted bipartite graph matching problem with topological constraints, and output the globally optimal matching scheme.

[0174] 6.1 Generation of Confidence Matrix C

[0175] Source image dataset of Individual elements and target map dataset of The comprehensive similarity score among the elements is integrated into... 5D confidence matrix The matrix elements are defined as follows:

[0176]

[0177] in, , The adaptive weighted fusion similarity calculated in the third step has a value range of [0,1]. To reduce computational complexity, the confidence matrix is ​​pre-filtered and pruned: when feature types do not match, the geometric distance exceeds 50 meters, or the overall similarity is below 0.3, it is directly set... By excluding this element from the candidate set, the effective candidate space can be reduced by more than 60%.

[0178] 6.2 Global Optimization Model with Topological Constraints

[0179] Incorporating topology preservation rate as a hard constraint into the global optimization framework aims to: (1) avoid the "local optimum trap": traditional methods may match two geometrically similar buildings, but destroy their connection with the surrounding roads; (2) ensure structural integrity: the road network remains connected after matching, and there will be no "dead ends". This ensures that the matching scheme meets the topology consistency requirement while achieving high geometric similarity.

[0180] (1) Definition of decision variables

[0181] definition 3D decision matrix Its elements :

[0182]

[0183] when Time indicates that the source graph node Matched target graph nodes ,when The time indicates a mismatch;

[0184] (2) Objective function

[0185] Establish a maximum weight bipartite graph matching model with the objective function:

[0186]

[0187] in, Represented as matching confidence, indicating the source graph node. With target graph nodes The confidence level of the match between them This represents the decision variable, taking values ​​of 0 or 1. , This represents the number of graph nodes in the target graph dataset and the source graph dataset;

[0188] (3) Constraints

[0189] Three types of constraints are set to ensure the rationality of matching and topological consistency: One-to-one matching constraint: ensures that each graph node is matched at most once, avoiding many-to-one or one-to-many conflicts.

[0190]

[0191]

[0192] Topology Preservation Hard Constraint: The proportion of topological edges in the source graph that are preserved in the target graph after matching is not less than a threshold. For each edge in the source graph ,like and Matched respectively and This requires that the target graph also contains corresponding edges. The optimized objective function formula is:

[0193]

[0194] Among them, indicator function The value is 1 if a corresponding edge exists in the target graph, and 0 otherwise. This represents the total number of edges in the source graph; The topology retention threshold is set to 0.85.

[0195] 6.3 Improved Hungarian Algorithm for Solving

[0196] The topology preservation constraint is transformed into a penalty term in the objective function using Lagrange relaxation techniques.

[0197]

[0198] in, These are Lagrange multipliers used to balance the trade-off between the objective function and topological constraints.

[0199] Iterative solution process: settings Iteration counting Iterative optimization and constraint verification: The Hungarian algorithm is used to solve the augmented objective function to obtain the candidate decision matrix. ; Calculate the number of topological edges preserved If satisfied Then output the optimal decision matrix. And terminate; Multiplier update: Update using subgradient method. ,limit Convergence criterion: If the degree is violated... If the number of iterations is ≥50, the process terminates; otherwise, return to step (2).

[0200] 6.4 Matching Scheme The generation

[0201] The optimal decision matrix obtained based on the improved Hungarian algorithm described above. Extract matching scheme Matching mapping function Defined as:

[0202]

[0203] If the source graph node In the decision matrix, all elements are 0 (i.e., ),but This indicates that no matching object was found for this graph node. Matching scheme The complete data structure is defined as an ordered set of triples:

[0204]

[0205] in: For identifying nodes in the source graph dataset; Identify the nodes of the target graph dataset; To match the confidence scores, we take the confidence score matrix C. The algorithm's time complexity is O(n log n). ,in The average number of iterations, This represents the number of features. For schemes with a topology retention rate below 0.80, the subsequent topology repair module will be automatically triggered.

[0206] 7. Construct a quantitative evaluation index based on the difference in topological invariants, namely, topological deviation measurement. Compare the topological deviation measurement value with the set topological deviation threshold to determine whether there is a significant topological conflict in the matching scheme. Based on the conflict situation, execute the corresponding hierarchical repair strategy.

[0207] 7.1 Extraction of Topological Invariants

[0208] Based on matching scheme Construct a mapping graph : Source graph nodes Replace with The corresponding target graph location retains the edge relationships of the source graph.

[0209] Three types of topological invariants were extracted for comparison: number of connected components. Count the number of disconnected subgraphs; count the number of loops. Degree distribution sequence .

[0210] 7.2 Topology Collision Detection

[0211] (1) Formula for measuring topological deviation

[0212] This invention proposes a quantitative evaluation index based on the difference in topological invariants, which comprehensively measures the degree to which the matching scheme preserves the topological structure of the source graph, and defines a topological deviation metric. for:

[0213]

[0214] In the formula, the first term is the relative rate of change of the connected components, and the numerator... Calculate the absolute difference in the number of connected components before and after matching, denominator The number of connected components in the source graph is normalized; the second term is the relative rate of change of the number of loops, and 1 is added to the denominator to avoid division by zero error in zero-loop graphs; This is a mapping graph formed by mapping the nodes of the source graph to the target graph space according to the matching scheme. The edge relationships of the source graph are preserved, but the node positions are replaced with the corresponding positions in the target graph.

[0215] in, To preserve the weights of connected components, To maintain the weighting coefficients of the loop, the weighting coefficients , Connectivity should be given higher priority because disruptions to connectivity have a more severe impact on applications such as road accessibility and river continuity.

[0216] (2) Conflict triggering threshold

[0217] Set topology deviation threshold ,when A significant topological conflict is detected in the matching scheme, triggering a subsequent tiered repair process. This threshold is determined through statistical analysis: in cross-validation experiments with 100 real datasets, when... The manual review pass rate for the matching results reached 98%. The pass rate dropped to 62%, therefore 0.15 is a reasonable quality control threshold.

[0218] 7.3 Graded Repair Strategy

[0219] This embodiment sets up a three-level repair strategy to achieve intelligent hierarchical and automatic repair, attempting each level from simple to complex. 90% of errors can be repaired at the first or second level, avoiding costly global recalculation.

[0220] Level 1 repair targets connectivity disruptions, traversing low-confidence matching pairs ( First, remove matches that cause connectivity disruption and replace them with suboptimal solutions. Second, for adjacency errors, search for alternative graph nodes that satisfy adjacency constraints within a 50-meter buffer. Third, for systematic errors, reduce the topology preservation rate threshold to 0.75 and re-execute global optimization.

[0221] 7.4 Final Matching Scheme and File Output

[0222] The final matching scheme will be generated after the repair. Expanded into a set of quintuples:

[0223]

[0224] in This refers to the identifier of the geographic map node in the source map dataset. Identify the matching graph nodes in the target graph dataset. To match confidence levels, Assuming a topology quality score, Mark the repair level. Based on Export a standard Shapefile file; geometric coordinates are derived from the nodes of the target graph. The geometric information includes the source map node identifier (SRC_ID), the target map node identifier (TGT_ID), the matching confidence (CONFIDENCE), the topology quality score (TOPO_QUAL), the repair level (REPAIR_LV), and the original attribute information. The coordinate system is set to CGCS2000, which is compatible with mainstream GIS platforms such as ArcGIS and QGIS.

[0225] Finally, it should be noted that the above is only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-source geographic vector data matching method based on topological correlation, characterized in that, The method includes: S1 collects multi-source basic geographic vector data and performs preprocessing, including topological cleaning and normalization, to form a preprocessed dataset. S2, perform multi-indicator data quality assessment on the preprocessed dataset based on local topological features to identify low-quality data areas, mark data below a predetermined threshold, and establish a tree space index for datasets that meet the quality standards. S3 maps the point, line, and polygon geographic features in the preprocessed dataset and the target graph dataset to graph nodes respectively. Based on the topological relationship of the preprocessed data, it constructs a mixed edge set containing directed and undirected edges to extract and classify spatial topological relationship edges. S4. Assign differentiated weights based on the type of topological relationship to quantify the degree of constraint of different spatial relationships on the matching process and construct a weighted adjacency matrix of the target graph. S5, for features in the source graph dataset and candidate features in the target graph dataset, construct a comprehensive similarity model with geometric-topological co-constraints, quantify the dynamic neighborhood constraints in the matching process into a computable metric, calculate the neighborhood topological consistency metric, and integrate adaptive weights, including: For nodes in the source graph dataset Candidate nodes in the target graph dataset Calculate the geometric similarity of three indicators: location distance metric, shape similarity metric, and scale ratio; Then, the three indicators mentioned above are weighted and fused together to obtain the comprehensive geometric similarity: In the formula, For location distance measurement, For shape similarity measurement, The scale ratio is used; the weight parameters are set as follows, depending on the feature type: Point features: focus on location, Line elements: Emphasizing shape. Surface elements: balancing shape and scale. ; and For node indexing, ; The neighborhood topological consistency is defined as the ratio of the intersection to the union of the neighborhoods of nodes in the source and target graphs under the matching mapping. The weights of geometry and topology are dynamically adjusted based on the centrality of nodes in the topological graph, and the overall similarity is calculated. in, For adaptive weighting coefficients, the range of values ​​is... , Source graph node With target graph nodes The comprehensive geometric similarity between them Source graph node With target graph nodes The neighborhood topological consistency metric between them, with a value range of [0,1]; S6. A global optimization objective function is established using the confidence matrix as input. The topology preservation rate is used as a hard constraint. An improved Hungarian algorithm is employed to solve the maximum weighted bipartite graph matching problem with topological constraints, outputting the globally optimal matching scheme. The optimized objective function formula is as follows: Among them, indicator function The value is 1 if a corresponding edge exists in the target graph, and 0 otherwise. This represents the total number of edges in the source graph; The topology retention rate threshold is set to 0.

85. For decision variables, take values ​​of 0 or 1. For nodes in the target graph Neighboring nodes that have topological relationships and , For the source graph and the source graph nodes Neighboring nodes that have topological relationships and ; S7. Construct a quantitative evaluation index based on the difference of topological invariants, topological deviation measurement, compare the topological deviation measurement value with the set topological deviation threshold, determine whether there is a significant topological conflict in the matching scheme, and execute the corresponding hierarchical repair strategy according to the conflict situation.

2. The multi-source geographic vector data matching method based on topological correlation according to claim 1, characterized in that, S2 The multi-indicator data quality assessment based on local topological features described in the article is: For each feature in the preprocessed dataset Calculate their topology quality scores respectively. : in, Represents the first in the preprocessed dataset A geographical element can be a point element, a line element, or a surface element. For connectivity quality indicators of road and pipeline elements, This is a geometric normalization index for all geographic vector features. The weighting is set as follows for the feature density consistency index based on neighborhood statistics: , , ; when When this occurs, it is marked as a low-quality feature, and its confidence weight is reduced in subsequent matching or it is given priority for manual review. Features with a confidence score of 0.6 or lower will be considered low-quality features. Elements with a value <0.8 are defined as qualified data. Elements with a value of ≥0.8 are defined as high-quality data.

3. The multi-source geographic vector data matching method based on topological correlation according to claim 1, characterized in that, The weighted adjacency matrix of the target graph described in S4 is: For the source image , build Weighted adjacency matrix Its element definition is as follows: The elements of the classification-weighted adjacency matrix are calculated as follows: In the formula, This represents the total number of nodes in the source graph. For nodes in the source graph, the weight parameter is set as follows: weight of line feature connectivity. As the strongest constraint, the adjacency weight of surface elements. This is a secondary constraint, with weights for point-to-surface inclusion relationships. This is a weak constraint. This represents the set of edges connecting all line features in the source graph. This represents the set of adjacency edges for all face features in the source graph. This represents the set of all points and faces in the source graph that contain relational edges.

4. The multi-source geographic vector data matching method based on topological correlation according to claim 1, characterized in that, The calculation of the neighborhood topological consistency metric described in S5 is specifically as follows: (1) Define the neighborhood set for nodes in the source graph. Its neighborhood set ( ) is defined in the topological association graph Zhongyu There exists a set of all nodes connected by an edge: In the formula, For the source graph and nodes Neighboring nodes that have topological relationships and For nodes in the target graph The neighborhood set is ( ), Let be the set of edges of the source graph; in, For nodes in the target graph Neighboring nodes that have topological relationships and , Let be the set of edges of the target graph; (2) Matching mapping function: During the matching process, the determined matching mapping function is maintained. Map the source graph nodes to the target graph nodes. For nodes that have not yet been matched, Return to the empty set. Represents the set of nodes in the source graph. To the target graph node set Matching mapping relationship; (3) Calculate topological similarity The neighborhood topology consistency is defined as the ratio of the intersection of the neighborhoods of nodes in the source graph to the union of the neighborhoods of nodes in the target graph under the matching mapping: in: Source graph node With target graph nodes The neighborhood topological consistency metric between them, with a range of values. ;molecular The denominator represents the number of nodes in the source graph's neighborhood that have successfully matched the target node's neighborhood, reflecting the degree of topology preservation; This represents the size of the union of two neighborhoods, used as a normalization factor; This indicates that the neighborhood of the target node is mapped back to the node space of the source graph through an existing matching mapping.

5. The multi-source geographic vector data matching method based on topological correlation according to claim 1, characterized in that, The calculation of the neighborhood topology consistency metric described in S5 also includes: Correcting edge weighting by introducing a weighted adjacency matrix The differentiated weights are calculated for neighboring nodes, and the corrected topological similarity is: For the source graph and the source graph nodes Neighboring nodes that have topological relationships and .

6. The multi-source geographic vector data matching method based on topological correlation according to claim 1, characterized in that, The calculation of the neighborhood topology consistency metric and the integration of adaptive weights, as described in S5, is as follows: Based on node degree centrality, the adaptive weights are calculated as follows: in, Source graph node The degree, i.e., the number of neighboring nodes. The threshold for the number of neighbors, To adjust the parameters, is the base of the natural logarithm.

7. The multi-source geographic vector data matching method based on topological correlation according to claim 1, characterized in that, The S6 objective function optimization method includes the following steps: First, define the decision variables, define... 3D decision matrix Its elements : when Time indicates that the source graph node Matched target graph nodes ,when The time indicates a mismatch; Secondly, a maximum weight bipartite graph matching model is established, with the objective function being: in, Represents source graph nodes With target graph nodes The confidence level of the match between them , This indicates the number of nodes in the target graph dataset and the source graph dataset; Finally, a hard constraint is set for the topology preservation rate: the percentage of topological edges in the source graph that are preserved in the target graph after matching is no less than a threshold. For each edge in the source graph ,like and Matched respectively and This requires that the target graph also contains corresponding edges. .

8. The multi-source geographic vector data matching method based on topological correlation according to claim 1, characterized in that, In S7, the set topology deviation threshold When topological deviation is measured If a significant topological conflict is detected in the matching scheme, a subsequent hierarchical repair process will be triggered.

9. The multi-source geographic vector data matching method based on topological correlation according to claim 1, characterized in that, The hierarchical repair described in S7 is as follows: Level 1 repair addresses connectivity disruptions and iterates through low-confidence matching pairs. Delete the matches that cause connectivity disruption and replace them with suboptimal solutions; Secondary repair addresses adjacency errors by searching for alternative features that satisfy adjacency constraints within a 50-meter buffer. Level 3 repair targets systemic errors, lowers the topology retention rate threshold to 0.75, and re-executes global optimization.

Citation Information

Patent Citations

  • Multi-source map fusion method, electronic equipment, storage medium and driving equipment

    CN117470255A

  • Intelligent old city boundary extraction and texture calculation system based on vector topology analysis and semantic segmentation

    CN120876855A