GIS multi-scale basic geographic space data efficient matching method
By adopting the model of relational database + spatial database engine and multi-forktree data structure in multi-scale basic geospatial data matching, combined with semantic change rules, the problem of low efficiency of multi-scale data matching is solved, and efficient and accurate spatial data matching and fusion is achieved.
Patent Information
- Application Number
- CN202510127243.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-01
- Publication Date
- 2025-06-03
AI Technical Summary
The existing technology is difficult to efficiently match multi-scale basic geospatial data, resulting in low spatial data fusion efficiency, low data quality and low map database update efficiency.
The spatial database model based on the relational database + spatial database engine is adopted, combined with multi-forktree data structure and semantic change rules, an object selection method and optimization scheme for multi-scale geospatial data matching is established to achieve efficient spatial data matching.
The matching efficiency and quality of multi-scale basic geospatial data is improved, and efficient fusion of spatial data and rapid update of map databases is promoted.
Smart Images

Figure CN120086250A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a method for matching geospatial multi-scale data, and particularly to an efficient method for matching GIS multi-scale basic geospatial data, belonging to the technical field of geospatial data matching. Background Art
[0002] Due to the increasing applications of geospatial technology in various industries, the demand for geospatial multi-scale data with high precision, good currency, multiple scales, and wide range is becoming more urgent. The key to maintaining data currency is to be able to obtain geospatial data in real time. When the scales of the earth space entities are different, they will show different spatial structures. The basic geographic information vector data are information on basic elements such as administrative regions, transportation, water systems, residential areas, and terrain; the geometric element types of these geographic information vector data are basically in the forms of points, lines, and surfaces, covering almost the entire land information range from macro to micro. Taking the basic geographic information resources as geospatial entity data, mining the basic information in the basic geographic information data, increasing its application scope, reducing resource waste, making the products of the basic geographic information data have richer information, and simplifying the process of basic geographic information sharing. Making the most of the existing spatial data information, reducing the production cost of spatial data, and establishing the association between basic geographic information element classes according to the production application requirements are of great significance for accelerating the update speed of the existing basic geographic information.
[0003] Due to the complexity of the objective world itself, the limitations in the human understanding and expression of the objective world, and the errors existing in the observation means of the real world, the computer's expression of geographic objects has limitations, and errors are inevitably generated when the computer processes data. Geospatial data can only be an approximation and abstraction of objective entities. Each expression of a geographic objective entity cannot be the same, which leads to differences in the data produced by different departments. In actual production, there are often different expression results for spatial objects of the same region and the same scale. Geospatial sharing is to study the essence of geographic information, explore the commonalities and differences in expressing geospatial entities at different scales, and then integrate and fuse the basic geographic data of the same region to find the geospatial data suitable for the needs of this department, and minimize the production cost of geospatial data as much as possible. Integrating and fusing multi-scale spatial data has become an urgent problem to be solved in GIS. Establishing the logical connection between spatial elements under different spatial data expressions of a unified geographic phenomenon is the prerequisite for effective spatial data integration and fusion, and the key lies in the integration and fusion of its position, spatial relationship, and attributes.
[0004] The key technical problems in spatial data integration, fusion or update are to find map elements of the same geographical features in the real world from different sources and establish logical connections between them, that is, the process of map data matching. Map data matching is used in many fields and plays an important role. Therefore, it is of great significance to match map data efficiently and accurately.
[0005] First, it is conducive to the efficient realization of spatial data fusion. The production and update of spatial data are time-consuming and costly tasks in GIS projects. Many information systems suffer from insufficient data quantity and long data update times, and some existing data is not effectively utilized, resulting in a waste of resources. Spatial data fusion uses certain methods and rules to effectively fuse spatial data of the same region from different sources and types, enriching the information of geographical features and facilitating a better understanding of things. This reduces problems such as ambiguity, incompleteness, and uncertainty in the data, improves the currency of the extracted information, increases the utilization efficiency of existing spatial data, and greatly reduces the cost of data collection. Map merging is a key technology in spatial information fusion and integration, and map data matching is the core of map merging. Therefore, map data matching is of great significance for the effective fusion of spatial data.
[0006] Second, it is conducive to improving the quality of spatial data. Map data matching can improve the quality of spatial data. The position accuracies of spatial data are different. By matching the data to find homonymous elements and establishing the conversion relationship between coordinates, the geometric positions of data with low position accuracy can be accurately corrected. In terms of data content, currency, and detail, data from different sources have different advantages and disadvantages in these aspects. Through data matching, the inconsistent problems in geometry, topology, and semantics of the data can be eliminated, thus improving the quality of spatial data.
[0007] Third, it is conducive to improving the update efficiency of the map database. At present, the technology of data collection has been rapidly improved, the data collection means are more advanced, and the acquisition cycle has been greatly shortened. GPS and RS technologies are effective methods for collecting spatial data with high accuracy and currency. When a large amount of such data emerges, the national basic geographic database has significant problems such as poor applicability and low currency. Therefore, it has become an urgent problem to update and maintain the existing geographic database with these large-scale topographic map data. Updating the existing vector data in the basic geographic database with new large-scale vector data is an important way of map database update. The scales of the new and old vector data can be the same, or the scale of the new vector data can be larger than that of the old vector data.
[0008] The problems to be solved in the multi-scale geospatial data matching of the prior art and the key technical difficulties of this application include: (1) The expression of geographical objects by computers has limitations. Errors are inevitably generated during computer data processing. Geospatial data can only be an approximation and abstraction of objective entities. Each expression of geographical objective entities cannot be the same, resulting in differences in data produced by different departments. There are different expression results for spatial objects of the same scale in the same area. Integrating and fusing multi-scale spatial data has become an urgent problem to be solved in GIS. Establishing the logical connection between spatial elements under different spatial data expressions of the same geographical phenomenon is the prerequisite for effective spatial data integration and fusion. The key lies in the integration and fusion of its location, spatial relationship, and attributes. The key technical problems in the integration, fusion, or update of spatial data in the existing technology are that it is impossible to find map elements of the same geographical features from different sources but reflecting the objective world, and it is impossible to establish a logical connection for them, that is, the process of map data matching. The efficiency and accuracy of map data matching are relatively low, which is not conducive to the efficient implementation of spatial data fusion, the improvement of spatial data quality, and the improvement of the update efficiency of the map database.
[0009] (2) With the continuous development and application of geospatial information technology, the demand for high-precision, up-to-date, multi-scale, and wide-range geospatial multi-scale data is becoming more urgent. The rapid update and integration of multi-scale basic geospatial data have become problems that need to be solved urgently, and the spatial data matching technology is the key to solving these problems. The existing technology lacks a spatial database construction model and a spatial data matching method based on it, and it is impossible to optimize multi-scale matching based on the geospatial database. First, there is a lack of a spatial database model based on a relational database + spatial database engine with fast storage speed and flexible application as a method for storing multi-scale basic geospatial data. Second, there is a lack of an object selection method for multi-scale geospatial data matching, and a lack of multi-scale semantic change rules based on geospatial data, making it impossible to effectively construct category matching relationships under multi-scale conditions. Third, there is a lack of an optimization scheme for multi-scale basic geospatial data matching, resulting in poor matching effects and low operating efficiency of geospatial multi-scale data matching.
[0010] (3) The existing technology has not solved the problem of multi-scale spatial data integration and fusion, has not established the logical connection between spatial elements under different spatial data expressions of the same geographical phenomenon, and lacks the integration and fusion of its location, spatial relationship, and attributes. There is a lack of a construction plan for multi-scale basic geospatial data, and the corresponding relationship between homonymous entities has not been established; the multi-scale semantic change rules of geospatial multi-scale data have not been analyzed, making it impossible to effectively deconstruct the complex relationships of geographical features. There is a lack of an object selection method for multi-scale basic geospatial data matching, and a multiple category matching relationship based on semantics between feature classes of different scales has not been established. There is a lack of an optimization scheme for multi-scale basic geospatial data matching, and it is impossible to achieve the matching of each homonymous entity in the map and geometric matching still needs to be carried out. The quality and efficiency of geospatial multi-scale data matching are relatively low. Summary of the Invention
[0011] The present application improves the method for matching multi-scale basic geospatial data, and optimizes the multi-scale matching scheme based on the geospatial database. First, a spatial database model based on a relational database + a spatial database engine is used as a method for storing multi-scale basic geospatial data. This method has a fast storage speed and flexible application. Second, a new object selection method for matching multi-scale geospatial data is established. Based on the multi-scale semantic change rules of geospatial data, basic geospatial entities are organized in the form of a multi-way tree data structure, and a category matching method for semantic changes of the same feature class at different scales is established to effectively construct a category matching relationship under multi-scale conditions. Third, an optimization scheme for matching multi-scale basic geospatial data is established, with good matching effect and high operating efficiency. It is beneficial to efficiently realize spatial data fusion, improve the quality of spatial data, and improve the update efficiency of the map database.
[0012] To achieve the above technical effects, the technical solutions adopted in the present application are as follows: An efficient matching method for GIS multi-scale basic geospatial data, which improves the method for matching geospatial multi-scale data based on a spatial database construction model and spatial data matching: First, a spatial database model based on a relational database + a spatial database engine is established as a method for storing multi-scale basic geospatial data; Second, a new object selection method for matching multi-scale geospatial data is established. Based on the multi-scale semantic change rules of geospatial data, basic geospatial entities are organized in the form of a multi-way tree data structure, and a category matching method for semantic changes of the same feature class at different scales is established to construct a category matching relationship under multi-scale conditions; Third, an optimization scheme for matching multi-scale basic geospatial data is established. On the basis of establishing a geospatial multi-scale database, pre-matching of spatial data is realized through a geospatial multi-scale data semantic object matching method, and then a multi-scale spatial data matching software is used for matching optimization to realize the matching of geospatial multi-scale data; Geospatial multi-scale data semantic object matching: (1) Among the basic geographic information at multiple scales, the national standard codes of the basic geographic entities at a scale of 1:5000 are sorted and organized according to semantic levels, and each major category contains sub-categories; (2) The organized national standard codes of the 1:5000 basic geographic entities are generated in the program in the form of a multi-way tree data structure. Each geographic entity code is a node of the multi-way tree, and the root node is the name of "feature", located at the 0th layer of the tree; the sub-categories of the feature are all sub-nodes of the "feature" node, located at the 1st layer of the tree; each sub-category under the water system and so on form a "multi-way tree" with a hierarchical structure; (3) If the code nodes of the two geographical entity feature classes to be matched are the same node in the multi-way tree, or the two nodes to be matched have the same ancestor node, the level number of the ancestor node is greater than or equal to 1 and the difference in the level numbers of the two nodes to be matched is less than or equal to 2, this application stipulates that these two feature classes can be matched, and vice versa, they cannot be matched.
[0013] Preferably, for the hierarchical coding of the basic geographical information data classification object: the data at each scale level is set to be associated. When obtaining data, the 1:1 million data is synthesized from the 1:250,000 data; the topographic features are divided into 8 major categories, namely positioning basis, water system, residential areas and facilities, transportation, pipelines, boundaries and administrative regions, landforms, soil and vegetation. The small class codes and sub-class codes are classified according to three scale segments of 1:500, 1:1000, 1:2000, 1:5000 - 1:10,000, and 1:25,000 - 1:100,000. Among them, new definitions and expansions are not allowed for the major and middle classes, and redefinitions are not allowed for the small and sub-classes. The topographic feature classification code consists of a major class code, a middle class code, a small class code, and a sub-class code, and is composed of 6 - digit decimal digital codes. The structure is: major class code + middle class code + small class code + sub-class code. Among them, the major class code is composed of 8 digits from 1 to 8. In the classification of basic geographical elements, the selection of the semantic relationship between the major and middle classes refers to the description form of the semantic relationship between the major and middle classes, and determines the existing semantic relationship between the major and middle classes. All the semantic relationships between the major and middle classes are an abstract conceptual - level semantic relationship. The semantic relationships between the small and sub - classes inherit from the major and middle classes and are the concretization of the semantic relationships between the major and middle classes.
[0014] Preferably, for the semantic object matching of geographical space multi - scale data: organize the basic geographical entities in the form of a "multi - way tree" data structure, and construct a method for all possible matches between geographical entities caused by cartographic generalization under multi - scale conditions. The specific process is as follows: (1) Among the basic geographic information at multiple scales, the national standard codes of the basic geographic entities at the scale of 1:5000 are sorted and organized according to the semantic hierarchy. Each major category contains subcategories. The first layer of the semantic classification tree is the major category of geographic features, which is divided into eight major categories in total: positioning basis (100000), water (200000), residential areas and facilities (300000), transportation (400000), pipelines (500000), boundaries and administrative regions (600000), landforms (700000), soil and vegetation (800000); the second layer of the semantic tree, the water system (200000), is divided into seven middle categories, including rivers (210000), ditches (220000), lakes (230000), reservoirs (240000), marine elements (250000), other water system elements (260000), water conservancy and affiliated facilities (270000); the third layer of the semantic classification tree, the river (210000), is divided into four small categories, including perennial rivers (210100), seasonal rivers (210200), dry rivers (210300), river annotations (219000); the fourth layer of the semantic classification tree, the perennial river (210100), is divided into four subcategories, including surface rivers (210101), underground river sections (210102), entrances and exits of underground river sections (210103), disappearing river sections (210104); (2) The organized national standard codes of the 1:5000 basic geographic entities are generated in the program in the form of a multi-way tree data structure. Each geographic entity code is a node of the multi-way tree, and the root node is the name of "geographic feature", located at the 0th layer of the tree; the subcategories of geographic features such as the water system (200000), residential areas and facilities (300000), and transportation (400000) are all child nodes of the "geographic feature" node, located at the 1st layer of the tree; and so on for each subclass under the water system, forming a "multi-way tree" with a hierarchical structure; (3) Scale changes cause changes in the semantic information of the same entity. The multi - fork tree logical structure of the 1:5000 basic geographic entities describes the most detailed semantic classification. Regardless of scale changes, the corresponding semantic information can be found in this multi - fork tree. If the code nodes of two geographic entity feature classes to be matched are the same node in the multi - fork tree, or the two nodes to be matched have the same ancestor node, the level number of the ancestor node is greater than or equal to 1 and the level difference between the two nodes to be matched is less than or equal to 2, this application stipulates that these two feature classes can be matched. In the first case, assuming that the two geographic entity feature classes to be matched are both rivers (210000), then these two feature classes can be matched. In the second case, if the first geographic entity feature class is a perennial river (210100) and the second geographic entity feature class is a perennial lake (230100), these two feature classes have a common ancestor, water system (200000). The level number of the water system (200000) node is greater than 1 and the level difference between the two nodes to be matched is 0. According to this application's regulations, these two feature classes can be matched. Similarly, the feature classes national highway (420100) and provincial highway (420200) have a common ancestor, inter - city highway (420000). The level number of the inter - city highway node is 2 and it satisfies that the level difference between the two nodes to be matched is less than or equal to 2, so the national highway and the provincial highway can also be matched. On the contrary, for the feature classes inter - city highway (420000) and lake (230000), although the level numbers of the nodes of these two feature classes are the same, their common ancestor "feature" node is the root node, so these two feature classes cannot be matched.
[0015] Preferably, the data preparation module includes spatial data loading, matching method selection, and parameter setting. Among them, the system supports two ways of loading, including local data files (in shapefile format) and feature sets in the spatial database. Corresponding matching methods are selected according to different types of features to be matched and corresponding matching parameters are set.
[0016] Preferably, the matching module is entity matching, including semantic matching and geometric matching. Semantic matching judges which features belong to the same major category as the source feature according to the attribute items of the data in the feature set through the semantic matching algorithm, and then obtains an alternative matching entity set. Then, entity geometric matching is carried out. According to the scale division, a total of 7 types of geometric matching are realized, namely, surface entity matching, surface - line entity matching, surface - point entity matching, line entity matching, line - point entity matching, point entity matching, and point - group surface entity matching.
[0017] Preferably, the surface entity matching method is based on considering the domain approximation degree and the global optimization matching strategy, and the architecture is as follows: ① The geometric approximation degree of the surface entity is based on the geometric features of a single surface entity. The parameters designed for the geometric approximation degree model include shape, size, and position weights. The index values are determined before matching, obtained by training and selecting a determined sample set, quantifying each index, and achieving optimized matching results through continuous training and feedback. ② Introduce the fixed grid space indexing technology for the alternative matching set of the surface entity to be matched, reducing the range of searching for the alternative matching set. At the same time, use the maximum position deviation and geometric similarity threshold to further narrow the scale of the alternative matching geometry, avoiding the omission of homonymous entities. ③ Construct the neighborhood environment space structure of the surface entity, laying the foundation for the evaluation of neighborhood approximation degree. For the recognition of one-to-one relationships of surface entities, the spatial approximation degree evaluation of surface entities combines geometric approximation degree and neighborhood approximation degree evaluation methods, taking into account both the approximation degree of the individual geometric features of the surface entity and the approximation degree of the corresponding spatial structure of the surface entity, ensuring the consistency of individual similarity and overall similarity. In the matching strategy, a global optimization method is adopted to ensure the consistency of local similarity and global similarity, the consistency of individual and overall similarity, and the consistency of local and global similarity, reducing the impact of spatial position offset on matching. ④ For the recognition of one-to-many relationships of surface entities, use the intersection points of the surface entity to be matched and its minimum bounding rectangle as marker points, obtain the alternative matching combinations of the surface entity to be matched through translation, and identify and construct the one-to-many matching relationship of the surface entity based on the overlapping similarity index.
[0018] Preferably, for line entity matching: ① Road network data preprocessing: First, perform stroke connection processing on the road network data, connect the roads according to the good coherence rule in spatial cognition, then, perform equidistant interpolation point encryption processing on the connected strokes, and finally, establish a grid space index for the strokes and process them using the grid space indexing algorithm for line features. ② Overall matching of roads: Use the approximate buffer space indexing algorithm to find the alternative matching set, then use the designed geometric similarity calculation method between strokes to measure the geometric approximation degree, use the improved spatial scene structure similarity evaluation method to enhance the matching effect of the algorithm, then use the forward and backward bidirectional matching algorithm to search for matching objects in reverse, and finally use multi-lane matching processing for the matching of multi-lane topological relationship changes caused by scale changes in the multi-scale road network map. ③ Partial matching of roads: Roads that do not find matching objects in the previous process need to enter the partial matching process. The partial matching process should handle the situations of changes and updates between homonymous road entities caused by inconsistent currency. The matching of circular roads should handle the matching situations of changes in the circular road topological structure to further enhance the matching effect.
[0019] Preferably, the line entity is matched with the point entity: The distance from the point to the line is used for judgment. The minimum value of the vertical distance from the point to the line segment and the minimum value of the distance from the point to each vertex of the broken line are calculated to obtain the minimum value between the two. The minimum distance value is compared with the critical value. If the distance is less than the critical value, the matching is successful; otherwise, the matching is unsuccessful. The surface is matched with the point entity: In two data sets with a large scale span, by searching for point targets in the large-scale surface features in the smaller scale, that is, according to the judgment algorithm of whether a point is inside a polygon, it is judged whether the point is in the surface target. If there are multiple points in the surface target, the point with the smallest distance from the centroid of the surface target is selected.
[0020] Preferably, for the matching of point entities: For the matching between single point targets, the method of comparing the spatial positions of two points is adopted. For the same-name point targets in different scale map spaces, their spatial positions should be the same, that is, their spatial coordinates should be the same. The position proximity is used for matching, that is, by calculating the Euclidean distance between two points. Suppose the reference point target and the point target to be matched are A and B respectively, and the plane coordinates of the two points are (x A , y A )(X B , y B ). The distance d AB between the two points is calculated by the following formula:
[0021] The closer the distance between points A and B is, the smaller d AB is, and the greater the possibility of matching. If the point entity closest to point A is B, and d AB is less than the critical value, and the point entity closest to point B is also A, then it is determined that the two points are matched; otherwise, they are not matched.
[0022] Preferably, for the matching of the point group on the large scale and the surface on the small scale: An extended minimum area circumscribed rectangle is established for the surface entity on the small scale, and the extension radius is the maximum error of the two data sets. If the point entity on the large scale falls within the rectangle, the point entity is matched with the surface entity on the large scale.
[0023] Compared with the prior art, the innovation points and advantages of this application are as follows: (1) The present application improves the method for matching multi-scale basic geospatial data, and optimizes the multi-scale matching scheme based on the geospatial database. First, a spatial database model based on a relational database + a spatial database engine is used as a method for storing multi-scale basic geospatial data. This method has a fast storage speed and flexible application. Second, a new object selection method for matching multi-scale geospatial data is established. Based on the multi-scale semantic change rules of geospatial data, basic geographical entities are organized in the form of a multi-way tree data structure, and a category matching method for semantic changes of the same feature class at different scales is established to effectively construct a category matching relationship under multi-scale conditions. Third, an optimization scheme for matching multi-scale basic geospatial data is established, with good matching effect and high operating efficiency. It is beneficial to efficiently realize spatial data fusion, improve the quality of spatial data, and improve the update efficiency of the map database.
[0024] (2) A database construction scheme for multi-scale basic geospatial data is established. For the merger and update optimization of map data in geospatial multi-scale data, the same-name entities in databases with different time phases and scales are identified based on data matching technology, and the corresponding relationship between the same-name entities is established. This involves the database construction scheme for vector data of each scale and the data storage problem. With the development of computer technology and database technology, there are also various modes for the basic geographic information database. The present application optimizes the two key technologies of spatial data model and spatial data engine in the construction of the basic geographic information database.
[0025] (3) The multi-scale semantic change rules of geospatial multi-scale data are analyzed. Due to different abstraction scales of geographical things, there are connections and differences in the expression results of the same geographical object. Geographical information has semantic connectivity and hierarchy. Starting from the hierarchical framework of map knowledge, the present application divides geospatial multi-scale data into a generic layer, an entity layer, and an attribute layer, and analyzes the multi-scale change rules of geographical information layer by layer, effectively deconstructing the complex relationships of geographical objects and providing a basis for finding a suitable method for data matching objects.
[0026] (4) An object selection method for matching multi-scale basic geospatial data is proposed. The basic geographical information is divided into 8 categories, namely positioning basis, water system, residential areas and facilities, transportation, pipelines, boundaries and administrative regions, landforms, soil and vegetation, and the element codes of each category are unique. Although the semantic information expressed by the same geographical object may be different due to cartographic generalization at different scales, this change also occurs within a large category. The present application establishes a semantic classification tree according to the national standard code and, based on this, establishes a multiple category matching relationship based on semantics for feature classes at different scales.
[0027] (5) An optimization scheme for multi-scale basic geospatial data matching is proposed. Selecting the matching objects of multi-scale geospatial multi-scale data only finds the alternative matching objects belonging to the same category. To achieve the matching of the same-name entities in the map, geometric matching is still required. This application uses the "Multi-scale Spatial Data Matching Software" to optimize the geometric matching and analyze the results, improving the quality and efficiency of multi-scale geospatial data matching. Brief Description of the Drawings
[0028] Figure 1 It is a schematic diagram of the organization of geographical entity names at a scale of 1:5000 according to semantic levels.
[0029] Figure 2 It is a schematic diagram of the "multi-way tree" logical structure with partial node code information.
[0030] Figure 3 It is a schematic diagram of the comparison of lakes in the 1:100,000 and 1:250,000 feature sets.
[0031] Figure 4 It is a schematic diagram of the matching data of settlements with similar scales.
[0032] Figure 5 It is a schematic diagram of the settlement matching result of the method of this application.
[0033] Figure 6 It is a schematic diagram of the matching data of roads with similar scales.
[0034] Figure 7 It is a schematic diagram of the overlay of road data with adjacent scales.
[0035] Figure 8 It is a schematic diagram of the statistics and evaluation of the line entity matching results.
[0036] Figure 9 It is a schematic diagram of the settlement data with adjacent scales.
[0037] Figure 10 It is a schematic diagram of the overlay of settlement data with adjacent scales.
[0038] Figure 11 It is a schematic diagram of the settlement data matching result. Detailed Embodiment
[0039] The following further describes the technical solution of the GIS multi-scale basic geospatial data efficient matching method provided by this application with reference to the drawings, so that those skilled in the art can better understand this application and be able to implement it.
[0040] As the digital representation of things, phenomena, and their interrelationships in the objective geographical world, basic geospatial data is the data source of basic geographic information systems and the core content jointly concerned by modern cartography and geographic information systems. With the continuous development and application of geographic information technology, the demand for geospatial multi-scale data with high precision, good currency, multiple scales, and wide scope is becoming more urgent. Therefore, the rapid update and integration of multi-scale basic geospatial data have become problems that need to be solved urgently, and spatial data matching technology is the key to solving these problems. This application improves the method of multi-scale basic geospatial data matching based on the spatial database construction model and spatial data matching method, and optimizes the multi-scale matching scheme based on the geospatial database: (1) The spatial database model based on the relational database + spatial database engine is used as a method for storing multi-scale basic geospatial data. This method has a fast storage speed and flexible application; (2) Establish a new object selection method for multi-scale geospatial data matching. Based on the multi-scale semantic change rules of geospatial data, organize basic geographical entities in the form of a multi-way tree data structure, and establish a category matching method for semantic changes of the same feature class at different scales, effectively constructing category matching relationships under multi-scale conditions; (3) Establish an optimization scheme for multi-scale basic geospatial data matching. On the basis of establishing a geospatial multi-scale database, realize the pre-matching of spatial data through the geospatial multi-scale data semantic object matching method, then use the multi-scale spatial data matching software to optimize the matching, and analyze the optimization results to improve the multi-scale basic geospatial data matching scheme, realize the matching of geospatial multi-scale data, with good matching effect and high operation efficiency.
[0041] I. Classification Object Hierarchical Coding of Basic Geographic Information Data The data of each scale of the national basic geographic information system is set to be associated. When obtaining data, the 1:1 million data is synthesized from the 1:250,000 data; topographic elements are divided into 8 categories, namely positioning basis, water system, residential areas and facilities, transportation, pipelines, boundaries and administrative regions, landforms, soil and vegetation. The minor class codes and subclass codes are classified according to three scale segments of 1:500, 1:1000, 1:2000, 1:5000 - 1:10,000, and 1:25,000 - 1:100,000. Among them, new definitions and expansions are not allowed for major classes and middle classes, and redefinitions are not allowed for minor classes and subclasses.
[0042] The classification code of topographic elements consists of major categories, middle categories, minor categories and sub-categories, which is composed of 6-digit decimal digital codes. Its structure is: major category code + middle category code + minor category code + sub-category code. Among them, the major category code consists of 8 digits from 1 to 8. In the classification of basic geographic elements, the selection of the semantic relationship between major and middle categories refers to the description form of the semantic relationship between major and middle categories, and determines the existing semantic relationship between major and middle categories. All semantic relationships between major and middle categories are an abstract conceptual-level semantic relationship. The semantic relationships between minor and sub-categories are all inherited from major and middle categories and are the concretization of the semantic relationships between major and middle categories.
[0043] The element code is used for category matching to reflect the classification characteristics of topographic elements. The category matching of elements is achieved by coding, but it does not truly reflect the matching of the graphic entity characteristics that make up the topographic elements, and it is impossible to find the matching relationship between basic geographic information entities. Entity geometric matching is still required.
[0044] II. Semantic Object Matching of Geospatial Multi-scale Data Scale is an important feature of spatial data. Due to cartographic generalization, the same feature will have different semantics in different scale representations. Specifically, the same entity changes from one element category to another, such as expressways and first-class highways becoming highways, paddy fields and dry land becoming cultivated land. However, these changes are all carried out under a large classification. This phenomenon of semantic multi-scale changes caused by spatial multi-scale changes determines that semantic matching needs to be carried out before geometric matching of multi-scale basic geographic information. The purpose of semantic matching is to unify the element classes of two different scale datasets to be matched, reduce the number of alternative matching classes, and improve the matching efficiency. However, these changes are the classification-level changes brought about by normal cartographic generalization, and these changes are under a large classification, and their classification codes still belong to the same parent node in the semantic classification tree. Regardless of how the scale changes, highway data will not be abstracted into river data due to cartographic generalization, and building data will not become lake data. That is, semantic matching excludes all situations that do not need to be matched, such as highway element classes and river element classes, before geometric matching of multi-scale basic geographic information; that is, under multi-scale, it determines all situations where matches may occur, such as first-class highways and highways.
[0045] Based on the above requirements, under multi-scale, the rule that the change of entity categories will not change from one major category to another major category and is all carried out under a large classification is used to organize basic geographic entities in the form of a "multi-way tree" data structure, and a method for all possible matches between geographic entities caused by cartographic generalization under multi-scale conditions is constructed. The specific process is as follows: (1) Among the basic geographic information at multiple scales, the national standard codes of the basic geographic entities at the scale of 1:5000 (the semantic description of the basic geographic information at this scale is the most detailed) are sorted and organized according to the semantic level. Each major category contains sub-categories. As shown in Figure 1 , the first layer of the semantic classification tree is the major category of geographic features, which is divided into eight major categories in total: positioning basis (100000), water (200000), residential areas and facilities (300000), transportation (400000), pipelines (500000), boundaries and administrative regions (600000), landforms (700000), soil and vegetation (800000); the second layer of the semantic tree, the water system (200000), is divided into seven middle categories, including rivers (210000), ditches (220000), lakes (230000), reservoirs (240000), marine elements (250000), other water system elements (260000), water conservancy and ancillary facilities (270000); the third layer of the semantic classification tree, the river (210000), is divided into four small categories, including perennial rivers (210100), seasonal rivers (210200), dry rivers (210300), and river annotations (219000); the fourth layer of the semantic classification tree, the perennial river (210100), is divided into four sub-categories, including surface rivers (210101), underground river sections (210102), entrances and exits of underground river sections (210103), and disappearing river sections (210104); (2) The organized national standard codes of the 1:5000 basic geographic entities are generated in the program in the form of a multi-way tree data structure. Each geographic entity code is a node of the multi-way tree. The root node is the name of "geographic feature" and is located at the 0th layer of the tree; the sub-categories of geographic features such as the water system (200000), residential areas and facilities (300000), and transportation (400000) are all sub-nodes of the "geographic feature" node and are located at the 1st layer of the tree; and so on for each sub-category under the water system, forming a "multi-way tree" with a hierarchical structure. As shown in Figure 2 is a schematic diagram of the logical structure of the "multi-way tree" with partial node code information; (3) Scale changes result in changes in the semantic information of the same entity. The multi-way tree logical structure of the 1:5000 basic geographic entities describes the most detailed semantic classification. Regardless of how the scale changes, the corresponding semantic information can be found in this multi-way tree. If the code nodes of two geographic entity feature classes to be matched are the same node in the multi-way tree, or the two nodes to be matched have the same ancestor node, the level number of the ancestor node is greater than or equal to 1 and the level numbers of the two nodes to be matched differ by less than or equal to 2. This application stipulates that these two feature classes can be matched. In the first case, assuming that the two geographic entity feature classes to be matched are both rivers (210000), then these two feature classes can be matched. In the second case, if the first geographic entity feature class is a perennial river (210100) and the second geographic entity feature class is a perennial lake (230100), these two feature classes have a common ancestor, water system (200000). The level number of the water system (200000) node is greater than 1 and the difference in the level numbers of the two nodes to be matched is 0. Therefore, according to the regulations of this application, these two feature classes can be matched. Similarly, the feature classes national highway (420100) and provincial highway (420200) have a common ancestor, intercity highway (420000). The level number of the intercity highway node is 2 and it satisfies that the difference in the level numbers of the two nodes to be matched is less than or equal to 2. The national highway and the provincial highway can also be matched. On the contrary, for the feature classes intercity highway (420000) and lake (230000), although the level numbers of the nodes of these two feature classes are the same, their common ancestor, the "feature" node, is the root node, so these two feature classes cannot be matched.
[0046] III. Optimization of Multi-scale Matching of Spatial Data In order to optimize the above multi-scale basic geographic spatial data semantic object matching method and matching scheme, this application uses the "Multi-scale Spatial Data Matching Software" to perform matching optimization on geographic spatial multi-scale data. First, use the semantic matching method of this application to perform pre-matching on geographic spatial multi-scale data, use the multi-scale spatial data matching software to perform precise matching based on geometric features on the results obtained after pre-matching, and finally evaluate the matching results.
[0047] (I) Configuration of Core Modules The configuration of core modules includes a data preparation module, a matching module, and a result analysis and visualization module.
[0048] 1) Data Preparation Module The content in the data preparation module includes spatial data loading, matching method selection, and parameter setting. Among them, the system supports two ways of loading, including local data files (in shapefile format) and feature sets in spatial databases; select the corresponding matching method and set the corresponding matching parameters according to the different types of features to be matched.
[0049] 2) Matching Module The matching module performs entity matching, including semantic matching and geometric matching. Semantic matching uses the attribute items of the data in the feature set and determines which features belong to the same major category as the source feature according to the semantic matching algorithm, thereby obtaining an alternative matching entity set. Then, entity geometric matching is performed. According to the scale division, a total of 7 types of geometric matching are realized, namely surface entity matching, surface-line entity matching, surface-point entity matching, line entity matching, line-point entity matching, point entity matching, and point group-surface entity matching. The principles of each matching algorithm are as follows.
[0050] ⑴ Surface Entity Matching At multiple scales, there are differences in shape, size, spatial position, and structural features between surface entities with the same name. The matching relationships between surface targets are mainly one-to-one and one-to-many, and there are also a small number of many-to-many cases. The surface entity matching method is based on considering the domain approximation degree and the global optimization matching strategy, and the architecture is as follows: ① The geometric approximation degree of the surface entity is based on the geometric features of a single surface entity. The parameters designed in the geometric approximation degree model include shape, size, and position weights. The index values are determined before matching and are obtained by training and selecting a determined sample set. Each index is quantified, and through continuous training and feedback, the matching result is optimized; ② Introduce a fixed grid spatial indexing technique for the alternative matching set of the surface entity to be matched, which reduces the range of searching for the matching alternative set and improves the matching efficiency. At the same time, the maximum position deviation and geometric similarity threshold are used to further reduce the scale of the alternative matching geometry, avoiding the omission of entities with the same name; ③ Construct the neighborhood environment spatial structure of the surface entity to lay the foundation for the evaluation of neighborhood approximation degree. For the recognition of the one-to-one relationship of surface entities, the spatial approximation degree evaluation of surface entities combines the geometric approximation degree and the neighborhood approximation degree evaluation methods, taking into account both the approximation degree of the individual geometric features of the surface entity and the approximation degree of the spatial structure corresponding to the surface entity, ensuring the consistency of individual similarity and overall similarity. The global optimization method is adopted in the matching strategy to ensure the consistency of local similarity and global similarity, the consistency of individual and overall similarity, and the consistency of local and global similarity, reducing the impact of spatial position offset on matching and improving the matching accuracy; ④ For the recognition of the one-to-many relationship of surface entities, the intersection points of the surface entity to be matched and its minimum circumscribed rectangle are used as marker points, and the alternative matching combinations of the surface entity to be matched are obtained by translation. Based on the overlapping similarity index, the one-to-many matching relationship of the surface entity is recognized and constructed.
[0051] (2) Line Entity Matching ① Road network data preprocessing: First, perform stroke connection processing on the road network data. Connect the roads according to the good coherence rule in spatial cognition. After that, perform equidistant interpolation point (vertex) encryption processing on the connected strokes. Finally, establish a grid spatial index for the strokes and process it using the grid spatial index algorithm for line features.
[0052] ② Overall matching of roads: Use the approximate buffer spatial index algorithm to find the alternative matching set. Then, use the designed geometric similarity calculation method between strokes to measure the geometric approximation degree, and use the improved spatial scene structure similarity evaluation method to enhance the matching effect of the algorithm. Subsequently, use the forward and backward two-way matching algorithm to search for matching objects in reverse. Finally, use multi-lane matching processing to match the changes in multi-lane topological relationships caused by scale changes in the multi-scale road network map. ③ Partial matching of roads: Roads that do not find matching objects in the previous process need to enter the partial matching process. The partial matching (stroke partial matching algorithm based on arc segment and vertex decomposition) process should handle the changes and updates that occur between homonymous road entities caused by inconsistent currency. The circular road matching should handle the matching situation of changes in the circular road topological structure to further enhance the matching effect.
[0053] (3) Matching of line entities and point entities Use the distance from the point to the line for judgment. Calculate the minimum value between the minimum value of the perpendicular distance from the point to the line segment and the minimum value of the distance from the point to each vertex of the polyline to obtain the minimum value between the two. Compare this minimum distance value with the critical value. If the distance is less than the critical value, the matching is successful; otherwise, the matching is unsuccessful.
[0054] (4) Matching of polygon and point entities In two data sets with a large scale span, search for point targets in the large-scale polygon features in the smaller scale, that is, judge whether the point is in the polygon target according to the algorithm for judging whether a point is inside a polygon. If there are multiple points in the polygon target, select the point with the minimum distance from the centroid of the polygon target.
[0055] (5) Matching of polygon entities and line entities The matching of polygon and line includes 1:1 matching and non-1:1 matching. Non-1:1 matching includes several situations such as 1:M, N:1, and N:M. First, extract the skeleton line of the polygon entity, calculate the length similarity of the part where the skeleton line matches the line to be matched, and judge whether the polygon matches the line according to the minimum and maximum similarity critical values.
[0056] (6) Matching of point entities For the matching between single point targets, a method of comparing the spatial positions of two points is adopted. For homologous point targets in the spatial maps of different scales, their spatial positions should be consistent, that is, their spatial coordinates should be the same. The position proximity is used for matching, that is, by calculating the Euclidean distance between two points. Suppose the reference point target and the point target to be matched are A and B respectively, and the plane coordinates of the two points are (x A , y A )(X B , y B ). The distance d AB between the two points is calculated by the following formula:
[0057] The closer the distance between points A and B is, the smaller d AB is, and the greater the possibility of matching. If the point entity closest to point A is B and d AB is less than the critical value, and the point entity closest to point B is also A, then the two points are considered to be matched; otherwise, they are not matched.
[0058] (7) Matching between point clusters and surface entities For the matching between point clusters on a large scale and surfaces on a small scale: An extended minimum-area circumscribed rectangle is established for the surface entities on the small scale, and the extension radius is the maximum error between the two data sets. If the point entities on the large scale fall within the rectangle, then the point entities are matched with the surface entities on the large scale.
[0059] (II) Optimizing system configuration 1. Selection of semantic matching objects for basic geospatial data When matching two geospatial multi-scale data sets from different sources, different time phases, or different scales, the following data preprocessing is performed before matching: Unify the geographic coordinate systems and projection coordinate systems of the two different data sets, and correct the spatial position errors of the two different data sets; In semantic matching, in the large-scale dataset of 1:100,000, there are four major categories of elements, namely water systems, residential areas and facilities, transportation, and boundaries and administrative regions, including 11 geographical elements such as natural villages (310108), prefecture-level boundaries (640201), ditches (220100), lakes (230101), double-line rivers (210101), village committee seats (311106), expressways (420801), provincial roads (420102), county roads (420301), rural roads (420400), and ramps (420600). In the small-scale dataset of 1:250,000, there are four major categories of elements, namely water systems, residential areas and facilities, transportation, and landforms, including 11 geographical elements such as agricultural and forestry units (310109), townships (310106), natural villages (310108), prefecture-level boundaries (640201), single-line rivers (210101), main canals (220200), lakes (230101), double-line rivers (210101), contour lines (710100), tractor roads (440100), provincial roads (420102), county roads (420301), rural roads (440200), and small paths (440300). The semantic matching process of the two element sets is as follows: (1) Check the basic geographic information DLG data to be stored in the same area in Oracle to ensure that the GB field of the national standard code and the CNAME field of the category name exist in the attribute layer of each element class to be stored; (2) Import each element class in the same area into the corresponding large- and small-scale element sets, and the naming format of the layer is "layer name code + category name + element class code + scale code"; (3) Load the data and connect to the SDE geographic database; (4) Through semantic matching between the large-scale 1:100,000 target object element set and the small-scale 1:250,000 matching object element set, the possible matching relationships of each target object are obtained; (5) Through semantic category matching, the elements in the small scale that match are obtained, including main canals (220200), lakes (230101), double-line rivers (210101), and single-line rivers (210101). After category matching, the target object and the alternative matching object are respectively geometrically matched; the target object lake in the large-scale 1:100,000 element set and the matching object lake in the small-scale 1:250,000 element set are as Figure 3 shown; (7) Analysis of the matching results: The matching objects selected according to semantic matching all belong to the same major category. From the geometric matching results of the target object selecting the matching object, it can be seen that if the two objects belong to different elements, there is no matching item; 2. Optimization of geometric matching of geospatial multi-scale data (1) Optimization of precise matching of surface entities Select two residential area data with similar scales in a certain area for matching optimization. Figure 4 (a) is the target data set with a scale of 1:10,000. This data set contains 160 residential area entities; Figure 4 (b) is the reference data set with a scale of 1:25,000. This data set contains 47 residential area entities. There are obvious position deviations between the homonymous entities, and the position deviations have the characteristic of non-uniformity. Some residential areas with small areas are completely separated from their homonymous entities, while some have smaller offsets. Visually, since the two data sets are of similar scales, there are one-to-one and one-to-many matching relationships, and there are a small number of residential area objects without homonymous entities.
[0060] Before the matching starts, a certain number of sample sets are selected for training, and the distance weight value is 0.2334, the shape weight value is 0.1211, the size weight is 0.6455, the geometric similarity critical value is 0.7216, and the maximum distance deviation is 14.069 meters. Figure 5 The figure is the matching result graph generated by using the method of this application. The two subgraphs on the left in the figure are respectively the enlarged views of the matching situation in a certain area in the figure.
[0061] From Figure 5 It can be seen that the relevant algorithms in the software correctly identify the one-to-one and one-to-many matching relationships in the case of obvious position offsets of the homonymous entities of the residential areas. The algorithm of this application correctly identifies all types of matching results. In terms of the matching speed, since there are fewer one-to-one matching relationships optimized this time, the one-to-many matching process in the algorithm of this application is used to identify the one-to-many matching relationships, so the matching speed is relatively fast, reaching 580 per second.
[0062] (2) Precise matching optimization of line entities Select the road data in a certain area for matching optimization. Figure 6 The left figure is the target data set with a scale of 1:2,000. This data set contains 196 road arcs; Figure 6 The right figure is the source data set with a scale of 1:50,000. This data set contains 112 road arcs. Figure 7 The figure is the superposition effect diagram of the target data set and the source data set. The upper left corner is the enlarged view of two roads in the figure. It can be seen from the figure that there are obvious position deviations between the two road entities. Visually, since the two data marts are of similar scales, there is a one-to-one matching relationship. In the matching parameters in the software interface, the Hausdorff distance is 150 meters, the interpolation point distance is 150 meters, and the buffer radii are all 150 meters.
[0063] When matching roads, their arcs are processed by stroke, and the matching results are displayed as whole roads. The target data consists of 196 arcs, and the source data consists of 112 arcs. After being processed by the stroke technology, the target data and the source data become 135 and 74 strokes respectively. By calculating the optimized results one by one under the spatial index of 30 by 30 grid index, it can be known that the one-to-one and one-to-many matching relationships can be correctly identified in the case of obvious position deviation of the roads. All types of matching results are correctly identified. The result analysis is as Figure 8 shown.
[0064] (3) Optimization of precise matching of point entities Select the entity data of residential locations with two similar scales in a certain area for matching optimization. Figure 9 The left figure is the target data set with a scale of 1:25,000, and this data set contains 288 residential location entities; Figure 9 The right figure is the reference data set with a scale of 1:50,000, and this data set contains 288 residential area entities. Figure 10 is the superimposed effect diagram of the target data set and the reference data set. The lower left corner is the enlarged view of a certain part of the figure. It can be seen from the enlarged view that there are obvious position deviations between the entities with the same name, and the position deviations have the characteristic of non-uniformity. Some residential areas with smaller areas are completely separated from their entities with the same name, and some have smaller offsets. Visually observed, since the two data sets have similar scales, there is a one-to-one matching relationship.
[0065] According to the maximum distance deviation between the point entities with the same name, the matching threshold of the point entities is set to 60 meters. The result obtained through geometric matching is as Figure 11 shown. The lower left corner is the partial enlarged view of the matching result.
Claims
1. An efficient matching method for GIS multi-scale basic geographic spatial data, characterized in that: Based on the spatial database construction model and spatial data matching, the method of geospatial multi-scale data matching is improved: first, a spatial database model based on relational database + spatial database engine is established as a method for storing multi-scale basic geospatial data; Instead, a new object selection method for multi-scale geospatial data matching is established. Based on the multi-scale semantic change rules of geospatial data, basic geographic entities are organized in the form of a multi-tree data structure, a category matching method for semantic changes of the same feature class at different scales is established, and a category matching relationship under multi-scale conditions is constructed; third, an optimization scheme for multi-scale basic geospatial data matching is established. On the basis of establishing a geospatial multi-scale database, spatial data pre-matching is achieved through a geospatial multi-scale data semantic object matching method, and then multi-scale spatial data matching software is used for matching optimization to achieve geospatial multi-scale data matching; Semantic object matching of geospatial multi-scale data: (1) In the basic geographic information of multiple scales, the national standard codes of basic geographic entities at a scale of 1:5000 are sorted and organized according to the semantic hierarchy, and each major category contains subcategories; (2) The organized 1:5000 basic geographic entity national standard codes are generated in the program in the form of a multi-branch tree data structure. Each geographic entity code is a node of the multi-branch tree. The root node is the name of the "land feature", which is located at the 0th level of the tree; the subclasses of the land feature are all subnodes of the "land feature" node, which are located at the 1st level of the tree; and the subclasses under the water system are analogous to each other, forming a "multi-branch tree" with a hierarchical structure; (3) If the code nodes of the two geographic entity feature classes to be matched are the same node in the multi-branch tree, or the two nodes to be matched have the same ancestor node, the level number of the ancestor node is greater than or equal to 1 and the level number difference between the two nodes to be matched is less than or equal to 2, this application stipulates that these two feature classes can be matched, otherwise they cannot be matched.
2. The GIS multi-scale basic geographic spatial data efficient matching method according to claim 1 is characterized in that: Classification and coding of basic geographic information data classification objects: data at all levels of scale are linked. When acquiring data, 1:1 million data is integrated from 1:250,000 data. Topographic elements are divided into 8 categories, namely positioning basis, water system, settlements and facilities, transportation, pipelines, boundaries and political divisions, landforms, soil and vegetation. Small and sub-class codes are divided into categories according to the three scales of 1:500, 1:1000, 1:2000, 1:5000-1:10,000, and 1:25,000-1:100,000. Among them, large and medium categories are not allowed to be newly defined or expanded, and small and sub-classes shall not be redefined. The terrain feature classification code is composed of major categories, medium categories, minor categories and subcategories, and consists of a 6-digit decimal code with the structure of major category code + medium category code + minor category code + subcategory code, where the major category code is composed of 8 numbers from 1 to 8. In the classification of basic geographic elements, the selection of major and medium category semantic relationships refers to the description form of the major and medium category semantic relationships, and determines the semantic relationships between the major and medium categories. All major and medium category semantic relationships are an abstract concept-level semantic relationship. The semantic relationships between minor and subcategories are inherited from the major and medium categories, and are the concretization of the major and medium category semantic relationships.
3. The GIS multi-scale basic geographic spatial data efficient matching method according to claim 1 is characterized in that: Geographic spatial multi-scale data semantic object matching: organize basic geographic entities in the form of a "multi-tree" data structure, and construct a method for matching all geographic entities that may occur due to cartographic generalization under multi-scale conditions. The specific process is as follows: (1) In the basic geographic information of multiple scales, the national standard codes of basic geographic entities at a scale of 1:5000 are sorted and organized according to the semantic hierarchy. Each major category contains subcategories. The first level of the semantic classification tree is the major category of land features, which is divided into eight major categories: positioning basis (100000), water (200000), residential areas and facilities (300000), transportation (400000), pipelines (500000), boundaries and political divisions (600000), landforms (700000), soil and vegetation (800000); the second level of the semantic tree is the water system (200000), which is divided into seven subcategories, including rivers (210000), ditches (220000), and water conservancy (230000). 00), lakes (230000), reservoirs (240000), ocean elements (250000), other water system elements (260000), water conservancy and ancillary facilities (270000); the third layer of the semantic classification tree, rivers (210000), is divided into four small categories, including perennial rivers (210100), seasonal rivers (210200), dry rivers (210300), and river notes (219000); the fourth layer of the semantic classification tree, perennial rivers (210100), is divided into four subcategories, including surface rivers (210101), underground river sections (210102), underground river section entrances and exits (210103), and disappeared river sections (210104); (2) The organized 1:5000 basic geographic entity national standard codes are generated in the program in the form of a multi-branch tree data structure. Each geographic entity code is a node of the multi-branch tree. The root node is the name of the "feature", which is located at the 0th level of the tree; the subclasses of the water system (200000), residential areas and facilities (300000) and transportation (400000) are all sub-nodes of the "feature" node, which are located at the 1st level of the tree; the subclasses under the water system are similar, forming a "multi-branch tree" with a hierarchical structure; (3) Scale changes cause changes in the semantic information of the same entity. The 1:5000 basic geographic entity multi-branch tree logical structure describes the most detailed semantic classification. Regardless of how the scale changes, the corresponding semantic information can be found in the multi-branch tree. If the code node of the two geographic entity feature classes to be matched is the same node in the multi-branch tree, or the two nodes to be matched have the same ancestor node, the level number of the ancestor node is greater than or equal to 1 and the level number difference between the two nodes to be matched is less than or equal to 2, this application stipulates that the two feature classes can be matched. In the first case, assuming that the two geographic entity feature classes to be matched are both rivers (210000), then the two feature classes can be matched; in the second case, if the first geographic entity feature class is a perennial river (210100), the second geographic entity feature class is a river. The class is perennial lake (230100). These two feature classes have a common ancestor water system (200000). The layer number of the water system (200000) node is greater than 1 and the difference in the number of layers of the two nodes to be matched is 0. According to the provisions of this application, these two feature classes can be matched; similarly, the feature classes national highway (420100) and provincial highway (420200) have a common ancestor intercity highway (420000). The layer number of the intercity highway node is 2 and the difference in the number of layers of the two nodes to be matched is less than or equal to 2. The national highway and the provincial highway can also be matched. On the contrary, the feature classes intercity highway (420000) and lakes (230000), although the nodes of these two feature classes have the same number of layers, their common ancestor "land feature" node is the root node, and these two feature classes cannot be matched.
4. The GIS multi-scale basic geographic spatial data efficient matching method according to claim 1 is characterized in that: Data preparation module: includes spatial data loading, matching method selection and parameter setting. The system supports two ways of loading: local data files (shapefile format) and feature sets in spatial databases. The corresponding matching method is selected and the corresponding matching parameters are set according to the different types of matched features.
5. The GIS multi-scale basic geographic spatial data efficient matching method according to claim 1 is characterized by: The matching module is entity matching, including semantic matching and geometric matching. Semantic matching uses the attribute items of the data in the feature set to determine which features belong to the same category as the source features according to the semantic matching algorithm, and then obtains the candidate matching entity set, and then performs entity geometric matching. According to the scale division, a total of 7 types of geometric matching are realized, namely, surface entity matching, surface and line entity matching, surface and point entity matching, line entity matching, line and point entity matching, point entity matching, and point group and surface entity matching.
6. The GIS multi-scale basic geographic spatial data efficient matching method according to claim 5 is characterized in that: The face entity matching method is based on the domain proximity and global optimization matching strategy, and the architecture is as follows: ① The geometric approximation of the surface entity is based on the geometric features of a single surface entity. The parameters of the geometric approximation model design include shape, size and position weights. The index value is determined before matching. It is obtained by selecting a determined sample set through training, quantifying each index, and achieving optimized matching results through continuous training and feedback; ② The fixed grid spatial index technology is introduced into the candidate matching set of matching face entities to reduce the scope of searching the candidate matching set. At the same time, the maximum position deviation and geometric similarity Zen value are used to further reduce the scale of candidate matching geometry to avoid missing entities with the same name. ③ Construct the neighborhood environment spatial structure of the face entity to lay the foundation for the neighborhood similarity evaluation. For the recognition of one-to-one relationship between face entities, the spatial similarity evaluation of face entities combines the geometric similarity and neighborhood similarity evaluation methods, taking into account both the similarity of individual geometric features of face entities and the similarity of the spatial structure corresponding to the face entities, ensuring the consistency between individual similarity and overall similarity. The global optimization method is used in the matching strategy to ensure the consistency between local similarity and global similarity, the consistency between individual and overall similarity, and the consistency between local and global similarity, reducing the impact of spatial position offset on matching; ④ For the recognition of one-to-many relationship between face entities, the intersection of the face entity to be matched and its minimum circumscribed rectangle is used as the marking point, and the alternative matching combination of the face entity to be matched is obtained by translation. The one-to-many matching relationship of the face entity is identified and constructed based on the overlapping similarity index.
7. The GIS multi-scale basic geographic spatial data efficient matching method according to claim 5 is characterized in that: Line entity matching: ① Road network data preprocessing: First, the road network data is processed by stroke connection. The roads are connected according to the good continuity rule in spatial cognition. Then, the connected strokes are encrypted by equidistant insertion points. Finally, a grid spatial index is established for the strokes, and the grid spatial index algorithm of line elements is used for processing. ② Overall matching of roads: The approximate buffer space index algorithm is used to find the candidate matching set, and then the geometric similarity calculation method between strokes is used to measure the geometric similarity. The improved spatial scene structure similarity evaluation method is used to enhance the matching effect of the algorithm. Then, the forward and reverse bidirectional matching algorithm is used to continue searching for matching objects in turn. Finally, multi-lane matching is used to handle the matching of multi-lane topological relationship changes caused by scale changes in the multi-scale road network map; ③ Partial matching of roads: Roads that have not found matching objects in the previous process need to enter the partial matching process. The partial matching process can deal with changes and updates between road entities with the same name due to inconsistent current conditions. Ring road matching can match changes in the topological structure of ring roads to further enhance the matching effect.
8. The GIS multi-scale basic geographic spatial data efficient matching method according to claim 5 is characterized in that: Matching of line entities and point entities: The distance from the point to the line is used for judgment. The minimum value of the vertical distance from the point to the line segment and the minimum value of the distance from the point to each vertex of the polyline are calculated to obtain the minimum value between the two. The minimum distance value is compared with the critical value. If the distance is less than the critical value, the match is successful; otherwise, the match is unsuccessful. Surface and point entity matching: In two data sets with large scale spans, point targets in smaller scale are searched for in large-scale surface features. That is, the algorithm for judging whether a point is in the surface target is used to determine whether the point is in the polygon. If there are multiple points in the surface target, the point with the smallest distance to the center of gravity of the surface target is selected.
9. The GIS multi-scale basic geographic spatial data efficient matching method according to claim 5 is characterized in that: Point entity matching: For the matching between single point targets, the spatial positions of the two points are compared. For point targets with the same name in different scale maps, their spatial positions should be consistent, that is, their spatial coordinates should be the same. The position proximity is used for matching, that is, the Euclidean distance between the two points is calculated. Assume that the reference point target and the point target to be matched are A and B, and the plane coordinates of the two points are (x A ,y A )(X B ,y B ), the distance d between the two points AB Use the following formula to calculate: The closer the distance between points A and B, the greater the AB The smaller the value, the greater the possibility of matching. If the point entity closest to point A is B, and d AB If it is less than the critical value and the point entity closest to point B is also A, the two points are considered to match, otherwise they are not matched.
10. The GIS multi-scale basic geographic spatial data efficient matching method according to claim 5 is characterized in that: Matching of large-scale point groups with small-scale surfaces: establish an expanded minimum area circumscribed rectangle for the surface entity on the small scale, and the expansion radius is the maximum error of the two data sets. If the point entity on the large scale falls within the rectangle, the point entity matches the large-scale surface entity.
Citation Information
Cited By
Satellite emergency communication rescue information grading generation and sending method, system and device
CN121984574A
School name normalization method based on multi-feature fusion and NLP semantic understanding
CN122287621A