Geographic information system (GIS) interest point space big data indexing method fused with semantic attributes

By combining Geohash encoding and attribute category encoding, the integration of grid and segmented hash indexes is solved, and efficient index construction and query performance is achieved.

CN120086199AInactive Publication Date: 2025-06-03孙智慧
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510127114.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-30
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When constructing spatial indexes of point-of-interest data, it is difficult to take into account both geographical coordinate information and attribute information, resulting in inefficient indexing performance and efficiency, and lack of effective classification and grading methods and coding schemes.

Method used

Geohash encoding method is used to convert the two-dimensional geographic coordinate information of the point of interest data into one-dimensional encoding, and combined with attribute category encoding, simplify encoding through Base32 encoding, integrate grid spatial index and segmented hash spatial index to build an efficient point of interest spatial index.

Benefits of technology

It improves the performance and execution efficiency of spatial indexes, shortens index construction time, reduces storage space usage, reduces the number of reads of external storage devices during query, and significantly improves the search and query speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086199A_ABST
    Figure CN120086199A_ABST
Patent Text Reader

Abstract

According to the semantic attribute fused GIS interest point space big data indexing method, starting from the three aspects of establishment of the space index, the size of the storage space occupied by the space index and retrieval and query of the interest point data based on the index, the existing space indexing method aiming at the interest point data is analyzed, and the retrieval and query efficiency of the interest point data is improved. On the basis of research and summarization of classification and grading of the interest point data, a set of spatial indexing method for coping with geographic spatial attributes and category information of the interest point data is established. The construction of the interest point spatial index improves the performance of managing and operating the interest point data in the geographic space database, and when the spatial index of the interest point data is constructed in the geographic space database, the geographic space attribute and the semantic attribute of the interest point data are considered as a whole, so that the spatial index performance of the interest point data is further improved; and the indexing speed and the accuracy are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a spatial indexing method for GIS points of interest, in particular to a spatial big data indexing method for GIS points of interest that integrates semantic attributes, and belongs to the technical field of spatial indexing of points of interest. Background Art

[0002] With the popularization of web map applications, the acquisition and application of spatial geographical data represented by points of interest have witnessed unprecedented development. Points of interest describe the point data of spatial geographical entities. Interest data includes two aspects: geographical coordinate information and attribute information. With the rise of location-based services, the application of point-of-interest data has received increasing attention. Abundant information can be abstracted from a large amount of interest data, providing important reference for users' decision-making, and even directly providing decisions, which has high application value and practical significance.

[0003] In a geographic information system, a spatial database undertakes the functions of storing and managing geographical information data. The spatial database is the sum of geographical spatial data related to applications stored on the computer physical storage medium of the geographic information system. Traditional database methods for data storage, management, and retrieval have many deficiencies when applied to spatial data, unable to meet the processing requirements of these data, thus forming the field of spatial databases.

[0004] With the continuous deepening of the construction of geographic information systems, the cumulative amount of point-of-interest data is also increasing continuously. In order to query and manage a large number and diverse types of point-of-interest data more efficiently, it is very necessary to apply a technology that can significantly improve the query performance of spatial databases, namely spatial indexing technology.

[0005] The spatial indexing technology for point-of-interest data organizes point-of-interest data in a certain order based on information such as the geographical coordinates of points of interest or based on the logical relationships between point-of-interest data. Due to the continuous growth of interest data volume and the increasingly wide application scope, traditional spatial indexing technologies show some deficiencies in operation efficiency. There is an urgent need to propose a more efficient spatial indexing method considering the geographical coordinate information and attribute information of point-of-interest data in terms of the establishment and operation efficiency of spatial indexing.

[0006] Spatial indexing of points of interest is a description of the storage location of point-of-interest data on hardware devices and is an effective method to improve the operation efficiency of spatial databases. For the spatial indexing technology of point-of-interest data, the existing technologies have achieved remarkable results in improving the operation efficiency of spatial databases, but so far, no optimal spatial indexing technology that adapts to various situations has been found. In particular, there is a lack of a spatial indexing that takes into account both the geographical coordinate information and attribute information of point-of-interest data.

[0007] The indexing methods of the prior art are all based on B-trees and hash indexes. These indexing methods are more suitable for linear structures. In spatial indexing methods, there are grid files and their variants, quadtrees and their variants, K-D trees and their variants, R-trees and their variants, etc. They can also be divided into two major categories: indexing structures based on point data and indexing structures based on non-point data. In addition to indexing structures, there are also some other indexing structures, such as bitmap indexes. The spatial indexing structure is an important part of the spatial database and also an important aspect affecting the performance of the spatial database. Due to the variety of application fields, it is difficult to find a general and perfect indexing structure to organize spatial data. Therefore, the spatial indexing structure is still a key research area in the current relevant fields.

[0008] After years of development, the spatial data indexing technology has developed a huge organizational system based on the original spatial indexing methods. There are a variety of indexing structures and indexing technologies for point-of-interest data, and different technical methods have significant differences in the operation efficiency of interest data. Some unpredictable factors will affect the efficiency of spatial indexing in practical applications, such as the hardware conditions adopted, the operating system status, the remaining size of the memory, and so on. In addition, there are also some direct influencing parameters, including the degree of data distribution uniformity, the applicability of the model, the spatial access method, the quantity of data, the distribution density of the data space, the degree of clustering, etc.

[0009] The problems to be solved by the prior art GIS point-of-interest spatial index and the key technical difficulties of this application include:

[0010] (1) Among the many interest data spatial indexing technical methods of the prior art, there are mainly the following problems: Some methods are simple but cannot handle relatively complex data objects; some methods can handle more complex spatial objects but greatly increase the number of disk accesses, and the indexing performance is greatly affected; some methods reduce the number of disk accesses but cannot adapt to high-dimensional data spaces; some methods can only handle the indexing of geographical space coordinates and cannot take into account the semantic attribute information of spatial objects; some methods are suitable for data spaces with relatively uniform distributions; some are more sensitive to the data order of the data. Therefore, it is difficult to find a spatial indexing method for point-of-interest data with excellent performance in all aspects. Therefore, it is still necessary to continue to conduct further research and discussion on the spatial indexing problem of point-of-interest data.

[0011] (2) The existing interest point data space indexing methods are carried out for the geospatial attributes of interest point data. However, when dealing with operations related to the semantic attributes of interest point data, the traditional database indexing methods are still used, which have problems such as slow speed and poor accuracy, and there is still room for improvement. Therefore, when constructing a space index for interest point data in a geospatial database, overall consideration of the geospatial attributes and semantic attributes of interest point data to further improve the space index performance of interest point data is a quite valuable and difficult problem. The existing technologies do not consider from three aspects: the establishment of the space index, the storage space size occupied by the space index, and the retrieval and query of interest point data based on the index. There is a lack of research on the classification and grading of interest point data, and there is no set of space index methods for dealing with the geospatial attributes and category information of interest point data.

[0012] (3) The existing technologies do not consider the geospatial characteristics of interest point data, lack the Geohash encoding method to encode the geographical coordinates of interest point data, do not transform the two-dimensional geographical coordinate information of interest point data into a one-dimensional space, and the performance and execution efficiency of the space index are low; lack hierarchical organization of the geographical coordinate information of interest point data and cannot retain the geospatial characteristics; the existing technologies lack a method for classifying and grading interest point data, do not combine the geographical coordinate encoding and attribute category encoding of interest point data, cannot obtain the encoding used to describe the geographical coordinate information and category attribute information of interest point data, lack an interest point space index method, the efficiency of the interest point space index is low, the index construction time is long, at the same time the index occupies a larger storage space, and the number of reads of the external storage device is large during query, and the retrieval and query speeds are slow. Summary of the Invention

[0013] Based on the geographical coordinate information and category attributes of interest point data, this application processes interest points based on Geohash geographical coding, transforms the two-dimensional geographical coordinate information of interest point data into a one-dimensional space, improves the performance and execution efficiency of the space index, and at the same time maintains the geospatial characteristics of interest point data; classifies and grades interest point data, and assigns unique encodings to each category of interest point data to improve the computer's management ability of the attribute information of interest point data, combines the geographical coordinate encoding and attribute category encoding of interest point data, and further simplifies the encoding based on Base32 encoding to obtain the encoding used to describe the geographical coordinate information and category attribute information of interest point data; integrates the grid space index and the segmented hash space index, constructs a space index using the newly obtained interest point data encoding, and at the same time uses online obtained interest point data to verify the new interest point space index. The new index method has a shorter index construction time, at the same time the index occupies a smaller storage space, the number of reads of the external storage device is less during index query, and the retrieval and query efficiency is higher.

[0014] To achieve the above technical effects, the technical solutions adopted in this application are as follows:

[0015] A semantic attribute integrated GIS point of interest (POI) spatial big data indexing method, starting from three aspects: the establishment of the spatial index, the storage space size occupied by the spatial index, and the retrieval and query of POI data based on the index. Based on the geographical coordinate information and category attributes of the POI data, the Geohash geocoding method is used to process the POI data, transforming the two-dimensional geographical coordinate information of the POI data onto a one-dimensional space while maintaining the geographical spatial characteristics of the POI data; the POI data is classified and graded, and each category of POI data is assigned a unique code;

[0016] The geographical coordinate encoding of the POI data is combined with the attribute category encoding, and the encoding is further simplified based on Base32 encoding to obtain an encoding for describing the geographical coordinate information and category attribute information of the POI data; the grid spatial index and the segmented hash spatial index are integrated, and the newly obtained POI data encoding is used to construct the spatial index; at the same time, online obtained POI data is used to cross-validate the new POI spatial index method.

[0017] Preferably, a set of spatial index methods for dealing with the geographical spatial attributes and category information of POI data is established:

[0018] 1) Based on the geographical spatial characteristics of the POI data, during the construction of the spatial index, the Geohash encoding method is used to encode the geographical coordinates of the POI data, transforming the two-dimensional geographical coordinate information of the POI data onto a one-dimensional space to improve the performance and execution efficiency of the spatial index; on the other hand, the geographical coordinate information of the POI data is hierarchically organized through the Geohash encoding method, retaining the geographical spatial characteristics;

[0019] 2) For the attribute characteristics of the POI data, based on the classification and grading methods of large Internet electronic map websites and national and industry standards, a method for classifying and grading the POI data is established;

[0020] 3) The geographical coordinate encoding of the POI data is combined with the attribute category encoding, and the encoding is further simplified using Base32 encoding to obtain an encoding for describing the geographical coordinate information and category attribute information of the POI data. Using the established encoding method, on the basis of integrating the advantages of the grid spatial index and the segmented hash spatial index, a POI spatial index method is constructed;

[0021] 4) Online obtained POI data is used to verify the new POI spatial index technical method.

[0022] Preferably, classify and grade the point-of-interest data: create a logical entropy for the point-of-interest data to accurately represent the point entities in the geospatial space, classify the point-of-interest data, summarize the entities with the same attributes and common features in the attribute information of the point-of-interest data, and separate the entities with different attributes and different features to form a multi-level and gradually expanding classification system;

[0023] Classify and grade the point-of-interest data. Classify the point-of-interest data according to different attribute categories of the point-of-interest data, divide the classified point-of-interest data into different levels, and determine the category, level, and subordination relationship between different point-of-interest data through the classification and grading of the point-of-interest data.

[0024] Preferably, encode the attributes of the point-of-interest data: encode based on the different attributes between different point-of-interest data, strengthen the organization of the point-of-interest data through attribute encoding, and provide convenience for retrieving the point-of-interest data at the same time. The encoding system includes the content of the encoding, the character format, and the character length. Assign corresponding attribute encodings to different levels and different types of point-of-interest data, and extract according to different encodings when retrieving the point-of-interest data. By encoding the attributes of the point-of-interest data, improve the operation efficiency of the point-of-interest data.

[0025] There are three forms of attribute encoding for the point-of-interest data: numeric attribute encoding, alphabetic attribute encoding, and mixed numeric and alphabetic attribute encoding.

[0026] Preferably, establish the rules for constructing the point-of-interest classification system:

[0027] 1) General rule: The point-of-interest classification starts from the basic needs of users. On the basis of meeting the general needs, further combine natural geographic information and human geographic information, and at the same time ensure that the final classification system is reasonable and clear;

[0028] 2) Consistency rule: The classification system of the point-of-interest data is consistent with the standards determined in the "Classification of Fundamental Geographic Information". Through classifying and encoding different types of geographic elements, a hierarchical geographic information classification structure is obtained, and finally a standard text for the classification of fundamental geographic information elements that covers all geographic information and is self-consistent is formed;

[0029] 3) Stability rule: The classification method of the point-of-interest data is based on the stable characteristics and attributes of different types of interest data, and a stable classification standard is obtained, so that the classification standard of the point-of-interest data remains stable;

[0030] 4) Completeness and expansion rule: Do not damage the integrity between the point-of-interest data sets. The point-of-interest classification system updates itself in a timely manner, and the classification system of the point-of-interest expands and updates to adapt to the update of the point-of-interest data;

[0031] 5) Public rules: POI data cannot involve sensitive content that endangers national public security.

[0032] Preferably, the geographic coordinates and the category attributes are jointly encoded: based on the spatial index based on the grid file and the segmented hash, the earth surface is segmented based on the Geohash encoding method, and the geographic spatial index of the point of interest data is constructed in combination with the classification and grading method of the point of interest data; firstly, the earth surface is alternately binary processed along the longitude and latitude directions, and the earth surface space is hierarchically divided into grids, and then a binary number (Geohash encoding) is used to represent the non-overlapping grids formed, and each string of Geohash binary codes represents a rectangular area surrounded by longitude and latitude lines on the earth's spherical surface;

[0033] On the other hand, the attribute information of the point of interest data within the rectangular area is encoded; in the classification system of the point of interest data, the number of categories under the first-level classification is set to 23, the number of categories under the second-level classification is set to 149, and the number of categories under the third-level classification is set to 537. The number of sub-categories under each individual classification does not exceed 32. Thirty binary digits are used to identify each level of sub-categories of the point of interest data. The two obtained codes are connected, and the generated new code simultaneously identifies the geographic coordinate code and attribute category of the point of interest data.

[0034] The 32 letters a, i, l, and o are removed from 0-9, bz for encoding. The letters i, l, and o are not included to avoid confusion with the numbers 1 and Ο. The letter a is also excluded to reduce the possibility of accidental confusion. The geographic coordinates of the point of interest data and the attribute category are mixed and encoded again using an encoding method to reduce the code length and further improve the indexing efficiency.

[0035] Compared with the prior art, the innovations and advantages of this application are:

[0036] (1) Based on the geographical coordinate information and category attributes of point-of-interest (POI) data, this application processes POIs using Geohash geocoding, transforms the two-dimensional geographical coordinate information of POI data into a one-dimensional space, improves the performance and execution efficiency of spatial indexing, and simultaneously maintains the geographical spatial characteristics of POI data; classifies and grades POI data, and assigns unique codes to each category of POI data to improve the computer's management ability of the attribute information of POI data. Combines the geographical coordinate coding and attribute category coding of POI data, and further simplifies the coding based on Base32 encoding to obtain a code used to describe the geographical coordinate information and category attribute information of POI data; integrates grid spatial indexing and segmented hash spatial indexing, constructs a spatial index using the newly obtained POI data code, and simultaneously uses online obtained POI data to verify the new POI spatial index. The new indexing method has a shorter index construction time, occupies less storage space, requires fewer read operations on external storage devices during index query, and has higher retrieval and query efficiency.

[0037] (2) Starting from three aspects: the establishment of spatial indexing, the storage space size occupied by spatial indexing, and the retrieval and query of POI data based on indexing, through the analysis of existing spatial indexing methods for POI data, combined with the research and summary of the classification and grading of POI data, this application establishes a set of spatial indexing methods for dealing with the geographical spatial attributes and category information of POI data. Constructing a POI spatial index improves the performance of geographical spatial database management and operation of POI data. When constructing a spatial index of POI data in a geographical spatial database, the geographical spatial attributes and semantic attributes of POI data are considered overall to further improve the spatial indexing performance of POI data, and both the indexing speed and accuracy are greatly improved.

[0038] (3) Based on the geospatial characteristics of the point-of-interest (POI) data, during the construction of the spatial index, the Geohash encoding method is used to encode the geographical coordinates of the POI data, transforming the two-dimensional geographical coordinate information of the POI data onto a one-dimensional space, thereby improving the performance and execution efficiency of the spatial index; on the other hand, the Geohash encoding method hierarchically organizes the geographical coordinate information of the POI data, preserving the geospatial characteristics; for the attribute characteristics of the POI data, based on the classification and grading methods of large-scale Internet electronic map websites and national and industry standards, a method for classifying and grading the POI data is established; the geographical coordinate encoding of the POI data is combined with the attribute category encoding, and the Base32 encoding is further used to simplify the encoding, obtaining the encoding for describing the geographical coordinate information and category attribute information of the POI data. Using the established encoding method, on the basis of integrating the advantages of the comprehensive grid spatial index and the segmented hash spatial index, a POI spatial index method is constructed, improving the efficiency of the POI spatial index; the POI data is obtained online to verify the new POI spatial index technology method. Experiments show that the new index method has a shorter index construction time, and at the same time, the index occupies less storage space. More importantly, when using this index method for querying, the number of reads from the external storage device is less, so the retrieval and query speed is faster. Brief Description of the Drawings

[0039] Figure 1 It is a schematic diagram of the first group of experiments for testing million-level POI data.

[0040] Figure 2 It is a schematic diagram of the test results of the second group of experiments on POI data.

[0041] Figure 3 It is a schematic diagram of the generation time (seconds) of the index in the second group of experiments.

[0042] Figure 4 It is a schematic diagram of the size (megabytes) of the index in the second group of experiments.

[0043] Figure 5 It is a schematic diagram of the number of index nodes traversed during the retrieval process in the second group of experiments. Detailed Embodiment

[0044] The following further describes the technical solution of the GIS POI spatial big data index method integrating semantic attributes provided by this application with reference to the drawings, so that those skilled in the art can better understand this application and be able to implement it.

[0045] Spatial indexing for point-of-interest (POI) data is an important part of the application field of POI data. POI data contains geographical spatial coordinate attribute information and semantic attribute information. The efficiency of the spatial indexing method directly affects the management operation performance of the geographical spatial database for POI data. To improve the query and retrieval speed, an efficient spatial indexing method must be found. Due to the characteristics of POI data itself, higher requirements are put forward for the technologies and methods of spatial indexing. How to efficiently construct the index of POI data and improve the retrieval and query processing algorithms of POI data are urgent problems to be solved.

[0046] Based on the geographical coordinate information and category attributes of POI data, this application uses the Geohash geographical coding method to process POI data, transforms the two-dimensional geographical coordinate information of POI data onto a one-dimensional space, improves the performance and execution efficiency of spatial indexing, and at the same time maintains the geographical spatial characteristics of POI data; and classifies and grades POI data, and assigns a unique code to each category of POI data to improve the computer's management ability of the attribute information of POI data.

[0047] Combining the geographical coordinate coding and attribute category coding of POI data, and further simplifying the coding based on Base32 coding, a code used to describe the geographical coordinate information and category attribute information of POI data is obtained; integrating the grid spatial index and the segmented hash spatial index, and constructing a spatial index using the newly obtained POI data code to improve the efficiency of the POI spatial index and enrich the POI data spatial index; at the same time, obtaining POI data online to verify the new POI spatial indexing method. Experiments show that the index construction time of this new indexing method is shorter, and at the same time the index occupies less storage space. More importantly, when using this indexing method for query, the number of reads of the external storage device is less, so the retrieval and query speed is faster.

[0048] I. Semantic Attribute Information of Point of Interest

[0049] In addition to the geometric information describing the geographical spatial location, geographical spatial data also includes a lot of text description information. In addition to the longitude and latitude coordinate information determining the geographical location of the point, POI data information also includes name, text address information, contact phone number, label, street view photo information, and corresponding attribute information for specific types of POI data. For example, general consumer places such as hotels and restaurants usually have per capita consumption attributes, and general business stores usually have contact phone number attributes.

[0050] Different users have different requirements when applying point-of-interest (POI) data, and geographical features such as the classification method of POI data and attribute coding are also different. Even for the same POI data, there are significant differences in the representation methods of its classification features and attribute information. Therefore, it is necessary to process according to the classification methods and attribute coding methods for different types of POI data.

[0051] 1. Classification and grading of POI data

[0052] Create a logical entropy for POI data to accurately represent point entities in the geographical space, classify the POI data, summarize entities with the same attributes and common features in the attribute information of POI data, and separate entities with different attributes and different features to form a multi-level and gradually expanding classification system.

[0053] Classify and grade POI data. Classify POI according to different attribute categories of POI data, divide the classified POI data into different levels, and determine the category, level, and subordination relationship between different POI data through the classification and grading of POI data.

[0054] 2. Attribute coding of POI data

[0055] Encode based on the different attributes between different POI data. Strengthen the organization of POI data through attribute coding, and at the same time facilitate the retrieval of POI data. The coding system includes the content of the code, character format, and character length. Assign corresponding attribute codes to different levels and different types of POI data, and extract according to different codes when retrieving POI data. Improve the operation efficiency of POI data by encoding the attributes of POI data.

[0056] There are three forms of attribute coding for POI data: numeric attribute coding, alphabetic attribute coding, and mixed numeric and alphabetic attribute coding.

[0057] Rules for constructing the POI classification system:

[0058] 1) General rule: POI classification starts from the basic needs of users, provides convenience for the daily life of the public, and on the basis of meeting universal needs, further combines natural geographical information and human geographical information, while ensuring that the final classification system is reasonable and clear;

[0059] 2) Consistency rule: The classification system of POI data is consistent with the standards determined in the "Classification of Fundamental Geographical Information". Through classifying and coding different types of geographical elements, a hierarchical geographical information classification structure is obtained, and finally a standard text for the classification of fundamental geographical information elements that covers all geographical information and is self-consistent is formed;

[0060] 3) Stable rules: The classification method of interest point data is based on the stable characteristics and attributes of different categories of interest data to obtain a stable classification standard, so that the classification standard of interest point data remains stable;

[0061] 4) Completeness and extension rules: The integrity of the POI data sets should not be destroyed. The POI classification system should update itself in a timely manner. The POI classification system should be expanded and updated to match the update of POI data.

[0062] 5) Public rules: POI data cannot involve sensitive content that endangers national public security.

[0063] 2. Joint encoding of geographic coordinates and category attributes

[0064] Based on the spatial index based on grid files and segmented hashing, the earth's surface is segmented based on the Geohash coding method, and the geographic spatial index of the point of interest data is constructed by combining the classification and grading method of the point of interest data; first, the earth's surface is alternately binary processed along the longitude and latitude directions, and the earth's surface space is layered into grids, and then a binary number (Geohash coding) is used to represent the formed non-overlapping grids, and each string of Geohash binary codes represents a rectangular area surrounded by longitude and latitude lines on the earth's spherical surface.

[0065] On the other hand, the attribute information of the point of interest data within the rectangular area is encoded; in the classification system of the point of interest data, the number of categories under the first-level classification is set to 23, the number of categories under the second-level classification is set to 149, and the number of categories under the third-level classification is set to 537. The number of sub-categories under each individual classification does not exceed 32. Thirty binary digits are used to identify each level of sub-categories of the point of interest data. The two obtained codes are connected, and the generated new code simultaneously identifies the geographic coordinate code and attribute category of the point of interest data.

[0066] The 32 letters a, i, l, and o are removed from 0-9, bz for encoding. The letters i, l, and o are not included to avoid confusion with the numbers 1 and Ο. The letter a is also excluded to reduce the possibility of accidental confusion. The geographic coordinates of the point of interest data and the attribute category are mixed and encoded again using an encoding method to reduce the code length and further improve the indexing efficiency.

[0067] 3. Process Framework and Experiments

[0068] 1. Data Acquisition

[0069] As the main providers of Internet electronic maps, Amap and Baidu Map possess extremely rich interest data information. Although these electronic map websites have abundant point-of-interest (POI) data information, their data is not fully open. Currently, most online electronic map manufacturers use tile transmission technology to display on the browser side and cannot directly obtain the vector data of POI information from the electronic map websites. However, large Internet electronic map manufacturers such as Amap and Baidu Map have specifically provided map development tools suitable for various terminals for developers. Therefore, the JavaScript APIs of these electronic map manufacturers are used to obtain the interest information in the database.

[0070] In addition to the POI data provided by online electronic map manufacturers, group-buying websites represented by Meituan also have a vast amount of POI data information and simultaneously provide JavaScript API service interfaces to obtain their POI data. Compared with the POI data provided by online electronic map manufacturers, the POI data of review websites is more real-time and also has richer attribute information, including the name of a store, address, whether there is group-buying information, supported transaction types, per capita price, per capita score, tag types, special features, member discounts, evaluation information, etc.

[0071] The POI data format extracted from the Internet online is relatively mixed and may have co-reference ambiguities. Therefore, before further processing, the obtained POI data is initially fused to obtain a POI data set without logical errors, and then it enters the formal stage.

[0072] 1. Obtain interest data from online maps

[0073] Call the API function of the electronic map manufacturer itself to obtain the POI data information. Taking Amap as an example below, the process of using the API function interface service released by it to obtain detailed POI data is described in detail.

[0074] The service class interface of the JavaScript API provided by Amap for developers is used to extract the POI data information of Amap online. The service class interface provides location-based retrieval services. The first step in extracting POI data is to determine the coordinate range of the extraction target area and retrieve relevant interests by specifying keywords or not specifying keywords for each coordinate according to the coordinate range, and obtain the returned POI data information.

[0075] The key to online extraction of point of interest (POI) data is to search block by block at a certain step size. The upper limit of the number of POIs returned in one search in Amap is 1,000. Therefore, when the step size is set too large, POI data in the target area may be missed; on the contrary, when the step size is set too small, the burden will be increased and the efficiency of online extraction will be reduced. Therefore, appropriate step sizes should be set for different target areas and different types of POI data respectively.

[0076] Amap provides two service class interfaces for developers to extract POI data online. One is to obtain POI data using the interest attributes of the ReGeocodeResult object of the Geocoder class interface. The Geocoder class returns the interest information near the point while the user is performing address parsing, and its input parameter is only the coordinates of a retrieval center point. Another method is to use the PlaceSearch interface provided by Amap. By using the search method of this interface, search for POI data of a specified type within a specified geographical range and return detailed information in the interest List object. The PlaceSearch function is used for location retrieval, surrounding area retrieval, and range retrieval, and its input parameters are the keyword for retrieving POI data and the center coordinates of the retrieval point. When retrieving using the PlaceSearch interface, the keyword type is selected from 20 first-level classifications of POI data in Amap, and the default classification is "Food & Beverage Service|Commercial & Residential|Life Service". The results returned by both methods are POI data near the retrieval point, but the POI data information returned by the first method is abbreviated information, and a large amount of attribute information has been deleted by Amap, only the longitude and latitude coordinates, the city where it is located, and the name information of the POI are retained; while in the second method, by setting the property value of the parameter extensions of the PlaceSearch interface to "all", the detailed attribute information of the POI provided by Amap can be obtained.

[0077] Build a system for scraping POI data information. The system uses the JavaScript API of Amap to scrape POI data block by block and save it to a JSON file.

[0078] Similar to Amap, Baidu Map, Sogou Map, and review websites such as Dianping, which have rich POI information, also provide JavaScript APIs for developers. The above methods can also be applied to scraping POI data to build a system to scrape POI data.

[0079] 2. POI Data Fusion

[0080] For POI data from different sources and with different standards, there are contradictions in the internal geospatial logic. Before storing it in the database, co-reference disambiguation needs to be used to fuse the co-referring objects in the POI data from different sources.

[0081] Co-referring objects are the same phenomenon or entity described in different materials and data, and co-reference ambiguity in names also exists in POI data. By initially fusing the POI data, co-reference ambiguity can be effectively eliminated.

[0082] Using the semantic matching method, the geographical coordinates and names of POI data from different data sources are matched for similarity, and the matched POI data is fused to eliminate co-reference ambiguity.

[0083] During the process of fusing POI data, only the internal logical contradictions in the POI data set are eliminated, and co-reference ambiguity is removed, while the consideration of the accuracy of POI data is ignored.

[0084] (2) Experiment

[0085] Based on obtaining a large amount of available POI data, first store it in the database for preparation, and use a B-tree to construct a joint index for the three attributes of longitude and latitude coordinates and POI classification as a control.

[0086] In the experiment, in order to verify the feasibility and rationality of the method proposed in this application, the experiment was carried out separately under different data volumes. At the same time, considering that in the actual production environment, there will be some POI data sets with relatively large data volumes, and the retrieval effect of the geospatial database is closely related to the size of the stored data volume. Since the increase in data may lead to a sharp change in the amount of computation during retrieval, which will affect the operation performance of the geospatial database for POI. Therefore, the data packets used in the experiment include POI data at three different data levels, namely POI with a data volume of one million, ten million, and one hundred million. Multiple experiments were carried out and analyzed and summarized from the time of establishing the spatial index and the retrieval efficiency of POI data. When evaluating the performance of the geospatial index, there are three important indicators. One of the most important performance indicators is the time of establishing the spatial index. Another important indicator is the number of accesses to the external storage device when using this spatial index for retrieval. There is also an important indicator which is the size of the storage space occupied by the spatial index. The above three aspects are important performance indicators for measuring the quality of the spatial index. Therefore, in the experiment, the time spent on constructing the index, the number of disk accesses when retrieving POI data, and the size of the storage space occupied by the spatial index are selected as the criteria for measuring the experimental results.

[0087] The first experiment is to test the POI data at the million level to verify the feasibility of the experimental index method. From Figure 1From the experimental results, it can be seen that without an index, when obtaining target data from a database storing millions of points of interest data, the entire dataset in the database needs to be traversed, resulting in relatively low efficiency. After building a query index for the database, the retrieval efficiency has been greatly improved.

[0088] The second experiment is a comparison between the experimental indexing method and the control indexing method, testing the indexing of points of interest data with data volumes of millions, tens of millions, and hundreds of millions. In the experiment, points of interest with different levels of data volume were used to build the index. Part of the data was extracted from the dataset to build it, and the average value of several times was used as the comparison data. The experimental results are as Figure 2 shown.

[0089] From Figure 3 and Figure 4 it can be seen that compared with the control group, the new indexing method has certain improvements in the index generation time and the storage space occupied by the index. Figure 5 The analysis in

[0090] shows that when using the new index for retrieval queries, the number of nodes traversed within the index structure is significantly reduced, which also indicates that the new indexing method organizes the points of interest data well. Therefore, the new indexing method also has a relatively high improvement in the data query speed. Process framework and experiment: First, use the Javascript API interface service provided by Internet electronic map manufacturers to capture the online points of interest data they provide. After capturing a large amount of points of interest data, fuse these points of interest data from different sources, enrich the data used as much as possible and eliminate the co-reference ambiguity between the data. Finally, store the fused points of interest data in the database for experiments. By comparing the retrieval efficiencies of different indexing methods, the feasibility of using the encoding method that combines geographical coordinate attributes and category attributes based on the Geohash encoding method to build a spatial index to further improve the operation efficiency of the geographical space database for points of interest data is verified.

Claims

1. A spatial big data indexing method for GIS points of interest integrating semantic attributes, characterized in that: Starting from the establishment of spatial index, the size of storage space occupied by spatial index, and the retrieval query of interest point data based on index, the Geohash geographic coding method is used to process interest point data based on the geographic coordinate information and category attributes of interest point data, transforming the two-dimensional geographic coordinate information of interest point data into one-dimensional space while maintaining the geographic spatial characteristics of interest point data; classifying and grading interest point data, and giving each category of interest point data a unique code; The geographic coordinate encoding of the point of interest data is combined with the attribute category encoding, and the encoding is further simplified based on Base32 encoding to obtain the encoding used to describe the geographic coordinate information and category attribute information of the point of interest data; the grid spatial index and the segmented hash spatial index are integrated, and the spatial index is constructed using the newly obtained point of interest data encoding; at the same time, the point of interest data obtained online is used to cross-validate the new point of interest spatial indexing method.

2. According to the method for indexing spatial big data of GIS points of interest with fusion of semantic attributes as described in claim 1, it is characterized in that: Establish a set of spatial indexing methods to deal with the geospatial attributes and category information of point of interest data: 1) Based on the geographic spatial characteristics of the POI data, the Geohash encoding method is used to encode the geographic coordinates of the POI data in the process of building the spatial index, and the two-dimensional geographic coordinate information of the POI data is converted into a one-dimensional space, thereby improving the performance and execution efficiency of the spatial index; on the other hand, the Geohash encoding method is used to hierarchically organize the geographic coordinate information of the POI data and retain the geographic spatial characteristics; 2) Based on the attribute characteristics of the POI data and the classification and grading methods of large Internet electronic map websites and national and industry standards, a method for classifying and grading the POI data is established; 3) The geographic coordinate coding of the point of interest data is combined with the attribute category coding, and the Base32 coding is used to further simplify the coding to obtain the coding used to describe the geographic coordinate information and category attribute information of the point of interest data. The established coding method is used to construct the point of interest spatial index method based on the advantages of the comprehensive grid spatial index and the segmented hash spatial index; 4) The new POI spatial indexing technology method is verified by using online POI data.

3. According to the method for indexing spatial big data of GIS points of interest with fusion of semantic attributes as described in claim 1, it is characterized in that: Classification and grading of POI data: Create logical entropy for POI data so that POI data can accurately express point entities in geographic space, classify information of POI data, summarize entities with the same attributes and common characteristics in the attribute information of POI data, separate entities with different attributes and characteristics, and form a multi-level and gradually expanding classification system; Classify and grade the interest point data, classify the interest points according to different attribute categories of the interest point data, divide the classified interest point data into different levels, and determine the categories, levels and subordinate relationships between different interest point data through classification and grading of the interest point data.

4. According to the method for indexing spatial big data of GIS points of interest with fusion of semantic attributes as claimed in claim 1, it is characterized in that: POI data attribute coding: Encoding is performed based on different attributes between different POI data. POI data organization is strengthened through attribute coding, and POI data retrieval is facilitated. The coding system includes the content, character format, and character length of the coding. Different levels and types of POI data are assigned corresponding attribute codes. When searching for POI data, different codes are used for extraction. By encoding the attributes of POI data, the efficiency of POI data operation is improved. There are three types of attribute encoding for POI data: digital attribute encoding, letter attribute encoding, and mixed attribute encoding of numbers and letters.

5. According to the method for indexing spatial big data of GIS points of interest with fusion of semantic attributes as claimed in claim 4, it is characterized in that: Rules for building a POI classification system: 1) General rules: The classification of points of interest starts from the basic needs of users. On the basis of meeting the general needs, it further combines the natural geographical information and the human geographical information, and at the same time ensures that the final classification system is reasonable and clear; 2) Consistent rules: The classification system of POI data is consistent with the standards set in the Basic Geographic Information Classification. By classifying and coding different types of geographic elements, a hierarchical geographic information classification structure is obtained, and finally a basic geographic information element classification standard text that covers all geographic information and is coordinated and unified is formed; 3) Stable rules: The classification method of interest point data is based on the stable characteristics and attributes of different categories of interest data to obtain a stable classification standard, so that the classification standard of interest point data remains stable; 4) Completeness and extension rules: The integrity of the POI data sets should not be destroyed. The POI classification system should update itself in a timely manner. The POI classification system should be expanded and updated to match the update of POI data. 5) Public rules: POI data cannot involve sensitive content that endangers national public security.

6. According to the method for indexing spatial big data of GIS points of interest with fusion of semantic attributes as claimed in claim 1, it is characterized in that: Joint encoding of geographic coordinates and category attributes: Based on the spatial index based on grid files and segmented hashing, the earth's surface is segmented based on the Geohash coding method, and the geospatial index of the point of interest data is constructed by combining the classification and grading method of the point of interest data; first, the earth's surface is binary processed alternately along the longitude and latitude directions, and the earth's surface space is layered into grids. Then, a binary number (Geohash coding) is used to represent the non-overlapping grids formed. Each string of Geohash binary codes represents a rectangular area surrounded by longitude and latitude lines on the earth's spherical surface; On the other hand, the attribute information of the point of interest data in the rectangular area is encoded; in the classification system of the point of interest data, the number of categories under the first-level classification is set to 23, the number of categories under the second-level classification is set to 149, and the number of categories under the third-level classification is set to 537. The number of sub-categories under each individual classification does not exceed 32. Thirty binary digits are used to identify each level of sub-categories of the point of interest data. The two obtained codes are connected, and the generated new code simultaneously identifies the geographic coordinate code and attribute category of the point of interest data; The 32 letters a, i, l, and o are removed from 0-9, bz for encoding. The letters i, l, and o are not included to avoid confusion with the numbers 1 and Ο. The letter a is also excluded to reduce the possibility of accidental confusion. The geographic coordinates of the point of interest data and the attribute category are mixed and encoded again using an encoding method to reduce the code length and further improve the indexing efficiency.

Citation Information

Patent Citations

  • Method and device for establishing incidence relation between point of interest and image corresponding to point of interest

    CN102080963A

  • GeoHash-based map visual range interest point retrieval method and system

    CN115309850A

  • Full-text retrieval method and system about points of interest

    CN115329035A

  • System and method for determining exact location results using hash encoding of multi-dimensioned data

    US20120226889A1