HBase spatio-temporal data storage indexing method and system based on spatio-temporal short codes
By adopting a storage indexing method based on space-time short codes in the HBase database, compact time slice short codes and space short codes are generated, which solves the problem of inefficient query of massive dynamic data and achieves efficient storage and query performance.
Patent Information
- Application Number
- CN202411966301.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-27
AI Technical Summary
The Rowkey design method of the existing HBase database cannot effectively support the time series calculation of massive dynamic data, resulting in inefficient storage query.
The HBase spatiotemporal data storage indexing method based on spatiotemporal short code is adopted. By generating time slice short codes and space short codes, the storage space occupation is reduced, the cache utilization efficiency is improved, and query efficiency is improved through various data partitioning models and indexing mechanisms.
It significantly improves storage query efficiency, reduces storage overhead, avoids data skew and hot issues, and realizes efficient management and fast access to massive spatio-temporal data.
Smart Images

Figure CN120045559A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular, to a method and system for storing and indexing spatio-temporal data of HBase based on spatio-temporal short codes. Background Art
[0002] In recent years, with the development of earth observation, satellite remote sensing, ecological assessment, and land supervision towards the macroscopic, dynamic, and refined directions, the update speed and accuracy requirements of basic geographic information data have become higher and higher, thus forming a large amount of multi-temporal and high-precision geographic information data. In the development and use of big data, high-performance and large-volume queries are essential. In the current environment, the database that meets the characteristics of high performance and large data volume is none other than Hbase. Hbase is a widely used nosql (not only sql) database, and it is widely used in the field of big data because of its large data storage capacity and the characteristics of accessing partition servers.
[0003] Chinese Patent with Publication No. CN113407518B discloses a method and device for designing the Rowkey of an Hbase database. The method includes: determining the number of partitions required for the data to be stored; determining the first half of the Rowkey corresponding to the data to be stored according to the current time and the number of partitions; generating a discrete random universally unique identifier and determining it as the second half of the Rowkey corresponding to the data to be stored; integrating the first half and the second half to obtain the designed value of the Rowkey. However, the solution provided by the above application only designs algorithms separately according to specific business requirements and data storage forms, and cannot provide time series calculation support for a large amount of dynamic data. Therefore, it is very necessary to provide a method and system for storing and indexing spatio-temporal data of HBase based on spatio-temporal short codes to improve the storage and query efficiency. Summary of the Invention
[0004] In view of this, the present invention proposes a method and system for storing and indexing spatio-temporal data of HBase based on spatio-temporal short codes. Through the generation mechanism of time slice short codes and space short codes, the occupation of storage space is greatly reduced, and shorter strings are used to form row keys, thereby saving a large amount of key value storage space and improving the utilization efficiency of the cache, thus improving the storage and query efficiency.
[0005] The present invention provides a method and system for storing and indexing spatio-temporal data of HBase based on spatio-temporal short codes, and the method includes:
[0006] Generating time slice short codes and space short codes respectively according to time information and spatial grids;
[0007] Combine the time slice short code, spatial short code, object identifier, and unique identifier to construct multiple data partitioning models, and pre-partition the amount of data required to be stored in the Hbase database according to the data partitioning models;
[0008] Perform spatial indexing or row key indexing according to the row key, time slice short code, and spatial short code in the multiple data partitioning models to obtain index results.
[0009] Based on the above technical solutions, preferably, the generation of the time slice short code and the spatial short code according to the time information and the spatial grid specifically includes:
[0010] Convert the time information into a binary code according to a preset number of digits, and group and map the binary code into visible characters;
[0011] Convert the spatial grid coordinates into visible characters by using an octal pyramid coding method, where the spatial grid coordinates include the x-axis coordinate value and the y-axis coordinate value.
[0012] Based on the above technical solutions, preferably, the multiple data partitioning models include a spatial data model, a time series data model, a spatio-temporal data model, and a cyclic spatio-temporal data model. The spatial data model includes a spatial short code and a unique identifier. The time series data model includes a time slice short code and an object identifier. The spatio-temporal data model includes a spatial short code, a time slice short code, and an object identifier. The cyclic spatio-temporal data model includes a time slice short code, a spatial short code, and an object identifier.
[0013] More preferably, the time slice short code represents a time code set based on the partition capacity and the data generation period. The object identifier represents the primary key corresponding to the object body that continuously and uniformly generates data in chronological order. The spatial short code represents a spatial code that encodes the grid where the spatial data is located after being divided according to the spatial grid. The unique identifier represents the primary key originally carried by the data information and having uniqueness.
[0014] More preferably, the conversion of the time information into a binary code according to a preset number of digits and the grouping and mapping of the binary code into visible characters specifically include:
[0015] Convert the time information into a binary code to obtain a binary sequence;
[0016] Group the binary sequence according to a preset grouping condition to obtain a first binary sequence;
[0017] Perform sequence conversion on the first binary sequence to obtain visible characters corresponding to the first binary sequence.
[0018] More preferably, the method for converting the spatial grid coordinates into visible characters by using the octal pyramid coding method specifically includes:
[0019] Confirm the spatial grid coding order, and obtain the x-axis coordinate value and y-axis coordinate value according to the spatial range data;
[0020] Convert the x-axis coordinate value and the y-axis coordinate value into octal respectively, determine the number of digits according to the maximum value and pad with zeros for alignment to obtain the octal spatial coding;
[0021] Cross-combine the octal spatial coding bit by bit to generate a second binary sequence, and perform sequence conversion on the second binary sequence to obtain visible characters corresponding to the second binary sequence.
[0022] More preferably, the method further includes:
[0023] Construct a decimation cache layer, and perform downsampling on the sampling data according to the scaling ratio of the spatio-temporal scale, where the sampling data includes high-precision map data, perception data collected by sensors, positioning data reported by intelligent driving vehicles, and corresponding spatio-temporal data;
[0024] Construct an aggregation cache layer, and perform statistical analysis on the sampling data within the specified time range and within the spatial grid.
[0025] In the second aspect of the present application, there is provided an HBase spatio-temporal data storage index system based on spatio-temporal short codes. The HBase spatio-temporal data storage index system includes a short code generation module, a data partitioning module, and an information indexing module, where
[0026] The short code generation module is used to generate a time slice short code and a spatial short code respectively according to time information and a spatial grid;
[0027] The data partitioning module is used to combine the time slice short code, the spatial short code, the object identifier, and the unique identifier to construct multiple data partitioning models, and pre-partition the data volume required to be stored in the Hbase database according to the data partitioning models;
[0028] The information indexing module is used to perform spatial indexing or row key indexing according to the row key, the time slice short code, and the spatial short code in the multiple data partitioning models to obtain an index result.
[0029] In the third aspect of the present application, there is provided an electronic device, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory.
[0030] In the fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and the computer program is executed by a processor to implement the steps of a method for storing and indexing spatio-temporal data in HBase based on spatio-temporal short codes.
[0031] The method and system for storing and indexing spatio-temporal data in HBase based on spatio-temporal short codes provided by the present invention have the following beneficial effects compared with the prior art:
[0032] (1) Through the generation mechanism of time-slice short codes and space short codes, the occupation of storage space is significantly reduced. The short code form is more compact than the complete spatio-temporal information, reducing the storage overhead. Using shorter strings to form row keys saves a large amount of key value storage space, improves the utilization efficiency of the cache, and thus improves the storage query efficiency. Using spatio-temporal short codes as indexes improves the retrieval efficiency of spatio-temporal data, and the multi-dimensional index mechanism makes the query more flexible, capable of meeting the query requirements of different scenarios. By constructing multiple data partitioning models, the problem of data skew is avoided, and the pre-partitioning mechanism ensures the uniform distribution of data in HBase, preventing hot spot problems in the HBase database.
[0033] (2) By constructing a two-layer cache structure of a thinning cache layer and an aggregation cache layer, the efficient management and rapid access of massive spatio-temporal data are realized. The thinning cache layer realizes the on-demand thinning of data through a downsampling mechanism, effectively reducing the storage and processing pressure, while the aggregation cache layer provides rapid data insight capabilities through pre-statistical analysis. The two layers work together not only significantly improves the system performance and resource utilization efficiency, but also provides a flexible data access method for upper-layer applications, which is especially suitable for processing the large-scale spatio-temporal data analysis requirements in intelligent transportation scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1 It is a schematic flowchart of the method and system for storing and indexing spatio-temporal data in HBase based on spatio-temporal short codes provided by the present invention;
[0036] Figure 2 It is a schematic diagram of the conversion of time-slice short codes provided by the present invention;
[0037] Figure 3 It is a schematic diagram of the conversion of an example of a time-slice short code provided by the present invention;
[0038] Figure 4 Coding schematic diagram of the TMS slice provided by the present invention;
[0039] Figure 5 Coding schematic diagram of the WMTS slice provided by the present invention;
[0040] Figure 6 Conversion schematic diagram of the octal space short code provided by the present invention;
[0041] Figure 7 Framework schematic diagram of the HBase spatio-temporal data storage index system provided by the present invention;
[0042] Figure 8 Structural schematic diagram of the electronic device provided by the present invention.
[0043] Explanation of reference numerals: 1, HBase spatio-temporal data storage index system; 11, short code generation module; 12, data partitioning module; 13, information indexing module; 2, electronic device; 21, processor; 22, communication bus; 23, user interface; 24, network interface; 25, memory. Detailed implementation manners
[0044] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] Before introducing the embodiments of the present invention, first, some terms and their abbreviations involved in the embodiments of the present invention are defined and explained.
[0046] Spatial data files in professional formats: such as shp files, cad format files, opendrive files, etc. Generally, static basic geographic information, building models, road models, etc. commonly use this method. This storage method is suitable for a small amount of data, focuses on the geometric shape of the data, and can also store other general types of attribute fields. Because it is a unique file format, it is difficult to exchange data with other software systems.
[0047] Relational spatial databases: such as Oracle spatial, Postgres (PostGIS), etc., are generally used for large-scale professional geographic information production or scenarios that require professional spatial statistics and analysis. This storage method can provide storage for a large amount of data and also provides storage for more abundant other attribute fields. Since the spatial database itself also provides algorithm support for professional spatial queries and spatial analysis, it can support relatively professional GIS industry applications.
[0048] NoSQL databases: such as HDFS files (geojson), HBase, mongoDB, etc., are mostly used for specific scenarios of massive data storage and query, such as the collection and storage of new energy vehicle positioning data. This storage method mainly solves the problem of storing massive spatial data that cannot be effectively processed by relational databases. However, the spatial information in the data is generally converted to strings for storage, lacking support for general-purpose spatial algorithms, and the query method is also greatly restricted according to the characteristics of the database.
[0049] The present invention discloses an HBase spatio-temporal data storage index method based on spatio-temporal short codes, referring to Figure 1 , and the steps of this method include S1 to S3.
[0050] Step S1, generate a time slice short code and a spatial short code respectively according to the time information and the spatial grid.
[0051] In this step, it also includes S11 to S12.
[0052] Step S11, convert the time information into a binary code according to a preset number of digits, and map the binary code into visible characters after grouping.
[0053] In this embodiment, convert the time information into a binary code to obtain a binary sequence; group the binary sequence according to a preset grouping condition to obtain a first binary sequence; perform sequence conversion on the first binary sequence to obtain visible characters corresponding to the first binary sequence.
[0054] Usually, in order to represent the year, month, and day, a string in the form of "yyyy-MM-dd" or "yyyyMMdd" is generally used, which requires 8 - 10 bytes. If the hour, minute, and second are also considered and a string in the form of "hh:mm:ss" or "hhmmss" is used, then another 8 - 10 bytes are required. This takes up too much space for the rowkey whose best design length should be within 16 bytes. For example Figure 2As shown, the design goal of the time slice short code is to use 3 bytes to represent the year, month, and day within a range of 500 years. Additionally, if further representation of hours, minutes, and seconds is required, one byte is extended for each level in sequence. That is, a maximum of 6 bytes can divide the time slice in seconds.
[0055] Year: 9 bits in binary, with the value being the upper limit of time (default is 2500 years) minus the current year, and it can go back a maximum of 511 years. For example, 2500 - 2024 = 476, which is 110011100 in binary.
[0056] Month: 4 bits in binary, with the value being the current month, ranging from 0 - 15. For example, May is 0101. When the time slice is in years, the month is set to 0.
[0057] Day: 5 bits in binary, with the value being the date of the current month, ranging from 0 - 31. For example, the 31st is 11111. When the time slice is in years and months, the date is set to 0.
[0058] Hour: Extended item, 6 bits in binary, with the value being the hour in 24 - hour format, ranging from 0 - 63. For example, 19:00 is 010011. This item is only extended when the time slice is in hours, minutes, and seconds.
[0059] Minute: Extended item, 6 bits in binary, with the value being the number of minutes, ranging from 0 - 63. For example, 46 minutes is 101110. This item is only extended when the time slice is in minutes and seconds.
[0060] Second: Extended item, 6 bits in binary, with the value being the number of seconds, ranging from 0 - 63. For example, 53 seconds is 110101. This item is only extended when the time slice is in seconds.
[0061] In this way, the time slice can use 18 bits of space to store the year, month, and day, and an extended 18 - bit space to store hours, minutes, and seconds.
[0062] As Figure 3 shown, for the binary time slice, we take 6 bits as a unit and convert it into visible characters. 6 bits can store 64 binary values. Referring to Base64 encoding, these 64 binary values are corresponded. 26 uppercase letters respectively correspond to 0 - 25, 26 lowercase letters respectively correspond to 26 - 51, 10 Arabic numerals respectively correspond to 52 - 61, and the symbol '+' corresponds to 62, the symbol ' / ' corresponds to 63. Thus, the three - byte short code '7i / ' represents May 31, 2024, the extended short code 'T' represents 19:00, 'u' represents 46 minutes, and '1' represents 53 seconds.
[0063] Step S12, convert the spatial grid coordinates into visible characters using the octal pyramid encoding method, where the spatial grid coordinates include the x - axis coordinate value and the y - axis coordinate value.
[0064] In this embodiment, the spatial grid coding order is confirmed, and the x-axis coordinate value and the y-axis coordinate value are obtained according to the spatial range data; the x-axis coordinate value and the y-axis coordinate value are respectively converted into octal, and the number of digits is determined according to the maximum value and padded with zeros for alignment to obtain the octal spatial coding; the octal spatial coding is combined bit by bit crosswise to generate a second binary sequence, and the second binary sequence is subjected to sequence conversion to obtain visible characters corresponding to the second binary sequence.
[0065] Furthermore, the spatial short code is the short code representing the spatial grid. Considering that the map visualization provides multi-level caching, the design of the spatial short code adopts an original octal pyramid coding method, and the spatial short code information at the lower level directly includes the spatial short code at the upper level. There are two common coding orders for spatial grids, and the main difference lies in the coding order of the y-axis. The y-axis adopted by the TMS slice is coded from bottom to top, and the specific example is as follows Figure 4 as shown. The y-axis adopted by the WMTS slice is coded from top to bottom, and the specific example is as follows Figure 5 as shown. However, no matter which coding method is adopted, the spatial grid uses integer subscripts on the x-axis and y-axis to mark the position of the grid. As a storage solution for spatial big data, it has nothing to do with the spatial grid coding order and only records the resulting spatial grid coding.
[0066] Generally, for data with a clear spatial range, when dividing the spatial grid, it starts from the relative origin (0, 0), and the spatial coding uses relative coding, and the smaller value is conducive to generating the spatial short code. For spatial data directly divided by the global grid, such as Tianditu tiles, Amap tiles, etc., up to 20 levels at most, while Google Maps tiles are up to 22 levels at most. The maximum grid coding in the x-axis direction is 4194303 (i.e.), and the maximum grid coding in the y-axis direction is 2097151 (i.e.). The larger value is not conducive to generating the spatial short code. Therefore, for spatial data using global coding, the spatial grid coding should subtract the spatial coding value of the starting point where the data spatial range is located. For example, the x-axis coding of a 20-level tile in Beijing is 1345973, and it should subtract 1344703 (the minimum x-axis coding value of the tiles in the Beijing area), and the result is 1270.
[0067] After obtaining the x-axis and y-axis codings of the spatial grid, first convert them into octal, and then pad the parts where the x-axis and y-axis do not have the same number of digits up to the maximum value with zeros. For example, if the spatial range is divided into 105x82 grids, and the number of a certain grid is (90, 59), it is converted into (132, 073) in octal. Please refer to Figure 6, Next, the octal code is placed into 6-bit binary numbers in the order from the high bit to the low bit, along the x-axis and y-axis respectively. Then, following the time slice short code pattern, visible characters are used to perform character conversion on each 6-bit binary number. In this way, the three-byte short code "IfT" represents the spatial grid cell numbered (90, 59). Moreover, from the high bit to the low bit of the spatial short code, each byte can determine the area of the grid within an 8x8 spatial range.
[0068] From the highest bit short code "I", it can be determined that the spatial grid is on the right side of the spatial range. Then, the next-level short code "f" can determine that the spatial grid is in the overall bottom area on the right side. Continuing like this, the position of the grid in space can be located all the way down. This coding method is also beneficial for cache loading during map zooming for visualization.
[0069] Step S2, combine the time slice short code, spatial short code, object identifier, and unique identifier to construct multiple data partitioning models, and perform pre-partitioning on the amount of data required to be stored in the Hbase database according to the data partitioning models.
[0070] Since after the amount of data in a partition grows to a certain extent, HBase will automatically perform partition splitting and merging. To avoid the frequent automatic partition splitting affecting data access efficiency and the situation where operations are concentrated in the same partition (hot spot) and resources cannot be evenly utilized, pre-partitioning processing needs to be performed on HBase, and a strategy for partitioning according to the front part of the rowkey is designed. At the same time, based on the distributed parallel processing architecture of HBase, to improve data access efficiency, the rowkey design should follow the principles of uniqueness, hashability, and being as short as possible. For example, for ordinary business data, the reverse of the business classification or business primary key can be used as the front part of the rowkey to hash the data into different partitions as much as possible while ensuring uniqueness.
[0071] In this embodiment, the multiple data partitioning models include a spatial data model, a time series data model, a spatio-temporal data model, and a cyclic spatio-temporal data model. The spatial data model includes a spatial short code and a unique identifier. The time series data model includes a time slice short code and an object identifier. The spatio-temporal data model includes a spatial short code, a time slice short code, and an object identifier. The cyclic spatio-temporal data model includes a time slice short code, a spatial short code, and an object identifier.
[0072] The characteristic of massive spatial data is that the data all carry spatial location information, and each record has the central coordinates of the spatial location, such as high-precision map data. A data partitioning scheme using spatial grid division is adopted. The following are the partition model and the composition of the rowkey:
[0073] X X X-Y Y Y Y Y Y Y Y
[0074] (Spatial Short Code)(Unique Identifier)
[0075] The Rowkey consists of two parts: the spatial short code and the unique identifier. The spatial short code is the result of encoding the grid where the spatial data is located after being divided according to the spatial grid. The unique identifier is the unique primary key originally carried by the data. If the original data does not have a unique primary key, a unique primary key needs to be regenerated.
[0076] Generally, the spatial grid is divided according to the spatial range of the data. Considering that it is impossible to predict whether the distribution of all data in space is uniform, a uniform grid is generally adopted by default. The size of the grid is estimated according to the data volume and the aspect ratio of the length and width of the spatial range. The specific estimation formula is as follows:
[0077]
[0078] where l is the side length of the square grid. a and b are the length and width of the spatial range respectively, C 分区容量 is the maximum capacity of a single designed partition, and C 总容量 is the estimated total data capacity.
[0079] For example, after the spatial data is divided using a uniform 32x32 grid, if it can be predicted that there are some grids with a relatively dense distribution, the data partition distribution can be further dispersed according to the quadtree division. After initially dividing the grid into 2x2, the upper right grid has concentrated three pieces of data such as 1, 2, and 4. Then, the upper right grid is further divided into 2x2 grids to make the data evenly distributed.
[0080] The characteristic of massive time-series data is that the data is continuously and evenly generated, and each record has a uniformly defined timestamp, such as the continuously uninterrupted sensing data collected by sensors. When processing time-series data, it is generally required to maintain this temporal order, such as time windows, etc. Therefore, the rowkey cannot be simply hashed. By using the time slice division as the data partitioning scheme, the data can be stored in order within the time slice, and at the same time, the data can be dispersed to different partitions on a larger scale of time slices. The following is the partitioning model and the composition of the rowkey:
[0081] X X X - Y Y Y Y Y Y Y Y
[0082] (Time Slice Short Code)(Object Identifier)
[0083] Time slices can be divided into units such as seconds, minutes, hours, days, and months. The length of the time slice cannot be completely matched according to the partition capacity, and it needs to be designed in combination with the data generation period. For example, the time slice of a sensor with a period of 10 milliseconds is in seconds, while the rain gauge with a period of 15 minutes is in days. In addition, the object identifier is the primary key of the object entity that sequentially generates data, rather than the primary key of the data record.
[0084] The characteristics of massive spatio-temporal data are that the data is continuously and evenly generated, and at the same time, the location of the object theme that generates the data is also constantly changing. Each record not only has a uniformly defined timestamp but also its location coordinates. For example, the continuously reported positioning data of intelligent driving vehicles.
[0085] The processing of spatio-temporal data not only requires maintaining this time sequence but also taking into account the spatial distribution. It is necessary to consider using time slices and spatial grids as the data partitioning scheme. Considering that spatio-temporal data services mainly focus on space and are supplemented by time synchronization conditions for processing, the following partitioning model and rowkey are used:
[0086] X X X X-Y Y Y Y-Z Z Z Z Z Z Z
[0087] (Spatial short code)(Time slice short code)(Object identifier)
[0088] Massive cyclic spatio-temporal data is based on the characteristics of massive spatio-temporal data. It also needs to consider recycling the storage space within a certain period. For example, only retain historical data for half a year, and delete data that exceeds half a year on a daily basis for storage space recycling. In this regard, priority should be given to the elastic use of space without affecting normal data storage, query, and other business operations. For this, the dynamic management of partitions by HBase can be used to complete. The following partitioning model and rowkey are used:
[0089] X X X X-Y Y Y Y-Z Z Z Z Z Z Z
[0090] (Time slice short code)(Spatial short code)(Object identifier)
[0091] For the time slice short code, generally, it is in days, and the inverted short code is not used to leave sufficient processing time. In the daily end batch task, regularly check the overdue partitions, and dynamically delete the overdue partitions after data archiving. At the same time, in the daily end batch task, it is also necessary to dynamically create partitions for the dates that will be used in advance for a certain period.
[0092] In this embodiment, the time slice short code represents a time code set based on the partition capacity and the data generation period. The object identifier represents the primary key corresponding to the object body that continuously and uniformly generates data in chronological order. The space short code represents a space code that encodes the grid where the space data is located after being divided according to the space grid. The unique identifier represents the primary key that the data information originally carries and is unique.
[0093] Step S3: Perform spatial indexing or row key indexing according to the row key, time slice short code, and space short code in multiple data partition models to obtain an index result.
[0094] In this embodiment, by constructing a secondary spatio-temporal index based on multiple engines (key-value storage engine, lucene search engine), rich spatio-temporal big data query capabilities are provided. For example, the quadtree space index constructed by redis can perform pre-screening on space queries and computational processing in high-concurrency scenarios, thereby providing high-concurrency space computing capabilities. The comprehensive index constructed by es can provide efficient comprehensive query capabilities for massive data. By constructing a double-layer cache structure of a thinning cache layer and an aggregation cache layer, efficient management and fast access to massive spatio-temporal data are realized. The thinning cache layer realizes on-demand reduction of data through a downsampling mechanism, effectively reducing storage and processing pressure, while the aggregation cache layer provides fast data insight capabilities through pre-statistical analysis. The two layers work together not only significantly improve the system performance and resource utilization efficiency, but also provide flexible data access methods for upper-layer applications, which is especially suitable for processing large-scale spatio-temporal data analysis requirements in intelligent transportation scenarios.
[0095] In one example, two modes of spatio-temporal indexes are provided, a short code-rowkey index based on key-value storage, and a spatio-temporal retrieval index based on the lucene search engine.
[0096] The short code-object index is used to quickly retrieve the object set under a specified time slice or space grid, and is generally used for spatio-temporal fusion calculation. For example, the set of vehicles and events that appear under a certain space grid during a certain time period, etc. The short code-object index structure is very simple. The time slice short code or space short code is used as the key, and the stored value is the set of object primary keys of the data within this time slice or space grid.
[0097] The spatio-temporal retrieval index is used to query spatial data based on each attribute field as a condition. At the same time, it can also utilize the characteristics of time-slice short codes and spatial short codes to retrieve data related to a specified spatial range or time range through the fuzzy calculation of a search engine. The spatio-temporal retrieval index is constructed using the lucene search engine. When each data record is written, keyword field segmentation statistics are performed according to the index. When searching, the data with the highest relevance is retrieved based on scoring statistics such as keyword frequencies. The spatio-temporal retrieval index is not equivalent to the join query in SQL, but it provides an effective means for the join retrieval of massive data.
[0098] Furthermore, a decimation cache layer is constructed, and the sampled data is decimated according to the scaling ratio of the spatio-temporal scale, where the sampled data includes high-precision map data, perception data collected by sensors, positioning data reported by intelligent driving vehicles, and the corresponding spatio-temporal data; a aggregation cache layer is constructed to perform statistical analysis on the sampled data within a specified time range and spatial grid.
[0099] The multi-level spatio-temporal decimation cache performs decimation calculations on the original data according to the scaling ratio of time or space and then caches it. It is generally used for the visualization of the overall structure of spatio-temporal data, such as roads and rivers. Without considering the extreme distribution of data, decimation by equidistant sampling can generally meet the requirements of visualization.
[0100] In an example, both the time-slice short code and the spatial short code design have a characteristic that the short code information with a larger scale is included in the short code with a smaller scale. For example, a five-byte time-slice short code in minutes, where the first three digits are the time-slice short code in days of the current day, and the first four digits are the time-slice short code in hours. Similarly, the first few digits of the spatial short code are also the spatial short code at a larger scale. Using this characteristic, a visualization cache can be constructed according to different levels of short codes when massive spatio-temporal data is stored in the database.
[0101] Decimation caching means that when the time or space scale is larger, a certain proportion of massive data is screened to reduce it to a reasonable amount of data. Generally, the screening proportion is the same as the scale factor. For example, when constructing a decimation cache for the original 3-bit spatial short code data according to the first 2-bit spatial short code, the extraction ratio is generally 1 / 64. For the original massive data divided by minutes and constructing a cache divided by hours, the extraction ratio is 1 / 60. Through decimation caching, only the overall trend of the data can generally be expressed, and the data distribution cannot be accurately reflected in terms of density. By using aggregation caching, the data within the same time slice and the same spatial grid is statistically analyzed, so that not only the location of the data can be displayed at a higher scale, but also the quantity represented by the data can be shown. Aggregation statistics generally adopt two methods: data volume statistics and numerical statistics. The former represents the number of times the data is generated, and the latter expresses a certain business meaning.
[0102] Through the multi-mode storage design, it provides the storage of massive spatio-temporal data far beyond what a general spatial database can handle. Especially for the requirement of recycling storage space, through a clever partition and rowkey design scheme, the elastic utilization and cyclic release of storage space can be achieved. Based on the unique time slice short code and spatial short code design, the constructed secondary spatio-temporal index can not only provide efficient spatio-temporal data query under the condition of massive data storage. At the same time, it can also achieve SQL-like data retrieval, and even further provide the general-purpose data retrieval ability based on spatial location and time range. Based on the general-purpose spatio-temporal query ability, the massive spatio-temporal data stored in it can be published as a personalized map service just like ordinary spatial data. At the same time, the multi-level spatio-temporal cache constructed based on the short code design provides more effective visualization optimization support for data display at a large scale.
[0103] Furthermore, through the generation mechanism of time slice short code and spatial short code, the occupation of storage space is greatly reduced. The short code form is more compact than the complete spatio-temporal information, reducing the storage overhead. Using shorter strings to form row keys saves a large amount of key value storage space, improves the utilization efficiency of the cache, thereby improving the storage query efficiency. Using the spatio-temporal short code as an index improves the retrieval efficiency of spatio-temporal data, and the multi-dimensional index mechanism makes the query more flexible, which can meet the query requirements of different scenarios. By constructing multiple data partition models, the data skew problem is avoided, and the pre-partition mechanism ensures the uniform distribution of data in HBase, preventing hot spot problems in the HBase database.
[0104] Based on the above method, an HBase spatio-temporal data storage index system based on spatio-temporal short codes, refer to Figure 7, the HBase spatio-temporal data storage index system 1 includes a short code generation module 11, a data partitioning module 12, and an information indexing module 13. Among them,
[0105] The short code generation module 11 is used to generate a time slice short code and a space short code according to time information and a spatial grid respectively;
[0106] The data partitioning module 12 is used to combine the time slice short code, the space short code, the object identifier, and the unique identifier to construct multiple data partitioning models, and pre-partition the data volume required to be stored in the Hbase database according to the data partitioning models;
[0107] The information indexing module 13 is used to perform spatial indexing or row key indexing according to the row key, the time slice short code, and the space short code in multiple data partitioning models to obtain an index result.
[0108] In one example, the short code generation module 11 is used to convert the time information into a binary code according to a preset number of digits, and group the binary code and then map it to a visible character; the spatial grid coordinates are converted into visible characters by using an octal pyramid coding method, and the spatial grid coordinates include the x-axis coordinate value and the y-axis coordinate value.
[0109] In one example, the multiple data partitioning models include a spatial data model, a time series data model, a spatio-temporal data model, and a cyclic spatio-temporal data model. The spatial data model includes a space short code and a unique identifier. The time series data model includes a time slice short code and an object identifier. The spatio-temporal data model includes a space short code, a time slice short code, and an object identifier. The cyclic spatio-temporal data model includes a time slice short code, a space short code, and an object identifier.
[0110] In one example, the time slice short code represents a time code set based on the partition capacity and the data generation period. The object identifier represents the primary key corresponding to the object body that generates data continuously and evenly in time order. The space short code represents a spatial code that encodes the grid where the spatial data is divided according to the spatial grid. The unique identifier represents the primary key that the data information originally carries and has uniqueness.
[0111] In one example, the short code generation module 11 is used to convert the time information into a binary code to obtain a binary sequence; group the binary sequence according to a preset grouping condition to obtain a first binary sequence; perform sequence conversion on the first binary sequence to obtain visible characters corresponding to the first binary sequence.
[0112] In one example, the short code generation module 11 is used to confirm the spatial grid coding order, and obtain the x-axis coordinate value and the y-axis coordinate value according to the spatial range data; convert the x-axis coordinate value and the y-axis coordinate value into octal respectively, determine the number of digits according to the maximum value and pad with zeros for alignment to obtain the octal spatial coding; cross-combine the octal spatial coding bit by bit to generate a second binary sequence, and perform sequence conversion on the second binary sequence to obtain visible characters corresponding to the second binary sequence.
[0113] In one example, a decimation cache layer is constructed, and the sampling data is decimated according to the scaling ratio of the spatio-temporal scale, where the sampling data includes high-precision map data, perception data collected by sensors, positioning data reported by intelligent driving vehicles, and corresponding spatio-temporal data; an aggregation cache layer is constructed to perform statistical analysis on the sampling data within a specified time range and spatial grid.
[0114] Please refer to Figure 8 , which provides a schematic structural diagram of an electronic device for an embodiment of the present application. As Figure 8 shown, the electronic device 2 may include: at least one processor 21, at least one network interface 24, a user interface 23, a memory 25, and at least one communication bus 22.
[0115] Among them, the communication bus 22 is used to realize the connection and communication between these components.
[0116] Among them, the user interface 23 may include a display screen (Display), a camera (Camera), and optionally the user interface 23 may further include a standard wired interface and a wireless interface.
[0117] Among them, the network interface 24 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0118] Among them, the processor 21 may include one or more processing cores. The processor 21 connects various parts within the entire server through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 25, and by calling the data stored in the memory 25, it performs various functions of the server and processes data. Optionally, the processor 21 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 21 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 21 and may be implemented separately through a single chip.
[0119] Among them, the memory 25 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 25 includes a non-transitory computer-readable storage medium. The memory 25 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 25 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store the data involved in the above-mentioned various method embodiments. Optionally, the memory 25 may also be at least one storage device located far from the aforementioned processor 21. As Figure 8 shown, the memory 25, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program of the HBase spatio-temporal data storage index method based on spatio-temporal short codes.
[0120] In Figure 8In the electronic device 2 shown, the user interface 23 is mainly used to provide an interface for the user to input and obtain the data input by the user; while the processor 21 can be used to call the application program stored in the memory 25 for the HBase spatio-temporal data storage index method based on spatio-temporal short codes. When executed by one or more processors, the electronic device is caused to execute one or more methods as in the above embodiments.
[0121] A computer-readable storage medium stores instructions. When executed by one or more processors, the computer is caused to execute one or more methods as in the above embodiments.
[0122] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0123] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0124] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some service interfaces. The indirect couplings or communication connections of the device or unit can be in electrical or other forms.
[0125] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0126] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0127] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. And the aforementioned memory includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0128] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and the disclosure of the practical truth, those skilled in the art will easily think of other implementation schemes of the present disclosure. The present application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include the common general knowledge or conventional techniques in the technical field not recorded in the present disclosure.
[0129] The above description is only the preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A HBase spatiotemporal data storage indexing method based on spatiotemporal short codes, characterized in that: The method comprises: According to the time information and the space grid, a time slice short code and a space short code are generated respectively; The time slice short code, space short code, object identifier and unique identifier are combined to construct multiple data partition models, and the amount of data required to be stored in the Hbase database is pre-partitioned according to the data partition model; Spatial indexing or row key indexing is performed according to the row keys, time slice short codes, and space short codes in the multiple data partitioning models to obtain index results.
2. The method according to claim 1, characterized in that The generating of a time slice short code and a space short code according to the time information and the space grid respectively specifically includes: Converting the time information into binary codes according to a preset number of bits, and grouping the binary codes and mapping them into visible characters; The spatial grid coordinates are converted into visible characters by using an octal pyramid encoding method, wherein the spatial grid coordinates include an x-axis coordinate value and a y-axis coordinate value.
3. The method according to claim 1, characterized in that The multiple data partition models include a spatial data model, a time series data model, a space-time data model and a cyclic space-time data model. The spatial data model includes a spatial short code and a unique identifier, the time series data model includes a time slice short code and an object identifier, the space-time data model includes a spatial short code, a time slice short code and an object identifier, and the cyclic space-time data model includes a time slice short code, a space short code and an object identifier.
4. The method according to claim 3, characterized in that The time slice short code represents the time code set based on the partition capacity and the data generation cycle, the object identifier represents the primary key corresponding to the object body that generates data continuously and evenly in time sequence, the space short code represents the spatial code for encoding the grid where the spatial data is located after being divided into spatial grids, and the unique identifier represents the primary key originally carried by the data information and is unique.
5. The method according to claim 2, characterized in that The step of converting the time information into binary codes according to the preset number of bits, and grouping the binary codes and mapping them into visible characters specifically includes: Converting the time information into binary code to obtain a binary sequence; Grouping the binary sequence according to a preset grouping condition to obtain a first binary sequence; Perform sequence conversion on the first binary sequence to obtain visible characters corresponding to the first binary sequence.
6. The method according to claim 2, characterized in that The method of converting the spatial grid coordinates into visible characters by using the octal pyramid encoding method specifically includes: Confirm the spatial grid coding order and obtain the x-axis coordinate value and y-axis coordinate value according to the spatial range data; The x-axis coordinate value and the y-axis coordinate value are converted into octal, and the number of bits is determined according to the maximum value and zero-filled for alignment to obtain the octal space code; The octal space codes are cross-combined bit by bit to generate a second binary sequence, and the second binary sequence is sequence-converted to obtain visible characters corresponding to the second binary sequence.
7. The method according to claim 1, characterized in that The method further comprises: Construct a sparse cache layer and downsample the sampled data according to the scaling ratio of the spatiotemporal scale, wherein the sampled data includes high-precision map data, sensor-collected perception data, positioning data reported by the intelligent driving vehicle, and corresponding spatiotemporal data; An aggregation cache layer is constructed to perform statistical analysis on the sampled data within a specified time range and the spatial grid.
8. An HBase spatiotemporal data storage indexing system based on spatiotemporal short codes, characterized in that: The HBase spatiotemporal data storage index system (1) comprises a short code generation module (11), a data partitioning module (12) and an information indexing module (13), wherein: The short code generation module (11) is used to generate a time slice short code and a space short code respectively according to the time information and the space grid; The data partitioning module (12) is used to combine the time slice short code, the space short code, the object identifier and the unique identifier to construct a plurality of data partitioning models, and pre-partition the amount of data required to be stored in the Hbase database according to the data partitioning model; The information indexing module (13) is used to perform spatial indexing or row key indexing according to the row keys, time slice short codes and space short codes in the multiple data partitioning models to obtain indexing results.
9. An electronic device, characterized in that: The electronic device (2) comprises a processor (21), a memory (25), a user interface (23) and a network interface (24), wherein the memory (25) is used to store instructions, the user interface (23) and the network interface (24) are used to communicate with other devices, and the processor (21) is used to execute the instructions stored in the memory (25) so that the electronic device (2) executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Rowkey design method and device for Hbase database
CN113407518B
Cited By
Geological survey surveying and mapping cloud platform data synchronization method and system
CN122173573A