Task scheduling method, device and storage medium for spatiotemporal big data
Through ETL cleaning and standardization processing, the construction of spatiotemporal indexes and multi-dimensional query indexes, combined with distributed storage architecture and caching mechanism, the problem of insufficient performance of traditional databases when processing spatiotemporal big data is solved, and efficient spatiotemporal data storage and query are achieved.
Patent Information
- Application Number
- CN202510259257.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-06
AI Technical Summary
Traditional databases face insufficient performance and obstacles in data integration and sharing when processing spatiotemporal big data, and existing big data processing technologies are not friendly to spatiotemporal data support.
Through ETL cleaning and standardization of raw spatiotemporal data, a spatiotemporal index is constructed, a spatial fill curve algorithm is used to reduce dimensionality, a multi-dimensional query index is established, and data storage and retrieval is carried out based on a distributed storage architecture and cache mechanism.
It realizes efficient storage and query of spatio-temporal big data, reduces the difficulty of data integration, improves data availability and interoperability, and supports sub-second spatial and attribute retrieval.
Smart Images

Figure CN119739745B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to a task scheduling method, device and storage medium for spatiotemporal big data. Background Art
[0002] With the rapid development of high-tech technologies such as mobile Internet and Internet of Things, as well as the technical iteration and update of software and hardware equipment in the geographic information industry, the data in the geographic information industry has shown an explosive growth trend. The sources of spatiotemporal data are becoming increasingly extensive, including satellite remote sensing, drones, sensor networks and social media. These heterogeneous data use different file formats and data standards, resulting in significant obstacles to data integration and sharing.
[0003] Traditional relational databases face severe challenges in processing spatiotemporal big data. Geographic spatiotemporal data includes vector data, raster data, spatiotemporal series and other types. Its data structure involves multi-dimensional, multi-resolution and multi-temporal information, and traditional databases are insufficient in performance when facing high-dimensional spatial queries. In addition, spatiotemporal data such as high-resolution remote sensing images, real-time IoT data and dynamic traffic information are collected frequently, the file size of a single collection is large, and the data volume grows rapidly, which places extremely high demands on the storage system. At present, although the massive data processing technology solutions such as hive, spark, and flink can efficiently process massive data, they are not friendly to spatiotemporal data. Although the big data solutions in the GIS industry such as Geomesa support the storage, indexing and query of spatiotemporal big data, their interface design is complex, and a large number of parameter configurations are required when using them. The related configurations are highly professional, requiring users to master both GIS professional knowledge and big data professional knowledge, which increases the threshold for technology use. Summary of the invention
[0004] The present invention provides a task scheduling method, device and storage medium for spatiotemporal big data, which realizes the reasonable allocation of computing resources, effectively reduces the performance pressure of each node, and improves the availability and interoperability of data.
[0005] In a first aspect, the present invention provides a task scheduling method for spatiotemporal big data, the task scheduling method for spatiotemporal big data comprising:
[0006] Perform ETL cleaning and standardization on the original spatiotemporal data to obtain spatiotemporal data in a unified format that complies with OGC specifications;
[0007] Performing a spatial dimension reduction operation on the unified format spatiotemporal data by using a space filling curve algorithm and constructing a spatiotemporal index to obtain a spatiotemporal index data set;
[0008] Determine a data partitioning strategy according to the spatiotemporal index data set, distribute and compress the unified format spatiotemporal data according to spatiotemporal features, and obtain partitioned compressed data;
[0009] Establishing a multi-dimensional query index for the partitioned compressed data to obtain a task allocation plan;
[0010] Execute distributed data retrieval based on the task allocation scheme to obtain original query results;
[0011] The original query result is subjected to data integrity verification and spatial relationship reconstruction to obtain target query data.
[0012] In a second aspect, the present invention provides a task scheduling device for spatiotemporal big data, the task scheduling device for spatiotemporal big data comprising:
[0013] The standardization processing module is used to perform ETL cleaning and standardization on the original spatiotemporal data to obtain spatiotemporal data in a unified format that complies with OGC specifications;
[0014] A construction module, used to perform a spatial dimension reduction operation on the unified format spatiotemporal data and construct a spatiotemporal index by using a space filling curve algorithm to obtain a spatiotemporal index data set;
[0015] A compression module is used to determine a data partitioning strategy according to the spatiotemporal index data set, distribute and compress the unified format spatiotemporal data according to spatiotemporal characteristics, and obtain partitioned compressed data;
[0016] A task allocation module, used to establish a multi-dimensional query index for the partition compressed data to obtain a task allocation plan;
[0017] An execution module, used for executing distributed data retrieval based on the task allocation scheme to obtain original query results;
[0018] The reconstruction module is used to perform data integrity verification and spatial relationship reconstruction on the original query results to obtain target query data.
[0019] A third aspect of the present invention provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned task scheduling method for spatiotemporal big data.
[0020] In the technical solution provided by the present invention, the problem of data heterogeneity is solved through ETL data cleaning and standardization processing, unified management of data from different sources and in different formats is achieved, and the difficulty of data integration is reduced; the space filling curve algorithm is used for dimensionality reduction processing, combined with the construction of Z index and S index, sub-second spatial retrieval and sub-second attribute retrieval of tens of millions of spatiotemporal data are achieved; based on the MMP distributed storage architecture and cache storage mechanism, high-speed data writing is achieved, and the writing rate reaches 120,000 records / second. At the same time, through the support of multiple compression algorithms, the maximum compression rate can reach 80%, which effectively saves storage space; through the establishment of multi-dimensional query indexes, the construction of redundant data tables is avoided, and the storage resource occupation and data maintenance complexity are reduced; based on the distributed task scheduling mechanism of dynamic load balancing, the reasonable allocation of computing resources is achieved, and the performance pressure of each node is effectively reduced; through integrity verification and spatial relationship reconstruction, the accuracy and consistency of query results are ensured, and data output in multiple standard formats is supported, which improves the availability and interoperability of data.
[0021] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0022] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A schematic diagram of an embodiment of a task scheduling method for spatiotemporal big data in an embodiment of the present invention;
[0024] Figure 2 Schematic diagram of an embodiment of a task scheduling device for spatiotemporal big data in an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device end including a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or device ends.
[0027] To facilitate understanding of this embodiment, a task scheduling method for spatiotemporal big data disclosed in an embodiment of the present invention is first described in detail. Figure 1 As shown, the method comprises the following steps:
[0028] 101. Perform ETL cleaning and standardization on the original spatiotemporal data to obtain spatiotemporal data in a unified format that complies with OGC specifications;
[0029] It is understandable that the execution subject of the present invention may be a task scheduling device for spatiotemporal big data, or a terminal or a server, which is not specifically limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0030] Specifically, the original spatiotemporal data is collected from satellite remote sensing equipment, unmanned aerial vehicle equipment and sensor network equipment, and the data source of the original spatiotemporal data is identified to determine the type and format of the data and extract the data source type information. According to the data source type information, the original spatiotemporal data is format parsed to extract its data format information, including WKB (Well-Known Binary), WKT (Well-Known Text), WMS (Web Map Service) and WMTS (Web Map Tile Service). Different data formats correspond to different data organization methods. For example, WKB and WKT are used to store and transmit vector data, while WMS and WMTS are suitable for image data and tile data of online map services. Based on the format parsing, a unified data structure conversion is performed to ensure that all data is organized according to the predefined data model to obtain preliminary conversion data. This conversion not only requires parsing the spatial information of the original data, but also processing its time and attribute information to ensure the integrity and consistency of the data. The spatial attributes of the preliminary conversion data are extracted, and the geographic coordinate information, timestamp and related attribute fields in the data are parsed into a standard format for subsequent query, calculation and analysis. Geographic coordinate information is expressed in longitude and latitude. In the process of extracting spatial attributes, different coordinate systems are converted to standard OGC-compatible coordinate reference systems to ensure cross-system compatibility and spatial consistency of data. At the same time, the time format is unified to ensure that all data are analyzed under the same time base. The analysis of attribute fields involves the conversion and standardization of data types, such as converting the units of numerical data into a unified format to ensure that errors are not generated due to inconsistent units in subsequent calculations. Based on spatial attribute data, a UDF SQL interface based on the ANSI SQL standard is constructed, allowing users to retrieve and analyze data using SQL queries. At the same time, for vector data, since it involves complex spatial relationship calculations, topological relationship calculations are performed to generate topological relationship data. Topological relationship calculations mainly include connection relationships, adjacency relationships, inclusion relationships, and intersection relationships between spatial objects. Through the calculation process, the data has stronger spatial analysis capabilities, thereby supporting complex geographic spatial computing tasks such as path analysis, regional clustering, and spatial proximity analysis. Outliers in topological relationship data are detected and corrected, missing values are processed using interpolation algorithms or other data filling methods, and duplicate data are removed through deduplication algorithms to obtain quality verification data. The quality verification data is reconstructed according to the OGC specification, including the unification of the spatial reference system, time format and attribute structure.In terms of spatial reference system, ensure that all data use a unified geographic coordinate system to ensure cross-platform compatibility and spatial consistency of data; in terms of time format, standardize all timestamps so that all data can be analyzed based on the same time base; in terms of attribute structure, organize data fields according to OGC specifications to ensure that data from different sources have a unified field structure to facilitate subsequent query, calculation and visualization. Through this step, the unified format spatiotemporal data that conforms to the OGC specification is finally obtained.
[0031] 102. Perform spatial dimension reduction operation on the unified format spatiotemporal data through a space filling curve algorithm and construct a spatiotemporal index to obtain a spatiotemporal index data set;
[0032] Specifically, the spatiotemporal data in a unified format are parsed by geographic vector spots to extract the spatial coordinate information and time information therein and obtain spatiotemporal feature data. The geometric information stored in the data, such as the coordinates of points, lines, and surfaces in the vector data, and the pixel coordinates in the raster data, are parsed and combined with the timestamp information to construct a spatiotemporal feature data set. When extracting spatiotemporal feature data, the coordinates are converted to a unified geographic coordinate system, and the time information is converted to a standardized time format to ensure the compatibility and comparability of different data. The spatiotemporal feature data is input into the Z-curve space filling curve algorithm, which is a spatial mapping method that maps three-dimensional spatial coordinates to one-dimensional data, thereby reducing the index complexity of high-dimensional data. Z-curve interleaves the binary encoding of the coordinates so that adjacent spatial points are kept as close as possible in the mapped Z sequence encoding data. Through Z sequence encoding, spatiotemporal data can be stored and processed in one-dimensional form without losing the original spatial information. The Z sequence encoding data is subjected to Hilbert-curve transformation. Hilbert-curve is a more optimized space filling curve algorithm. Unlike Z-curve, it can maintain the spatial proximity of data while reducing the dimension. Through Hilbert-curve transformation, the spatial proximity relationship of data is recalculated, and a spatial coding sequence is generated to obtain one-dimensional spatial index data. The Hilbert curve allows adjacent spatial points to maintain relatively close index values after mapping, thereby improving the efficiency of index query. Especially in the scenarios of range query and nearest neighbor query, the Hilbert-curve index can locate the target data more accurately. A spatial multi-layer index tree is constructed based on one-dimensional spatial index data to improve the performance of index query. Index information of different granularities is stored in a hierarchical relationship, that is, coarser granularity index data is stored at a higher level, and more refined index data is stored at a lower level, so that the spatial range of the data can be quickly located during query, and specific data points can be further accurately retrieved when needed. On the basis of constructing the spatial index, the hierarchical index structure is extended in the time dimension to establish a time index tree, and it is associated with the spatial index to form a spatiotemporal combined index. The time index tree is constructed using a B+ tree or time interval index, which allows efficient querying of data according to the time dimension. By combining the time index tree with the spatial index structure, joint queries on spatiotemporal data can be implemented, such as querying data for a specific area within a certain time period, or querying the changes in a certain spatial location at different time points. Targeted optimization is performed on the spatiotemporal combined index, and the Z index and S index are constructed separately.The Z index is used for range query. Its principle is to quickly filter out data that meets the query conditions based on the index value range after Z-curve mapping, thereby reducing the cost of data scanning; while the S index is mainly used for nearest neighbor query. Its optimization method is to use the index value generated after Hilbert-curve transformation to quickly find the data point closest to the query target through the calculation of spatial proximity relationship. Through the optimization of these two indexes, the retrieval efficiency of spatiotemporal data in different query scenarios is improved respectively. Through the above steps, the spatiotemporal index dataset is obtained.
[0033] 103. Determine the data partitioning strategy according to the spatiotemporal index data set, distribute and compress the unified format spatiotemporal data according to the spatiotemporal characteristics, and obtain partitioned compressed data;
[0034] Specifically, data partition planning is performed based on the spatiotemporal index data set, and data storage and query performance are optimized through reasonable data partitioning strategies, so that data with similar spatial locations and time characteristics can be stored in the same partition to reduce the cross-region access cost during data retrieval. In this process, the spatial coordinate information and timestamp information of the data are comprehensively analyzed to ensure that the data is efficiently retrieved according to the proximity of the geographical location and the continuity of the time series during query. Therefore, data clustering is performed based on the index information generated by the space filling curve algorithm, and adjacent data is aggregated into the same storage partition to form an initial partitioning scheme. The size threshold is detected for each partition in the initial partitioning scheme. According to the set sharding rules, the partitions that exceed the threshold are further divided so that each data block meets the optimal load requirements of system storage and query, and the partition sharding data is obtained. The partition sharding data is input into the MMP distributed storage architecture, and independent cache space is configured on the storage node to optimize the cache during data writing and reading. Each storage node maintains a part of the data cache locally to reduce the frequency of disk I / O, thereby improving the data access speed. The MMP architecture supports dynamic expansion of data and multi-node distributed management, so that data storage not only has high throughput, but also has good load balancing characteristics, so that data can be dynamically adjusted between different storage nodes according to the load situation. In the process of writing data into the cache, the access frequency of the data is analyzed, and the cache strategy is intelligently adjusted according to the heat distribution of the data to ensure that the data with high frequency access can be cached in high-speed storage media first, thereby improving query efficiency. The compression algorithm of the distributed cache data is analyzed to optimize storage occupancy and improve access efficiency. Since different types of spatiotemporal data have different efficiencies when compressed, the optimal compression algorithm is selected according to the characteristics of the data to compress the data and obtain node compressed data. For example, for vector data, the lightweight snappy compression algorithm is selected to ensure that the data can be quickly decompressed and used for calculation, while for high-resolution remote sensing image data, efficient compression algorithms such as bzip2, lz4 or zstd are used to minimize storage space occupancy. In this process, the content of the data block is analyzed and the selected compression strategy is dynamically adjusted to achieve the best balance between storage space usage and data decompression performance, ensuring that the data is compressed to the maximum extent when stored, and can be quickly decompressed and returned when queried. In order to improve the reliability and disaster recovery capabilities of the data, a master-slave replication mechanism is established based on node compression data, and the compressed data is backed up and synchronized to build a complete backup data set. Data is replicated between different storage nodes to ensure that even if a node fails, data can still be restored from other nodes to ensure high data availability.Under the management of the master-slave replication mechanism, the data of each storage partition will be stored in at least two or more copies, so that when reading data, the optimal data source is selected from multiple copies through the load balancing strategy to access, thereby improving the query response speed and reducing the pressure on a single storage node. At the same time, during the data synchronization process, a consistency check is performed to ensure that all backup data is consistent with the primary data to avoid incomplete data due to data replication delays or partial loss. After completing the data backup, a partition index table is established for the backup data set to record the data distribution information and storage location mapping relationship to form the final partition compressed data.
[0035] 104. Establish a multi-dimensional query index for the partitioned compressed data to obtain a task allocation plan;
[0036] Specifically, query conditions are parsed for partitioned compressed data, and SQL query statements are converted into structured query trees to obtain query parsing data. Parsing of SQL query statements includes extracting query fields, filter conditions, sorting rules, and aggregate functions, and converting them into a tree structure so that the query optimizer can intuitively analyze the query logic. As query complexity increases, query statements contain multiple JOIN operations, subqueries, or nested queries. Recursive parsing methods are used in the parsing process to ensure that all query conditions and data associations are fully parsed to generate accurate query parsing data. Based on the query parsing data, time dimension indexes, spatial dimension indexes, and attribute dimension indexes are constructed to form a multi-dimensional index set to improve the efficiency of data query. The construction of the time dimension index adopts a B+ tree or time interval index structure, so that the query can quickly retrieve the target data within a given time range; the spatial dimension index is constructed based on the Z index or S index to support efficient range queries and nearest neighbor queries, so that spatial data can be quickly retrieved according to geographic location; the attribute dimension index uses an inverted index or hash index to support efficient attribute filtering queries. In this process, different types of data indexes are independent of each other but interrelated. Through reasonable index combinations, the query system can dynamically select the most appropriate index according to the query conditions when executing the query, thereby improving retrieval efficiency and reducing the amount of data that needs to be traversed during the query process. After the multi-dimensional index set is built, a cost calculation model is established to calculate the execution cost of different query paths and obtain query cost data. The core purpose of the cost calculation model is to quantify the computational overhead of different query solutions and evaluate the pros and cons of each query path so as to select the optimal execution solution in the query optimization stage. When calculating the query cost, multiple factors are considered, including the distribution of data storage, the computing resources required for the query, the index hit rate, the data scanning range, and the I / O overhead. Usually, the system simulates different query execution plans and calculates the expected execution time of the query based on historical query execution or statistical information, and then selects the query solution with the lowest execution cost to optimize the query performance. In this process, the query cost calculation needs to be combined with the query tree structure to analyze the dependencies between different query subtasks and evaluate the computational complexity of each subtask to ensure that the query path finally selected can complete the computational task in the shortest time. The optimal query path is selected based on the query cost data, and an optimized execution plan is generated. At the same time, the optimized execution plan is task-splitting to divide the query task into multiple subtask units to form a subtask sequence. According to the query dependency and computing resource conditions, the query task is disassembled into independent computing units that can be executed in parallel, so as to make full use of computing resources in a distributed computing environment. For large-scale data scanning tasks involved in the query, it is split into parallel processing tasks of multiple data blocks, and for complex computing tasks, it is hierarchically split according to the computing dependency to ensure that the computing task can be executed in the best way.In the process of task splitting, the location of data storage is considered to minimize cross-node data transmission, thereby improving computing efficiency. A task scheduling strategy is established based on the subtask sequence to achieve resource allocation and parallelism setting for each subtask, forming a complete task allocation solution. The design of the task scheduling strategy needs to comprehensively consider the utilization of computing resources, the priority of task execution, the real-time requirements of the query, and the load of the computing nodes, so as to achieve optimal resource scheduling during the execution of the query task.
[0037] 105. Perform distributed data retrieval based on the task allocation scheme to obtain the original query results;
[0038] Specifically, based on the task allocation scheme, task reception and parsing are performed, and each subtask of the query task is assigned to different processing nodes in the distributed computing cluster to obtain node task data. The task allocation scheme is parsed to identify the computing requirements of each subtask, and the tasks are reasonably assigned to the computing nodes in combination with the data storage location and computing resources to ensure load balancing of data processing. At the same time, the task parsing process also needs to consider the principle of data locality, that is, to assign computing tasks to nodes that store data as much as possible to reduce data transmission overhead and improve task execution efficiency. In order to ensure that tasks can be stably executed in the cluster, the execution dependencies of tasks are analyzed to ensure that tasks can be executed in the correct order, thereby avoiding data competition and computing conflicts. After completing the task allocation, a task execution queue is constructed based on the node task data, and the load status of each processing node in the cluster is monitored in real time to generate a task scheduling sequence. The construction of the task execution queue needs to comprehensively consider the priority of the task, the data access mode, the availability of computing resources, and the estimated execution time of the task to ensure that the task can be executed in the optimal order. At the same time, the load monitoring system will continuously collect the CPU utilization, memory usage, I / O load and network bandwidth usage of each computing node, so that the task allocation plan can be dynamically adjusted during the task scheduling process to prevent some nodes from becoming performance bottlenecks due to excessive computing load. In order to improve the flexibility of task scheduling, the task execution time is estimated in combination with historical query records, and the scheduling strategy is dynamically adjusted according to the task execution history to improve the efficiency of task execution. After determining the task scheduling sequence, data prefetching operations are performed according to the sequence to reduce the performance degradation caused by data I / O delay during the calculation process. The core goal of data prefetching is to load the data required for the task into the memory cache in advance, thereby reducing the overhead of disk access and improving the efficiency of data retrieval. In this process, the query conditions of the task are analyzed, the required data range is identified, and the relevant data is loaded from the storage system to the high-performance cache through the distributed cache mechanism to speed up the query execution. After preloading the data, the computing node will use the parallel computing framework to perform distributed retrieval operations to achieve efficient data query. During the distributed retrieval process, each computing node performs query operations based on preloaded data and uses index optimization strategies to accelerate data filtering to ensure that the query task can be completed in the shortest time. During the query execution process, data is exchanged between computing nodes to optimize computing efficiency and reduce the overhead of repeated calculations. Exception monitoring is performed based on the node retrieval results to ensure that the query task can be completed correctly. The exception monitoring system analyzes the execution status of each computing node in real time and detects task execution anomalies, such as task timeouts, abnormal calculation results, or missing data. When an abnormal task is detected, the system automatically triggers the task reallocation mechanism to reallocate the abnormally executed task to other nodes with idle computing resources to ensure that the task can continue to execute and generate reallocation results.During the task redistribution process, priority is given to computing nodes with lower load and optimal data locality to reduce data transmission overhead and ensure that the restarted computing tasks can be completed as soon as possible. After all nodes have completed computing tasks, the redistribution results are merged with the normal execution results to obtain the complete original query results. In the process of merging data, data integrity is ensured, and operations such as deduplication, sorting, or aggregation are performed to ensure the correctness and consistency of the query results. At the same time, according to the query requirements, the original query results are formatted or optimized to meet the subsequent data analysis or visualization requirements.
[0039] 106. Perform data integrity verification and spatial relationship reconstruction on the original query results to obtain target query data.
[0040] Specifically, the original query results are checked for data integrity to ensure that the returned data is not lost, duplicated or damaged during the calculation and storage process, and the verification result data is obtained. The query results are checked for data integrity, including checking the data integrity, the correctness of the data fields, and the spatial consistency of the data. The integrity check uses hash verification, data comparison, and statistical analysis to ensure that all query data meet the expected format and content requirements, and that data is not missing or duplicated due to distributed storage or computing node failures. At the same time, the spatial consistency check will check whether the geometric structure in the data is complete, such as ensuring that the point, line, and surface data are not broken or missing, and the topological relationship conforms to the geographic space rules. The topological relationship of the geographic elements is reconstructed based on the verification result data to restore the spatial adjacency relationship between spatial objects and construct complete spatial topological data. In this process, the topological rules of vector data are used to perform relationship analysis on the point, line, and surface elements in the query results to determine the connection, intersection, inclusion, adjacency, and other relationships between spatial objects, and to construct a spatial adjacency relationship graph. The topological relationship calculation adopts the spatial index technology based on R-tree or Quadtree, so that the calculation of adjacency relationship can be performed efficiently on large-scale data sets and ensure the spatial consistency of data. The spatial topological data is input into the result cache mechanism to optimize the access efficiency of query results and reduce the overhead of repeated calculations. The historical access pattern of query requests is analyzed, and based on factors such as data access frequency, similarity of query parameters, and computing overhead, which data needs to be cached is intelligently selected and cache structure data is constructed. The cache mechanism adopts a distributed cache architecture, such as Redis or Memcached, to ensure the access speed of query results and support large-scale concurrent access. After the cache index is established, the query results are paging managed according to the cache structure data to ensure that the returned data is segmented according to the preset paging size and form a paging result set. In the paging management process, the required paging size is calculated according to the parameters of the query request, and the data segmentation strategy is dynamically adjusted according to the index information of the query result to ensure that the data volume of each page is balanced and can be efficiently loaded on the client. After obtaining the paginated result set, the data is formatted based on the preset second data format information to meet the needs of different application scenarios and generate format conversion data. Since different GIS applications and data analysis platforms use different data formats, the goal of format conversion is to convert the query results into the standard format required by the user, such as GeoJSON, Shapefile, or KML.GeoJSON is a lightweight format widely used in Web GIS applications, suitable for online map display and front-end data interaction, while Shapefile is a common format used by traditional GIS software (such as ArcGIS) and is suitable for local data storage and analysis, while KML is mainly used for visualization of three-dimensional geographic information systems such as Google Earth. After completing the data format conversion, a data security control mechanism is established based on the format conversion data to ensure that the query results can be safely accessed under different user permissions and prevent unauthorized users from obtaining sensitive data. In this process, the query results are filtered according to the user's access rights, and only the data that the user has permission to access is returned to the user. The data security control mechanism is based on role access control or attribute-based access control models to ensure that users of different levels access data according to predefined security policies. In order to enhance data security, the sensitive fields in the query results are fuzzy processed in combination with data desensitization technology to prevent unauthorized users from obtaining detailed geographic information or business data. Generate target query data to meet the query needs of spatiotemporal big data in fields such as GIS, remote sensing, and smart cities, and ensure the efficiency, reliability, and security of data queries.
[0041] In the embodiment of the present invention, through ETL data cleaning and standardization processing, the problem of data heterogeneity is solved, unified management of data from different sources and in different formats is achieved, and the difficulty of data integration is reduced; the space filling curve algorithm is used for dimensionality reduction processing, combined with the construction of Z index and S index, sub-second spatial retrieval and sub-second attribute retrieval of tens of millions of spatiotemporal data are achieved; based on the MMP distributed storage architecture and cache storage mechanism, high-speed data writing is achieved, and the writing rate reaches 120,000 / second. At the same time, through the support of multiple compression algorithms, the maximum compression rate can reach 80%, which effectively saves storage space; through the establishment of multi-dimensional query indexes, the construction of redundant data tables is avoided, and the storage resource occupation and data maintenance complexity are reduced; based on the distributed task scheduling mechanism of dynamic load balancing, the reasonable allocation of computing resources is achieved, and the performance pressure of each node is effectively reduced; through integrity verification and spatial relationship reconstruction, the accuracy and consistency of query results are ensured, and data output in multiple standard formats is supported, which improves the availability and interoperability of data.
[0042] In a specific embodiment, the process of executing step 101 may specifically include the following steps:
[0043] Collect original spatiotemporal data from satellite remote sensing equipment, UAV equipment and sensor network equipment, and identify the data source of the original spatiotemporal data to obtain data source type information;
[0044] Format-parse the original spatiotemporal data according to the data source type information to obtain first data format information, and perform unified data structure conversion according to the first data format information to obtain preliminary conversion data, where the first data format information includes WKB, WKT, WMS and WMTS;
[0045] Extract spatial attributes from the preliminary converted data, parse the geographic coordinates, timestamps and attribute fields into a standard format, and obtain spatial attribute data;
[0046] Based on spatial attribute data, a UDF SQL interface in the ANSI SQL standard is constructed to calculate the topological relationship of vector data and obtain topological relationship data.
[0047] Correct outliers, missing values, and duplicate values in topological relationship data to obtain quality verification data;
[0048] The quality verification data is reconstructed according to the OGC specification, and the spatial reference system, time format and attribute structure are unified to obtain unified format spatiotemporal data that conforms to the OGC specification.
[0049] Specifically, raw spatiotemporal data are collected from a variety of data sources, including satellite remote sensing equipment, drone equipment, and sensor network equipment. The data collected by these data sources have large differences in format, coordinate system, time accuracy, etc. The data source is identified to determine its source and characteristics and obtain data source type information. The process of data source identification is based on metadata analysis, file format parsing, and data content feature matching, wherein metadata analysis determines the data format by reading file header information, and data content feature matching classifies the data based on statistical analysis or machine learning models. For example, satellite remote sensing data is stored in the form of raster images, and its file format is GeoTIFF, while data collected by drone equipment is stored in the form of vector point clouds, and its format is LAS or Shapefile, and sensor network data is stored in a time series database, and its data format is JSON or CSV. The raw spatiotemporal data is format parsed according to the data source type information to extract the first data format information, and a unified data structure conversion is performed based on the format information. The first data format information includes WKB (Well-Known Binary), WKT (Well-Known Text), WMS (Web Map Service) and WMTS (Web Map Tile Service). WKB and WKT are standard formats for storing and transmitting vector data, while WMS and WMTS are used to provide online map services, supporting dynamic map rendering and tile data caching, respectively. During the format parsing process, data from different sources are converted into a standard format, such as converting Shapefile to WKT or WKB, so that subsequent storage and query operations use a unified data structure. For example, given a vector data in Shapefile format, which contains multiple geographic objects (such as points, lines, and surfaces), convert it to WKT format:
[0050] POINT(120.5,31.2);
[0051] Among them, the WKT format is an easy-to-read text representation, while the WKB format is a binary encoding format suitable for efficient storage and transmission. During the format conversion process, the spatial reference system of the data is unified to ensure that data from different sources are calculated in the same coordinate system. The spatial attributes of the preliminary converted data are extracted, and the geographic coordinates, timestamps and attribute fields are parsed into a standard format to obtain spatial attribute data. The geographic coordinate information is represented by longitude and latitude, but different data sources use different projection coordinate systems. All data are converted to a unified geographic coordinate system. The timestamp information is standardized and uniformly converted to the standard time representation in ISO 8601 format. The parsing of attribute fields involves data type conversion, such as converting string type values to floating point numbers and ensuring that the units are consistent. After the spatial attribute data is standardized, the ANSI SQL standard UDF SQL interface is constructed based on these data to support the topological relationship calculation of vector data and obtain topological relationship data. The core of topological relationship calculation is to analyze the mutual relationship between spatial objects, such as the inclusion relationship, adjacency relationship, and intersection relationship between points, lines, and surfaces. Topological relationship calculation relies on spatial index structures, such as R-trees or quadtrees, to improve query efficiency. For any two spatial objects and , calculate their spatial relationship:
[0052] ;
[0053] For example, given a building polygon area and a polyline segment of a road , calculate whether they intersect:
[0054]
[0055] Indicates that the road passes through the building area. Similarly, calculate spatial operations such as nearest neighbor query and containment relationship to support subsequent data analysis. After completing the topological relationship calculation, perform data quality detection on the topological relationship data to correct the outliers, missing values and duplicate values to obtain quality verification data. Outlier detection uses statistical methods or machine learning methods, such as outlier detection based on standard deviation or outlier identification based on clustering methods. For example, if a data point of a GPS sensor The Euclidean distance of is far from its adjacent points before and after:
[0056]
[0057] if If the value is greater than the set threshold, the point is considered an outlier and is corrected or deleted. Missing value filling uses interpolation methods, such as linear interpolation:
[0058]
[0059] To ensure the integrity of the data. For duplicate values, the hash deduplication method is used to ensure that redundant data is not stored in the database. After completing the data quality test, the quality verification data is reconstructed according to the OGC (Open Geospatial Consortium) specification to ensure that all data meets the standardization requirements. The data reconstruction process includes the standardization of the spatial reference system, the unification of the time format, and the normalization of the attribute structure. For example, all spatial data uses a unified EPSG:4326 coordinate system, all timestamps are converted to ISO 8601 format, and the attribute fields are organized according to the standard data model so that different systems can be seamlessly connected.
[0060] In a specific embodiment, the process of executing step 102 may specifically include the following steps:
[0061] Perform geographic vector spot analysis on unified format spatiotemporal data, extract spatial coordinate information and time information, and obtain spatiotemporal feature data;
[0062] Input the spatiotemporal feature data into the Z-curve space filling curve algorithm, perform one-dimensional mapping calculation on the three-dimensional space coordinates, and obtain Z sequence encoding data;
[0063] Perform Hilbert-curve transformation on the Z sequence coded data, calculate the spatial proximity relationship and generate a spatial coding sequence to obtain one-dimensional spatial index data;
[0064] Based on the one-dimensional spatial index data, a spatial multi-layer index tree is constructed to store index information of different granularities in layers to obtain a layered index structure.
[0065] The hierarchical index structure is expanded in time dimension, a time index tree is constructed and associated with the spatial index to obtain a time-space combined index;
[0066] The spatiotemporal combined index is optimized according to range query and nearest neighbor query, and the Z index and S index are constructed respectively to obtain the spatiotemporal index dataset.
[0067] Specifically, the unified format of spatiotemporal data is parsed by geographic vector spot analysis to extract spatial coordinate information and time information to obtain spatiotemporal feature data. Vector data exists in the form of points, lines, surfaces, etc. and is stored in standard formats such as WKB or WKT. Spatial coordinate information is parsed from vector data And the corresponding time information ,in and represents geographic coordinates (such as longitude and latitude), and Represents height or other specific dimensions, such as temperature or population density, time information It is used to support temporal queries. The spatiotemporal feature data is input into the Z-curve space filling curve algorithm to perform one-dimensional mapping calculations on the three-dimensional space coordinates in order to reduce the index dimension and optimize query performance. Z-curve is a space filling curve algorithm that converts multi-dimensional coordinates into one-dimensional index values by interleaving bit encoding. Suppose the coordinates of a point are , whose binary representation is:
[0068]
[0069] Then the calculation method of Z sequence encoding is:
[0070]
[0071] For example, for (Binary: 0011), (Binary: 0101), (binary: 0010), The sequence is encoded as:
[0072]
[0073] This encoding method ensures that similar spatial points maintain the closest order as possible in the one-dimensional index space. However, although the Z-curve can effectively reduce dimensionality, its mapping process will destroy some spatial proximity relationships. Therefore, the Z-sequence encoded data is transformed by Hilbert-curve to optimize spatial proximity and generate a spatial encoding sequence to obtain one-dimensional spatial index data. The Hilbert curve is a continuous curve, which is characterized by being able to better maintain the adjacency relationship of spatial objects, especially better than the Z-curve in high-dimensional data indexing. The calculation method of the Hilbert transform is relatively complicated. It recursively divides the space into sub-grids and numbers each sub-grid according to predefined rules. For example, in 2D space, the order of Hilbert encoding is as follows:
[0074]
[0075] In 3D space, the calculation of Hilbert coding is generalized as:
[0076]
[0077] The numbers of points with similar spatial positions in the Hilbert sequence are arranged more closely, improving query performance. Based on the one-dimensional spatial index data after Hilbert transformation, a spatial multi-layer index tree is constructed so that index information of different granularities can be stored in layers, thereby improving query efficiency. The spatial multi-layer index tree uses an R-tree or a quadtree to store data at different levels. For example, in an R-tree, each leaf node stores the actual data object, while the non-leaf node stores the minimum bounding rectangle of the child node, so that the search range can be quickly narrowed down by filtering layer by layer during query. Assume that the minimum bounding rectangle of a spatial data object is:
[0078]
[0079] When querying, only check whether the query range intersects with the MBR:
[0080]
[0081] It can be determined whether the object needs further calculation. On the basis of establishing the spatial index, the hierarchical index structure is expanded in the time dimension to build a time index tree, and it is associated with the spatial index to obtain a spatiotemporal combined index. The time index uses a B+ tree or time interval index to support efficient time range queries. For example, suppose a data point has a time range:
[0082]
[0083] Query time range The judgment condition for whether it intersects with it is:
[0084]
[0085] The construction method of the time index tree is similar to that of the spatial index, that is, time segment information is stored at different levels, and the time index is associated with the spatial index through an index adapter, so that data objects that meet the time conditions can be filtered out at the same time during the query. In order to optimize the query performance, the spatiotemporal combined index is optimized according to the range query and the nearest neighbor query, and the Z index and S index are constructed respectively to obtain a complete spatiotemporal index data set. The Z index is used for range queries. Its principle is to quickly filter out data that meets the query conditions based on the index value range encoded by the Z-curve, thereby reducing the overhead of data scanning. The S index is used for nearest neighbor queries. Its optimization method is to use the index value generated after the Hilbert-curve transformation to quickly find the data point closest to the query target by calculating the spatial proximity relationship. For example, given a query point , calculate the Hilbert index of its nearest neighbor data point:
[0086]
[0087] Through this optimization strategy, we ensure that the query performance can meet the high efficiency requirements in the massive data environment, so that both range queries and nearest neighbor queries can be completed in sub-second time, and finally form a complete spatiotemporal index data set.
[0088] Among them, the spatiotemporal combined index is optimized according to the range query and the nearest neighbor query, and the Z index and the S index are constructed respectively to obtain the spatiotemporal index data set, including: performing incremental clustering processing on the spatial index data in the spatiotemporal combined index, dynamically allocating the newly added index data to the existing spatial clusters or creating new spatial clusters to obtain the initial clustering results; constructing a recursive and extensible aggregation function based on the initial clustering results, dynamically updating the cluster boundaries to obtain cluster boundary data; performing overlap analysis on the cluster boundary data, calculating the spatial similarity between adjacent clusters, and obtaining cluster overlap data; setting a fusion threshold according to the cluster overlap data, An incremental fusion operation is performed on clusters whose overlap exceeds a threshold to obtain a fused cluster set; the spatial mapping relationship of the Z index is updated based on the fused cluster set, and the index hierarchical structure is dynamically adjusted to obtain an optimized Z index; the optimized Z index is input into the nearest neighbor comparator, the distance relationship between the center points of each cluster is calculated, and a spatial proximity network is constructed to obtain an optimized S index; a double-layer index structure is established based on the optimized Z index and the optimized S index, and the index data of the range query and the nearest neighbor query are stored separately to obtain hierarchical index data; the hierarchical index data is indexed and merged, and the query results of the Z index and the S index are unified to obtain a spatiotemporal index dataset.
[0089] In a specific embodiment, the process of executing step 103 may specifically include the following steps:
[0090] Data partition planning is performed based on the spatiotemporal index dataset, and data with similar spatial location and time characteristics are divided into the same partition to obtain the initial partition scheme;
[0091] Perform size threshold detection on each partition in the initial partitioning scheme, and perform sharding processing on the data according to the preset threshold to obtain partition sharding data;
[0092] Input the partition and shard data into the MMP distributed storage architecture, configure independent cache space on the storage node, and obtain distributed cache data;
[0093] Analyze the compression algorithm of distributed cache data, select the optimal compression algorithm to compress data according to data characteristics, and obtain node compressed data;
[0094] A master-slave replication mechanism is established based on node compressed data, and the compressed data is backed up and synchronized to obtain a backup data set. A partition index table is established for the backup data set to record data distribution information and storage location mapping to obtain partition compressed data.
[0095] Specifically, data partition planning is performed based on the spatiotemporal index data set to ensure that data with similar spatial locations and time characteristics are divided into the same storage partition, thereby reducing the cross-region access cost during query. In the process of data partitioning, the spatial proximity relationship and time range of each data point are calculated based on the spatial index (such as the Hilbert Curve index) and the time index (such as the B+ tree time index) to maintain a high degree of spatial locality when storing data. Assume that the three-dimensional coordinates of a data point are represented as , whose timestamp is , then through the Hilbert mapping function and time index functions Calculate its partition identifier:
[0096]
[0097] in, represents the partition step of the Hilbert index, Indicates the time window size of the time index, Represents the partition number to which the data belongs. Through this method, all data points that are close in space and time are mapped to the same partition number to form an initial partitioning scheme. The size threshold of each partition in the initial partitioning scheme is tested to ensure that the amount of data in a single partition is not too large or too small to affect storage efficiency or query performance. The data is sharded according to the preset threshold to generate balanced storage data blocks. Assume that the storage data size of each partition is , the system detects whether the storage threshold is met :
[0098]
[0099] The sharding method is based on the local density of data points. For example, for data points gather , according to the data density Divide into sub-areas:
[0100] ;
[0101] When the data density is too high, subdivide the spatial region to keep the data size of each shard close to After sharding is completed, the partitioned data is input into the MMP (Multi-node Memory Parallelism) distributed storage architecture, and independent cache space is configured on the storage node to increase the speed of data access. In the MMP architecture, data is numbered according to the partition. Distributed to multiple storage nodes, and maintain independent cache pools on each storage node to reduce frequent access to the disk. Suppose a query request needs to access partition , the system first checks whether the data of the partition has been cached:
[0102] ;
[0103] If the cache hits, the data is read directly from the memory cache. Otherwise, the data needs to be loaded from the disk and stored in the cache pool to speed up access during subsequent queries. In order to further optimize storage resources, the compression algorithm of the distributed cache data is analyzed, and the optimal compression algorithm is selected according to the data characteristics to reduce storage occupancy. The selection of the compression algorithm depends on the type and redundancy of the data. For example, a lossless compression algorithm (such as LZ4, Zstd) is used for high-resolution remote sensing image data, and a lightweight compression algorithm (such as Snappy) is used for vector data. Suppose a data block With original size , whose compression ratio is defined as:
[0104]
[0105] in, is the compressed data size. If Greater than a certain minimum compression ratio threshold , then the compression algorithm is applied, otherwise a more efficient compression method is used. After completing data compression, a master-slave replication mechanism is established based on node compressed data to improve the data disaster tolerance capability, and the compressed data is backed up and synchronized to generate a backup data set. The master-slave replication strategy adopts asynchronous or semi-synchronous mode to ensure that the data can be copied to the backup node in time after being written to the master node, thereby avoiding data loss. Assume that a data block If replication is required on multiple storage nodes, the replication strategy is as follows:
[0106]
[0107] in, Represents the storage node that stores the data block, ensuring that even if a node fails, the data can still be restored from the backup node. Synchronization delay between the master node and the slave node Calculated as:
[0108] ;
[0109] if If the data synchronization time is too long, optimize the data synchronization strategy, such as increasing concurrency or using incremental synchronization technology to reduce data synchronization time. After the data backup is completed, a partition index table is established for the backup data set to record the data distribution information and storage location mapping relationship to form the final partition compressed data. The role of the partition index table is to speed up data location during query, avoid global scanning, and improve retrieval efficiency. The partition index table contains the following information:
[0110] ;
[0111] When querying, the system first searches the partition index table to determine the storage node where the query data is located, and quickly loads the data based on the storage location. For example, a query request needs to access a data block , the system looks up its storage location based on the index table:
[0112]
[0113] Then directly from the node Read data to optimize query performance.
[0114] Among them, data partition planning is carried out based on the spatiotemporal index data set, and data with similar spatial location and time characteristics are divided into the same partition to obtain an initial partitioning scheme, including: performing feature decomposition operation on the spatiotemporal index data set, separating the spatiotemporal features into global dynamic features and local time-varying features, and obtaining feature decomposition data; constructing a two-layer feature representation model based on the feature decomposition data, performing spatiotemporal stability analysis on the global dynamic features, and obtaining a global feature vector; performing temporal correlation calculation on the local time-varying features, extracting the local spatiotemporal change pattern, and obtaining a local feature vector; inputting the global feature vector and the local feature vector into the mutual information constraint model, calculating the feature independence score, and obtaining feature decoupled data; establishing a partition mapping relationship based on the feature decoupling data, clustering and grouping data with similar feature patterns, and obtaining a partition clustering result; optimizing the spatial continuity of the partition clustering result, merging adjacent partitions and eliminating fragmented partitions, and obtaining an optimized partitioning scheme; constructing a partition adaptive mechanism based on the optimized partitioning scheme, dynamically adjusting the partition boundaries and sizes, and obtaining dynamic partition data; performing spatiotemporal consistency verification on the dynamic partition data, and verifying the rationality of the partition division results to obtain an initial partitioning scheme.
[0115] In this embodiment, the partition and shard data are input into the MMP distributed storage architecture, and independent cache space is configured on the storage node. Before obtaining the distributed cache data, it also includes: performing delay sensitivity analysis on the partition and shard data, dividing the spatiotemporal data jobs into high-sensitivity jobs and low-sensitivity jobs, and obtaining job sensitivity classification data; constructing a multi-granularity time slot scheduling model based on the job sensitivity classification data, allocating different time slot offsets to jobs with different sensitivities, and obtaining a time slot allocation scheme; performing bandwidth-aware mapping on the storage nodes according to the time slot allocation scheme, allocating priority transmission time slots to data with high bandwidth transmission requirements, and obtaining a node mapping strategy; setting topology reconstruction thresholds for the node mapping strategy, and hierarchically configuring the reconstruction thresholds of ports with different priorities to obtain reconstruction configuration data; performing adaptive topology reconstruction based on the reconstruction configuration data, preferentially adjusting the bandwidth of low-priority ports, and obtaining topology optimization data; establishing a bandwidth dynamic allocation mechanism based on the topology optimization data, adjusting the port transmission bandwidth in real time, and obtaining a bandwidth configuration scheme; dynamically monitoring the throughput of the bandwidth configuration scheme, recording bandwidth utilization and transmission delay indicators, and obtaining performance monitoring data; adjusting port configuration parameters based on the performance monitoring data, balancing network delay and bandwidth utilization, and obtaining optimized configuration parameters.
[0116] In a specific embodiment, the process of executing step 104 may specifically include the following steps:
[0117] Parse query conditions on the partition compressed data, convert the SQL query statement into a query tree structure, and obtain query parsing data;
[0118] Based on the query and parsing data, a time dimension index, a space dimension index, and an attribute dimension index are constructed to obtain a multi-dimensional index set;
[0119] Establish a cost calculation model for the multi-dimensional index set, calculate the execution costs of different query paths, and obtain query cost data;
[0120] Select the optimal query path according to the query cost data, generate an optimized execution plan, and split the optimized execution plan into tasks, dividing the query task into multiple subtask units to obtain a subtask sequence;
[0121] A task scheduling strategy is established based on the subtask sequence, and resources are allocated and the degree of parallelism is set for each subtask to obtain a task allocation plan.
[0122] Specifically, query conditions are parsed for the partitioned compressed data, and the SQL query statement is converted into a query tree structure to obtain query parsing data. During the SQL query parsing process, the structure of the SQL statement is parsed, including clauses such as SELECT, FROM, WHERE, GROUP BY, and the logical representation of the query is constructed. Each node of the query tree represents an operation, such as filtering, aggregation, and connection. Among them, each leaf node represents a query condition, and the internal nodes represent logical operations. The construction of the query tree facilitates subsequent query optimization and index selection. After generating the query parsing data, a time dimension index, a space dimension index, and an attribute dimension index are constructed based on the query conditions to optimize query efficiency and reduce the amount of data scanning. The time dimension index uses a B+ tree or a time interval index, so that the query can quickly lock the target time range. For example, for a time range query:
[0123]
[0124] B+ trees efficiently locate time segments that meet the conditions. Spatial dimension indexes are based on Hilbert curves or R-trees, allowing queries to efficiently perform spatial range searches, such as:
[0125]
[0126] The hierarchical structure of the R-tree allows the query to recursively narrow the search range from the root node, while the Hilbert index optimizes the access order of adjacent data through one-dimensional encoding. The attribute dimension index uses a hash index or an inverted index to speed up attribute-based screening. For example, for the hash index of attribute A:
[0127]
[0128] in, is the hash value, is the size of the hash table, and this method is used to quickly find data that meets the attribute query conditions. After building the index, a cost calculation model is established to evaluate the execution cost of different query paths and obtain query cost data. The query cost is determined by factors such as data scanning volume, index utilization, calculation complexity, and 1 / O overhead. Assume that the cost calculation formula for a query is as follows:
[0129]
[0130] in, represents the total cost of the query, Indicates the amount of data scanned, Represents the effectiveness of index usage. represents the computational complexity, Represents the amount of data transferred. is the weight coefficient. The query optimizer calculates the cost of different query solutions and selects the execution path with the lowest cost. Based on the query cost data, the optimal query path is selected and an optimized execution plan is generated. In this process, the query task is split into multiple subtask units according to the index hit situation, query cost and parallelism of the query task to obtain a subtask sequence. Assuming that the query task needs to access data in different partitions, the query task is split into multiple parallel subtasks:
[0131]
[0132] Among them, each Represents a subtask, corresponding to a specific spatial, temporal or attribute range. For example, if the query involves multiple time segments:
[0133]
[0134] The query is split into three subtasks, which are executed in parallel on different computing nodes to improve query efficiency. Based on the subtask sequence, a task scheduling strategy is established to allocate resources and set the degree of parallelism for each subtask to obtain the final task allocation solution. The task scheduling strategy mainly considers the load balancing of computing resources, task dependencies, and data locality to ensure that the query task can be executed in the best way. Assume that the query task Need to be on the compute node If it is executed on, the task scheduling is expressed as:
[0135]
[0136] in, Representation Node The current load, Representative query task The system will select the computing node with the lowest execution cost for scheduling. In order to improve the parallelism of query tasks, a task splitting strategy is adopted to further divide larger query tasks into smaller computing units to be executed simultaneously on multiple computing nodes. For example, a complex GIS spatial query is split into parallel queries of multiple regions, and each computing node is only responsible for processing the corresponding spatial region, thereby reducing computing pressure and accelerating query execution.
[0137] In a specific embodiment, the process of executing step 105 may specifically include the following steps:
[0138] Receive and parse the task allocation plan, assign subtasks to cluster processing nodes, and obtain node task data;
[0139] Build a task execution queue based on node task data, monitor the load status of each processing node, and obtain the task scheduling sequence;
[0140] Perform data pre-fetching operations according to the task scheduling sequence, load the data required for the task into the memory cache, obtain pre-loaded data, perform parallel calculations on the pre-loaded data, perform distributed retrieval operations, and obtain node retrieval results;
[0141] Anomaly monitoring is performed based on the node retrieval results, and computing resources are reallocated to the tasks with abnormal execution to obtain the reallocation results. The reallocation results and normal execution results are merged to obtain the original query results.
[0142] Specifically, the task allocation scheme is received and parsed, the query task is split into multiple subtasks, and they are assigned to appropriate cluster processing nodes to obtain node task data. In the task parsing process, the spatial scope, time scope, and computational complexity of the query task are analyzed, and the subtasks are assigned to the nodes with the best computing resources based on the load balancing strategy. Assume a query task Need access to dataset The query is split into multiple subtasks for the records that meet the conditions. , where each Responsible for processing specific sub-datasets ,Right now:
[0143]
[0144] The task scheduling system is based on the current load of the computing node. And data locality To optimize the task allocation scheme, ensure that the computing tasks are as close as possible to the data storage location to reduce data transmission overhead and optimize computing efficiency. After the tasks are assigned to each computing node, a task execution queue is built based on the node task data to ensure that the tasks can be executed in order according to priority and computing resources. The construction process of the task execution queue involves task priority calculation and load balancing management. Assuming that the task The execution cost is determined by the data size , computational complexity and current node load The execution priority of the task is calculated as:
[0145] ;
[0146] in, is a weight coefficient used to balance the relationship between data volume, computational complexity, and node load. The system sorts tasks by priority and dynamically adjusts the order of task execution to ensure that high-priority tasks can be processed as quickly as possible. Data prefetching is performed based on the task execution queue to load the data required for the task into the memory cache in advance, thereby reducing the 1 / O overhead during query and improving query speed. The data prefetching process involves data sharding, cache strategy optimization, and data consistency management. Assume that the data block number that the computing node needs to access is , the system checks whether these data blocks are already cached in memory:
[0147] ;
[0148] If the data is already cached, it is read directly from the cache. Otherwise, the system will perform a data load operation and store it in the cache for subsequent queries. In order to optimize the cache hit rate, the LRU (Least Recently Used) or LFU (Least Frequently Used) strategy is adopted to ensure that frequently accessed data is kept in the cache first. After completing the data pre-fetching, the computing nodes perform parallel computing on the pre-loaded data and use the distributed computing framework to perform efficient data retrieval. Assume that the query involves a spatial range , the computing node accelerates the query based on the spatial index (such as R-tree or Z-index) to retrieve the data points that meet the conditions:
[0149]
[0150] in, represents a data point, Representative Inquiry The generated sub-result set. Since the computing tasks are executed in parallel, the query response time is reduced through distributed computing, that is:
[0151]
[0152] in, Represents the actual query time, represents the serial execution time, Represents the parallelism of the computing node. During the computing process, due to problems such as uneven load, data anomalies, or network interruptions on the computing nodes, anomaly monitoring is performed based on the node retrieval results, and computing resources are reallocated to the abnormally executed tasks to obtain the reallocation results. The anomaly detection method is based on the task execution time threshold and result consistency check. For example, if the execution time of a computing node is Exceeded the maximum allowed time , then the task execution is judged to be abnormal:
[0153] ;
[0154] The system reallocates the abnormal task to other computing nodes for execution and ensures that the task can be completed within the specified time. At the same time, in order to ensure data consistency, check whether the results returned by each computing node meet the query constraints and compare the results, such as:
[0155]
[0156] If the hash values do not match, it means that the data is lost or damaged and needs to be recalculated or restored from a backup node. After all tasks are completed, the redistribution results and normal execution results are merged to obtain the complete original query results. In the process of merging data, data deduplication, sorting, aggregation and other operations are performed to ensure that the final result data returned is complete, accurate and non-duplicate. For example, if the query involves multiple data partitions , the final query result is expressed as:
[0157]
[0158] in, Represents the query result that is finally returned. Represents the total number of data partitions. In order to improve the efficiency of data merging, a distributed merge sort or distributed hash aggregation method is used to reduce the computational burden and optimize the data merging speed.
[0159] In a specific embodiment, the process of executing step 106 may specifically include the following steps:
[0160] Perform data integrity check on the original query results to obtain verification result data, and reconstruct the topological relationship of geographic elements based on the verification result data, build a spatial adjacency relationship graph, and obtain spatial topological data;
[0161] Input the spatial topology data into the result cache mechanism, establish cache indexes for frequently accessed query results, and obtain cache structure data;
[0162] Perform query result paging management based on cache structure data, split data according to preset paging size, and obtain paging result set;
[0163] Based on preset second data format information, performing data format conversion on the paginated result set to obtain format conversion data, where the second data format information includes GeoJSON, Shapefile or KML;
[0164] A data security control mechanism is established based on the format conversion data, and the query results are filtered for permissions to obtain the target query data.
[0165] Specifically, the original query results are tested for data integrity to verify whether the data is lost, duplicated, or erroneous during storage and calculation. The integrity test is based on hash verification, data comparison, and topological consistency analysis. Hash verification is used to detect the consistency of data content, while data comparison detects anomalies by cross-validating the results of different data sources. Assume that the original query result dataset is , which contains Records , the integrity check is verified by calculating the hash value:
[0166]
[0167] If the hash value of the query result does not match the expected checksum, that is, , it indicates that the data has been tampered with or lost and needs further repair. After completing the data integrity check, the topological relationship of the geographic elements is reconstructed based on the verification result data to build a complete spatial adjacency relationship graph and obtain spatial topological data. The topological relationship mainly includes the connection, adjacency, inclusion and intersection relationships between spatial objects, which are calculated using topological rules. Assume that the query result contains multiple polygon objects , then define their adjacency relationship for:
[0168] ;
[0169] By traversing all spatial objects and calculating their adjacency relationships, a spatial adjacency graph is constructed, where each node represents a geographic feature and each edge represents the adjacency relationship between two objects. For example, in road network data, a topological graph of nodes (intersections) and edges (roads) is constructed to support path calculation and spatial analysis. After obtaining the spatial topological data, it is input into the result cache mechanism to optimize query efficiency and reduce repeated calculations. The core goal of the cache mechanism is to establish a cache index for frequently accessed data, so that subsequent identical or similar queries can directly obtain results from the cache without re-calculating. Assume that the cache system contains cache blocks, each cache block stores a query result , then the storage structure of the cache index is expressed as:
[0170]
[0171] For new queries , calculate its hash value And look for a match in the cache:
[0172] ;
[0173] If the cache hits (i.e. If the query result exists, the cache result is returned directly. Otherwise, the query result is stored in the cache and cache management is performed according to the LRU (Least Recently Used) or LFU (Least Frequently Used) strategy to ensure that frequently accessed data is kept in the cache first. After the cache index is established, the query result is paging managed according to the cache structure data to ensure that the returned data can be divided according to the preset paging size, thereby improving data transmission efficiency and optimizing query response time. The key to paging management is to determine the appropriate paging size. , based on the size of the data record and the maximum amount of data requested by the client Perform the calculation:
[0174] ;
[0175] in, is the total number of records in the query results. The paginated data set is represented as:
[0176] ;
[0177] Among them, each Include The data is collected and managed according to the paging index so that users can efficiently perform page-turning queries. After obtaining the paging result set, the data is formatted based on the preset second data format information to meet the needs of different application scenarios and generate format conversion data. Since different GIS platforms and data analysis tools use different data formats, the goal of data format conversion is to convert the query results into the standard format required by users, such as GeoJSON, Shapefile or KML. For example, GeoJSON is a lightweight format suitable for Web GIS applications, while the Shapefile format is suitable for traditional GIS software (such as ArcGIS), and its data is stored in multiple files (.shp, .shx, dbf), which requires coordinate projection conversion and field mapping, while the KML format is used for Google Earth for three-dimensional visualization. Therefore, during the format conversion process, the consistency of the spatial reference system, attribute fields and geometric structure of the data is ensured to avoid data loss or distortion. After completing the data format conversion, a data security control mechanism is established based on the format conversion data to ensure that the query results are securely accessed under different user permissions and prevent unauthorized users from obtaining sensitive data. Data security control is implemented based on access control lists (ACLs) or role-based access control (RBAC). Assume that the user Need to access query results , then its access rights Determined by the following conditions:
[0178] ;
[0179] Desensitize the data to prevent the leakage of sensitive fields, reduce the exposure of accurate data, and improve data security. Through the above steps, the target query data is obtained.
[0180] The above describes the task scheduling method of spatiotemporal big data in the embodiment of the present invention. The following describes the task scheduling device of spatiotemporal big data in the embodiment of the present invention. Figure 2 An embodiment of a task scheduling device for spatiotemporal big data in an embodiment of the present invention includes:
[0181] The standardization processing module 201 is used to perform ETL cleaning and standardization processing on the original spatiotemporal data to obtain spatiotemporal data in a unified format that complies with the OGC specification;
[0182] A construction module 202 is used to perform a spatial dimension reduction operation on the unified format spatiotemporal data by using a space filling curve algorithm and construct a spatiotemporal index to obtain a spatiotemporal index data set;
[0183] The compression module 203 is used to determine the data partitioning strategy according to the spatiotemporal index data set, distribute and compress the spatiotemporal data in the unified format according to the spatiotemporal characteristics, and obtain partitioned compressed data;
[0184] The task allocation module 204 is used to establish a multi-dimensional query index for the partitioned compressed data to obtain a task allocation plan;
[0185] An execution module 205 is used to execute distributed data retrieval based on the task allocation scheme to obtain original query results;
[0186] The reconstruction module 206 is used to perform data integrity verification and spatial relationship reconstruction on the original query results to obtain target query data.
[0187] Through the coordinated cooperation of the above components, through ETL data cleaning and standardization processing, the problem of data heterogeneity is solved, unified management of data from different sources and in different formats is achieved, and the difficulty of data integration is reduced; the space filling curve algorithm is used for dimensionality reduction processing, combined with the construction of Z index and S index, sub-second spatial retrieval and sub-second attribute retrieval of tens of millions of spatiotemporal data are achieved; based on the MMP distributed storage architecture and cache storage mechanism, high-speed data writing is achieved, with a writing rate of 120,000 records per second. At the same time, with the support of multiple compression algorithms, the maximum compression rate can reach 80%, effectively saving storage space; through the establishment of multi-dimensional query indexes, the construction of redundant data tables is avoided, and the storage resource occupation and data maintenance complexity are reduced; based on the distributed task scheduling mechanism of dynamic load balancing, the reasonable allocation of computing resources is achieved, effectively reducing the performance pressure of each node; through integrity verification and spatial relationship reconstruction, the accuracy and consistency of query results are ensured, and data output in multiple standard formats is supported, which improves the availability and interoperability of data.
[0188] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the task scheduling method for spatiotemporal big data.
[0189] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0190] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0191] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A task scheduling method for spatiotemporal big data, characterized in that: The method comprises: Perform ETL cleaning and standardization on the original spatiotemporal data to obtain spatiotemporal data in a unified format that complies with OGC specifications; Performing a spatial dimension reduction operation on the unified format spatiotemporal data by using a space filling curve algorithm and constructing a spatiotemporal index to obtain a spatiotemporal index data set; Determine a data partitioning strategy based on the spatiotemporal index data set, distribute and compress the unified format spatiotemporal data according to the spatiotemporal features to obtain partitioned compressed data; specifically include: perform data partitioning planning based on the spatiotemporal index data set, divide data with similar spatial location and time features into the same partition, and obtain an initial partitioning scheme; wherein obtaining the initial partitioning scheme includes: perform feature decomposition operation on the spatiotemporal index data set, separate the spatiotemporal features into global dynamic features and local time-varying features, and obtain feature decomposition data; construct a two-layer feature representation model based on the feature decomposition data, perform spatiotemporal stability analysis on the global dynamic features, and obtain a global feature vector; perform time series correlation calculation on the local time-varying features, extract the local spatiotemporal change pattern, and obtain a local feature vector; input the global feature vector and the local feature vector into the mutual information constraint model, calculate the feature independence score, and obtain feature decoupled data; establish a partition mapping relationship based on the feature decoupling data, cluster and group data with similar feature patterns, and obtain partition clustering results; The clustering results are optimized for spatial continuity, adjacent partitions are merged and fragmented partitions are eliminated to obtain an optimized partitioning scheme; a partition adaptive mechanism is constructed based on the optimized partitioning scheme, and the partition boundaries and sizes are dynamically adjusted to obtain dynamic partition data; the dynamic partition data is checked for spatiotemporal consistency, and the rationality of the partition division results is verified to obtain an initial partitioning scheme; a size threshold detection is performed on each partition in the initial partitioning scheme, and the data is sliced according to a preset threshold to obtain partition slice data; the partition slice data is input into the MMP distributed storage architecture, and an independent cache space is configured on the storage node to obtain distributed cache data; a compression algorithm analysis is performed on the distributed cache data, and the optimal compression algorithm is selected according to the data characteristics to perform data compression to obtain node compressed data; a master-slave replication mechanism is established based on the node compressed data, and the compressed data is backed up and synchronized to obtain a backup data set, and a partition index table is established for the backup data set to record the data distribution information and the storage location mapping relationship to obtain partition compressed data; Establishing a multi-dimensional query index for the partitioned compressed data to obtain a task allocation plan; Execute distributed data retrieval based on the task allocation scheme to obtain original query results; The original query result is subjected to data integrity verification and spatial relationship reconstruction to obtain target query data.
2. The task scheduling method for spatiotemporal big data according to claim 1 is characterized in that: The ETL cleaning and standardization of the original spatiotemporal data to obtain spatiotemporal data in a unified format that complies with OGC specifications includes: Collecting original spatiotemporal data from satellite remote sensing equipment, unmanned aerial vehicle equipment and sensor network equipment, and identifying the data source of the original spatiotemporal data to obtain data source type information; Performing format parsing on the original spatiotemporal data according to the data source type information to obtain first data format information, and performing unified data structure conversion according to the first data format information to obtain preliminary conversion data, wherein the first data format information includes WKB, WKT, WMS and WMTS; Extracting spatial attributes from the preliminary converted data, parsing geographic coordinates, timestamps and attribute fields into a standard format, and obtaining spatial attribute data; Based on the spatial attribute data, a UDF SQL interface of the ANSI SQL standard is constructed to calculate the topological relationship of the vector data to obtain topological relationship data; Correcting abnormal values, missing values and duplicate values in the topological relationship data to obtain quality verification data; The quality verification data is reconstructed according to the OGC specification, and the spatial reference system, time format and attribute structure are unified to obtain unified format spatiotemporal data that conforms to the OGC specification.
3. The task scheduling method for spatiotemporal big data according to claim 1 is characterized in that: The method of performing a spatial dimension reduction operation on the unified format spatiotemporal data by using a space filling curve algorithm and constructing a spatiotemporal index to obtain a spatiotemporal index data set includes: Perform geographic vector spot analysis on the unified format spatiotemporal data, extract spatial coordinate information and time information, and obtain spatiotemporal feature data; Input the spatiotemporal feature data into a Z-curve space filling curve algorithm, perform one-dimensional mapping calculation on the three-dimensional space coordinates, and obtain Z sequence encoding data; Performing Hilbert-curve transformation on the Z sequence coded data, calculating the spatial proximity relationship and generating a spatial coding sequence to obtain one-dimensional spatial index data; Building a spatial multi-layer index tree based on the one-dimensional spatial index data, storing index information of different granularities in layers, and obtaining a layered index structure; Expanding the hierarchical index structure in time dimension, constructing a time index tree and associating it with the spatial index to obtain a time-space combined index; The spatiotemporal combined index is optimized according to range query and nearest neighbor query, and a Z index and an S index are constructed respectively to obtain a spatiotemporal index data set.
4. The task scheduling method for spatiotemporal big data according to claim 1, characterized in that: The step of establishing a multi-dimensional query index for the partitioned compressed data to obtain a task allocation scheme includes: Parsing query conditions on the partition compressed data, converting the SQL query statement into a query tree structure, and obtaining query parsing data; Based on the query parsing data, a time dimension index, a space dimension index and an attribute dimension index are constructed to obtain a multi-dimensional index set; Establishing a cost calculation model for the multi-dimensional index set, calculating the execution costs of different query paths, and obtaining query cost data; Selecting an optimal query path according to the query cost data, generating an optimized execution plan, and performing task splitting on the optimized execution plan, dividing the query task into multiple subtask units, and obtaining a subtask sequence; A task scheduling strategy is established based on the subtask sequence, and resources are allocated and the degree of parallelism is set for each subtask to obtain a task allocation plan.
5. The task scheduling method for spatiotemporal big data according to claim 1, characterized in that: The performing of distributed data retrieval based on the task allocation scheme to obtain original query results includes: Receiving and parsing the task allocation scheme, allocating subtasks to cluster processing nodes, and obtaining node task data; Building a task execution queue based on the node task data, monitoring the load status of each processing node, and obtaining a task scheduling sequence; Execute a data pre-fetch operation according to the task scheduling sequence, load the data required for the task into the memory cache to obtain the pre-loaded data, perform parallel calculation on the pre-loaded data, execute a distributed search operation, and obtain a node search result; Based on the node retrieval result, abnormality monitoring is performed, computing resources are reallocated to the abnormally executed tasks to obtain reallocation results, and the reallocation results are merged with the normal execution results to obtain the original query results.
6. The task scheduling method for spatiotemporal big data according to claim 1, characterized in that: The performing of data integrity verification and spatial relationship reconstruction on the original query result to obtain target query data includes: Performing data integrity check on the original query result to obtain verification result data, and reconstructing the topological relationship of the geographic elements based on the verification result data, constructing a spatial adjacency relationship graph, and obtaining spatial topological data; Inputting the spatial topological data into a result cache mechanism, establishing a cache index for frequently accessed query results, and obtaining cache structure data; Perform query result paging management according to the cache structure data, split the data according to the preset paging size, and obtain a paging result set; Based on preset second data format information, performing data format conversion on the paginated result set to obtain format conversion data, wherein the second data format information includes GeoJSON, Shapefile or KML; A data security control mechanism is established based on the format conversion data, and the query results are filtered for permissions to obtain target query data.
7. A task scheduling device for spatiotemporal big data, characterized in that: The method for scheduling spatiotemporal big data tasks according to any one of claims 1 to 6, wherein the task scheduling device for spatiotemporal big data comprises: The standardization processing module is used to perform ETL cleaning and standardization on the original spatiotemporal data to obtain spatiotemporal data in a unified format that complies with OGC specifications; A construction module, used to perform a spatial dimension reduction operation on the unified format spatiotemporal data and construct a spatiotemporal index by using a space filling curve algorithm to obtain a spatiotemporal index data set; A compression module is used to determine a data partitioning strategy according to the spatiotemporal index data set, distribute and compress the unified format spatiotemporal data according to spatiotemporal characteristics, and obtain partitioned compressed data; A task allocation module, used to establish a multi-dimensional query index for the partition compressed data to obtain a task allocation plan; An execution module, used for executing distributed data retrieval based on the task allocation scheme to obtain original query results; The reconstruction module is used to perform data integrity verification and spatial relationship reconstruction on the original query results to obtain target query data.
8. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the task scheduling method for spatiotemporal big data as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Space-time data management method
CN112328583A
Data compression and fusion method and device for digital earth space and medium
CN119202324A