Distributed mass spatio-temporal data rapid retrieval method

By constructing a unified metadata model and composite RowKey based on the ISO-19115-2 standard, and combining it with HBase BulkLoad and the Observer coprocessor, the problem of low efficiency in spatiotemporal data retrieval in existing technologies is solved, achieving efficient retrieval and load balancing of massive spatiotemporal data.

CN120873034APending Publication Date: 2025-10-31HENAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510982229.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

The existing "relational database + shared file system" architecture has inherent bottlenecks in horizontal scaling capabilities, concurrent query processing, and multidimensional index construction, making it difficult to meet the processing requirements of high throughput and low latency. Furthermore, it lacks a unified metadata extraction model and an automated data entry mechanism, resulting in poor efficiency in spatiotemporal data retrieval.

Method used

A unified metadata model conforming to the ISO-19115-2 standard is constructed. MapReduce is used for parallel parsing and field normalization of multi-source heterogeneous formats to generate composite RowKeys. Load-balanced storage is achieved through HBase's BulkLoad mechanism. Spatial relationship is refined by combining the Observer coprocessor, and a retrieval process of "coarse screening - fine sorting - local deduplication - global merging" is adopted.

Benefits of technology

It enables rapid retrieval of massive spatiotemporal data, reduces the cost of accessing and maintaining new data sources, improves retrieval efficiency, enhances concurrent write throughput, and reduces network and I/O overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873034A_ABST
    Figure CN120873034A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of spatio-temporal data retrieval, in particular to a distributed massive spatio-temporal data quick retrieval method, which comprises the following steps of: analyzing spatio-temporal data; constructing RowKey, and creating a column family; constructing a metadata object; metadata storage is achieved through a BulkLoad mechanism, and user retrieval conditions are analyzed; the method comprises the following steps of: screening through a RegionServer, and carrying out space filtering through an Observer coprocessor deployed at a RegionServer end; and returning the local screening results returned by the RegionServers back to the driving node in a unified manner, executing data de-duplication and time sorting operations in a global dimension, and constructing a final retrieval result set. According to the method, RowKey indexing and coprocessor end-side filtering are utilized, a two-stage retrieval mechanism of global screening and local fine detection is achieved, and it is ensured that quick retrieval is achieved in a massive spatio-temporal data scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spatiotemporal data retrieval technology, specifically to a distributed method for rapid retrieval of massive spatiotemporal data. Background Technology

[0002] With the widespread deployment of high-resolution remote sensing satellites, vehicle-to-everything (V2X) positioning systems, cellular communication signaling, and Internet of Things (IoT) sensor networks, the frequency and volume of spatiotemporal data acquisition are growing exponentially. Typical data types include remote sensing imagery, trajectory data, and environmental observation records, with daily new data volumes reaching tens to hundreds of terabytes (TB), and the overall volume has entered the petabyte (PB) level. This type of data is characterized by heterogeneous sources, complex formats (such as GeoTIFF, NetCDF, CSV, JSON, etc.), and inconsistent spatial reference systems, posing significant challenges to centralized storage and efficient retrieval.

[0003] The existing "relational database + shared file system" architecture has inherent bottlenecks in horizontal scalability, concurrent query processing, and multidimensional index construction, making it difficult to meet the demands for high throughput and low latency. Although distributed databases such as HBase have advantages in storage scalability, directly using auto-incrementing or random RowKeys as primary keys can easily lead to uneven data distribution across Regions, resulting in write hotspots and cross-Region full table scans, severely impacting query performance. Furthermore, the current lack of a unified metadata extraction model and automated data insertion mechanism when dealing with multi-source heterogeneous data results in poor system maintainability, high access costs, and difficulty in adapting to complex and ever-changing spatiotemporal data scenarios, leading to poor spatiotemporal data retrieval efficiency. Summary of the Invention

[0004] To address the technical problem of poor efficiency in spatiotemporal data retrieval, this invention proposes a distributed method for rapid retrieval of massive spatiotemporal data.

[0005] This invention provides a distributed method for rapid retrieval of massive spatiotemporal data. The method includes: acquiring the original spatiotemporal data file; using MapReduce to perform parallel parsing and field normalization on multi-source heterogeneous formats (such as XML, JSON, TXT, and TIF) to construct a unified metadata model conforming to the ISO-19115-2 standard and generate a spatiotemporal parsing dataset; generating a 5-bit GeoHash code based on latitude and longitude, concatenating it with satellite identifiers and shooting dates, and obtaining a three-bit prefix through MD5 (Message DigestAlgorithm5) hashing to construct a composite RowKey; importing the data into HDFS using a pre-partitioning strategy, and utilizing HBase's BulkLoad mechanism to achieve parallel loading of HFile files and load-balanced storage; receiving user search conditions, parsing the spatial and temporal ranges, generating corresponding GeoHash lists and time segments, and constructing multiple Scan start and end keys; firstly, the RegionServer node performs initial screening in parallel based on the RowKey prefix and column value filter; then, by deploying an Observer coprocessor, the JTS spatial computation library is called on the server side to complete the spatial relationship fine-tuning; finally, the local results are returned in parallel to the driving node and globally aggregated to obtain the retrieval results. This invention utilizes RowKey indexing and coprocessor-side filtering to implement a two-level retrieval mechanism of first global screening and then local fine-tuning, ensuring fast retrieval in scenarios with massive spatiotemporal data. Specifically, a distributed method for fast retrieval of massive spatiotemporal data may include the following steps:

[0006] Step 1: Analyze the spatiotemporal data acquired from different sources;

[0007] Step 2: Based on the parsed spatiotemporal data, construct the RowKey, create column families, and extract metadata information;

[0008] Step 3: Based on the constructed RowKey and the extracted metadata information, construct a metadata object;

[0009] Step 4: Based on the constructed metadata object, the metadata is loaded into the database using the BulkLoad mechanism, and the user's search criteria are parsed.

[0010] Step 5: Verify the validity of the parameters based on the parsed user search criteria;

[0011] Step 6: If the verification parameters are valid, construct the RowKey to generate the scan range, and filter it through the RegionServer, thereby performing spatial filtering through the Observer coprocessor deployed on the RegionServer.

[0012] Step 7: The local filtering results returned by multiple RegionServers are uniformly sent back to the driver node to perform data deduplication and time sorting operations on a global dimension, and to build the final set of search results.

[0013] Optionally, constructing the RowKey includes:

[0014] For the parsed spatiotemporal data, the latitude and longitude of the center point are first encoded into a five-bit GeoHash, and then concatenated with the satellite identifier and the shooting date. The three-bit prefix is ​​obtained by MD5 hashing, and a composite RowKey with spatial locality and hash dispersion is constructed to support subsequent partition parallel access and load balancing.

[0015] Optionally, the construction of the metadata object based on the constructed RowKey and the extracted metadata information includes:

[0016] Based on the constructed RowKey and the extracted metadata information, the corresponding metadata items are extracted according to the unified metadata model conforming to the ISO-19115-2 standard, and a metadata object is generated.

[0017] Optionally, the unified metadata model includes at least UUID, GeoHash, SatelliteID, SensorID, SceneID, Resolution, CloudCover, BeginTime, EndTime, CenterLong, CenterLat, image four-corner coordinates, and thumbnails. Figure 2 Number base field.

[0018] Optionally, the formula for constructing a composite RowKey with spatial locality and hash dispersion is:

[0019] RowKey=PrefixID3+SatelliteID2+GeoHash(r)5+Time6

[0020] GeoHash(r) is the encoding after latitude and longitude coordinate transformation, with a length of 5 bits. GeoHash encoding stores spatially adjacent data in adjacent physical locations; Time represents the acquisition time of the remote sensing data, in the format YYMMDD.

[0021] PrefixID=MD5[(GeoHash(r)4+SatelliteID2+Time_Month)]%nums

[0022] Among them, PrefixID represents the prefix identifier, which is the first four bits of the GeoHash encoding of the r space region; SatelliteID represents the type identifier of the data source; Time_Month represents the year and month of the shooting time; and nums is the total number of partitions.

[0023] Optionally, nums determines the optimal number of Regions on each RegionServer based on a formula, which is:

[0024]

[0025] Rs memory is the amount of memory allocated by the server for HBase;

[0026] Set hbase.regionserver.global.memstore.size to 0.4;

[0027] hbase.hregion.memstore.flushsize represents the size of each Memstore, which is 128MB;

[0028] columnfamilies represents the number of column families.

[0029] Optionally, the metadata loading process via the BulkLoad mechanism includes:

[0030] The metadata object is written to HDFS, and then generated into a partitioned HFile file in the cluster through a MapReduce task. Finally, the HFile is atomically loaded into multiple Regions in the table using the HBase BulkLoad mechanism.

[0031] Optionally, the retrieval process includes:

[0032] By combining the Scan operation of the HBase database with column family filters, the data is retrieved and the remote sensing data that meets the search criteria is initially selected.

[0033] By implementing the RegionObserver.preGetOp method of the Observer coprocessor, the JTS spatial calculation engine is integrated on the RegionServer side. Through this integration, spatial relationship judgment is performed on the RegionServer side where the data is located.

[0034] Optionally, parsing user search criteria includes:

[0035] The query range and time interval of the user space are parsed to generate the corresponding GeoHash prefix list and time segment. All possible PrefixIDs are derived by combining the satellite type. Multiple start and end RowKey ranges are constructed on the client, mapped to multiple Scan tasks, and distributed in parallel to each RegionServer node of the HBase cluster.

[0036] Optionally, the filtering via RegionServer includes:

[0037] Each HBase RegionServer node independently receives Scan requests and filters candidate records based on the RowKey range in its local Region. Combined with column value filters, it completes the matching of sensor identifiers, realizing a distributed and parallel preliminary filtering process.

[0038] The present invention has the following beneficial effects:

[0039] First, by constructing a unified spatiotemporal metadata model based on the ISO-19115-2 standard, and using template-mapping rules to automatically extract data fields in various formats such as XML, JSON, TXT, and GeoTIFF, format normalization and field validation can be completed simultaneously during the data entry stage without manual intervention. This mechanism ensures the structural consistency of spatiotemporal data from different sources and significantly reduces the cost of accessing and maintaining new data sources, thereby improving the retrieval efficiency of spatiotemporal data.

[0040] Second, constructing a composite RowKey not only ensures spatial locality but also eliminates write hotspots through MD5 prefixes and pre-partitioning strategies, thereby improving concurrent write throughput.

[0041] Third, the standardized metadata is exported as an HFile using the MapReduce method and then atomically loaded into the target table using the HBaseBulkLoad tool, avoiding the large amount of random write overhead of WAL (Write Ahead Log) and MemStore.

[0042] Fourth, the retrieval end expands the user-given spatial polygon / circle range using MBR (Minimum Bounding Rectangle) and an eight-neighbor GeoHash grid, combines it with time slices, and then sends out Scan tasks in parallel. The RegionServer first performs a coarse screening by row key range, and then the Observer coprocessor calls JTS (Java TopologySuite) spatial operators to perform precise intersection or containment filtering. This "coarse screening—fine sorting—local deduplication—global merging" process minimizes network and I / O overhead. Attached Figure Description

[0043] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a flowchart of a distributed method for rapid retrieval of massive spatiotemporal data according to the present invention;

[0045] Figure 2 The flowchart below is a further detailed representation of the distributed, high-volume spatiotemporal data retrieval method of the present invention.

[0046] Figure 3 This is a flowchart of spatiotemporal data parsing in a distributed, high-volume spatiotemporal data fast retrieval method of the present invention;

[0047] Figure 4 This is a flowchart illustrating the metadata import process in a distributed, high-volume spatiotemporal data fast retrieval method of the present invention.

[0048] Figure 5 This is a flowchart illustrating the retrieval process in a distributed, high-volume spatiotemporal data retrieval method according to the present invention. Detailed Implementation

[0049] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the specific implementation methods, structures, features, and effects of the technical solution proposed according to the present invention are described in detail below with reference to the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0051] refer to Figure 1 The diagram illustrates the flow of some embodiments of a distributed, high-volume spatiotemporal data fast retrieval method according to the present invention. (Reference) Figure 2 The diagram illustrates a further detailed flowchart of a distributed, high-speed retrieval method for massive spatiotemporal data according to the present invention. Specifically, a distributed, high-speed retrieval method for massive spatiotemporal data may include the following steps:

[0052] Step 1: Analyze the spatiotemporal data obtained from different sources.

[0053] In some embodiments, spatiotemporal data parsing includes: using MapReduce to perform parallel parsing and normalization of spatiotemporal source data from different satellite platforms, sensors, and various formats such as XML, JSON, TXT, and TIF.

[0054] Specifically, the spatiotemporal data analysis in step 1 is as follows: Figure 3 As shown: First, data from different sources is read, and the basic structure information of the metadata is obtained according to the data type table. Through the mapping rule table, the system automatically parses and transforms the fields of the remote sensing data, mapping them to standard template fields. The field standards in the template library provide clear guidance for parsing and transformation, ensuring that the metadata conforms to a predetermined unified model. After completing the metadata standardization process, the system constructs a metadata object based on the parsing results. Only after constructing the RowKey can the complete metadata object be stored in HBase.

[0055] Step 2: Based on the parsed spatiotemporal data, construct the RowKey, create column families, and extract metadata information.

[0056] In some embodiments, RowKey construction includes: encoding the latitude and longitude of the center point into a five-digit GeoHash (with an accuracy of approximately 2.4km) from the parsed spatiotemporal data, concatenating it with the satellite identifier and the shooting date, and obtaining a three-digit prefix through MD5 hashing to construct a composite RowKey with spatial locality and hash dispersion, to support subsequent partitioned parallel access and load balancing. Creating column families includes: defining the column family MetaCF (core geographic and temporal attributes) and UsageCF (business extension attributes and image storage path). Each column family enables compression and BlockCache, combined with the region distribution characteristics, to improve read / write efficiency and data locality in a distributed environment.

[0057] As an example, the latitude and longitude of the spatiotemporal data center are obtained and a 5-bit GeoHash code is generated. The formula for constructing a composite RowKey with spatial locality and hash dispersion is as follows:

[0058] RowKey=PrefixID3+SatelliteID2+GeoHash(r)5+Time6

[0059] GeoHash(r) is the 5-bit encoding of latitude and longitude coordinates. GeoHash encoding stores spatially adjacent data in adjacent physical locations, facilitating efficient range lookups. Time represents the time the remote sensing data was captured, in YYMMDD format.

[0060] PrefixID=MD5[(GeoHash(r)4+SatelliteID2+Time_Month)]%nums

[0061] Among them, PrefixID represents the prefix identifier, which is the first four bits of the GeoHash encoding of the r space region; SatelliteID represents the type identifier of the data source; Time_Month represents the year and month of the shooting time; and nums is the total number of partitions.

[0062] As another example, nums allocates the optimal number of Regions across each RegionServer according to a formula to ensure a balance between MemStore capacity and RegionServer memory. The corresponding formula is:

[0063]

[0064] Rs memory is the amount of memory allocated by the server for HBase;

[0065] According to the official HBase recommendation, hbase.regionserver.global.memstore.size should be set to 0.4;

[0066] hbase.hregion.memstore.flushsize represents the size of each Memstore, which is typically 128MB;

[0067] columnfamilies represents the number of column families.

[0068] As another example, the MetaCF column family stores the core geographic and shooting attributes of the imagery, while the UsageCF column family stores the imagery's business attributes and storage path. The UsageCF column family is designed with greater flexibility, allowing new column qualifiers to be added dynamically according to business needs without modifying the table structure.

[0069] Step 3: Based on the constructed RowKey and the extracted metadata information, construct a metadata object.

[0070] In some embodiments, based on the constructed RowKey and extracted metadata information, corresponding metadata items are extracted according to a unified metadata model conforming to the ISO-19115-2 standard to generate a metadata object. The unified metadata model includes at least UUID, GeoHash, SatelliteID, SensorID, SceneID, Resolution, CloudCover, BeginTime, EndTime, CenterLong, CenterLat, image four-corner coordinates, and thumbnails. Figure 2 Number base field.

[0071] Step 4: Based on the constructed metadata object, the metadata is loaded into the database through the BulkLoad mechanism, and the user's search criteria are parsed.

[0072] In some embodiments, metadata import via the BulkLoad mechanism includes: writing metadata objects to HDFS, generating partitioned HFile files in the cluster using MapReduce tasks, and atomically loading the HFiles into multiple Regions in the table using the HBase BulkLoad mechanism, avoiding MemStore and WAL write bottlenecks and achieving efficient distributed import of massive amounts of data. Parsing user search conditions includes: parsing the user's spatial query range and time interval, generating a corresponding GeoHash prefix list and time segment, and deriving all possible PrefixIDs based on satellite type; constructing multiple start and end RowKey ranges on the client, mapping them to multiple Scan tasks, and distributing them in parallel to each RegionServer node in the HBase cluster.

[0073] As an example, the parsed metadata unified model is written to HDFS, and the MapReduce framework is used to process this data and generate HFile format files. The generated HFile files are then written to HDFS again. Through the BulkLoad operation, the HFile files in HDFS are loaded into the HBase cluster, and the data is distributed to the various Regions under the corresponding RegionServers according to predefined Region partitioning rules.

[0074] Specifically, the flowchart for metadata ingestion is as follows: Figure 4As shown: The parsed metadata is written to HDFS as a data source. The MapReduce framework is used to process this data and generate HFile format files. The generated HFile files are then written to HDFS again. Through the BulkLoad operation, the HFile files in HDFS are loaded into the HBase cluster, and the data is distributed to the corresponding RegionServers according to predefined Region partitioning rules.

[0075] Step 5: Verify the validity of the parameters based on the parsed user search criteria.

[0076] Step 6: If the verification parameters are valid, construct the RowKey to generate the scan range, and filter it through the RegionServer, thereby performing spatial filtering through the Observer coprocessor deployed on the RegionServer.

[0077] In some embodiments, precise filtering by the coprocessor includes: integrating the JTS spatial operation library into the Observer coprocessor deployed on the RegionServer side, and performing spatial relationship judgments (such as intersection and containment) on the nearest data node. Specifically, coarse filtering by the RegionServer includes: each HBase RegionServer node independently receives Scan requests and filters candidate records based on the RowKey range in the local Region; combining column value filters to complete precise matching of sensor identifiers, realizing a distributed and parallel preliminary filtering process, reducing the burden of cross-node scanning and network transmission.

[0078] Step 7: The local filtering results returned by multiple RegionServers are uniformly sent back to the driver node to perform data deduplication and time sorting operations on a global dimension, and to build the final set of search results.

[0079] In some embodiments, deduplication and sorting include: uniformly transmitting the local filtering results returned by multiple RegionServers back to the driver node, performing data deduplication and time sorting operations on a global dimension, and constructing the final set of search results.

[0080] As an example, the retrieval process includes two steps: First, a rapid coarse search is performed on the data using HBase's Scan operation combined with column family filters, reducing unnecessary data scanning and initially filtering out remote sensing data that meets the search criteria. Then, by implementing the RegionObserver.preGetOp method of the Observer coprocessor, the JTS (Java Topology Suite) spatial computation engine is integrated on the RegionServer side. This integration allows spatial relationship determinations (such as intersection and containment) to be performed directly on the RegionServer where the data resides, avoiding the high latency of transmitting data to the client before computation.

[0081] Specifically, the retrieval flowchart is as follows: Figure 5 As shown: After the user selects the search area, polygon or administrative region range, time interval, and satellite type on the front end, the back end verifies the parameters and extracts the polygon vertices, calculates the minimum bounding rectangle (MBR) and time range; generates a five-bit GeoHash for the MBR and expands the adjacent grid, extracts the first four bits to construct a GeoHash prefix for calculating the MD5 prefix PrefixID; combines the GeoHash prefix, satellite identifier, and time information to calculate all possible start and end RowKeys, and batches HBase Scan requests to send to the RegionServer; each RegionServer first performs coarse screening based on the RowKey range and column value filter, and then calls the JTS (Java Topology Suite) library in the coprocessor to perform precise intersection / containment judgment; finally, the driver node deduplicates the candidate records returned by each partition and sorts them by time to generate the final paginated search results.

[0082] In summary, this invention constructs a unified spatiotemporal metadata model based on the ISO-19115-2 standard and automatically extracts data fields in various formats such as XML, JSON, TXT, and GeoTIFF using template-mapping rules. This allows for format normalization and field validation during the data ingestion phase without manual intervention. This mechanism ensures structural consistency of spatiotemporal data from different sources while significantly reducing the cost of accessing and maintaining new data sources, thereby improving the retrieval efficiency of spatiotemporal data. Constructing a composite RowKey not only guarantees spatial locality but also eliminates write hotspots through MD5 prefixes and pre-partitioning strategies, improving concurrent write throughput. The standardized metadata is exported as an HFile using MapReduce and atomically loaded into the target table using the HBaseBulkLoad tool, avoiding the significant random write overhead of WAL and MemStore. The retrieval end expands the user-given spatial polygon / circular range using MBR and an eight-neighbor GeoHash grid, combines it with time slices, and then sends out Scan tasks in parallel. The RegionServer first performs a coarse screening by row key range, and then the Observer coprocessor calls JTS (Java TopologySuite) spatial operators to perform precise intersection or containment filtering. This "coarse screening—fine sorting—local deduplication—global merging" process minimizes network and I / O overhead.

[0083] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A distributed method for rapid retrieval of massive spatiotemporal data, characterized in that, Includes the following steps: Step 1: Analyze the spatiotemporal data acquired from different sources; Step 2: Based on the parsed spatiotemporal data, construct the RowKey, create column families, and extract metadata information; Step 3: Based on the constructed RowKey and the extracted metadata information, construct a metadata object; Step 4: Based on the constructed metadata object, the metadata is loaded into the database using the BulkLoad mechanism, and the user's search criteria are parsed. Step 5: Verify the validity of the parameters based on the parsed user search criteria; Step 6: If the verification parameters are valid, construct the RowKey to generate the scan range, and filter it through the RegionServer, thereby performing spatial filtering through the Observer coprocessor deployed on the RegionServer. Step 7: The local filtering results returned by multiple RegionServers are uniformly sent back to the driver node to perform data deduplication and time sorting operations on a global dimension, and to build the final set of search results.

2. The distributed method for rapid retrieval of massive spatiotemporal data according to claim 1, characterized in that, The construction of RowKey includes: For the parsed spatiotemporal data, the latitude and longitude of the center point are first encoded into a five-bit GeoHash, and then concatenated with the satellite identifier and the shooting date. The three-bit prefix is ​​obtained by MD5 hashing, and a composite RowKey with spatial locality and hash dispersion is constructed to support subsequent partition parallel access and load balancing.

3. The distributed method for rapid retrieval of massive spatiotemporal data according to claim 1, characterized in that, The metadata object is constructed based on the constructed RowKey and the extracted metadata information, including: Based on the constructed RowKey and the extracted metadata information, the corresponding metadata items are extracted according to the unified metadata model conforming to the ISO-19115-2 standard, and a metadata object is generated.

4. The distributed method for rapid retrieval of massive spatiotemporal data according to claim 3, characterized in that, The unified metadata model includes at least UUID, GeoHash, SatelliteID, SensorID, SceneID, Resolution, CloudCover, BeginTime, EndTime, CenterLong, CenterLat, image corner coordinates, and thumbnail binary fields.

5. The distributed method for rapid retrieval of massive spatiotemporal data according to claim 2, characterized in that, The formula for constructing a composite RowKey with spatial locality and hash dispersion is as follows: RowKey=PrefixID3+SatelliteID2+GeoHash(r)5+Time6 GeoHash(r) is the encoding after latitude and longitude coordinate transformation, with a length of 5 bits. GeoHash encoding stores spatially adjacent data in adjacent physical locations; Time represents the acquisition time of the remote sensing data, in the format YYMMDD. PrefixID=MD5[(GeoHash(r)4+SatelliteID2+Time_Month)]%nums Among them, PrefixID represents the prefix identifier, which is the first four bits of the GeoHash encoding of the r space region; SatelliteID represents the type identifier of the data source; Time_Month represents the year and month of the shooting time; and nums is the total number of partitions.

6. The distributed method for rapid retrieval of massive spatiotemporal data according to claim 5, characterized in that, nums determines the optimal number of Regions on each RegionServer based on a formula, which is: Rs memory is the amount of memory allocated by the server for HBase; Set hbase.regionserver.global.memstore.size to 0.4; hbase.hregion.memstore.flushsize represents the size of each Memstore, which is 128MB; columnfamilies represents the number of column families.

7. The distributed method for rapid retrieval of massive spatiotemporal data according to claim 1, characterized in that, The method of loading metadata into the database via the BulkLoad mechanism includes: The metadata object is written to HDFS, and then generated into a partitioned HFile file in the cluster through a MapReduce task. Finally, the HFile is atomically loaded into multiple Regions in the table using the HBase BulkLoad mechanism.

8. The distributed method for rapid retrieval of massive spatiotemporal data according to claim 1, characterized in that, The retrieval process includes: By combining the Scan operation of the HBase database with column family filters, the data is retrieved and the remote sensing data that meets the search criteria is initially selected. By implementing the RegionObserver.preGetOp method of the Observer coprocessor, the JTS spatial calculation engine is integrated on the RegionServer side. Through this integration, spatial relationship judgment is performed on the RegionServer side where the data is located.

9. The distributed method for rapid retrieval of massive spatiotemporal data according to claim 1, characterized in that, The parsing of user search criteria includes: The query range and time interval of the user space are parsed to generate the corresponding GeoHash prefix list and time segment. All possible PrefixIDs are derived by combining the satellite type. Multiple start and end RowKey ranges are constructed on the client, mapped to multiple Scan tasks, and distributed in parallel to each RegionServer node of the HBase cluster.

10. A distributed method for rapid retrieval of massive spatiotemporal data according to claim 1, characterized in that, The filtering via RegionServer includes: Each HBase RegionServer node independently receives Scan requests and filters candidate records based on the RowKey range in its local Region. Combined with column value filters, it completes the matching of sensor identifiers, realizing a distributed and parallel preliminary filtering process.