Remote Sensing Image Data Search Engine System and Method Based on High-Speed Caching Technology
By employing high-speed caching technology and a unified data access model, the problem of inconvenient multi-source remote sensing image data querying has been solved, enabling efficient and unified image data querying and management. It supports querying by attribute and field from multiple data sources, improving query efficiency and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING AEROSPACE SPACE VIEW INFORMATION TECH
- Filing Date
- 2022-12-26
- Publication Date
- 2026-07-17
AI Technical Summary
In the current technology, major satellite data query websites lack a unified search engine, which makes it inconvenient to query remote sensing image data and makes it impossible to achieve efficient and unified query and management of multi-source remote sensing data.
Employing high-speed caching technology, a unified data access model, and an image search protocol, high-concurrency image data queries are achieved. Combining hash mapping and a two-level caching mechanism, the differences in image attributes across multiple platforms are masked, and unified results are returned. Furthermore, multi-threaded connections to major image websites are supported, enabling efficient filtering and storage of image query results.
It enables efficient querying and browsing of multi-source remote sensing image data, improves query efficiency, supports filtering by attribute and querying by field from multiple data sources, reduces search time, enhances user experience, and supports visual querying of multi-type, multi-source spatial data.
Smart Images

Figure CN116244456B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image data search engine technology, specifically to a remote sensing image data search engine system and method based on high-speed caching technology, which is suitable for efficient and unified retrieval of archived remote sensing images on various satellite websites in the industry. Background Technology
[0002] In recent years, satellite remote sensing technology has developed rapidly, ushering in a new era of multi-layered, three-dimensional, multi-angle, all-round, and all-weather Earth observation. Satellite remote sensing offers advantages unmatched by traditional methods, including high viewpoints, wide fields of view, rapid data acquisition, and repeatable and continuous observation. A wide range of resolution remote sensing satellites have been launched, featuring multi-source, multi-temporal, multi-spectral, and multi-resolution capabilities. Satellite remote sensing manufacturers offer a variety of products including Gaofen satellites (SV1-01, SV1-02, SV1-03, SV1-04, SV2-01, SV2-02, GFDM01), MAXAR satellites (GE01, QB02, WV01, WV02, WV03, WV04), and the KOMPSAT series (KOMPSAT-3, K...). The archived imagery data is growing rapidly every day. With its large volume of information and short update cycles, the need for convenient data retrieval and management based on spatial and attribute information is an objective requirement for enterprises and industries. To expedite the delivery of this expensive, up-to-date, multi-source archived data to the market, major satellite manufacturers have launched their own imagery query systems. These systems include OMPSAT-3A and KOMPSAT2, Airbus series (Pleiades (PHR), SPOT, etc.), the Resource Satellite Center series (GF1, GF2, GF3, etc.), Resource Satellite (ZY3, ZY3-1, ZY3-2), Beijing 2 (01, 02), COSMO series (CSKS1, CSKS2), ALOS series (ALOS, PALSAR, PRISM), Tianhui series (TH01, TH02, SXZ), IKONOS-2, TerraSAR-X1, Deimos 2, and LandSat series (LANDSAT_8), etc.
[0003] Domestic and international industry satellite data query websites such as MAXAR, Airbus, KOMPSAT, Gaofen, China Resources Satellite Center, and Orbita have scattered remote sensing image archives, making querying inconvenient. There is no unified industry-wide search engine, nor a unified definition and standard for remote sensing data models. To better utilize industry remote sensing data and move beyond the limitations of relying on proprietary data, this system integrates domestic and international remote sensing data into a unified platform, creating a unified industry data query system. Driven by data, this remote sensing data search engine facilitates comprehensive and unified exchange of industry remote sensing data, enabling unified querying and browsing of data from various remote sensing vendors, while also supporting flexible integration of data from new vendors. Summary of the Invention
[0004] The purpose of this invention is to provide a remote sensing image data search engine system and method using high-speed caching technology, which solves the problems mentioned in the background art.
[0005] To achieve the above objectives, this invention provides the following technical solution: a remote sensing image data search engine system and method using high-speed caching technology, comprising a unified data access model, establishing a hash mapping between image data information (including attribute information, bounding boxes, thumbnails, and attribute details) and image data information items (oid) based on shared entity IDs and their respective data types; implementing a high-concurrency image search protocol based on a network framework; encapsulating query requests in the background according to an interface parameter configuration table, connecting to the data search APIs of major image websites via multiple threads, retrieving in real time, and separately obtaining their respective image data query information, including bounding boxes, attribute details, etc.; acquiring data information and implementing image data caching through a two-level high-speed cache of modeled configuration information and query results, including local caching and KV caching database caching. Efficient filtering and storage of query results; rapid extraction of standardized basic attribute information from cached image query results through image attribute key-value pair hash mapping and keyword matching modules, masking differences in image attribute fields across multiple platforms, returning unified results, and adding them to the high-speed cache; unified retrieval of thumbnails based on entity IDs, achieving standardized thumbnail processing through thumbnail deflection configuration and processing modules; analysis of remote sensing data query results through data information selection linkage and data coverage analysis; unified retrieval of standardized thumbnails, attributes, and bounding boxes stored in the cache by large object OIDs, and through the task system, organizing the results by range and fields into common formats such as Shapefile required by the user, along with thumbnails, to achieve the download of key information from image query results.
[0006] One-stop service, a single interface for querying images from multiple data sources, comprehensively covering satellite data and greatly improving the efficiency of querying industry satellite data sources; at the same time, real-time retrieval achieves the effect of trading time for space, exchanging the time spent collecting data from various websites for the space to store large-scale industry images, and using high-speed caching and matching mechanisms to save search time as much as possible, such as key-value pair hash mapping and matching, to achieve the optimal performance of the industry image search system in terms of time and space.
[0007] The beneficial effects of this invention are:
[0008] 1. Directly search industry remote sensing data from major satellite data websites, uniformly support multiple commonly used satellite website data sources in the industry, and store and browse image query results in a unified manner, making it convenient to expand the query of industry satellite data sources (including stereo data);
[0009] 2. Real-time retrieval: The industry remote sensing data query system supports retrieval and query modules for multi-source satellite data such as MAXAR, Pleiades, KOMPSAT, and Gaofen. It supports filtering by attribute and querying by field (supporting multiple scene numbers or product numbers, etc.) from multiple industry data sources, enabling efficient query and browsing of multi-source remote sensing image data. It supports query result analysis (one-time loading of thumbnails, selection, coverage) and saving and exporting, greatly improving the ease of use. The pressure on supplier websites can be compensated by the conversion rate of commercial orders.
[0010] 3. The remote sensing data search engine provides backend support and is deployed externally as a service, supporting deployment in Windows and Linux environments. The engine abstracts and defines the data query APIs of industry vendors, forming its own complete set of query APIs, which can serve Worldview Technology and also support third-party access, thereby realizing one-stop, multi-type, multi-source spatial data visualization and efficient query, with data types including rich optical, radar, and satellite imagery;
[0011] 4. Integrates with satellite source query interfaces such as MAXAR, KOMPSAT, Pleiades, Gaofen, China Resources Satellite Center, Changguang Satellite Jilin-1, and Capella, supporting queries from multiple industry data sources by attribute, spatial range, and field (scene number, product number, etc.), enabling one-click query of industry remote sensing data; supports the return of stereo image pairs from the MAXAR official website with complete attribute information;
[0012] 5. Improved functionality to meet user needs, such as supporting vector / thumbnail / vector + thumbnail download options for the original and results channels, ensuring that thumbnails from KOMPSAT, Gaofen, Jilin-1, and the Resource Satellite Center are correctly rotated, and meeting the different needs of users for downloading data information;
[0013] 6. Optimize the left and right linkage of the query and browsing windows and multi-level data picking to facilitate data analysis such as coverage rate, and improve the analysis and export of query results, as well as vector saving of covered / uncovered ranges;
[0014] 7. To enable and protect the industry data query capabilities of beta testers, a role application system has been added during the user registration phase. Approvers can view the applicant's company, department, position, email, mobile phone number, and other information, and obtain platform query permissions upon approval. Attached Figure Description
[0015] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0016] Figure 1 This is a service framework diagram of an industry remote sensing image data search engine system and method based on a high-speed caching technology, as described in this invention.
[0017] Figure 2 This is a schematic diagram of the module structure of an industry remote sensing image data search engine system and method based on a high-speed caching technology of the present invention.
[0018] Figure 3 This invention provides a high-speed caching technology for remote sensing image data search engine systems and methods, including a core algorithm diagram for keyword matching in an industry-specific remote sensing image data search engine system.
[0019] Figure 4 This invention relates to a remote sensing image data search engine system and method using high-speed caching technology, which uses Redis to cache various image information maps. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention are within the scope of protection of the present invention.
[0021] A remote sensing image data search engine system and method using high-speed caching technology includes the following steps:
[0022] S1: A unified data access model that establishes a hash mapping from image data information, including attribute information, bounding box, thumbnail, and attribute details, to image data information item OID based on shared entity IDs and their respective data types;
[0023] S2: A high-concurrency image search protocol based on a network framework;
[0024] S3: Based on the interface parameter configuration table, encapsulate the query request in the background, connect to the data search APIs of major image websites in multiple threads, retrieve the image data query information of each website in real time, including the bounding box, attribute details, etc.
[0025] S4: Acquire data information and achieve efficient filtering and storage of image query results through a two-level high-speed cache of modeled configuration information and query results, including local cache and KV cache database cache;
[0026] S5: Through the hash mapping of image attribute key-value pairs and keyword matching module, standardized basic attribute information is quickly extracted from the cached image query results, the differences in image attribute fields across multiple platforms are masked, the results are returned after unification, and added to the cache.
[0027] S6: Thumbnails are retrieved uniformly based on entity IDs, and standardized thumbnail processing is achieved through thumbnail deflection configuration and processing modules;
[0028] S7: Through data information selection and linkage and data coverage analysis, the query results analysis of remote sensing data is realized;
[0029] S8: Unified retrieval of standardized thumbnails, attributes, and bounding boxes stored in the cache by large object OID. Through the task system, the results are organized by range and field into common formats such as Shapefile required by the user, along with thumbnails, to achieve the download of key information of image query results.
[0030] Remote sensing data is spatial data, and its two fundamental characteristics are spatial features and attribute features. Remote sensing imagery is stored as a unified raster layer, ensuring each layer is unique. Furthermore, each provider has its own unique channel. Image queries must be performed by channel; cross-channel queries are not possible in a single session. This facilitates multi-threaded retrieval of data from various imagery query websites. Therefore, data queries, data information retrieval, and thumbnail retrieval all require specifying a specific channel. For example, when querying Gaofen satellite imagery, all related API calls need to set the channel parameter. The following table illustrates the channel references for each provider.
[0031]
[0032] The entity ID is a unique 64-bit ID generated for each image result in the query result set, based on the layer and channel where the data is located. It is a globally unique image identifier. Numerous related information about the image, such as its basic attributes, detailed attributes, geometry, and overview (including address, size, and data), are all attributed to this image entity ID and mapped accordingly.
[0033] Furthermore, as a search engine for remote sensing imagery data in the industry, it abstracts and defines the data query APIs of industry vendors to form its own complete set of query APIs. This is the search protocol. The image search protocol is a custom protocol based on the HTTP application layer, providing lightweight, high-concurrency image query services, including compatibility with OGC standards. This protocol is implemented based on spatial operation protocols, making it convenient, flexible, and easily extensible. It can serve existing company users and also supports third-party user access, thereby achieving one-stop, multi-type, multi-source spatial data visualization and efficient querying in the industry. Data types include rich optical and radar satellite imagery; it also supports new satellites.
[0034] Image search is primarily divided into four main interfaces: attribute query interface, field query interface, thumbnail interface, and login interface. All image query interfaces in the industry invariably fall into these four categories. Therefore, based on this pattern, an industry data interface configuration table is created, specifying fields such as query interface, field query interface, thumbnail interface, login interface, authentication type, username, and password. Since the image query APIs of various data providers in the industry differ, compatibility, simplicity, reusability, and minimal modification are carefully considered. By establishing interface specifications, a unified interface is provided for developers to connect to external APIs when extending image queries. Under the constraints of the interface configuration table specifications, the entire engine system defines protocols for image attribute query, total image count retrieval, image thumbnail retrieval, image bounding box retrieval, and image attribute details retrieval, as well as field-based queries. Implementing these protocols is sufficient to meet the search needs of the suppliers we encounter. Therefore, a concise and lightweight protocol is designed using an appropriate number of parameter dimensions, employing the fewest key parameter definitions to reduce redundancy in protocol design and achieve search requirements across multiple scenarios.
[0035]
[0036]
[0037] The system defines several key request parameters, such as the parameter name "channel," which distinguishes different suppliers. `channel=517` indicates a query for data from supplier MAXAR. The parameter "category" indicates the satellite names belonging to the supplier specified by "channel," such as MAXAR's WV01; WV02; WV03; WV04; QB02; GE01; IKONOS-2, with multiple satellite names separated by semicolons. The `mode` parameter, `attr`, indicates an attribute query for image data; `mode`, `field`, indicates a query by scene number. The `datatype` parameter, `none`, returns only the number of records; `attr`, returns the attributes. The `output` parameter takes the values `json` or `xml`, indicating the data format of the response. The `page_size` parameter indicates the page size of the returned data, and the `page_no` parameter indicates the page number, starting from 0. The parameter design for each attribute feature of the image data is simple, directly using the common English names of the attribute names. At the same time, the atomicity requirement of the protocol design must not be violated. Following this principle ensures that users can invoke various flexible protocol combinations for their requests. For example, the protocol can be obtained by first calling the total number of images, and then querying page by page based on the total number of data and the user-defined page size. The definition of the `page_size` parameter can satisfy user requests for different page sizes. Similarly, after specifying `page_no` to retrieve images at a specified page number, image thumbnails, image bounding boxes, or image details can be retrieved based on the `entityId` of each data entry. Therefore, a well-designed protocol, in addition to atomicity, must also consider the logical relationships between protocols to better adapt to changes in application requirements. The HTTP packet body is organized according to parameter configuration information and packet type. For simple cases, parameters can be quickly organized directly based on the parameter configuration information, and then packaged according to the required HTTP packet body data type. However, fields with mismatched data types need to be converted. For example, the system query range uniformly uses WKT format, but some industry data source query interfaces use GeoJSON format, which requires prior format conversion. Meanwhile, for some irregularities, you can use a template to directly replace the placeholders in the template. For example, for a template with an irregular time range, "date=between{date0)and{date1)", you can directly replace date0 and date1 with the upper and lower limits of the time range respectively.
[0038] Furthermore, during data queries, the corresponding industry interface configuration information is obtained based on the industry's layer and channel (channel1), and then the query interface information is obtained based on the query type. However, the query parameter field names and formats vary across industry query interfaces, lacking a unified standard. We need to convert the unified query parameters into parameter names and formats specific to the corresponding industry query interfaces. Here, we reuse the 7.37.3 key-value pair mapping and matching submodule and the parameter configuration table. We use a common image attribute field name (attr_name) to store the unified query field name, a data type (attr_type) to store the target data conversion type, a target field name (field_name) to store the target field name, and a field value conversion rule (field_value) to store the conversion relationship. Through the industry data interface configuration table and parameter configuration table, we can flexibly, dynamically, and in real-time adjust the settings to respond to any changes in the data source interfaces of various industries simply by updating the relevant configurations, thus avoiding repeated compilation and modification of the program and achieving a "one-size-fits-all" approach.
[0039] Furthermore, the characteristics of query interfaces for various industries were summarized, and the final definition of the interface parameter configuration table is as follows:
[0040]
[0041] Furthermore, interface configuration and parameter configuration information are obtained from the industry data source channel. This information changes infrequently; retrieving largely unchanged information from the database each time would increase database load and reduce concurrency. Therefore, configuration information is cached for faster processing. If the interface requires authentication information, this information is also cached during the first configuration retrieval, with an expiration time set to reduce the frequency of authentication-related interfaces. When subsequent authentication information expires from the cache, new authentication information is retrieved and added to the cache again. The libcurl library is used to call image query APIs from multiple data providers, satisfying existing query and browsing requirements and completing API integration with various satellite data providers.
[0042] Furthermore, the entity ID is a key generated for each image based on its layer and channel. Other image information is stored as a field of this key, using a caching mechanism to cache information such as bounding boxes, thumbnails, and attributes. The entity ID is mapped to various data types (oids) such as thumbnails, bounding boxes, and image attribute details. The caching core uses a key-value structure, supporting two-level caching, including a combination of local caching and a KV cache database. This significantly reduces resource contention and network requests, achieving high speed while meeting the requirements of high concurrency and multiple data types, ensuring an efficient user experience. Locally, the entity ID is stored to establish relationships with various data types (oids). The KV cache database stores basic image attributes, detailed attributes, bounding boxes, thumbnails (including address, size, and data), and other information using hash data types. Figure 3 As shown, the basic attributes of the image correspond to the attribute fields that are most important to ordinary users, such as resolution, cloud cover, satellite name, and side view. Here, they are all packaged into a JSON string and stored as a single field called "attributes" for the entity ID. Although storing all the attributes separately as entity ID fields is very convenient for retrieval and use, in actual applications, the frequency of use is not that high, which actually increases the storage burden (the memory usage of the merged storage is about 1 / 10 of that of storing them all as HASH fields). High-speed caching is a core requirement of this data search engine's caching core, and its high-speed caching characteristics are mainly reflected in the following design: It is entirely based on memory, because memory has extremely fast read and write speeds, making it the fastest storage device for computational read and write operations; it is implemented in a single thread, because being memory-based, the CPU is not the bottleneck of the entire cache, thus saving a lot of time spent on context switching threads; it uses I / O multiplexing technology, enabling a single thread to efficiently handle multiple connections, non-blocking I / O, and is based on epoll, registering the file descriptors (FDs) corresponding to multiple socket connections into epoll. epoll listens for read, write, connection, and close events, and epoll's event notification is much faster than select's polling.
[0043] Furthermore, Redis, a high-performance key-value database, supports high concurrency and multiple data types, which are precisely the characteristics required by industry remote sensing image data search engines. Therefore, Redis was chosen as the core storage of the system. Satellite imagery contains a wealth of information, including basic attributes, detailed attributes, geometry, and overviews (containing address, size, and data). To store this information and establish relationships, appropriate data types must be selected, and Redis's hash data type perfectly meets these requirements. Based on Redis caching, industry archive data retrieved from external APIs is cached, allowing industry archive data to be retrieved from the cache at certain intervals, rather than calling the external API every time. This achieves call efficiency comparable to directly accessing the original website. Figure 4 As shown, the entity ID is a key generated for each image based on its layer and channel. Other image information is stored as a field of this key. The image's basic attributes correspond to the fields most relevant to ordinary users, such as resolution, cloud cover, satellite name, and side view. Here, these are uniformly packaged into a JSON string and stored as an attribute field within the entity ID. While storing all attributes separately as entity ID fields would be very convenient, in practice, their usage frequency is not that high, and this would actually increase the storage burden.
[0044] Furthermore, data compression can help the system effectively utilize limited memory. Data compression can be used to reduce memory usage for fields with low access frequency. Two compression algorithms are used: zstd and lz4. zstd offers the highest compression ratio and is suitable for cold storage. lz4 offers the fastest compression and decompression speeds and is suitable for OLAP query scenarios. Different compression policies can be configured for the basic attributes of the image on devices with varying memory capacities. On devices with ample memory, no compression can be selected; on devices with insufficient memory, zstd compression can be chosen (saving approximately 1 / 5 of the memory); and in scenarios where memory is limited but speed is critical, lz4 compression can be selected.
[0045] Furthermore, using appropriate cache eviction allows the system to utilize memory more effectively. Redis handles memory reclamation in two ways: expiration time and memory eviction policies. For expiration time, thumbnail data is stored separately using Redis's string data type (HASH types do not support setting expiration time for individual fields). The key uses a combination of entity ID and overview, ensuring key uniqueness while maintaining relationships. For the memory eviction policy, allkeys-lru is selected, using the LRU algorithm to evict all keys.
[0046] Furthermore, since the image attributes returned by various image query systems in the industry are not entirely the same, the names of fields with the same meaning vary, and even the character values with the same meaning are different. Therefore, this system needs a keyword matching module to return unified image attribute field names and values, thereby masking the differences in attribute field names and values among various systems in the industry.
[0047] Furthermore, the core of the key field matching module is the key field matching table, which contains five fields: data source (cat_id), general image attribute field name (attr_name), data type (attr_type), target field name (field name), and field value conversion rule (field value). Each industry generates a unique data source identifier (cat_id) based on its layer and channel. Each cat_id corresponds to all key field matching details. The data type (attr_type) is the storage type. When the industry's corresponding field does not match the storage field, conversion is required. For values that cannot be directly converted, the field value conversion rule (field value) can be used for matching and conversion.
[0048] Furthermore, the core of key field matching is the parameter configuration table, which contains five fields: data source, common image attribute field name, data type, target field name, and field value conversion rule. To quickly retrieve all information matching a specific key field from the configuration table, an index is created on the data source (cat_id) field. During initialization, the parameter configuration table is loaded into memory, thus eliminating the need to read the database during field matching. The specific field definitions of the attribute parameter configuration table are as follows:
[0049] Field Name Data types length Remark cat_id int4 32 Data source attr_name varchar 32 Attribute Name attr_type int4 32 Attribute data type field_name varchar 256 Corresponding attribute name field_value varchar 256 Corresponding attribute value mapping rules
[0050] Furthermore, given the issue of image data retrieved from supplier websites having thumbnails that don't align with the vector range, the engine has been configured to handle whether the supplier data thumbnails need to be deflected. If deflection is configured, it will be applied during thumbnail display or export.
[0051] The system uses a geometric polynomial as the transformation function for deflection processing, takes the four boundaries of the data vector range as the GCP control points, uses the least squares method for data fitting, and performs correction processing on the data thumbnails. Essentially, it corrects the unprojected thumbnails of the data to the projection of the range box, thereby achieving the purpose of fitting under the same projection.
[0052] The number of control points (N) required for the aforementioned deflection processing increases with the polynomial degree (n), as follows: N = (n+1)(n+2) / 2. In fact, most of the raw image data from the suppliers we've encountered uses regular quadrilateral vector ranges and thumbnails, making the actual implementation of deflection processing simpler. Regular shapes make it easier to determine accurate GCP control points based on thumbnails and bounding boxes, while irregular shapes generally make it difficult to determine the row and column numbers corresponding to any point on the vector. Furthermore, this regularity ensures that the polynomial degree used for deflection processing is not excessively high, making the computational load acceptable and essentially achieving real-time deflection processing when users browse thumbnails. Responding to user needs, download options in the raw and results channels support vector / thumbnail / vector + thumbnail, ensuring correct thumbnail rotation for KOMPSAT, Gaofen, Jilin-1, and resource satellite centers, meeting diverse user needs for downloaded data.
[0053] Furthermore, image queries include industry-wide multi-data source queries by attribute and by field. Field queries support queries by scene number and product number, and support multiple numbers. The system organizes image query parameters, performs request and parsing, returns the total number of query results, and provides complete attribute information. Pagination is controlled by the `page_no` and `page_size` parameters.
[0054] This includes the following steps:
[0055] S1: Send the authentication information (cookie or header) and parameter body using the network library according to the permission type.
[0056] S2: Obtain image field parameter configuration information based on the industry data source channel.
[0057] S3: Convert each image attribute into an attribute storage field under the unified storage model based on the image field parameter configuration information.
[0058] S4: Organize thumbnails, attribute details request URLs, and use the original attribute information as attribute details information.
[0059] S5: Generate a unique identifier entity ID for the image, and add the converted attribute information, thumbnail address, and attribute details to the cache.
[0060] S6: Returns the transformed query results, including basic image attributes and bounding boxes.
[0061] Furthermore, the image query service provides a set of image attributes, where the `entityid` field is a globally unique image identifier, specifying the image ID from which to retrieve thumbnails or attribute details. The data acquisition service retrieves the file address and data of the thumbnail or attribute details from the cache based on the image's unique identifier, the image ID. If the image exists, it is returned directly; if retrieving a thumbnail corresponding to the image cache, and if multiple thumbnails exist for the image, the first one is returned. If the image does not exist, the data is retrieved based on the thumbnail or attribute details address, including retrieving file attribute values, obtaining thumbnail size, format, and other attributes, adding the thumbnail to the cache, and then returning the retrieved thumbnail. During image browsing, all thumbnails can be selected at once and loaded.
[0062] Furthermore, query result analysis includes data information selection and linkage, and data coverage analysis, including vector saving and export of covered / uncovered ranges.
[0063] Furthermore, through data picking, it is convenient to cooperate to achieve that when the source data of a certain industry (such as MAXAR) cannot fully cover the user's AOI range, the AOI coverage can be supplemented by data from another industry (such as KOMPSAT). Check the data range box in the map display area, and pick the range box with the left mouse button to enable linked browsing. When picking with the left mouse button, according to the user's current operation and mouse position, the matching geometry and image ID are filtered out from the data set of the current operation and fed back to the current operation interface such as the image query result column or the order image list column. It realizes the linked browsing of the range box, attributes, and thumbnails among multiple industries and multiple data. When multiple features are selected, the selected items are presented. When one of them is selected, it can be positioned to the corresponding data for quick browsing. After the user selects a certain vector point V, the search result data set Ps is traversed, and it is calculated whether the vector polygon of each data contains the vector point V. The data that contains the point V forms a new data set, called the picking data set Pi. The picking data set Pi is the data set picked after the user's click operation. Its feature is that the vector polygon of each data element in the set contains the vector point V selected by the user. This inclusion relationship is a vector inclusion relationship in geometric space, not a set inclusion relationship. At this time, when the user browses each data in the picking data set Pi, it has an inclusion relationship with the point V selected by the user in terms of spatial features. The data in the set Ps that does not contain the point V selected by the user cannot be browsed. As the vector point V selected by the user changes, the composition of the obtained picking data set Pi also changes. In terms of process, the picking linkage is presented in a fast and convenient manner for these two features of the query result data, and the spatial features and attribute features are associated and displayed in real time from the view, facilitating the analysis and picking of search results. This mechanism fully considers the user's intelligent and automated browsing requirements for data after search. When the user clicks on a certain spatial point vector V in the map area, all the data whose spatial features contain the spatial point vector V will be picked again from the search result set Ps (referred to as the first query search), and the obtained results are presented in another list (the picking data set Pi, referred to as the second screening search). As the clicked spatial point V changes, the secondary screening result list Pi will be updated synchronously. The picking linkage is completed based on the above two searches. When clicking on a certain data for browsing from the second screening search result list Pi, the spatial feature vector range box and attribute features (cloud amount, resolution, time phase, side view angle, azimuth angle, band type, product type, product level, etc.) of this data will be displayed in the map area in a synchronous manner, facilitating the user to browse the data in a foolproof and efficient all-round and all-element way. As different data are switched from Pi, the spatial feature range and attribute feature information of the data will also be updated synchronously, realizing the all-element spatio-temporal linkage display of the data.The picking linkage, based on the initial query search, provides users with a quick and easy way to fully and thoroughly interpret each piece of data from a micro perspective; it also provides a simple and feasible pre-selection method for finally selecting a piece of data that closely matches the ideal characteristics for data interpretation.
[0064] Furthermore, it provides analytical functions, supporting the calculation of coverage rate for query-matched data within the query bounding box. Data coverage analysis, based on the search, provides preliminary criteria for user decision-making from a macro perspective. After selecting target matching data from the query results, coverage analysis can be automatically completed with one click, displaying the coverage percentage and automatically loading thumbnails and bounding boxes of the selected data. After the engine system searches, selecting all or part of the search results data allows for automated calculation of data coverage through data coverage analysis. Sometimes users need to know how much data within the searched spatial range meets the attribute characteristics and is retrieved; here, "how much" refers to area, not number. For this purpose, the system provides data coverage analysis. The search range vector polygon is defined as Vs. The search results dataset Ps is traversed, and the union of all data range vector polygons in Ps is calculated, resulting in the set Vu (the union of data vectors). The effective data range is calculated as follows:
[0065] Vv=Vu ∩ Vs
[0066] Therefore, data coverage (Cp) = area of effective data range (Vv) / area of search range (Vs) * 100%. This indicator can guide users to quantitatively analyze the image data of a certain area (search range) that meets the attribute characteristics (cloud cover, resolution, time phase, side view, azimuth, band type, product type, product level, etc.), providing data support for industry decision-making or planning.
[0067] Furthermore, scattered files on the internet cannot be directly downloaded by users from their browsers. Therefore, the data needs to be packaged into a tar archive for download. The archive contains bounding box files in the commonly used Shapefile format and thumbnail files with coordinates. The bounding box, thumbnail address, and thumbnail data are retrieved from the cache based on the image entity ID; if the image thumbnail does not exist, it is retrieved from the thumbnail address and added to the cache. The bounding box is converted from WKT (the system uniformly stores data in WKT format) to the more commonly used Shapefile format, and thumbnail files with coordinates in JPEG or other formats are generated based on the bounding box. When a user selects or clicks to download image data in the search results, the task system automatically executes the task, exports the data, packages the bounding box (Shapefile file) and thumbnails into a tar archive, returns the data address for download, and completes the tar archive download after user confirmation.
[0068] In summary, by utilizing the technical solution of this invention, by establishing a common definition and standard for a remote sensing data model, unifying the data from various manufacturers in the remote sensing industry, and providing backend support through a remote sensing data search engine, real-time querying and browsing of remote sensing data is achieved, while also supporting flexible access to data from new manufacturers.
[0069] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A remote sensing image data search engine method based on high-speed caching technology, comprising the following steps: S1: A unified data access model that establishes a hash mapping from image data information, including attribute information, bounding box, thumbnail, and attribute details, to image data information item OID based on shared entity IDs and their respective data types; S2: A high-concurrency image search protocol based on a network framework; S3: Based on the interface parameter configuration table, encapsulate the query request in the background, connect to the data search API of major image websites in multiple threads, retrieve in real time, and obtain the image data query information of each website separately, including the bounding box and attribute details; S4: Acquire data information and achieve efficient filtering and storage of image query results through a two-level high-speed cache of modeled configuration information and query results, including local cache and KV cache database cache; S5: Through the hash mapping of image attribute key-value pairs and keyword matching module, standardized basic attribute information is quickly extracted from the cached image query results, the differences in image attribute fields across multiple platforms are masked, the results are returned after unification, and added to the cache. S6: Thumbnails are retrieved uniformly based on entity IDs, and standardized thumbnail processing is achieved through thumbnail deflection configuration and processing modules; S7: Through data information selection and linkage and data coverage analysis, the query results analysis of remote sensing data is realized; S8: Unified retrieval of standardized thumbnails, attributes, and bounding boxes stored in the cache by large object OID. Through the task system, the results are organized by range and field into the user-required Shapefile general format, along with the thumbnails, to achieve the download of key information from the image query results.
2. The remote sensing image data search engine method based on high-speed caching technology according to claim 1, characterized in that: The high-concurrency image search protocol based on the network framework includes a protocol mechanism, an interface matching mechanism, an interface parameter configuration table structure, and the acquisition of interface configuration information and parameter configuration information. The protocol mechanism provides a unified interface for developers to connect to external APIs by defining interface specifications. The interface matching mechanism obtains the corresponding industry interface configuration information based on the industry's layer and channel during data query, and then obtains the query interface information based on the query type. The interface parameter configuration table structure organizes the characteristics of query interfaces for each industry. The acquisition of interface configuration information and parameter configuration information is based on the industry data source channel.
3. The remote sensing image data search engine method based on high-speed caching technology according to claim 1, characterized in that: The hash mapping established by sharing entity IDs and their respective data types to image data information items (oids) includes a caching mechanism, a key-value (KV) cache database and data types, data compression, and cache eviction. The caching mechanism is used to cache bounding boxes, thumbnails, and attribute type information using a high-speed caching mechanism, mapping entity IDs to thumbnails, bounding boxes, and image attribute details data types (oids). The KV cache database and data types support high concurrency and multiple data type selections, which are precisely the characteristics required by industry remote sensing image data search engines. Therefore, Redis is chosen as the storage core of the system. The data compression helps the system effectively utilize limited memory, reducing memory usage for fields with low access frequency. The cache eviction mechanism uses allkeys-lru, employing the LRU algorithm to evict all keys.
4. The remote sensing image data search engine method based on high-speed caching technology according to claim 1, characterized in that: Returning a unified set of image attribute field names and values requires a keyword matching module to mask the differences in attribute field names and values across various industry players. The core of this keyword matching module is a key field matching table, which contains five fields: data source, general image attribute field name, data type, target field name, and field value conversion rule. To quickly retrieve all information matching a specific key field from the configuration table, an index is created on the data source field of this key field matching table.
5. The remote sensing image data search engine method based on high-speed caching technology according to claim 1, characterized in that: The thumbnail deflection configuration uses a geometric polynomial as the transformation function for deflection processing. The vector bounding box of the data and the thumbnail are used as GCP control points. The least squares method is used to fit the data and correct the thumbnail. Essentially, it corrects the thumbnail of the data that has no projection to the projection of the bounding box, so as to achieve the purpose of fitting under the same projection and meet the different needs of users to download data information.
6. The remote sensing image data search engine method based on high-speed caching technology according to claim 1, characterized in that: The picking linkage, through data picking, can easily supplement AOI coverage by using data from another industry when source data from one industry cannot fully cover the user's AOI range. By selecting a data range box in the map display area and left-clicking the picking range box, users can browse the data in a linked manner. The data coverage analysis provides analytical functions, supporting the calculation of coverage rate of the data coverage query range box. Based on the search, data coverage analysis provides preliminary criteria for user decision-making from a macro perspective. The search range vector polygon is defined as Vs. The search result dataset Ps is traversed, and the union of all data range vector polygons in Ps is calculated to obtain the set Vu, where the effective data range is calculated. Vv = Vu ∩ Vs; Data coverage = Area of effective data range / Area of search range * 100%.
7. The remote sensing image data search engine method based on high-speed caching technology according to claim 1, characterized in that: When a user selects or clicks to download image data in the search results, the task system automatically executes the task, exports the data, packages the bounding box and thumbnails into a tar archive, returns the data address for download, and completes the tar archive download after the user confirms.
8. A remote sensing image data search engine system based on high-speed caching technology, characterized in that: The system includes modules for accessing industry data sources, configuring and storing databases, interface configuration, parameter matching, a unified data access model, a high-speed caching module, a data export task module, a search protocol module, a high-concurrency service framework, image query APIs, attribute retrieval APIs, data download, result analysis, and query result browsing. The unified data access model establishes a hash mapping between image data information (including attribute information, bounding boxes, thumbnails, and attribute details) and image data information items (oids) based on shared entity IDs and their respective data types. The high-concurrency service framework is a network-based image search protocol. The system encapsulates query requests in the background according to the interface parameter configuration table, connects to the data search APIs of major image websites using multiple threads, retrieves image data query information (including bounding boxes and attribute details) separately, and obtains data information through a modeled approach. A two-level high-speed cache for configuration information and query results, including local cache and KV cache database cache, enables efficient filtering and storage of image query results. Through image attribute key-value pair hash mapping and keyword matching modules, standardized basic attribute information is quickly extracted from cached image query results, masking differences in image attribute fields across multiple platforms. The unified results are then added to the high-speed cache. Thumbnails are retrieved uniformly based on entity IDs. Through thumbnail deflection configuration and processing modules, standardized thumbnail processing is achieved. Through data information selection linkage and data coverage analysis, remote sensing data query result analysis is realized. Standardized thumbnails, attributes, and bounding boxes stored by large object OIDs in the cache are retrieved uniformly. Through the task system, the results are organized by range and field into the user-required Shapefile general format, along with thumbnails, to achieve the download of key information from image query results.