Ephemeris data storage indexing method and equipment based on distributed database

Through the method based on dual-table collaborative storage structure and dynamic index update, combined with distributed query optimization of open source distributed relational databases, the problem of index delay of distributed databases under high throughput data processing is solved, and efficient star ephemeris data management and query are achieved.

CN120296018AActive Publication Date: 2025-07-11JIANGSU HEHAI SCIENCE & TECHNOLOGY PARK CO LTD

Patent Information

Application Number
CN202510787880.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-11
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

In the case of large-scale high-throughput data processing, the data index delay of distributed databases is large, which cannot meet the needs of high concurrent query, and the cross-node query latency increases.

Method used

The ephemeris data storage method based on a dual-table collaborative storage structure is adopted, including ephemeris data table and ephemeris index table. The index table is updated in real time through a dynamic maintenance mechanism, and the distributed query optimization technology of open source distributed relational databases is used to quickly locate data sharding and parallel query.

Benefits of technology

It effectively reduces the data index delay, improves query efficiency, supports efficient management and dynamic expansion of massive data, and maintains stable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296018A_ABST
    Figure CN120296018A_ABST
Patent Text Reader

Abstract

The invention discloses an ephemeris data storage indexing method and equipment based on a distributed database, belongs to the technical field of ephemeris data management, and remarkably improves the query performance of mass spatio-temporal data through a double-table collaborative storage and dynamic index optimization mechanism. The method comprises the following steps: constructing a double-table collaborative storage structure of an ephemeris data table and an ephemeris index table; preprocessing ephemeris data, converting the ephemeris data into structured data, arranging and storing the structured data in the double-table collaborative storage structure according to the sequence of satellite numbers and global positioning system weeks, and optimizing the continuity and the distribution rule of the data; updating the ephemeris index table in real time based on a dynamic maintenance mechanism, and triggering self-adaptive index recombination based on a data fragment splitting event; a two-stage retrieval strategy is adopted, a target data primary key range and data fragments are quickly positioned through joint indexing, parallel processing query is carried out to realize cross-node accurate retrieval, and finally, a result is fed back to a client, so that query delay of massive ephemeris data is reduced, and query efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ephemeris data management, and in particular to an ephemeris data storage and indexing method and device based on a distributed database. Background Art

[0002] With the rapid development of the Global Navigation Satellite System (GNSS), multi-constellation systems including the Global Positioning System and Beidou continue to expand. As the core spatio-temporal reference information, ephemeris data plays an irreplaceable role in positioning, navigation, and timing (PNT) services. High-precision ephemeris data is not only a key input for real-time navigation systems but also widely used in fields such as geophysical monitoring, precision agriculture, intelligent transportation, and aerospace. Modern global satellite navigation systems are characterized by multiple satellites, multiple frequency bands, and high update frequencies, resulting in dynamic, massive, and highly time-sensitive ephemeris data. This poses great challenges to traditional data management technologies. On the one hand, ephemeris data needs to support high-concurrency queries with millisecond-level responses to meet the requirements of real-time applications. On the other hand, its spatio-temporal correlation requires the storage system to have efficient multi-dimensional indexing capabilities. How to achieve efficient storage, fast retrieval, and dynamic expansion of ephemeris data has become a key issue in the construction of satellite navigation data infrastructure.

[0003] To address the significant technical bottlenecks faced by traditional databases in the field of ephemeris data management, open-source distributed databases can be adopted. Open-source distributed relational databases combine the advantages of traditional relational databases and non-relational databases (NoSQL), enabling horizontal expansion of large-scale data while ensuring transaction consistency. Open-source distributed relational databases solve challenges such as high concurrency, low latency, and large-scale data storage through automatic sharding mechanisms and efficient distributed query optimization. They can horizontally expand infinitely to store massive data and support online transaction processing (OLTP) and online analytical processing (OLAP), effectively improving query efficiency. However, ephemeris data is frequently updated, and when updating data, it is usually necessary to recalculate and update the index, resulting in increased overhead. In addition, data is distributed across multiple nodes, and query latency often increases when cross-node queries are involved. How to reduce the latency of data indexing under the condition of meeting large-scale high-throughput data processing is a key issue for distributed databases. Summary of the Invention

[0004] The present invention aims to overcome the problem of large latency of data indexing in large-scale high-throughput data processing in the prior art, and provides an ephemeris data storage and indexing method and device based on a distributed database.

[0005] To achieve the above object, the technical solution of the present invention is: providing an ephemeris data storage and indexing method based on a distributed database, including the following steps, Step S1: Construct a dual-table collaborative storage structure for the ephemeris data table and the ephemeris index table; Step S2: Preprocess the ephemeris data to convert it into structured data, and store the structured data in the dual-table collaborative storage structure according to the sorting of satellite numbers and GPS weeks; Step S3: Based on the dynamic maintenance mechanism, update the ephemeris index table in real time during the storage stage of the structured data; Step S4: Obtain the primary key range of the target data through the ephemeris index table, and quickly locate the data shard where the target data is located based on the primary key range and query it.

[0006] In one embodiment, in the above step S4, it specifically includes the following steps: Step S41: Based on the query request containing the satellite number and the GPS week, parse and extract the filtering conditions; Step S42: Initiate a query to the ephemeris index table based on the filtering conditions, and establish a combined index of the satellite number and the GPS week in the ephemeris index table to match the target index records that meet the conditions and obtain the starting primary key and the ending primary key of the target data; Step S43: Combine the starting primary key and the ending primary key of the target data obtained from the query and the filtering conditions in the query request to generate a query condition, and locate the data shard number and the storage node range where the target data is located based on the range of the starting primary key and the ending primary key of the target data to obtain the index range; Step S44: Query the index range in the ephemeris data table based on the query condition to obtain the output result.

[0007] In one embodiment, in the above step S44, when the target data spans multiple primary key ranges, the query task will be automatically split and the relevant data shards will be accessed in parallel.

[0008] In one embodiment, in the above step S4, the query adopts an index cost model, and the specific formula is as follows: TC = ITC 1 + NC 1 + SC 1 + RLC + TSC + NC 2 + SC 2 + MC + RC Among them, TC is the total index cost, ITC 1 , NC 1 , SC 1 are the query cost models in step S42; RLC , TSC , NC 2 and SC 2 are the query cost models in step S43; MC and RC are the cost models in step S44; Each cost factor is dynamically calibrated through the statistical information of the execution plan of the open-source distributed relational database.

[0009] In one embodiment, in the above step S2, it specifically includes the following steps. Step S21, parse and standardize the ephemeris data, perform hash grouping according to the satellite number and the GPS week, filter redundant data and complete verification, and then generate structured data; Step S22, the user or the data access layer submits a write request to the open-source distributed relational database based on the structured data, specifies the storage target of the structured data as the ephemeris data table, and completes operations including data sharding strategy selection and transaction consistency verification; Step S23, store the structured data into the ephemeris data table, and trigger the update condition judgment mechanism of the ephemeris index table; Step S24, dynamically determine whether to update the ephemeris index table or update the index record according to the GPS week and the satellite number.

[0010] In one embodiment, in the above step S3, it specifically includes the following steps. Step S31, query the ephemeris index table according to the satellite number and the GPS week of the structured data; Step S32, if there is a matching index record, update the end primary key value of the index record; Step S33, if there is no matching index record, insert a new index record.

[0011] In one embodiment, in the above step S1, the ephemeris data table includes the primary key, satellite number, GPS week, reference time, and orbit parameter set of the ephemeris data; the ephemeris index table includes a composite primary key, a start primary key, and an end primary key.

[0012] In one embodiment, in the above step S42,ITC 1 = iRC × sF represents the cost of scanning the ephemeris index table to locate records that meet the query conditions. iRC represents the number of records hit in the ephemeris index table. sF represents the cost of single-line scanning. NC 1 = iRC × nF represents the network transmission cost of transmitting the index query result, i.e., the start primary key or the end primary key, from the distributed key-value storage component node to the open-source distributed relational database server. nF represents the network transmission factor. SC 1 = iRC × selF represents the cost of calculating the filtering condition. selF represents the condition calculation factor; in step S43, RLC = log 2 (N_Regions) × rF represents the positioning cost of the open-source distributed relational database to determine the data shard where the data is located according to the primary key range. rF represents the data shard positioning factor. TSC = dRC × sF represents the data scanning cost at the distributed key-value storage component node, where dRC represents the number of records returned by the data table. NC 2 = dRC × nF represents the network transmission cost of transmitting the data table query result from the distributed key-value storage component node to the open-source distributed relational database server. SC 2 = dRC × selF represents the calculation cost of executing the final filtering condition of the data table; in step S44, MC = dRC × mF represents the cost of operations such as sorting the results returned from multiple data shards, where mF represents the merge operation factor.

[0013] In one embodiment, the method optimizes the data shard positioning path according to the scheduler metadata cache of the distributed scheduling server.

[0014] The present invention also provides an ephemeris data storage and indexing device based on a distributed database. The device executes the method described above. The device includes a number of servers, which are respectively used to build a distributed key-value storage component, a distributed scheduling server, and an open-source distributed relational database server cluster node. A data visualization tool and a system monitoring and alarm toolkit are deployed on the servers of the distributed key-value storage component nodes, and the 2-replica Raft protocol is used to ensure data consistency. The data shard splitting threshold is 96MB to 144MB, and the key ranges of the sub-data shards are continuous after splitting.

[0015] In summary, the present invention discloses an ephemeris data storage and indexing method based on a distributed database. The ephemeris data storage and indexing method of the present invention based on an open-source distributed relational database utilizes a dual-table collaborative storage architecture, an ephemeris index table dynamic maintenance mechanism, and a two-level retrieval strategy to achieve efficient management of massive spatio-temporal data. By designing an ephemeris index table and combining it with the distributed data shard storage mechanism of an open-source distributed relational database, the present invention realizes efficient retrieval of complex ephemeris data under multi-condition queries. By storing ephemeris data in a structured manner and innovatively designing an ephemeris index table, and utilizing the distributed horizontal expansion ability of an open-source distributed relational database, it supports dynamically adding storage nodes to cope with data volume growth while maintaining stable query performance and ensuring long-term stable operation of the system.

[0016] To make the above features and advantages of the invention more obvious and understandable, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of the method in the present invention.

[0018] Figure 2 It is a flowchart of ephemeris data storage in the present invention.

[0019] Figure 3 It is a diagram of the ephemeris index update strategy in the present invention.

[0020] Figure 4 It is a flowchart of the ephemeris data secondary retrieval strategy query in the present invention.

[0021] Figure 5 It is a node topology diagram of the open-source distributed relational database cluster in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the objectives and technical solutions of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0023] In order to achieve efficient storage and query optimization of ephemeris data in an open-source distributed relational database (TiDB), and to solve the problem of overcoming large indexing delays in the case of massive growth of ephemeris data, the present invention provides a method for storing and indexing ephemeris data based on a distributed database. Figure 1 This is a flowchart of the method of the present invention, as Figure 1 shown, the steps include: Step S1, constructing a dual-table collaborative storage structure for the ephemeris data table and the ephemeris index table; Step S2, preprocessing the ephemeris data to convert it into structured data, and storing the structured data in the dual-table collaborative storage structure according to the sorting of satellite numbers and GPS week numbers; Step S3, based on a dynamic maintenance mechanism, updating the ephemeris index table in real time during the storage stage of the structured data; Step S4, obtaining the primary key range of the target data through the ephemeris index table, quickly locating the data shard (Region) where the target data is located based on the primary key range and querying, completing the indexing and outputting the result.

[0024] In step S1, the ephemeris data table is a structured database for storing the data body of the broadcast ephemeris file. The ephemeris data table corresponds to each field of the broadcast ephemeris data. The ephemeris data table includes the primary key, satellite number, GPS week, reference time, and orbit parameter set of the ephemeris data. The satellite number field is stored using a fixed-length string data type (CHAR format), and the GPS week field is stored using an integer data type (INT) to store the week count.

[0025] The ephemeris index table is used to support and optimize the indexing of the ephemeris data table, and is based on the storage and indexing characteristics of the open-source distributed relational database and the structural characteristics of the ephemeris data. The ephemeris index table includes a composite primary key, a starting primary key, and an ending primary key. The starting primary key and ending primary key fields are of the large integer type (BIGINT) to record the primary key interval corresponding to the ephemeris data table.

[0026] In step S2, the dual-table collaborative storage model makes use of the gradual continuity of data on the storage medium in the open-source distributed relational database and the storage characteristics of data sharding in units of the data shards. Specifically, the preprocessing method and storage steps of the ephemeris data are as Figure 2 shown and include: Step S21: Parse and standardize the ephemeris data in the original Receiver Independent Exchange Format (RINEX), extract information such as satellite orbit parameters, time parameters, satellite numbers, and GPS weeks, and perform hash grouping according to the satellite numbers. Multiple groups of ephemeris data of the same satellite are sorted in ascending order according to the GPS week. After filtering redundant data and completing format verification and logical verification, structured data that conforms to the database storage specification is generated; Step S22: The user or the data access layer submits a transactional write request to the open-source distributed database based on the structured data, specifying that the storage target of the structured data is the ephemeris data table, and operations such as data sharding strategy selection and transaction consistency verification are completed to ensure the atomicity and traceability of the write operation; Step S23: Persistently store the structured data into the ephemeris data table according to preset rules, and trigger the ephemeris index table update condition judgment mechanism; Step S24: Dynamically determine whether to update the ephemeris index table or update the index record according to features such as the GPS week and the satellite number.

[0027] In step S21, it is necessary to convert the Coordinated Universal Time (UTC) of the ephemeris data into the second format within the GPS week.

[0028] In step S3, when the structured data is stored in the ephemeris data table, the update strategy of the ephemeris index table is as Figure 3 shown and includes the following steps: Step S31: Query the ephemeris index table according to the satellite number and the GPS week of the structured data; Step S32: If there is a matching index record, update the end primary key value of the index record; Step S33: If there is no matching index record, insert a new index record.

[0029] First, check whether there is already an index record in the ephemeris index table that is the same as the current satellite number and the GPS week. If there is a corresponding record, it indicates that the inserted data belongs to an existing ephemeris data group. Then, only update the end primary key value in this index record to avoid generating redundant index records, and the updated value is the data primary key of the data currently written into the ephemeris data table. If there is no corresponding record in the ephemeris index table, it means that the inserted data belongs to a new ephemeris data group. Then, insert a new index record, and the record fields, GPS week, start primary key, and end primary key values in the ephemeris index table are respectively the satellite number, GPS week, start primary key, and end primary key values in the current ephemeris data table.

[0030] In step S4, a two-level retrieval strategy will be adopted, and its flowchart is as Figure 4 shown, including the following steps: Step S41, based on a query request that includes the satellite number, time parameter, the GPS week, etc., parse and extract filtering conditions such as the satellite number and the GPS week. Step S42, initiate a query to the ephemeris index table based on the filtering conditions, and establish a combined index of the satellite number and the GPS week in the ephemeris index table to quickly match the target index record that meets the conditions and obtain the start primary key and the end primary key of the target data. Step S43, the open-source distributed relational database server will generate accurate query conditions by combining the start primary key and the end primary key of the target data obtained from the query and the filtering conditions in the query request, and locate the data shard number and the storage node range where the target data is located according to the range of the start primary key and the end primary key of the target data to obtain a reduced index range, avoiding a full database scan. Step S44, query the ephemeris data table based on the query conditions. If the target data spans multiple primary key ranges, the query task will be automatically split and access the relevant data shards in parallel to make full use of the performance advantages of the distributed architecture and reduce the index latency; sort the query results according to time and output the results in the form of formatted data.

[0031] In addition, an index cost model is adopted in step S4, and the specific formula is as follows: TC = ITC 1 + NC 1 + SC 1 (ephemeris index table stage) + RLC + TSC +NC 2 + SC 2 (Ephemeris data table stage) + MC + RC (Result stage) Wherein, TC is the total index cost, ITC 1 , NC 1 , SC 1 are the query cost models in step S42; RLC , TSC , NC 2 and SC 2 are the query cost models in step S43; MC and RC are the cost models in step S44. Each cost factor is dynamically calibrated through the statistical information of the execution plan of the open-source distributed relational database.

[0032] In step S42, ITC 1 is the index scan cost, and the cost model ITC 1 corresponds to the process of determining the data shard where the target data is located by using the combined index of the satellite number and the GPS week and filtering out irrelevant data shards to quickly locate the target index record. ITC 1 is positively correlated with the number of hit records in the index table. Specifically, ITC 1 = iRC × sF , representing the cost of scanning the ephemeris index table to locate the records that meet the query conditions, iRC represents the number of hit records in the ephemeris index table, sF represents the single-row scan cost. The cost model NC 1 corresponds to the process of transmitting the query result from the distributed key-value storage component node to the open-source distributed relational database server. Specifically, NC 1 = iRC × nF , representing the network transmission cost of transmitting the index query result (starting primary key / ending primary key) from the distributed key-value storage component node to the open-source distributed relational database server, where nF represents the network transmission factor. The cost model SC 1 corresponds to the process of verifying conditions such as the time range boundary. Specifically,SC 1 = iRC × selF represents the cost of calculating the filtering condition. selF represents the conditional calculation factor. In step S42, the compressible scan range can be used to avoid full table traversal.

[0033] In step S43, RLC is the data shard localization overhead, which determines the localization cost of the data shard where the data is located according to the primary key range. The cost model RLC corresponds to the process of locating the target data shard based on the primary key range, and its time consumption is logarithmically positively correlated with the number of data shards. RLC can reduce the probability of cross-node queries. Specifically, RLC = log 2 (N_Regions) × rF represents the localization cost of the data shard where the data is located determined by the open-source distributed relational database according to the primary key range. rF represents the data shard localization factor. The cost models TSC and SC 2 correspond to the processes of optimizing the scanned data and performing time interpolation calculation through the Log-Structured Merge Tree (LSM-Tree) within the target data shard respectively. Specifically, TSC = dRC × sF represents the cost of scanning data at the distributed key-value storage component node, where dRC represents the number of records returned by the data table; NC 2 = dRC × nF represents the network transmission cost of transferring the query result of the data table from the distributed key-value storage component node to the open-source distributed relational database server; the cost model NC 2 corresponds to the process of transmitting the result set to the open-source distributed relational database server. Specifically, SC 2 = dRC × selF represents the cost of calculating the final filtering condition of the data table. Step S43 achieves distributed acceleration by querying multiple data shards in parallel.

[0034] In step S44, MC is the result merging cost, and the cost model MCThe parallel merge algorithm using an open-source distributed relational database performs in-memory sorting on the results of multiple data shards, feeds the query results back to the open-source distributed relational database server for the corresponding multiple data shards, and performs in-memory merge sorting according to the ephemeris reference time, and the time consumption has a linear logarithmic relationship with the data volume. Specifically, MC = dRC × mF represents the cost of operations such as sorting the return results of multiple data shards, where mF represents the merge operation factor. The cost model RC corresponds to the process of serializing the results into the format specified by the client; RC represents returning the final result to the client, which is related to the network bandwidth.

[0035] This index cost model effectively shows that by compressing the scan range through the ephemeris index table to reduce ITC 1 and RLC and using distributed parallel queries to reduce the cost MC can effectively improve the index query performance.

[0036] Furthermore, in step S44, the data shard where the target data is located will be located according to the range of the start primary key and the end primary key of the target data, and parallel queries will be performed on multiple distributed key-value storage component nodes; in step S45, each data shard will return the query results to the open-source distributed relational database server. After receiving the data stream from multiple distributed key-value storage component nodes, the open-source distributed relational database server stores it in the memory buffer, sorts it in ascending order according to the ephemeris reference time field, converts the final result into the MySQL protocol format, and returns it to the client for reception.

[0037] In addition, the data shard location path can be optimized according to the scheduler metadata cache of the distributed scheduling server.

[0038] Figure 5This is the node topology diagram of the open-source distributed relational database cluster of the present invention. The present invention also includes an ephemeris data storage and indexing device based on a distributed database. The device can adopt the cluster architecture of an open-source distributed relational database (TiDB-v7.5.1), and use 7 servers to build an open-source distributed relational database cluster. Among them, 3 independent servers are used as cluster nodes of the distributed key-value storage component, each equipped with an Intel Xeon Platinum 8374C processor, 64GB of memory, and 20TB of solid-state drive (SSD) storage; 2 independent servers are used as cluster nodes of the distributed scheduling server, equipped with an Intel Xeon Silver 4210R processor, 64GB of memory, and 1.2TB of solid-state drive storage; 2 independent servers are used as cluster nodes of the open-source distributed relational database server, equipped with an Intel Xeon Silver 4210R processor, 64GB of memory, and 1.2TB of solid-state drive storage; and a data visualization tool (Grafana) and a system monitoring and alarm toolkit (Prometheus) are simultaneously deployed on the servers of the distributed key-value storage component nodes to monitor the nodes, for real-time monitoring of the cluster status, namely the queries per second (QPS), latency (the time interval from when a request is sent to when a response is received), and resource utilization metrics. Further, the distributed key-value storage component nodes adopt a 2-replica consensus algorithm (Raft protocol) to ensure data consistency; in the open-source distributed relational database, the granularity of each data shard is 96MB to 144MB. When the data volume continues to grow and reaches the threshold, the distributed key-value storage component nodes will automatically trigger a split operation according to the data distribution characteristics, dividing the original data shard into two sub-data shards, but the key ranges of the newly generated sub-data shards and the original data shard always remain globally ordered and non-overlapping, that is, the key ranges of the split sub-data shards are continuous. In addition, the servers are connected through a gigabit network to ensure the efficiency of data transmission within the cluster. Each server in the data center is interconnected through a high-speed local area network, ensuring low-latency and high-bandwidth communication between nodes.

[0039] In summary, the present invention discloses an ephemeris data storage and indexing method and device based on a distributed database. The present invention's method for storing and indexing ephemeris data based on an open-source distributed relational database utilizes a dual-table collaborative storage architecture, an ephemeris index table dynamic maintenance mechanism, and a two-level retrieval strategy to achieve efficient management of massive spatio-temporal data. By designing an ephemeris index table and combining it with the distributed data sharding storage mechanism of the open-source distributed relational database, the present invention realizes efficient retrieval of complex ephemeris data under multi-condition queries; by storing ephemeris data in a structured manner and innovatively designing an ephemeris index table, and utilizing the distributed horizontal expansion ability of the open-source distributed relational database, it supports dynamically adding storage nodes to cope with data volume growth while maintaining stable query performance, ensuring the long-term stable operation of the system.

[0040] Although the present invention has been disclosed above by way of embodiments, it is not intended to limit the present invention. Any person with ordinary knowledge in the relevant technical field may make some modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to that defined by the appended patent application scope.

Claims

1. A method for storing and indexing ephemeris data based on a distributed database, characterized in that, It includes the following steps: Step S1, constructing a double-table collaborative storage structure for the ephemeris data table and the ephemeris index table; Step S2, preprocessing the ephemeris data to convert it into structured data, and storing the structured data in the double-table collaborative storage structure according to the sorting of satellite numbers and GPS weeks; Step S3, updating the ephemeris index table in real time during the storage stage of the structured data based on the dynamic maintenance mechanism; Step S4, obtaining the primary key range of the target data through the ephemeris index table, and quickly locating the data shard where the target data is located and querying based on the primary key range.

2. The ephemeris data storage indexing method based on a distributed database according to claim 1, characterized in that In the above step S4, it specifically includes the following steps: Step S41, based on the query request containing the satellite number and the GPS week, parsing and extracting the filtering conditions; Step S42, initiating a query to the ephemeris index table based on the filtering conditions, and establishing a joint index of the satellite number and the GPS week in the ephemeris index table to match the target index records that meet the conditions and obtain the start primary key and end primary key of the target data; Step S43, combining the start primary key and end primary key of the target data obtained from the query and the filtering conditions in the query request to generate a query condition, and positioning the data shard number and storage node range where the target data is located based on the range of the start primary key and the end primary key of the target data to obtain the index range; Step S44, querying the index range in the ephemeris data table based on the query condition to obtain the output result.

3. The ephemeris data storage and indexing method based on a distributed database according to claim 2, wherein, In the above step S44, when the target data spans multiple primary key ranges, the query task will be automatically split and the relevant data shards will be accessed in parallel.

4. The ephemeris data storage indexing method based on a distributed database according to claim 3, characterized in that, The above step S4 adopts an index cost model, and the specific formula is as follows: TC = ITC 1 + NC 1 + SC 1 + RLC + TSC + NC 2 + SC 2 + MC + RC Among them, TC is the total index cost, ITC 1 , NC 1 , SC 1 is the query cost model in step S42; RLC , TSC , NC 2 and SC 2 is the query cost model in step S43; MC and RC is the cost model in step S44; Each cost factor is dynamically calibrated through the statistical information of the execution plan of the open-source distributed relational database.

5. A method for storing and indexing ephemeris data based on a distributed database according to claim 1, characterized in that, In the above step S2, it specifically includes the following steps: Step S21, parsing and standardizing the ephemeris data, performing hash grouping according to the satellite number and the GPS week, filtering redundant data and completing verification, and generating structured data; Step S22, the user or the data access layer submits a write request to the open-source distributed relational database based on the structured data, specifying that the storage target of the structured data is the ephemeris data table, and completing operations including data sharding strategy selection and transaction consistency verification; Step S23, storing the structured data in the ephemeris data table, and triggering the update condition judgment mechanism of the ephemeris index table; Step S24, dynamically determining whether to update the ephemeris index table or update the index records according to the GPS week and the satellite number.

6. The ephemeris data storage and indexing method based on a distributed database according to claim 1, wherein, In the above step S3, it specifically includes the following steps: Step S31, querying the ephemeris index table according to the satellite number and the GPS week of the structured data; Step S32, if there are matching index records, updating the end primary key value of the index records; Step S33, there is no matching index record, insert a new index record.

7. A ephemeris data storage index method based on a distributed database according to claim 1, characterized in that In the above step S1, the ephemeris data table includes the primary key of the ephemeris data, satellite number, GPS week, reference time, and orbit parameter set; the ephemeris index table includes a composite primary key, start primary key, and end primary key.

8. A method for storing and indexing ephemeris data based on a distributed database according to claim 4, characterized in that, In the above step S42, ITC 1 = iRC × sF , which represents the cost of scanning the ephemeris index table to locate records that meet the query conditions, iRC represents the number of records hit in the ephemeris index table, sF represents the cost of single-row scanning; NC 1 = iRC × nF , which represents the network transmission cost of transferring the index query result, i.e., the start primary key or the end primary key, from the distributed key-value storage component node to the open-source distributed relational database server, nF represents the network transmission factor; SC 1 = iRC × selF , which represents the cost of calculating the filtering conditions, selF represents the condition calculation factor; In step S43, RLC = log 2 (N_Regions) × rF , indicating the positioning cost of the data shard where the data is located determined by the open-source distributed relational database according to the primary key range. rF Indicates the data shard positioning factor; TSC = dRC × sF , indicating the data scanning cost of the distributed key-value storage component node, where dRC Indicates the number of records returned by the data table; NC 2 = dRC × nF , indicating the network transmission cost of transferring the query result of the data table from the distributed key-value storage component node to the open-source distributed relational database server; SC 2 = dRC × selF , indicating the calculation cost of executing the final filtering condition of the data table; In step S44, MC = dRC × mF , indicating the cost of operations such as sorting the returned results of multiple data shards, where mF Indicates the merge operation factor.

9. The ephemeris data storage and indexing method based on a distributed database according to claim 8, wherein The method optimizes the data shard location path according to the scheduler metadata cache of the distributed scheduling server.

10. An ephemeris data storage index device based on a distributed database, characterized in that, The device executes the method according to any one of claims 1-9. The device includes several servers, which are respectively used to build a distributed key-value storage component, a distributed scheduling server, and an open-source distributed relational database server cluster node; a data visualization tool and a system monitoring and alarm toolkit are deployed on the server of the distributed key-value storage component node, and the 2-replica Raft protocol is used to ensure data consistency; the data shard splitting threshold is 96MB to 144MB, and the key ranges of the sub-data shards are continuous after splitting.

Citation Information

Patent Citations

  • Storage and retrieval method for GPS satellite broadcast ephemeris data

    CN104035976A

  • Method and device for extracting mass data

    CN104112011A

  • Distributed calculating platform facing to spatio-temporal data k neighbor query and query method

    CN105893605A

  • A satellite remote sensing big data optimization inquiry method based on a mixed index

    CN109284338A

  • Massive astronomical data query processing method

    CN119862211A

Cited By

  • GNSS (Global Navigation Satellite System) trajectory data storage and query method oriented to quick retrieval of position trajectory

    CN120561388A