A method for querying and processing massive astronomical data
By constructing a distributed spatial database and adopting LB-ANN indexing technology, combining K-Means clustering and index connectivity loss function, the problems of high computational complexity, low storage efficiency and insufficient index optimization in traditional astronomical data query processing methods are solved, and efficient and accurate query processing of massive astronomical data is achieved.
Patent Information
- Application Number
- CN202510352290.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Traditional astronomical data query processing methods have problems such as high computational complexity, low storage efficiency and insufficient index optimization when processing massive astronomical data. Especially when processing high-dimensional data and multimodal data, the query efficiency has significantly decreased.
A distributed spatial database is used to combine time series databases, column storage databases and distributed databases, and the index structure and query load balancing are optimized through load balancing-approximate nearest neighbor search (LB-ANN index) technology. Using the index subset division strategy and index connectivity loss function of K-Means clustering, the allocation method of vector data is dynamically adjusted to balance query efficiency and recall.
It significantly improves the query response speed and data retrieval efficiency of massive astronomical data, solves the problems of low query efficiency and insufficient recall in traditional methods, and achieves efficient and accurate data retrieval.
Smart Images

Figure CN119862211B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of database optimization, and particularly to a method for querying and processing massive astronomical data. Background Art
[0002] With the rapid development of modern astronomy, the improvement of observation technology and the diversification of data collection means, astronomy has entered the big data era; currently, scientific devices such as major observatories, space telescopes, and radio telescopes around the world generate massive amounts of astronomical data every day, including spectral data, image data, electromagnetic spectrum data, star catalog data, time series data, etc.; however, traditional astronomical data query and processing methods have many deficiencies in practical applications, mainly reflected in the following aspects: First, when traditional relational databases process large-scale astronomical data, limited by the storage structure and indexing mechanism, the query speed is restricted by the growth of the data scale; at the same time, the multi-modal characteristics of astronomical data make it difficult to unify the storage model, increasing the complexity of data management; second, astronomical data usually has high-dimensional characteristics, for example, spectral data and time series data often contain hundreds to thousands of dimensions; traditional indexing structures such as B-Tree, R-Tree, KD-Tree, etc. face the "curse of dimensionality" problem when processing high-dimensional data, resulting in a significant decline in indexing efficiency; finally, traditional astronomical database systems usually adopt a centralized architecture, with low utilization rate of computing resources and difficulty in meeting the high-concurrency query requirements of large-scale data. Summary of the Invention
[0003] The present invention proposes an efficient query and processing method for massive astronomical data. Aiming at the multi-modal, multi-time series, large-scale, and high-dynamic characteristics of astronomical data, the present invention constructs an efficient distributed spatial database using a time series database, a column store database, and a distributed database, and optimizes the index structure and query load balance through LB-ANN index (Load Balancing - Approximate Nearest Neighbor Search Index); during the index construction process, an index subset partitioning strategy based on K-Means clustering is proposed, and an index connectivity loss function is introduced to dynamically adjust vector allocation through gradient optimization of relaxation parameters to balance query efficiency and recall rate; in addition, the present invention constructs a SQL-ANN hybrid query to achieve the efficient integration of structured queries and unstructured queries, optimizes query execution through JIT compilation, avoids redundant calculations, and improves distributed query performance; the overall method breaks through the bottlenecks of traditional astronomical data queries in terms of computational complexity, storage efficiency, and index optimization, providing an efficient and accurate solution for the intelligent processing of large-scale astronomical data.
[0004] The present invention provides a method for querying and processing massive astronomical data, which specifically includes the following steps:
[0005] Step S1: Collect astronomical data sets, which include spectral data, image data, electromagnetic spectrum data, star catalog data, sound field data, time series data, and virtual data;
[0006] Step S2: Construct a distributed spatial database. Analyze the data types, data scales, query patterns, data update frequencies, data access methods, and data sharing requirements of the astronomical data sets, and store them in the distributed spatial database. The distributed spatial database includes a time series database, a column store database, and a distributed database;
[0007] Step S3: Construct an LB-ANN index through an overload-aware adaptive strategy and a distributed ANNS index construction technology. Optimize the index structure, data allocation, and query load balancing of the distributed spatial database through the LB-ANN index;
[0008] Step S4: Optimize the query execution of the distributed spatial database through SQL-ANN hybrid queries, reduce redundant calculations, improve query efficiency, and accelerate data retrieval.
[0009] Further, in step S3, the process of constructing the LB-ANN index specifically includes the following:
[0010] Step S31: Obtain the vector data set of the distributed spatial database, set the maximum storage capacity and the maximum overlap factor of the subsets in the vector data set, and calculate the minimum number of index partitions;
[0011] Step S32: Generate an initial cluster center set according to the minimum number of index partitions, and construct index subsets according to the initial cluster center set;
[0012] Step S33: Optimize the basic structure of the index subsets to obtain overload-aware index subsets, allocate the vector data set to the overload-aware index subsets, and perform quantization encoding and storage in the distributed spatial database;
[0013] Step S34: After the allocation of the vector data set is completed, perform a distributed parallel construction of the ANNS subgraph, form a global index through subgraph aggregation, and obtain the LB-ANN index to support efficient retrieval.
[0014] Further, step S32 specifically includes the following:
[0015] Step S321: Randomly extract a sampling subset from the vector data set;
[0016] Step S322: Calculate the initial cluster centers corresponding to the minimum number of index partitions based on the sampling subset using the K-Means clustering algorithm to obtain the initial cluster center set;
[0017] Step S323: Construct the basic structure of the index subset based on the initial cluster center set, that is, determine the partitioning method and the initial center points of the index subset.
[0018] Further, in step S33, an overload-aware adaptive strategy is adopted to optimize the basic structure of the index subset and construct an overload-aware index subset.
[0019] Further, the process of allocating the vector data set to the overload-aware index subset specifically includes the following:
[0020] Step Q1: Calculate the Euclidean distance from each vector in the vector data set to the initial cluster center set, and preliminarily determine the index subset to which the vector belongs as the initial candidate set.
[0021] Step Q2: Based on the initial candidate set, define a relaxation parameter, introduce an index connectivity loss function, dynamically update the relaxation parameter through an adaptive optimization method to generate a gradient-optimized relaxation parameter, and dynamically control the allocation of vectors to the index subset through the gradient-optimized relaxation parameter to preliminarily evaluate the load status of the index subset; the process of dynamically controlling the allocation of vectors through the gradient-optimized relaxation parameter optimizes the redundant partitioning degree of vectors in the index subset, ensures the connectivity of the index structure, and balances query performance and storage redundancy at the same time.
[0022] When the relaxation parameter is small, the vector data set strictly belongs to the nearest index subset, reducing redundancy and improving the index query efficiency.
[0023] When the relaxation parameter is large, the vector data set can belong to multiple index subsets, enhancing the connectivity between subsets and improving the recall rate, which is suitable for cases with complex data distributions.
[0024] Step Q3: Set a maximum load limit, dynamically monitor the load status of the index subset, and when the maximum load limit is reached, optimize the vector attribution through angle constraints and allocate the vector data set to the overload-aware index subset to keep the index structure balanced and improve the query efficiency.
[0025] Further, step S4 specifically includes the following:
[0026] Step S41: Obtain an astronomical query statement, parse and semantically analyze the astronomical query statement to generate a structured query semantic result.
[0027] Query parsing includes identifying the structured data and unstructured data involved in the astronomical query statement, and semantic analysis identifies the type of the astronomical query statement, including Top-K nearest neighbor query, distance range query, distance Join query, and KNN-Join query.
[0028] Step S42: Construct an SQL-ANN balance formula through complexity analysis, query optimization cost minimization, and data distribution modeling. Combine the structured query semantic results to perform SQL-ANN hybrid queries and generate the optimal execution order; ensure that SQL filtering and LB-ANN indexing are in the optimal execution positions;
[0029] Step S43: According to the optimal execution order, use the LB-ANN index to limit the search range, then execute JOIN to improve the calculation efficiency, avoid full-table JOIN, and reduce the calculation complexity;
[0030] Step S44: Based on the structured query semantic results, SQL-ANN balance formula, and optimal execution order in Steps S41 - S43, determine whether Just-In-Time (JIT) compilation of astronomical query statements is required to avoid the additional overhead of function calls, improve the execution efficiency, and thus optimize the query execution of the distributed spatial database.
[0031] Adopting the above solution, the beneficial effects achieved by the present invention are as follows:
[0032] The present invention provides an efficient query processing method for massive astronomical data. By constructing an efficient distributed spatial database that combines a time series database, a column store database, and a distributed database, the efficient storage and management of large-scale astronomical data are realized; this architecture can effectively cope with the multi-modal, multi-temporal, large-scale, and high-dynamic characteristics of astronomical data, significantly improving the data access efficiency and query response speed, and providing a solid technical foundation for the intelligent processing of massive astronomical data;
[0033] The present invention introduces the load balancing - approximate nearest neighbor search (LB-ANN index) technology, innovatively optimizes the index structure of astronomical data, and significantly improves the query load balancing and data retrieval efficiency; through the index subset partitioning strategy of K-Means clustering and the gradient optimization relaxation parameter based on the index connectivity loss function, the allocation method of vector data is dynamically adjusted, thus solving the problems of low query efficiency and insufficient recall rate existing in traditional methods; the optimized LB-ANN index method greatly enhances the index optimization ability of the distributed spatial database, ensures efficient and accurate query processing, effectively avoids redundant calculations, and significantly improves the response speed of data retrieval;
[0034] In addition, through SQL-ANN hybrid query, the present invention realizes the efficient integration of structured query and unstructured query; combined with JIT compilation technology, it optimizes the query execution order, avoids unnecessary full-table scans and redundant calculations, and significantly improves the query efficiency; this method not only effectively solves the computational complexity problem in traditional astronomical data queries, but also enhances the flexibility and intelligence of the system through reasonable query optimization; finally, the query processing ability of the distributed spatial database in a large-scale astronomical data environment has been greatly improved, ensuring efficient and accurate data retrieval. Brief Description of the Drawings
[0035] Figure 1 It is a schematic structural diagram of a method for querying and processing a large amount of astronomical data proposed by the present invention;
[0036] Figure 2 It is a schematic flow diagram of step S4 proposed in Embodiment 7. Detailed Embodiments
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0038] Embodiment 1, according to Figure 1 , the present invention provides a method for querying and processing a large amount of astronomical data, which specifically includes the following steps:
[0039] Step S1: Collect astronomical data sets, which include spectral data, image data, electromagnetic spectrum data, star catalog data, acoustic field data, time series data, and virtual data;
[0040] Step S2: Construct a distributed spatial database, analyze the data types, data scales, query modes, data update frequencies, data access methods, and data sharing requirements of the astronomical data sets, and store them in the distributed spatial database. The distributed spatial database includes a time series database, a column store database, and a distributed database;
[0041] Step S3: Through an overload-aware adaptive strategy and a distributed ANNS index construction technology, construct an LB-ANN index, and optimize the index structure, data distribution, and query load balancing of the distributed spatial database through the LB-ANN index;
[0042] Step S4: Introduce SQL filtering. According to the LB-ANN index, optimize the query execution of the distributed spatial database through SQL-ANN hybrid query, reduce redundant calculations, improve query efficiency, and accelerate data retrieval.
[0043] Embodiment 2. This embodiment is based on Embodiment 1. In this embodiment, in Step S3, the process of constructing the LB-ANN index specifically includes the following contents:
[0044] Step S31: Obtain the vector data set of the distributed spatial database, set the maximum storage capacity and the maximum overlap factor of the subsets in the vector data set, and calculate the minimum number of index partitions;
[0045] Step S32: Generate an initial cluster center set according to the minimum number of index partitions, and construct index subsets according to the initial cluster center set;
[0046] Step S33: Optimize the basic structure of the index subsets to obtain overload-aware index subsets, allocate the vector data set to the overload-aware index subsets, and perform quantization encoding and storage in the distributed spatial database;
[0047] Step S34: After the allocation of the vector data set is completed, perform distributed parallel construction of the ANNS subgraph, form a global index through subgraph aggregation, and obtain the LB-ANN index to support efficient retrieval.
[0048] Embodiment 3. This embodiment is based on Embodiment 2. In this embodiment, Step S32 specifically includes the following contents:
[0049] Step S321: Randomly extract a sampling subset from the vector data set;
[0050] Step S322: Calculate the initial cluster centers corresponding to the minimum number of index partitions based on the sampling subset using the K-Means clustering algorithm to obtain an initial cluster center set;
[0051] Step S323: Construct the basic structure of the index subsets according to the initial cluster center set, that is, determine the partitioning method and the initial center points of the index subsets.
[0052] Embodiment 4. This embodiment is based on Embodiment 3. In this embodiment, in Step S33, an overload-aware adaptive strategy is used to optimize the basic structure of the index subsets to construct overload-aware index subsets.
[0053] Embodiment 5. This embodiment is based on Embodiment 4. In this embodiment, the process of allocating the vector data set to the overload-aware index subsets specifically includes the following contents:
[0054] Step Q1: Calculate the Euclidean distance from each vector in the vector dataset to the initial cluster center set, and preliminarily determine the index subset to which the vector belongs as the initial candidate set;
[0055] Step Q2: Based on the initial candidate set, define a relaxation parameter, introduce an index connectivity loss function, dynamically update the relaxation parameter through an adaptive optimization method to generate a gradient-optimized relaxation parameter. Through the gradient-optimized relaxation parameter, dynamically control the allocation of vectors to the index subset and preliminarily evaluate the load status of the index subset; The process of dynamically controlling the vector allocation through the gradient-optimized relaxation parameter optimizes the redundant partitioning degree of the vectors in the index subset, ensures the connectivity of the index structure, and balances the query performance and storage redundancy at the same time;
[0056] When the relaxation parameter is small, the vector dataset strictly belongs to the nearest index subset, reducing redundancy and improving the index query efficiency;
[0057] When the relaxation parameter is large, the vector dataset can be assigned to multiple index subsets, enhancing the connectivity between subsets and improving the recall rate, which is suitable for the case of complex data distribution;
[0058] ;
[0059] Among them, represents the index connectivity loss function, represents the vector index, represents the total number of vectors in the vector dataset, represents the index subset index, represents the vector, represents the vector is assigned to the set of index subsets, represents the index subset of the cluster center, represents the vector belongs to the number of index subsets, represents the calculation vector to the index center weighted average distance; represents the balance coefficient; represents the global average size of the index subset, represents the current vector of the redundant partitioning degree;
[0060] ;
[0061] Among them, represents the iteration index, represents the relaxation parameter, represents the updated relaxation parameter after iteration; represents the relaxation parameter of the current iteration, denotes the learning rate, denotes the index connectivity loss function the partial derivative with respect to the relaxation parameter;
[0062] ;
[0063] wherein, denotes the vector belongs to the index subset ; denotes the gradient-optimized relaxation parameter, denotes the average distance of all currently assigned vectors; denotes the load balancing factor;
[0064] Step Q3: Set the maximum load limit, dynamically monitor the load status of the index subset. When the maximum load limit is reached, optimize the vector attribution through angle constraints, and allocate the vector dataset to the overload-aware index subset to keep the index structure balanced and improve the query efficiency.
[0065] Example 6. This example is based on Example 4. In this example, the process of allocating the vector dataset to the overload-aware index subset specifically includes the following content:
[0066] Step T1: Calculate the Euclidean distance from each vector in the vector dataset to the initial cluster center set, and preliminarily determine the index subset to which the vector belongs as the initial candidate set;
[0067] Step T2: Based on the initial candidate set, define the relaxation parameter, and dynamically control the redundant partitioning degree of vector allocation to the index subset through the relaxation parameter to preliminarily evaluate the load status of the index subset;
[0068] Step T3: Set the maximum load limit, dynamically monitor the load status of the index subset. When the maximum load limit is reached, optimize the vector attribution through angle constraints to ensure that the vectors are reasonably allocated to the index subset, keep the index structure balanced, and improve the query efficiency.
[0069] Example 7. According to Figure 2 , this example is based on Example 5. In this example, Step S4 specifically includes the following content:
[0070] Step S41: Query analysis: Obtain the astronomical query statement, parse and semantically analyze the astronomical query statement to generate a structured query semantic result;
[0071] Query parsing includes identifying the structured data and unstructured data involved in the astronomical query statement, and semantic analysis identifies the types of astronomical query statements, including Top-K nearest neighbor query, distance range query, distance Join query, and KNN-Join query;
[0072] Step S42: Generate the execution order: Construct an SQL-ANN balance formula through complexity analysis, query optimization cost minimization, and data distribution modeling, combine the structured query semantic results, perform SQL-ANN hybrid queries, and generate the optimal execution order; ensure that SQL filtering and LB-ANN indexing are in the optimal execution positions;
[0073] ;
[0074] Among them, represents the optimal SQL filtering ratio, represents the total number of data points in the distributed spatial database, represents the target number of neighbors of the LB-ANN index, represents the probability that the target data is still included after SQL filtering, represents the index scan calculation factor, represents the SQL filtering full table scan cost, represents the LB-ANN index calculation cost, represents the Top-K calculation cost of the LB-ANN index, represents the index complexity coefficient of the LB-ANN index, represents the correction term, represents logarithmic growth, represents power growth;
[0075] Step S43: Execute JOIN: According to the optimal execution order, use the LB-ANN index to limit the search range, and then execute JOIN to improve the calculation efficiency, avoid full table JOIN, and reduce the calculation complexity;
[0076] Step S44: JIT compilation: According to the structured query semantic results, SQL-ANN balance formula, and optimal execution order in Steps S41 - S43, determine whether JIT compilation is required for the astronomical query statement, avoid the additional overhead of function calls, improve the execution efficiency, and thus optimize the query execution of the distributed spatial database.
[0077] Example 8, this example is based on Example 5. In this example, Step S4 specifically includes the following content:
[0078] Step R1: Receive the astronomical query statement, parse and semantically analyze the astronomical query statement to generate structured query semantic results; query parsing includes identifying the structured and unstructured data involved in the astronomical query statement, and semantic analysis identifies the types of astronomical query statements, including Top-K nearest neighbor queries, distance range queries, distance Join queries, and KNN-Join queries;
[0079] Step R2: Map the structured query semantic results to an execution strategy suitable for the current query type and generate an optimal execution order;
[0080] Step R3: According to the optimal execution order, use the LB-ANN index to limit the search scope, then perform JOIN to improve the computing efficiency, avoid full-table JOIN, and reduce the computing complexity;
[0081] Step R4: Determine whether JIT compilation of the astronomical query statement is required based on the structured query semantic results and the optimal execution order, avoid the additional overhead of function calls, improve the execution efficiency, and thus optimize the query execution of the distributed spatial database.
[0082] The above describes the present invention and its implementation manners. Such a description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the spirit of the present invention, they should fall within the protection scope of the present invention.
Claims
1. A method for querying and processing massive astronomical data, characterized in that: The method specifically comprises the following steps: Step S1: Collect astronomical data sets; Step S2: construct a distributed spatial database, analyze the characteristics of the astronomical data set, and store it in the distributed spatial database; Step S3: construct a load-balanced approximate nearest neighbor index, i.e., LB-ANN index, through an overload-aware adaptive strategy and distributed ANNS index construction technology, and optimize the index structure, data allocation, and query load balancing of the distributed spatial database through the LB-ANN index; Step S4: Introduce SQL filtering and optimize the query execution of distributed spatial database through SQL-ANN hybrid query; The step S4 specifically includes the following contents: Step S41: obtaining an astronomical query statement, parsing and semantically analyzing the astronomical query statement, and generating a structured query semantic result; Step S42: constructing a SQL-ANN balance formula through complexity analysis, query optimization cost minimization and data distribution modeling, calculating the optimal SQL filtering ratio, combining the structured query semantic results, performing SQL-ANN hybrid query, and generating the optimal execution order to ensure that SQL filtering and LB-ANN indexing are in the optimal execution position; ; in, Indicates the optimal SQL filtering ratio, Represents the total number of data points in a distributed spatial database, represents the number of target neighbors indexed by LB-ANN, Indicates the probability of still containing the target data after SQL filtering. Indicates the index scan calculation factor, Indicates the SQL filtering full table scan cost, represents the LB-ANN index calculation cost, represents the Top-K computation cost of the LB-ANN index, represents the LB-ANN index complexity coefficient, Represents a correction item, represents logarithmic growth, represents power growth; Step S43: Execute JOIN according to the optimal execution order; Step S44: According to steps S41 to S43, it is determined whether it is necessary to perform JIT compilation on the astronomical query statement to optimize the query execution of the distributed spatial database.
2. A method for querying and processing massive astronomical data according to claim 1, characterized in that: The process of building the LB-ANN index includes the following: Step S31: obtaining a vector data set of a distributed spatial database, and calculating the minimum number of index partitions; Step S32: Generate an initial cluster center set according to the minimum number of index partitions and construct an index subset; Step S33: optimizing the index subset, constructing the overload-aware index subset, and assigning the vector data set to the overload-aware index subset; Step S34: After the vector data set is allocated, ANNS subgraphs are constructed in distributed parallel, and then the subgraphs are aggregated to form a global index to construct a LB-ANN index.
3. A method for querying and processing massive astronomical data according to claim 2, characterized in that: In step S33, an overload-aware adaptive strategy is used to optimize the basic structure of the index subset to construct an overload-aware index subset.
4. A method for querying and processing massive astronomical data according to claim 2, characterized in that: The process of assigning a vector dataset to an overload-aware index subset includes the following: Step Q1: Calculate the Euclidean distance from each vector in the vector data set to the initial cluster center set, and preliminarily determine the index subset to which the vector belongs as the initial candidate set; Step Q2: Based on the initial candidate set, dynamically control the vector allocation and preliminarily evaluate the load status of the index subset; Step Q3: Dynamically monitor the load status of the index subset, optimize the vector attribution through angle constraints, and assign the vector data set to the overload-aware index subset.
5. A method for querying and processing massive astronomical data according to claim 4, characterized in that: The dynamic control of vector allocation specifically includes: defining relaxation parameters based on the initial candidate set, introducing the index connectivity loss function, dynamically updating the relaxation parameters, generating gradient optimized relaxation parameters, and dynamically controlling the allocation of vectors to index subsets through gradient optimized relaxation parameters.