Data range query method and device for Cassandra distributed key-value storage system

Through the Cassandra-oriented data range query method, dynamically adjusting the number of packets and query range, the performance bottleneck of Cassandra in scope query is solved, and more efficient data retrieval and query response are achieved.

CN119862209BActive Publication Date: 2025-06-06HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510346080.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-06
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The Cassandra distributed key-value storage system has performance bottlenecks in scope query. It is mainly because the data is not organized in the order of keys, so the scope query needs to read a large amount of irrelevant data across multiple partitions, which reduces query efficiency.

Method used

A data range query method for Cassandra is proposed. By obtaining the range query request and analyzing the query range and data quantity, dynamically adjusting the number of packets, grouping data according to the key prefix, optimizing the query range and data quantity, and reducing the reading and transmission of irrelevant data.

Benefits of technology

It effectively improves the scope query performance of Cassandra, reduces the overhead of cross-partition access, improves data retrieval efficiency, reduces the system I/O burden, optimizes query response speed, and adapts to data query needs of different scales and complexities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862209B_ABST
    Figure CN119862209B_ABST
Patent Text Reader

Abstract

The present invention discloses a data range query method and device for a Cassandra distributed key-value storage system, and relates to the field of data query. The method groups key-value pairs with the same key prefix according to the key prefix to realize an efficient grouping and partitioning strategy. When executing a range query, the number of groups to be accessed is accurately determined according to the query range, and a batch scheduling strategy is adopted to send the query request. During the query process, the number of groups for the next subsequent round is dynamically adjusted according to the query results returned each time, thereby optimizing data access efficiency. The present invention effectively reduces the reading of invalid data in the traditional range query of Cassandra, and solves the problem of low performance of Cassandra range query.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data query, and in particular to a data range query method and device for a Cassandra distributed key-value storage system. Background Art

[0002] With the rapid development of information technology, the amount of global data is growing exponentially, and the demand for data storage is also rising sharply. According to the forecast of International Data Corporation (IDC), the scale of global data storage will continue to expand in the future, especially unstructured data, which will still dominate the entire data storage system. Although the proportion of structured data is expected to gradually increase, unstructured data is still the core component of big data applications, and it is widely present in the storage and processing scenarios of various data types such as text, images, audio, and video.

[0003] In the traditional data storage model, relational databases (such as MySQL, Oracle, and SQL Server) are widely used in various business systems due to their mature structured data models, transaction support, and data consistency advantages. However, with the rapid development of the Internet and the increasing complexity of data types, relational databases have gradually shown their shortcomings of limited scalability and prominent performance bottlenecks when facing massive data and high-concurrency requests. Especially when processing unstructured data, the limitations of traditional relational databases are more obvious. In order to solve this challenge, NoSQL databases came into being and became a new generation of data storage solutions. Unlike traditional relational databases, NoSQL databases do not require fixed data models, have stronger scalability and flexibility, and can handle a variety of data types more efficiently.

[0004] In the NoSQL system, key-value storage is an important storage method and is widely used in large-scale data management scenarios with its simple and efficient architecture. Many modern key-value storage systems (such as LevelDB, RocksDB, and Redis) use Log-Structured Merge-Tree (LSM tree) as the core data structure. LSM tree not only significantly improves write performance by converting random write operations into sequential writes, but also optimizes data reading efficiency with the help of hierarchical storage and memory cache. Therefore, storage systems based on LSM trees perform well under high-throughput and low-latency application requirements, making them an ideal choice for large-scale data processing.

[0005] However, as the scale of data continues to expand, traditional centralized key-value storage systems have gradually exposed performance bottlenecks and limited scalability. In order to meet the needs of large-scale data storage and high concurrent access, distributed key-value storage systems have emerged. The distributed architecture ensures high availability, reliability and consistency of data through data sharding, load balancing and multi-copy mechanisms, and can dynamically expand according to business needs to adapt to the growing data load. At present, distributed key-value storage systems such as BigTable, TiDB, HBase, and FoundationDB have been widely used in the Internet, e-commerce and other fields, and have become the core supporting technology for modern big data storage.

[0006] Among many distributed key-value storage systems, Cassandra, as a typical representative based on the LSM tree architecture, has shown significant advantages in large-scale data processing scenarios with its excellent scalability and high availability. Figure 1 Cassandra uses a consistent hashing algorithm for data partitioning and combines it with a random partitioner to optimize load balancing. Through a multi-copy storage mechanism, it can ensure redundant storage of data between different nodes, thereby enhancing the system's high availability and fault tolerance. In terms of data organization, Cassandra uses a row storage model and manages data through column families, making it particularly good in time-series data storage and frequent update application scenarios.

[0007] Although Cassandra has made remarkable achievements in the field of distributed storage, it still has performance bottlenecks in range queries. Since Cassandra uses a random partitioner to ensure balanced distribution of data, although this design optimizes the load balancing of the system, it also leads to inefficient range queries. The main reason is that the data is not organized in the order of the key when it is stored. Therefore, when performing range queries, it is often necessary to read a large amount of irrelevant data across multiple partitions, which reduces the query efficiency. Especially on large-scale data sets, the performance problem of range queries becomes more and more prominent, becoming a major challenge for Cassandra in efficient query scenarios.

[0008] Therefore, under the premise of ensuring the high availability and scalability of Cassandra, how to optimize the partitioner design and query strategy to improve its range query efficiency has become an important direction of current research. Summary of the invention

[0009] The purpose of this application is to propose a data range query method and device for the Cassandra distributed key-value storage system in response to the above-mentioned technical problems.

[0010] In a first aspect, the present invention provides a data range query method for a Cassandra distributed key-value storage system, comprising the following steps:

[0011] S1, obtain the range query request and parse it to obtain the corresponding entire query range and the total data volume, determine the length of the key prefix and the initial number of groups, use the entire query range as the query range of the current round, use the total data volume as the data volume required for the current round, and use the initial number of groups as the number of groups for the current round;

[0012] S2, taking the left boundary of the query range of the current round as the starting point of grouping, grouping the query range of the current round according to the number of groups of the current round and the key prefix, and obtaining the grouping result of the current round, the number of all grouping ranges in the grouping result of the current round is the number of groups of the current round, and the data in the same grouping range in the grouping result of the current round has the same key prefix;

[0013] S3, perform range query according to the grouping result of the current round, obtain the query result of the current round and store it in the result set;

[0014] S4, determine whether the amount of data in the result set reaches the amount of data required for the current round and / or whether the last grouping range in the grouping results of the current round reaches the right boundary of the query range of the current round. If so, take all the query results in the result set as the final query results; otherwise, determine whether the amount of data in the query results of the current round is greater than or equal to 1. If so, calculate the updated number of groups based on the amount of data in the query results of the current round, the amount of data required for the current round and the number of groups in the current round; otherwise, increase the number of groups based on the number of groups in the current round to obtain the updated number of groups; determine the number of groups for the next round based on the updated number of groups;

[0015] S5, based on the query range of the current round, remove all grouping ranges in the grouping results of the current round, and update the query range of the next round; based on the amount of data required for the current round, remove the amount of data in the query results of the current round, and update the amount of data required for the next round. The number of groups in the next round is used as the number of groups in the current round, the query range of the next round is used as the query range of the current round, and the amount of data required for the next round is used as the amount of data required for the current round. Repeat steps S2-S5 until the final query result is obtained.

[0016] As a preference, it also includes:

[0017] Determine whether each grouping range in the current round of grouping results exceeds the right boundary of the current round of query range. If so, modify the right boundary of the first grouping range that exceeds the right boundary of the current round of query range to the right boundary of the current round of query range, and discard the subsequent grouping ranges of the first grouping range that exceeds the right boundary of the current round of query range; otherwise, perform range query according to the current round of grouping results.

[0018] Preferably, the updated number of groups is calculated according to the amount of data in the query result of the current round, the amount of data required for the current round, and the number of groups of the current round, specifically including:

[0019] The updated number of groups is calculated using the following formula:

[0020] ;

[0021] in, Indicates the number of groups in the current round, Indicates the amount of data in the query results of the current round, Indicates the amount of data required for the current round, Indicates the number of groups after update. Indicates rounding up.

[0022] Preferably, the number of groups in the current round is increased to obtain an updated number of groups, which specifically includes:

[0023] The updated number of groups is calculated using the following formula:

[0024] ;

[0025] in, Indicates the number of groups in the current round, Indicates the increase, , Indicates the number of groups after update.

[0026] Preferably, determining the number of groups for the next round according to the updated number of groups specifically includes:

[0027] Determine the minimum and maximum number of groups;

[0028] The updated number of groups is compared with the minimum number of groups and the maximum number of groups. If the updated number of groups is greater than or equal to the minimum number of groups and less than or equal to the maximum number of groups, the updated number of groups is used as the number of groups for the next round; if the updated number of groups is less than the minimum number of groups, the minimum number of groups is used as the number of groups for the next round; if the updated number of groups is greater than the maximum number of groups, the maximum number of groups is used as the number of groups for the next round.

[0029] Preferably, each storage node in the Cassandra distributed key-value storage system is provided with a group partitioner; the data storage process is as follows:

[0030] The data to be written is grouped by the grouping partitioner according to the fixed key prefix to obtain the grouping key;

[0031] Perform hash calculation on the grouping key to obtain the hash value corresponding to the grouping key; determine the target storage node according to the hash value corresponding to the grouping key;

[0032] Write the data to be written into the LSM tree in the target storage node.

[0033] In a second aspect, the present invention provides a data range query device for a Cassandra distributed key-value storage system, comprising:

[0034] The data parsing module is configured to obtain a range query request and parse the corresponding entire query range and the entire data volume, determine the length of the key prefix and the initial number of groups, use the entire query range as the query range of the current round, use the entire data volume as the data volume required for the current round, and use the initial number of groups as the number of groups for the current round;

[0035] The grouping module is configured to use the left boundary of the query range of the current round as the grouping starting point, group the query range of the current round according to the number of groups of the current round and the key prefix, and obtain the grouping result of the current round, the number of all grouping ranges in the grouping result of the current round is the number of groups of the current round, and the data in the same grouping range in the grouping result of the current round has the same key prefix;

[0036] A query module is configured to perform a range query based on the grouping result of the current round, obtain the query result of the current round and store it in a result set;

[0037] The judgment module is configured to judge whether the amount of data in the result set reaches the amount of data required for the current round and / or whether the last grouping range in the grouping result of the current round reaches the right boundary of the query range of the current round. If so, all query results in the result set are used as the final query results; otherwise, it is judged whether the amount of data in the query result of the current round is greater than or equal to 1. If so, the updated number of groups is calculated based on the amount of data in the query result of the current round, the amount of data required for the current round and the number of groups of the current round; otherwise, the number of groups is increased based on the number of groups in the current round to obtain the updated number of groups; the number of groups for the next round is determined based on the updated number of groups;

[0038] The repetition module is configured to remove all grouping ranges in the grouping results of the current round based on the query range of the current round, and update the query range of the next round; remove the data amount in the query results of the current round based on the data amount required for the current round, and update the data amount required for the next round; based on the number of groups in the next round as the number of groups in the current round, the query range of the next round is used as the query range of the current round, and the data amount required for the next round is used as the data amount required for the current round, and the grouping module to the repetition module are repeatedly executed until the final query result is obtained.

[0039] In a third aspect, the present invention provides an electronic device comprising one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.

[0040] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.

[0041] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] (1) The data range query method for the Cassandra distributed key-value storage system proposed in the present invention can effectively improve the range query performance of Cassandra. Key-value pairs with the same key prefix are grouped according to the key prefix, making the physical storage of data more compact, thereby reducing the overhead of cross-partition access during range query and improving the efficiency of data retrieval. At the same time, the method can reduce the reading and transmission of irrelevant data, reduce the system I / O burden, and further optimize the query response speed.

[0044] (2) The data range query method for the Cassandra distributed key-value storage system proposed in the present invention adopts a dynamic batch scheduling strategy, which flexibly adjusts the number of groups in the next round according to the query range and the returned results, making the query process more efficient and intelligent, avoiding unnecessary data access and waste of system resources. This optimization not only improves the throughput of the system, but also better adapts to data query requirements of different scales and complexities.

[0045] (3) The data range query method for the Cassandra distributed key-value storage system proposed in the present invention helps to balance the distribution of query requests, relieve the pressure on storage nodes, and improve the stability of Cassandra in a high-concurrency environment. At the same time, the method can be seamlessly integrated into the existing Cassandra system without affecting other storage and query functions, providing a more efficient and flexible solution for large-scale data storage and processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0047] Figure 1 A schematic diagram of the data storage process of the existing Cassandra distributed key-value storage system;

[0048] Figure 2 A flow chart of a data range query method for a Cassandra distributed key-value storage system according to an embodiment of the present application;

[0049] Figure 3 A schematic diagram of a data storage process of a data range query method for a Cassandra distributed key-value storage system according to an embodiment of the present application;

[0050] Figure 4 A schematic diagram of a data range query device for a Cassandra distributed key-value storage system according to an embodiment of the present application;

[0051] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0053] Figure 2 A data range query method for a Cassandra distributed key-value storage system provided by an embodiment of the present application is shown, comprising the following steps:

[0054] S1, obtain the range query request and parse it to obtain the corresponding entire query range and total data volume, determine the length of the key prefix and the initial number of groups, use the entire query range as the query range of the current round, use the total data volume as the data volume required for the current round, and use the initial number of groups as the number of groups for the current round.

[0055] In a specific embodiment, each storage node in the Cassandra distributed key-value storage system is provided with a group partitioner; the data storage process is as follows:

[0056] The data to be written is grouped by the grouping partitioner according to the fixed key prefix to obtain the grouping key;

[0057] Perform hash calculation on the grouping key to obtain the hash value corresponding to the grouping key; determine the target storage node according to the hash value corresponding to the grouping key;

[0058] Write the data to be written into the LSM tree in the target storage node.

[0059] Specifically, refer to Figure 3 The Cassandra distributed key-value storage system in the embodiment of the present application includes a plurality of storage nodes. A group partitioner is introduced on each storage node in the Cassandra distributed key-value storage system to group the data to be written according to the length of a fixed key prefix, so that data with the same key prefix are grouped in the same group, thereby optimizing the locality of data storage.

[0060] The data writing process of the Cassandra distributed key-value storage system in the embodiment of the present application is as follows:

[0061] S11, obtain the user's write request and obtain the primary key of the data to be written, group the data to be written according to the same key prefix to obtain the grouping key, perform hash calculation on the grouping key to obtain the hash value, and enter S12;

[0062] S12, writing the data to be written into the target storage node according to the hash value for partition storage, and organizing and managing the data in groups, and then entering S13;

[0063] S13, write the data to be written into the LSM tree of the target storage node to ensure efficient storage and access performance, and enter S14;

[0064] S14, completing the user's write request and returning confirmation information of successful write to the user.

[0065] The above storage method can realize grouping and partitioning storage of data. After grouping and partitioning storage, data range query can be performed.

[0066] First, obtain the range query request, which can exist in the form of a CQL statement. By parsing the range query request, the entire query range and the total amount of data required for this range query can be obtained. Determine the length of the key prefix, initialize the number of groups, and determine the initial number of groups. In one embodiment, the length of the fixed key prefix is ​​3, that is, the first 3 bits of the data in the query range of all rounds are the key prefix; the initial number of groups is set to 2. Use the entire query range as the query range of the current round, the total amount of data as the amount of data required for the current round, and the initial number of groups as the number of groups for the current round, and enter S2 for grouping.

[0067] S2, taking the left boundary of the query range of the current round as the starting point of grouping, groups the query range of the current round according to the number of groups of the current round and the key prefix, and obtains the grouping result of the current round. The number of all grouping ranges in the grouping result of the current round is the number of groups of the current round, and the data in the same grouping range in the grouping result of the current round has the same key prefix.

[0068] Specifically, in the range query process, the left boundary of the query range of the current round is used as the grouping starting point to group the query range of the current round. During the grouping process, it is necessary to consider that the number of grouping ranges in the grouping result of the current round is the number of groups in the current round, and the same grouping range in the grouping result of the current round has the same key prefix. When the current round is the first round, the entire query range is used as the query range of the first round, the total amount of data is used as the amount of data required for the first round, and the initial number of groups is used as the number of groups in the first round. Assuming that the entire query range is [11000, 15000], the total amount of data is 800, the initial number of groups G=2, and the length of the fixed key prefix is ​​3, the grouping results obtained after grouping are: the first grouping range: 11000-11999 (key prefix is ​​110), the second grouping range: 12000-12999 (key prefix is ​​120). The grouping ranges in the grouping results are sorted from small to large, and are continuously segmented in sequence without intervals in between.

[0069] In a specific embodiment, it also includes:

[0070] Determine whether each grouping range in the current round of grouping results exceeds the right boundary of the current round of query range. If so, modify the right boundary of the first grouping range that exceeds the right boundary of the current round of query range to the right boundary of the current round of query range, and discard the subsequent grouping ranges of the first grouping range that exceeds the right boundary of the current round of query range; otherwise, perform range query according to the current round of grouping results.

[0071] Specifically, during the grouping process, it is necessary to determine whether the grouping range will exceed the right boundary of the query range of the current round. If it exceeds, the right boundary of the first grouping range that exceeds the right boundary of the query range of the current round needs to be modified to the right boundary of the query range of the current round, and the subsequent grouping ranges that exceed the right boundary of the query range of the current round are discarded. In other words, when it is detected that the right boundary of one of the grouping ranges exceeds the right boundary of the query range of the current round, the right boundary of the grouping range is modified to the right boundary of the query range of the current round, and the subsequent grouping is stopped.

[0072] S3, perform range query according to the grouping result of the current round, obtain the query result of the current round and store it in the result set.

[0073] Specifically, a range query is performed on each group range in the group result of the current round in parallel. When all group ranges have completed the range query, the query result of the current round can be obtained, and the query result of the current round is stored in the result set. For example, the amount of data queried by the first group range is 300, and the amount of data queried by the second group range is 200, and the total amount of data is 500. Stored in the result set.

[0074] S4, determine whether the amount of data in the result set reaches the amount of data required for the current round and / or whether the last grouping range in the grouping results of the current round reaches the right boundary of the query range of the current round. If so, take all the query results in the result set as the final query results; otherwise, determine whether the amount of data in the query results of the current round is greater than or equal to 1. If so, calculate the updated number of groups based on the amount of data in the query results of the current round, the amount of data required for the current round and the number of groups in the current round; otherwise, increase the number of groups based on the number of groups in the current round to obtain the updated number of groups; determine the number of groups for the next round based on the updated number of groups.

[0075] In a specific embodiment, the updated number of groups is calculated based on the amount of data in the query results of the current round, the amount of data required for the current round, and the number of groups of the current round, specifically including:

[0076] The updated number of groups is calculated using the following formula:

[0077] ;

[0078] in, Indicates the number of groups in the current round, Indicates the amount of data in the query results of the current round, Indicates the amount of data required for the current round, Indicates the number of groups after update. Indicates rounding up.

[0079] In a specific embodiment, the number of groups in the current round is increased to obtain the updated number of groups, which specifically includes:

[0080] The updated number of groups is calculated using the following formula:

[0081] ;

[0082] in, Indicates the number of groups in the current round, Indicates the increase, , Indicates the number of groups after update.

[0083] In a specific embodiment, determining the number of groups for the next round according to the updated number of groups specifically includes:

[0084] Determine the minimum and maximum number of groups;

[0085] The updated number of groups is compared with the minimum number of groups and the maximum number of groups. If the updated number of groups is greater than or equal to the minimum number of groups and less than or equal to the maximum number of groups, the updated number of groups is used as the number of groups for the next round; if the updated number of groups is less than the minimum number of groups, the minimum number of groups is used as the number of groups for the next round; if the updated number of groups is greater than the maximum number of groups, the maximum number of groups is used as the number of groups for the next round.

[0086] Specifically, determine whether the amount of data in the result set reaches the amount of data required for the current round, and determine whether the last grouping range in the grouping results of the current round reaches the right boundary of the query range of the current round. If at least one of the above two situations is met, all query results in the result set are used as the final query results. If the amount of data in the result set exceeds the amount of data required for the current round, discard the excess data at the end of the result set. If neither of the above two situations is met, dynamically adjust the number of groups for the next round. Specifically, if the query results of the current round have queried data, that is, whether the amount of data in the query results of the current round is greater than or equal to 1, then according to the formula , calculate the updated number of groups, this formula compares the amount of data in the query results of the current round The amount of data required for the current round , dynamically calculate the updated number of groups , thereby adjusting the query range; if the query results of the current round do not find any data, that is, the amount of data in the query results of the current round is 0, then according to the formula , calculate the updated number of groups. If the amount of data in the query results of the current round is insufficient, increase the number of groups to expand the query range; if the amount of data in the query results of the current round is too much, reduce the number of groups to improve the query accuracy. The embodiments of the present application also set reasonable upper and lower limits of the score groups to prevent extreme situations. In one embodiment, the minimum number of groups is 2 and the maximum number of groups is 20. If the updated number of groups is between the minimum number of groups and the maximum number of groups, the updated number of groups is used as the number of groups for the next round; if the updated number of groups is less than the minimum number of groups, the minimum number of groups is used as the number of groups for the next round; if the updated number of groups is greater than the maximum number of groups, the maximum number of groups is used as the number of groups for the next round. In this way, the number of groups for the next round can be dynamically adjusted to ensure that data access is more accurate and efficient.

[0087] S5, based on the query range of the current round, remove all grouping ranges in the grouping results of the current round, and update the query range of the next round; based on the amount of data required for the current round, remove the amount of data in the query results of the current round, and update the amount of data required for the next round. The number of groups in the next round is used as the number of groups in the current round, the query range of the next round is used as the query range of the current round, and the amount of data required for the next round is used as the amount of data required for the current round. Repeat steps S2-S5 until the final query result is obtained.

[0088] Specifically, after obtaining the number of groups for the next round, the query range needs to be updated, that is, all grouping ranges in the grouping results of the current round are removed on the basis of the query range of the current round to obtain the query range of the next round. For example, if the query range of the first round is [11000, 15000], the first grouping range is: 11000-11999, and the second grouping range is: 12000-12999; then the query range of the second round is [13000, 15000]. It is also necessary to update the data volume, that is, on the basis of the data volume required for the current round, the data volume in the query results of the current round is removed to obtain the data volume required for the next round. For example, if the data volume required for the first round is 800 and the data volume in the query results of the first round is 500, then the data volume required for the second round is 300. The next round is taken as the current round and steps S2-S5 are repeated until the data volume required for the current round is met or the right boundary of the query range of the current round is reached. The final query result is returned to the client.

[0089] In order to verify the effect of the group partitioner designed in the embodiment of the present application, the embodiment of the present application deployed a Cassandra cluster on two 2-core 4G lightweight cloud hosts, and used YCSB to import 10G data for testing. The experiment set the scan length to 100 and the number of operations to 1000, and compared the range query performance of the traditional random partitioner and the group partitioner proposed in the embodiment of the present application. The results are shown in Table 1.

[0090] Table 1 Benchmark table of two partitioners:

[0091]

[0092] The results show that the running time of the group partitioner is 899 seconds, which is significantly better than the 1515 seconds of the random partitioner; the throughput is increased from 0.6597 to 1.1121 ops / sec (number of operations per second); the average latency is reduced from 1.51 seconds to 0.89 seconds. This shows that the group partitioner can effectively improve the performance of range queries. The reason is that the random partitioner will randomly scatter the data, and the data cannot be read in order, and the query cost is high; while the group partitioner ensures the order of data within the group and ensures good load balancing, thereby improving the performance of range queries.

[0093] The larger the group, the more data in the group, and the better the range query performance. However, this will affect the load balancing, causing some nodes to store more data and some nodes to store less data. Therefore, an appropriate group size should be designed.

[0094] Further references Figure 4 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a data range query device for a Cassandra distributed key-value storage system. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0095] The embodiment of the present application provides a data range query device for a Cassandra distributed key-value storage system, which is characterized by comprising:

[0096] The data parsing module 1 is configured to obtain a range query request and parse the corresponding entire query range and the entire data volume, determine the length of the key prefix and the initial number of groups, use the entire query range as the query range of the current round, use the entire data volume as the data volume required for the current round, and use the initial number of groups as the number of groups for the current round;

[0097] The grouping module 2 is configured to use the left boundary of the query range of the current round as the grouping starting point, group the query range of the current round according to the number of groups of the current round and the key prefix, and obtain the grouping result of the current round, the number of all grouping ranges in the grouping result of the current round is the number of groups of the current round, and the data in the same grouping range in the grouping result of the current round has the same key prefix;

[0098] Query module 3, configured to perform range query according to the grouping result of the current round, obtain the query result of the current round and store it in the result set;

[0099] The judgment module 4 is configured to judge whether the amount of data in the result set reaches the amount of data required for the current round and / or whether the last grouping range in the grouping result of the current round reaches the right boundary of the query range of the current round. If so, all the query results in the result set are used as the final query results; otherwise, it is judged whether the amount of data in the query result of the current round is greater than or equal to 1. If so, the updated number of groups is calculated based on the amount of data in the query result of the current round, the amount of data required for the current round and the number of groups of the current round; otherwise, the number of groups in the current round is increased to obtain the updated number of groups; the number of groups in the next round is determined according to the updated number of groups;

[0100] The repetition module 5 is configured to remove all grouping ranges in the grouping results of the current round based on the query range of the current round, and update the query range of the next round; remove the data amount in the query results of the current round based on the data amount required for the current round, and update the data amount required for the next round; based on the number of groups in the next round as the number of groups in the current round, the query range of the next round is used as the query range of the current round, and the data amount required for the next round is used as the data amount required for the current round, and repeatedly execute the grouping module 2 to the repetition module 5 until the final query result is obtained.

[0101] Figure 5 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Figure 5 As shown, the electronic device of this embodiment includes: a processor 501 and a memory 502; wherein the memory 502 is used to store computer-executable instructions; the processor 501 is used to execute the computer-executable instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant description in the above method embodiment.

[0102] Optionally, the memory 502 may be independent or integrated with the processor 501 .

[0103] When the memory 502 is independently provided, the electronic device further includes a bus 503 for connecting the memory 502 and the processor 501 .

[0104] The embodiment of the present invention further provides a computer storage medium, in which computer execution instructions are stored. When the processor 501 executes the computer execution instructions, the above method is implemented.

[0105] The embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 501, the above method is implemented.

[0106] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0107] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to implement the solution of this embodiment.

[0108] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each module may exist physically separately, or two or more modules may be integrated into one unit. The unit formed by the above modules may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0109] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 501 to perform some steps of the methods of various embodiments of the present application.

[0110] It should be understood that the processor 501 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or the processor 501 may be any conventional processor 501, etc. The steps of the method disclosed in the invention may be directly embodied in the execution of the hardware processor 501, or may be executed by a combination of hardware and software modules in the processor 501.

[0111] The memory 502 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disk.

[0112] The bus 503 may be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 503 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 503 in the drawings of the present application is not limited to only one bus 503 or one type of bus 503.

[0113] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special purpose computer.

[0114] An exemplary storage medium is coupled to the processor 501, so that the processor 501 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 501. The processor 501 and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor 501 and the storage medium can also exist as discrete components in an electronic device or a main control device.

[0115] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data range query method for a Cassandra distributed key-value storage system, characterized in that: The following steps are involved: S1, obtain the range query request and parse it to obtain the corresponding entire query range and the total data volume, determine the length of the key prefix and the initial number of groups, use the entire query range as the query range of the current round, use the total data volume as the data volume required for the current round, and use the initial number of groups as the number of groups for the current round; S2, taking the left boundary of the query range of the current round as the grouping starting point, grouping the query range of the current round according to the number of groups of the current round and the key prefix, and obtaining the grouping result of the current round, the number of all grouping ranges in the grouping result of the current round is the number of groups of the current round, and the data in the same grouping range in the grouping result of the current round has the same key prefix; S3, performing a range query according to the grouping result of the current round, obtaining the query result of the current round and storing it in a result set; S4, judging whether the amount of data in the result set reaches the amount of data required for the current round and / or whether the last grouping range in the grouping results of the current round reaches the right boundary of the query range of the current round. If so, taking all the query results in the result set as the final query results; Otherwise, determine whether the data volume of the query result of the current round is greater than or equal to 1. If so, calculate the updated number of groups according to the data volume in the query result of the current round, the data volume required for the current round, and the number of groups in the current round. Otherwise, increase the number of groups based on the number of groups in the current round to obtain the updated number of groups; determine the number of groups in the next round according to the updated number of groups; S5, based on the query range of the current round, remove all grouping ranges in the grouping results of the current round, and update the query range of the next round; based on the amount of data required for the current round, remove the amount of data in the query results of the current round, and update the amount of data required for the next round, based on the number of groups in the next round as the number of groups in the current round, use the query range of the next round as the query range of the current round, use the amount of data required for the next round as the amount of data required for the current round, repeat steps S2-S5 until the final query result is obtained.

2. The data range query method for the Cassandra distributed key-value storage system according to claim 1 is characterized in that: Also includes: Determine whether each grouping range in the grouping results of the current round exceeds the right boundary of the query range of the current round. If so, modify the right boundary of the first grouping range that exceeds the right boundary of the query range of the current round to the right boundary of the query range of the current round, and discard the subsequent grouping ranges of the first grouping range that exceeds the right boundary of the query range of the current round; otherwise, perform a range query according to the grouping results of the current round.

3. The data range query method for the Cassandra distributed key-value storage system according to claim 1 is characterized in that: The updated number of groups is calculated based on the amount of data in the query results of the current round, the amount of data required for the current round, and the number of groups of the current round, specifically including: The updated number of groups is calculated using the following formula: ; in, Indicates the number of groups in the current round, Indicates the amount of data in the query results of the current round, Indicates the amount of data required for the current round, Indicates the number of groups after update. Indicates rounding up.

4. The data range query method for the Cassandra distributed key-value storage system according to claim 1, characterized in that: The updated number of groups is obtained by increasing the number of groups in the current round, including: The updated number of groups is calculated using the following formula: ; in, Indicates the number of groups in the current round, Indicates the increase, , Indicates the number of groups after update.

5. The data range query method for the Cassandra distributed key-value storage system according to claim 1, characterized in that: Determining the number of groups for the next round according to the updated number of groups specifically includes: Determine the minimum and maximum number of groups; The updated number of groups is compared with the minimum number of groups and the maximum number of groups. If the updated number of groups is greater than or equal to the minimum number of groups and less than or equal to the maximum number of groups, the updated number of groups is used as the number of groups for the next round; if the updated number of groups is less than the minimum number of groups, the minimum number of groups is used as the number of groups for the next round; if the updated number of groups is greater than the maximum number of groups, the maximum number of groups is used as the number of groups for the next round.

6. The data range query method for the Cassandra distributed key-value storage system according to claim 1, characterized in that: Each storage node in the Cassandra distributed key-value storage system is provided with a group partitioner; the data storage process is as follows: The data to be written is grouped according to a fixed key prefix by the group partitioner to obtain a group key; Performing hash calculation on the grouping key to obtain a hash value corresponding to the grouping key; Determine the target storage node according to the hash value corresponding to the grouping key; The data to be written is written into the LSM tree in the target storage node.

7. A data range query device for a Cassandra distributed key-value storage system, characterized in that: include: The data parsing module is configured to obtain a range query request and parse the corresponding entire query range and the entire data volume, determine the length of the key prefix and the initial number of groups, use the entire query range as the query range of the current round, use the entire data volume as the data volume required for the current round, and use the initial number of groups as the number of groups for the current round; a grouping module configured to use the left boundary of the query range of the current round as a grouping starting point, group the query range of the current round according to the number of groups of the current round and the key prefix, and obtain a grouping result of the current round, wherein the number of all grouping ranges in the grouping result of the current round is the number of groups of the current round, and data in the same grouping range in the grouping result of the current round has the same key prefix; A query module, configured to perform a range query according to the grouping result of the current round, obtain the query result of the current round and store it in a result set; A judgment module is configured to judge whether the amount of data in the result set reaches the amount of data required for the current round and / or whether the last grouping range in the grouping results of the current round reaches the right boundary of the query range of the current round, and if so, take all the query results in the result set as the final query results; Otherwise, determine whether the data volume of the query result of the current round is greater than or equal to 1. If so, calculate the updated number of groups according to the data volume in the query result of the current round, the data volume required for the current round, and the number of groups in the current round. Otherwise, increase the number of groups based on the number of groups in the current round to obtain the updated number of groups; determine the number of groups in the next round according to the updated number of groups; The repetition module is configured to remove all grouping ranges in the grouping results of the current round based on the query range of the current round, and update the query range of the next round; remove the data amount in the query results of the current round based on the data amount required for the current round, and update the data amount required for the next round; based on the number of groups in the next round as the number of groups in the current round, the query range of the next round is used as the query range of the current round, the data amount required for the next round is used as the data amount required for the current round, and repeatedly execute the grouping module to the repetition module until the final query result is obtained.

8. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Distributed expandable quadtree indexing mechanism oriented to Cassandra and query method based on mechanism

    CN105630968A

  • Method, device, medium for retrieving common prefixes of keys in key-value database

    CN109388641A