Data processing and query method of a key-value storage system and related device
Patent Information
- Application Number
- CN202510288644.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2026-09-11
AI Technical Summary
若上述事务的执行出现时序错误,一个事务提交者先将底层存储中的键值对更新为版本B,另一个事务提交者再将版本B的键值对替换为版本A,从而造成底层存储的数据错误
[0045] Seventhly, this application also provides a computer program product, including computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the data processing method of the key-value storage system provided in the first aspect or the query method of the key-value storage system provided in any possible implementation of the second aspect. Any of the service layers, computing devices, computing device clusters, computer storage media, or computer program products provided above are used to execute the methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects of the corresponding solutions in the corresponding methods provided above, and will not be repeated here.
Smart Images

Figure CN122734136A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular to a data processing and querying method and related apparatus for a key-value storage system. Background Technology
[0002] In Kubernetes (k8s) clusters, a distributed key-value (KV) storage system is typically used to achieve service discovery, state storage, and cluster configuration. The key-value storage system consists of multiple compute nodes and underlying storage. The compute nodes handle data processing requests from tenants, while the underlying storage stores the key-value pairs belonging to that tenant.
[0003] Existing key-value stores, when processing tenant data processing requests, first rely on an election mechanism to elect a transaction committer from multiple compute nodes. This transaction committer then creates a transaction based on the data processing request. For example, if the data processing request is an update request for a key-value pair, the transaction may include the new version of the key-value pair to be updated, along with a series of comparison and swap operations related to that key-value pair; the underlying storage contains the old version of the key-value pair. Then, the transaction committer can submit the transaction to other compute nodes. Once more than half of the compute nodes agree to the transaction, the transaction committer finally performs a series of comparison and swap operations on the new version of the key-value pair, updating the old version of the key-value pair in the underlying storage based on the keys in the new version.
[0004] Because key-value stores require only one transaction committer to participate in transaction execution at any given time, adding or deleting a compute node may trigger a re-election process, resulting in two transaction committers simultaneously. In this scenario, if a tenant sends two data update requests for the same key-value pair with different versions (e.g., version A and version B, where version A < version B), the two transaction committers might execute the transactions corresponding to the two update requests separately. If a timing error occurs during the execution of these transactions, one transaction committer might update the key-value pair in the underlying storage to version B, while the other transaction committer might replace the key-value pair of version B with version A, causing data errors in the underlying storage. Therefore, existing key-value stores do not support improving scalability and utilization by adding compute nodes. Summary of the Invention
[0005] This application provides a data processing and querying method and related apparatus for a key-value storage system, which improves the scalability and utilization of the key-value storage system.
[0006] Firstly, this application provides a data processing method for a key-value storage system. The method is applied to a service layer deployed on at least one computing node in the key-value storage system, which also includes a storage cluster. At least one computing node and the storage cluster are communicatively connected.
[0007] The method may include receiving a data processing request sent by a first tenant. The service layer can then create a corresponding transaction based on the data processing request and generate a pushdown operator related to the data processing request. Each pushdown operator includes at least one sub-request, and each sub-request requests the storage cluster to perform a corresponding operation. The service layer can then send the pushdown operator to the storage cluster and submit the transaction to the storage cluster. This transaction instructs the storage cluster to process at least one sub-request in the pushdown operator. Finally, the service layer can receive the processing result of the transaction returned by the storage cluster and send the processing result back to the first tenant.
[0008] In the above scheme, the computing cluster undertaking computational tasks is separated from the storage cluster undertaking storage tasks (storage-compute separation), and the service layer is deployed on at least one computing node. The service layer constructs the data processing requests of the first tenant into corresponding transactions and pushdown operators, and submits these transactions and pushdown operators to the storage cluster for execution. This eliminates the need for the transaction execution of the key-value storage system to rely on the transaction committer obtained through an election mechanism. This method supports improving the scalability and utilization of the key-value storage system by adding computing nodes (horizontal scaling).
[0009] In some embodiments, the data processing request includes a tenant identifier of a first tenant. The method includes: the service layer determining the traffic quota of the first tenant based on the tenant identifier of the first tenant; if the traffic quota of the first tenant exceeds a threshold, the service layer may return a processing result indicating that the data processing request has failed to the first tenant; if the traffic quota of the first tenant does not exceed the threshold, the service layer may create a transaction corresponding to the data processing request based on the key or key-value pair entered by the first tenant.
[0010] In the above scheme, the service layer can manage the traffic quota of each tenant based on the tenant identifier corresponding to each tenant, support tenant-level traffic control, realize multi-tenant sharing of the key-value storage system, and improve the utilization of the key-value storage system.
[0011] In some embodiments, when the data processing request includes a first key-value pair, which is a key-value pair input by a first tenant, and the storage cluster pre-stores a second key-value pair including the first tenant's lease information, which is an older version of the key-value pair associated with the first key-value pair, the method further includes: the service layer can create a write transaction corresponding to the data processing request based on the first key-value pair, and send the key in the first key-value pair to the storage cluster, so that the storage cluster can return the first tenant's lease information based on the key in the first key-value pair. Then, the service layer can generate a pushdown operator related to the write transaction based on the first tenant's lease information. Finally, the service layer can send the pushdown operator related to the write transaction to the storage cluster, so that the storage cluster can use the pushdown operator related to the write transaction to execute at least one sub-request in the pushdown operator to update the second key-value pair to the first key-value pair.
[0012] In the above scheme, the service layer can create a corresponding write transaction based on the first key-value pair. This write transaction can be executed independently by the storage cluster. When the write transaction involves a lease, this method constructs the write transaction based on the lease information of each tenant, generates pushdown operators related to the write transaction, and distributes the pushdown operators and the write transaction to the storage system. When the write transaction is executed by the storage system, it can use the pushdown operators to perform atomic operations on the first key-value pair, thereby updating the first key-value pair. Compared to key-value storage systems based on election mechanisms, which require the transaction committer to synchronize the changed data to other compute nodes before executing the write transaction, this method ensures the separation of storage and computation in the key-value storage system, achieves horizontal scaling of the key-value storage system, and improves the utilization rate of the key-value storage system.
[0013] In some embodiments, the pushdown operator related to write transactions includes: a first sub-request requesting the storage cluster to modify the revision identifier field and version identifier field recorded in the first key-value pair; a second sub-request requesting the storage cluster to replace the lease index related to the first key-value pair, wherein the lease index is determined based on the key in the first key-value pair and the lease identifier of the first tenant; a third sub-request requesting the storage cluster to store the first key-value pair and generate a multi-version record of the first key-value pair, wherein the key of the first key-value pair is determined based on the version identifier of the write transaction; a fourth sub-request requesting the storage cluster to generate monitoring information related to the first key-value pair and generate a version index of the monitoring information, wherein the key of the monitoring information is determined based on the version identifier of the write transaction; and a fifth sub-request requesting the storage cluster to compress the second key-value pair and generate a compressed index of the second key-value pair, wherein the key of the compressed index is determined based on the version identifier of the write transaction.
[0014] In the above scheme, by generating pushdown operators related to write requests, the storage cluster can support changes to the revision identifier, version identifier, and lease index of the first key-value pair. This generates multi-version records and monitoring information for the first key-value pair, updates the second key-value pair to the first key-value pair, and achieves atomic modification of data in the storage cluster. Simultaneously, by determining the version of the key based on the version identifier of the write transaction, monotonically increasing the key-value pair version identifier is supported, enabling concurrent execution of multi-version write requests.
[0015] In some embodiments, after sending the pushdown operator to the storage cluster and committing the transaction to the storage cluster, the method further includes: obtaining monitoring information of the first key-value pair from the storage cluster; and sending the monitoring information of the first key-value pair to a first compute node, so that the first compute node updates the index file of the first key-value pair according to the monitoring information of the first key-value pair. The first compute node is at least one compute node that has subscribed to the monitoring information of the first key-value pair.
[0016] In the above scheme, the storage cluster can atomically write the first key-value pair and notification information to multiple storage nodes in the storage cluster, and generate monitoring information related to the first key-value pair, so that other computing nodes can perform data synchronization based on the monitoring information. This method ensures the separation of storage and computation in the key-value storage system and the consistency of data writing in the storage cluster, realizes the horizontal scaling of the computing cluster, and improves the utilization rate of the key-value storage system.
[0017] In some embodiments, the data processing request includes a key input by a first tenant. The aforementioned storage cluster pre-stores a third key-value pair including the first tenant's lease information. This third key-value pair is a specific range of key-value pairs related to the data processing request. The method further includes: creating a deletion transaction corresponding to the data processing request based on the key input by the first tenant, and sending the key input by the first tenant to the storage cluster, so that the storage cluster returns the first tenant's lease information based on the key input by the first tenant. Then, the service layer can generate a pushdown operator related to the deletion transaction based on the first tenant's lease information. Finally, the service layer can send the pushdown operator related to the deletion transaction to the storage cluster, so that the storage cluster can use the pushdown operator related to the deletion transaction to execute at least one sub-request in the pushdown operator to delete the third key-value pair.
[0018] In the above scheme, the service layer can create a corresponding deletion transaction based on the key input by the first tenant. This deletion transaction can be executed independently by the storage cluster. When the deletion transaction involves a lease, this method constructs the deletion transaction based on the lease information of each tenant, generates pushdown operators related to the deletion request, and sends the pushdown operators and deletion transaction to the storage system. When the deletion transaction is executed by the storage system, the pushdown operators can be used to perform atomic operations on the third key-value pair, thereby deleting the third key-value pair. Compared to key-value storage systems based on election mechanisms, which require the transaction committer to synchronize the deletion transaction to other computing nodes before executing it, this method ensures the separation of storage and computation in the key-value storage system, achieves horizontal scaling of the key-value storage system, and improves the utilization rate of the key-value storage system.
[0019] In some embodiments, the pushdown operator related to the deletion transaction includes: a sixth sub-request requesting the storage cluster to add the deletion record to the third key-value pair, wherein the key of the third key-value pair is determined based on the version identifier of the deletion transaction; a seventh sub-request requesting the storage cluster to replace the lease index related to the third key-value pair, wherein the lease index is determined based on the key in the third key-value pair and the lease identifier of the first tenant; an eighth sub-request requesting the storage cluster to compress the third key-value pair and generate a compressed index of the third key-value pair, wherein the key of the compressed index is determined based on the version identifier of the deletion transaction; and a ninth sub-request requesting the storage cluster to generate monitoring information related to the third key-value pair and generate a version index of the monitoring information, wherein the key of the monitoring information is determined based on the version identifier of the deletion transaction.
[0020] In the above scheme, by generating pushdown operators related to deletion requests, the storage cluster can add deletion identifiers for third key-value pairs. Changes to the lease index generate monitoring information and compressed indexes for the third key-value pairs, completing the deletion of the third key-value pairs and achieving atomic modification of data in the storage cluster. Simultaneously, the version of the key-value pair is determined based on the identifier of the deletion transaction, supporting monotonically increasing key-value pair version identifiers, enabling concurrent execution of multiple versions of deletion requests and improving the utilization rate of the key-value storage system.
[0021] In some embodiments, after the storage cluster executes a sub-request related to a data processing request using a push operator, the method further includes: obtaining monitoring information of a third key-value pair from the storage cluster; sending the monitoring information of the third key-value pair to a second compute node so that the second compute node updates the index file of the third key-value pair based on the monitoring information of the third key-value pair; wherein the second compute node is at least one compute node that has subscribed to the monitoring information of the third key-value pair.
[0022] In the above scheme, the storage cluster can atomically write the third key-value pair and notification information to multiple storage nodes in the storage cluster, and generate monitoring information related to the third key-value pair, so that other computing nodes can perform data synchronization based on the monitoring information. This method ensures the separation of storage and computation in the key-value storage system and the consistency of data writing in the storage cluster, realizes the horizontal scaling of the computing cluster, and improves the utilization rate of the key-value storage system.
[0023] In some embodiments, the method can also handle data processing requests from multiple tenants. First, the service layer can receive a data processing request sent by a second tenant. The second tenant is any one of multiple tenants other than the first tenant, and the data processing request includes the tenant identifier of the second tenant. Then, the service layer can create a transaction corresponding to the data processing request based on the tenant identifier of the second tenant, and generate a pushdown operator related to the data processing request. The service layer can then send the pushdown operator to the storage cluster and submit the transaction to the storage cluster. Finally, the service layer can receive the transaction processing result returned by the storage cluster and send the processing result to the second tenant based on the tenant identifier of the second tenant.
[0024] In the above scheme, the service layer can distinguish multiple tenants based on the tenant identifier corresponding to each tenant, enabling the key-value storage system to support tenant-level traffic control, realizing multi-tenant sharing of the key-value storage system, and improving the utilization rate of the key-value storage system.
[0025] Secondly, this application provides a query method for a key-value storage system. This method is applied to a service layer deployed on at least one compute node in the key-value storage system, which also includes a storage cluster. At least one compute node is communicatively connected to the storage cluster. The method includes: the service layer receiving a first query request. The first tenant is one of multiple tenants accessing the service layer; the query request includes a key input by the first tenant, which includes a first version identifier corresponding to that key; the service layer sending the key input by the first tenant to the storage cluster to retrieve a second version identifier from multiple versions of key-value pairs. The second version identifier is the version identifier of the higher version key-value pair among the multiple versions. If the first version identifier is less than the second version identifier, the service layer can retrieve a third version identifier from the storage cluster. The third version identifier is the version identifier of the lower version key-value pair among the multiple versions. If the first version identifier is greater than or equal to the third version identifier, the service layer can create a read transaction corresponding to the query request based on the key input by the first tenant.
[0026] In the above scheme, before constructing a read transaction, the service layer can determine whether the storage cluster contains key-value pairs within a specific range corresponding to the first version identifier based on the first version identifier, and create a read transaction based on the above operation so that the read transaction can be executed from the storage cluster. Compared with the key-value storage system based on the election mechanism, which requires the computing nodes to participate in the version maintenance of key-value pairs and the execution of read transactions, this method ensures the separation of storage and computing in the key-value storage system, realizes the horizontal scaling of the computing cluster, and improves the utilization rate of the key-value storage system.
[0027] In some embodiments, the method further includes: if the first version identifier is equal to the second version identifier, the service layer can create a read transaction corresponding to the query request based on the key entered by the first tenant.
[0028] In some embodiments, the method further includes: the service layer may submit a read transaction to the storage cluster, so that the storage cluster returns the processing result of the read transaction. The processing result of the read transaction includes a fourth key-value pair, which is a key-value pair in the storage cluster that conforms to the first version identifier.
[0029] In some embodiments, the method further includes: if the first version identifier is less than the third version identifier, the service layer may return a processing result indicating that the query request failed to the first tenant.
[0030] Thirdly, embodiments of this application provide a service layer for a key-value storage system. This service layer can be applied to at least one computing node on which the service layer is deployed. The service layer includes a data processing module. The data processing module is used to receive a data processing request sent by a first tenant among multiple tenants, and send the processing result of the data processing request to the first tenant. The data processing module is also used to create a corresponding transaction according to the data processing request and generate a pushdown operator related to the data processing request. The pushdown operator includes at least one sub-request, each of the at least one sub-request being used to request the storage cluster to perform a corresponding operation. The data processing module is used to send the pushdown operator to the storage cluster and submit a transaction to the storage cluster. The transaction is used to instruct the storage cluster to process at least one sub-request using the pushdown operator. After the storage cluster completes the above transaction, the data processing module can also receive the transaction processing result returned by the storage cluster and send the result to the first tenant.
[0031] In some embodiments, the data processing module is further configured to determine the traffic quota of the first tenant based on the tenant identifier of the first tenant. If the traffic quota of the first tenant exceeds the threshold, the data processing module may return a processing result indicating that the data processing request has failed to the first tenant; if the traffic quota of the first tenant does not exceed the threshold, the data processing module may create a transaction corresponding to the data processing request based on the key or key-value pair entered by the first tenant.
[0032] In some embodiments, the service layer further includes a lease management module for obtaining lease information of the first tenant from the storage cluster.
[0033] In some embodiments, when the data processing request is a write request, the data processing request includes a first key-value pair, which is a key-value pair input by a first tenant. The storage cluster pre-stores a second key-value pair including the first tenant's lease information. The second key-value pair is an older version of the key-value pair associated with the first key-value pair. In this case, when the data processing module creates a write transaction, it is also used to send the key in the first key-value pair to the storage cluster through the lease management module, so that the storage cluster returns the first tenant's lease information based on the key in the first key-value pair; and generates a pushdown operator related to the write transaction based on the first tenant's lease information. The pushdown operator related to the write transaction is sent to the storage cluster, so that the storage cluster uses the pushdown operator related to the write transaction to execute at least one sub-request in the pushdown operator to update the second key-value pair to the first key-value pair.
[0034] In some embodiments, the data processing module generates pushdown operators related to write transactions based on the lease information of the first tenant, specifically for: requesting the storage cluster to modify the revision identifier field and version identifier field recorded in the first key-value pair; requesting the storage cluster to replace the lease index related to the first key-value pair, wherein the lease index is determined based on the key in the first key-value pair and the lease identifier of the first tenant; requesting the storage cluster to store the first key-value pair and generate a multi-version record of the first key-value pair, wherein the key of the first key-value pair is determined based on the version identifier of the write transaction; requesting the storage cluster to generate monitoring information for the first key-value pair and generate a version index for the monitoring information, wherein the key of the monitoring information is determined based on the version identifier of the write transaction; and requesting the storage cluster to compress the second key-value pair and generate a compressed index for the second key-value pair, wherein the key of the compressed index is determined based on the version identifier of the write transaction.
[0035] In some embodiments, the data processing request includes a key input by a first tenant, and the storage cluster pre-stores a third key-value pair including the first tenant's lease information. This third key-value pair is a specific range of key-value pairs related to the data deletion request. In this case, when creating a deletion transaction, the data processing module is also used to: send the key input by the first tenant to the storage cluster via the lease management module, so that the storage cluster returns the first tenant's lease information based on the key input by the first tenant. The data processing module can generate a pushdown operator related to the deletion transaction based on the first tenant's lease information. The data processing module can send the pushdown operator related to the deletion transaction to the storage cluster, so that the storage cluster uses the pushdown operator related to the deletion transaction to execute at least one sub-request in the pushdown operator to delete the third key-value pair.
[0036] In some embodiments, the data processing module generates pushdown operators related to the deletion transaction based on the lease information of the first tenant, specifically for: requesting the storage cluster to add a deletion record to the third key-value pair. The key of the third key-value pair is determined based on the version identifier of the deletion transaction. The module requests the storage cluster to replace the lease index related to the third key-value pair. The lease index is determined based on the key in the third key-value pair and the lease identifier of the first tenant. The module requests the storage cluster to compress the third key-value pair and generate a compressed index for the third key-value pair. The key of the compressed index is determined based on the version identifier of the deletion transaction. The module requests the storage cluster to generate monitoring information related to the third key-value pair and generate a version index for the monitoring information. The key of the monitoring information is determined based on the version identifier of the deletion transaction.
[0037] In some embodiments, the service layer further includes a monitoring module. When the content of the first key-value pair changes, the monitoring module obtains monitoring information for the first key-value pair from the storage cluster, and also obtains monitoring information for the third key-value pair from the storage cluster. The monitoring information for the first key-value pair is sent to a first compute node, so that the first compute node updates its index file for the first key-value pair based on the monitoring information. The first compute node is at least one compute node that has subscribed to the monitoring information for the first key-value pair. When the content of the third key-value pair changes, the monitoring module also obtains monitoring information for the third key-value pair from the storage cluster. The monitoring information for the third key-value pair is sent to a second compute node, so that the second compute node updates its index file for the third key-value pair based on the monitoring information. The second compute node is at least one compute node that has subscribed to the monitoring information for the third key-value pair.
[0038] In some embodiments, the service layer can also process data processing requests sent by other tenants. In this case, the data processing module is used to receive a data processing request sent by a second tenant. The second tenant is any one of a plurality of tenants other than the first tenant, and the data processing request includes the tenant identifier of the second tenant. The data processing module is used to create a transaction corresponding to the data processing request based on the tenant identifier of the second tenant, and generate a pushdown operator related to the data processing request. The data processing module is also used to send the pushdown operator to the storage cluster and submit the transaction to the storage cluster. The data processing module is used to receive the processing result of the above transaction returned by the storage cluster, and send the processing result to the second tenant based on the tenant identifier of the second tenant.
[0039] Fourthly, embodiments of this application provide a service layer for a key-value storage system. This service layer can be applied to at least one computing node on which the service layer is deployed. The service layer includes a data processing module, which receives a query request sent by a first tenant. The query request includes a key input by the first tenant, and the key input by the first tenant includes a first version identifier corresponding to the key. The data processing module is also used to send the key input by the first tenant to a storage cluster to obtain a second version identifier from the storage cluster. The storage cluster pre-stores multiple versions of key-value pairs, which are key-value pairs within a specific range related to the query request. The second version identifier is the version identifier of the higher version key-value pair among the multiple versions. If the first version identifier is less than the second version identifier, the data processing module can obtain a third version identifier from the storage cluster. The third version identifier is the version identifier of the lower version key-value pair among the multiple versions. When the first version identifier is greater than or equal to the third version identifier, the data processing module can create a read transaction corresponding to the query request based on the key input by the first tenant.
[0040] In some embodiments, if the first version identifier is equal to the second version identifier, the data processing module can also create a read transaction corresponding to the query request based on the key entered by the first tenant.
[0041] In some embodiments, after a read transaction is created, the data processing module is further configured to submit the read transaction to the storage cluster so that the storage cluster returns the processing result of the read transaction. The processing result of the read transaction includes a fourth key-value pair, which is a key-value pair in the storage cluster that conforms to the first version identifier.
[0042] In some embodiments, if the first version identifier is less than the third version identifier, the data processing module may return a processing result indicating that the query request failed to the first tenant.
[0043] Fifthly, this application also provides a computing device cluster, which includes at least one computing device, each computing device including a processor and a memory, wherein the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the data processing method of the key-value storage system provided in the first aspect or the query method of the key-value storage system provided in any possible implementation of the second aspect.
[0044] Sixthly, this application also provides a computer-readable storage medium including computer program instructions, which, when executed by a computing device cluster, enable the computing device cluster to execute the data processing method of the key-value storage system provided in the first aspect or the query method of the key-value storage system provided in any possible implementation of the second aspect.
[0045] Seventhly, this application also provides a computer program product, including computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the data processing method of the key-value storage system provided in the first aspect or the query method of the key-value storage system provided in any possible implementation of the second aspect. Any of the service layers, computing devices, computing device clusters, computer storage media, or computer program products provided above are used to execute the methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects of the corresponding solutions in the corresponding methods provided above, and will not be repeated here. Attached Figure Description
[0046] Figure 1 This is a schematic diagram of the structure of a known key-value storage system;
[0047] Figure 2 This is a schematic diagram of the structure of a key-value storage system provided in an embodiment of this application;
[0048] Figure 3 This is a diagram illustrating the data transformation of key-value pairs corresponding to a specific tenant;
[0049] Figure 4 This is a flowchart illustrating a data processing method for a key-value storage system proposed in this application;
[0050] Figure 5 This is a flowchart illustrating a query method for a key-value storage system proposed in this application;
[0051] Figure 6 This is a schematic diagram illustrating the interaction scenario of a key-value storage system executing a leaseless Put request and a Watch request.
[0052] Figure 7 This is a flowchart illustrating the process of a key-value store system executing a Put request carrying a lease.
[0053] Figure 8 This is a flowchart illustrating the process of a key-value store system executing a Delete request carrying a lease;
[0054] Figure 9 This is a flowchart illustrating the process of a key-value store system executing a range request;
[0055] Figure 10 This is a schematic diagram of the service layer structure of the key-value storage system proposed in this application;
[0056] Figure 11 This is a schematic diagram of another service layer structure of the key-value storage system proposed in this application;
[0057] Figure 12This is a schematic diagram of a computing device proposed in this application;
[0058] Figure 13 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;
[0059] Figure 14 This application provides an embodiment of a computing device cluster deployment Figure 10 or Figure 11 The diagram shows a service layer. Detailed Implementation
[0060] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0061] To facilitate understanding of the technical solution of this application, the relevant terms used in this document are explained below.
[0062] A key-value (KV) storage system refers to a database system that stores, indexes, deletes, and retrieves data using key-value pairs as the basic unit. A key-value storage system typically includes a client, a compute cluster, and underlying storage. The client receives key-value pair-related data processing requests from tenants. The compute cluster creates corresponding transactions based on these requests and then collaborates with the storage cluster to execute the transactions, thereby achieving key-value pair storage, indexing, deletion, and retrieval. The storage implementation of a key-value storage system consists of two parts: a storage module and an index module. The index module is deployed on top of the compute cluster. The index module records index files, which contain the location information of keys on the storage devices. The data in the index files is an additional structure derived from the main data and does not affect the data content, but it often determines the data access efficiency and therefore plays a crucial role. The underlying storage deploys the storage module, which determines the storage location and write format of the key-value pairs.
[0063] A tenant refers to a set of accessible storage resources in a key-value storage system. A tenant's set of storage resources can be used by multiple users bound to that tenant.
[0064] A transaction is the smallest unit of operation in a key-value storage system, comprising a series of operations performed when a data processing request is requested. These operations are packaged together at the beginning of a transaction and executed as a whole throughout the transaction. If any operation within a transaction fails, the entire transaction is rolled back to its state before the transaction began, ensuring data consistency within the key-value storage system.
[0065] Horizontal scaling (also known as scale-out) refers to expanding a system's capacity by adding more compute nodes. This approach typically involves a distributed cluster of multiple compute nodes working together to handle high-concurrency traffic. The opposite is vertical scaling (scale-up).
[0066] Vertical scaling, also known as vertical scaling, refers to increasing the processing power of a system by adding hardware resources (such as CPU, memory, and hard drive) to a single compute node. This method is suitable for single-node systems; when the system needs to expand, performance can only be enhanced by upgrading the hardware. Vertical scaling is generally simpler because it only requires adding hardware, but it cannot be further expanded when the system's concurrency exceeds the hardware limits of a single compute node.
[0067] In Kubernetes (k8s) clusters, a distributed key-value store system is typically used to achieve service discovery, state storage, and cluster configuration. For example, commonly used key-value stores can include distributed key-value stores implemented using Zookeeper (zk) or ETCD protocol interfaces.
[0068] Figure 1 This is a schematic diagram of a known key-value storage system, such as... Figure 1 As shown, the key-value storage system includes a client layer, an application programming interface (API) layer, a distributed consensus algorithm (Raft) layer, a logic layer, and a storage layer.
[0069] The Client layer provides clients to tenants, enabling them to access the key-value store system. Illustratively, this client can be either an implementation of the ZooKeeper protocol interface or an implementation of the ETCD protocol.
[0070] The API layer provides a communication interface between the Client layer and the logic layer. The API layer can receive data processing requests from the client through this interface and then forward these requests to the logic layer.
[0071] The logic layer can create data processing requests as transactions and send them to the Raft layer for processing, so that the Raft layer returns Raft logs related to the Put request. The Raft logs consist of a series of Raft log entries, each corresponding to an operation within the transaction.
[0072] The Raft layer is implemented using the Raft algorithm. The Raft algorithm will be introduced below.
[0073] The Raft distributed consensus algorithm is a protocol used to manage and maintain the consistency of distributed systems. This algorithm elects a transaction committer (also known as the Leader node) from multiple nodes in a key-value storage system to handle all client data processing requests and replicates the corresponding Raft log entries to other compute nodes (also known as Follower nodes).
[0074] The Leader node is used to receive transactions sent by the KV server module, generate Raft logs based on the transactions, and persist the Raft logs to the storage layer.
[0075] When a leader node commits a transaction, it needs to rely on the voting results of the leader node and other follower nodes. Therefore, key-value stores typically include an odd number of compute nodes. The number of compute nodes can be three, five, or seven. Taking three compute nodes as an example, each compute node has its own membership state, such as one leader node and two follower nodes.
[0076] The Leader node is also used to send Raft logs to the logic layer and other Follower nodes.
[0077] The logic layer and storage layer together implement the multi-version concurrency control (MVCC) module of the key-value storage system. Based on the Raft log mentioned above, it can perform a series of comparison and exchange operations on the key-value pairs in the storage layer to achieve data consistency of the key-value storage system.
[0078] Based on the above, it can be concluded that key-value storage systems require only one Leader node to participate in transaction execution at all times. Adding or deleting a compute node may trigger a re-election process, resulting in two Leader nodes simultaneously. In this scenario, if a tenant sends two data update requests for the same key-value pair with different versions (e.g., version A and version B, where version A < version B), the two Leader nodes may execute the transactions corresponding to these requests separately. If a timing error occurs during the execution of these transactions, one Leader node might update the underlying storage with the key-value pair of version B first, while the other Leader node then replaces version B with version A, causing data errors in the underlying storage. Therefore, existing key-value storage systems do not support improving scalability and utilization by adding compute nodes.
[0079] Meanwhile, in key-value storage systems that rely on the Raft algorithm, each write operation requires a majority of nodes in the system to successfully write the Raft log entry to disk before the leader node modifies its internal state machine and returns the result to the client. Therefore, the more nodes in a key-value storage system, the lower its performance, and the probability of errors increases linearly, resulting in a linear performance degradation.
[0080] In summary, key-value storage systems based on Raft and multi-version concurrency modules have poor scalability and are difficult to support horizontal scaling of compute and storage nodes.
[0081] A known optimization scheme for key-value storage systems separates the Raft layer from the Server layer and interfaces the storage layer with an open-source KV storage instance. This solves problems such as slow system startup due to file reads and memory index rebuilding by multiple concurrent modules in scenarios with large data volumes, memory overflow when node memory is insufficient, long traversal times, and latency jitter when committing transactions. Furthermore, this key-value storage system supports vertical scaling.
[0082] However, this scheme still relies on the leader election module of the algorithm. For example, the three data replicas (one leader node and two follower nodes) implemented using the Raft algorithm do not achieve statelessness of the compute nodes, making it difficult to achieve horizontal scaling of the storage system.
[0083] Another known optimization scheme for key-value stores involves modifying the server layer to move the underlying storage to distributed storage instances, thus expanding storage resources through these instances. However, this scheme still relies on the Raft algorithm and does not achieve horizontal scaling of compute nodes.
[0084] We hope to find an improved solution that can support the horizontal scaling of the key-value storage system, thereby improving its scalability and utilization.
[0085] To improve the scalability and utilization of key-value storage systems, this application provides a data processing and querying method and related apparatus for key-value storage systems. By separating the computing cluster that undertakes computing tasks from the storage cluster that undertakes storage tasks, a first computer node creates a transaction and generates a sinking operator related to the data processing request based on the user's input data processing request. The transaction and sinking operator are then submitted to the storage cluster for execution, enabling the key-value storage system to execute transactions using the sinking operator without relying on an election mechanism to determine the transaction committer. This ensures the statelessness of multiple computing nodes in the computing cluster and achieves horizontal scaling of both the computing and storage clusters. Furthermore, by adding a tenant identifier as a prefix to the key, the key-value storage system supports data isolation among tenants, achieving logical multi-tenancy and improving the utilization of the key-value storage system.
[0086] For example, Figure 2 This is a schematic diagram of the structure of a key-value storage system provided in an embodiment of this application. For example... Figure 2 As shown, the key-value storage system includes a client layer 100, a server layer 200, and a storage cluster 300.
[0087] Client layer 100 includes at least one client for receiving data processing requests from multiple tenants and forwarding these requests to computing cluster 200. These data processing requests include key-value pair-related write (Put) requests, watch (Watch) requests, delete (Delete) requests, and range query (Range) requests. Illustratively, Client layer 100 includes clients 110, 120, and 130 with interfaces compatible with both zk and ETCD protocols.
[0088] The following sections will introduce Put requests, Watch requests, Delete requests, and Range requests.
[0089] A Put request is used to request the storage of key-value pairs to storage cluster 300. The Put request includes a key-value pair input by a tenant. If the key-value pair already exists in storage cluster 300, executing the Put request will update the older version of the key-value pair data in storage cluster 300.
[0090] A Watch request is used to request a compute node to asynchronously monitor changes to key-value pairs, which are specified by a tenant through a client.
[0091] The Delete request is used to request the removal of one or more key-value pairs from storage cluster 300.
[0092] As an illustration, a tenant can use a Delete request to delete multiple key-value pairs within a specific range related to a key in storage cluster 300; the same tenant can also use a Delete request to delete a single key-value pair corresponding to that key in storage cluster 300.
[0093] A Range request is used to retrieve a specific range of key-value pairs from storage cluster 300. A tenant can retrieve all key-value pairs within this range by specifying a start key (Key) and an end key (range_end).
[0094] In some possible implementations, the key-value store system may randomly assign clients 110, 120, and 130 to tenants 10, 20, and 30, respectively.
[0095] As one possible implementation, tenants 10, 20, and 30 may designate one or more clients among clients 110, 120, and 130 to access at least one computing node in the computing cluster.
[0096] As an example and not a limitation, the aforementioned clients 110, 120, and 130 can be deployed on separate computing devices, or they can be deployed together on a single computing device at the granularity of virtual machines or containers. In other words, this application does not limit the deployment method of multiple clients in the Client layer 100.
[0097] Server layer 200 can be deployed on at least one compute node in the key-value storage system. Illustratively, the at least one compute node may include compute node 210, compute node 220, and compute node 230.
[0098] Taking Server layer 200 deployed on compute node 210 as an example, Server layer 200 can receive data processing requests sent by at least one tenant through a client.
[0099] To achieve horizontal scaling of the computing cluster 200 in the key-value storage system, in this application, the Server layer 200 can receive data processing requests from at least one tenant through a client, create corresponding transactions based on these requests, and submit the transactions to the storage cluster 300. This allows the storage cluster 300 to execute related operations for key-value pairs, including Put, Delete, and Range requests, based on the type of the transaction. These transactions include write transactions corresponding to Put requests, delete transactions corresponding to Delete requests, and range query transactions corresponding to Range requests. Each transaction includes a corresponding transaction version identifier.
[0100] In some possible implementations, when the data processing request includes a Put or Delete request carrying a lease, the Server layer 200 can also generate a corresponding pushdown operator based on the data processing request, and send the pushdown operator to the storage cluster 300, based on the creation of a transaction. The pushdown operator includes at least one sub-request, and each sub-request requests the storage cluster 300 to perform a corresponding operation.
[0101] Server layer 200 sends the aforementioned pushdown operator to storage cluster 300 and submits a transaction to storage cluster 300 so that storage cluster 300 can use the pushdown operator to execute sub-requests related to the data processing request.
[0102] As one possible implementation, when the Server layer 200 connects to multiple tenants, the Server layer 200 is also used to manage the traffic quotas of the multiple connected tenants.
[0103] Server layer 200 is also used to receive the processing results of the above transactions returned by the storage node, and send the processing results to at least one client accessed by a tenant.
[0104] As an example rather than a limitation, the Server layer 200 described above can be deployed independently on a single physical device, or it can be deployed collectively on a single computing device at the granularity of virtual machines or containers. In other words, this application does not limit the deployment method of the Server layer 200.
[0105] Storage cluster 300 includes multiple storage instances implemented using Globally Distributed Component Services (GDCS).
[0106] To ensure data consistency in the key-value store system, each storage instance in the storage cluster 300 includes a complete copy of the data. When performing Put, Delete, and Range operations on key-value pairs, the storage cluster 300 can use distributed transactions to force the same data operations to be performed on the data copies of each storage instance, thereby achieving multi-version concurrency control in the key-value store system.
[0107] Storage cluster 300 is used to provide a data interface (not shown in the figure) to multiple computing nodes in the computing cluster. Through the above data interface, storage cluster 300 can receive transactions sent by any computing node through its own data processing module, and perform Put, Delete and Range operations on key-value pairs according to the type of transaction.
[0108] The storage cluster 300 can also receive a query request from a tenant sent by any computing node through its own data processing module via the aforementioned data interface, and respond to the query request by returning the tenant's authentication information and / or lease information to the aforementioned computing node.
[0109] The storage nodes 310, 320, and 330 in storage cluster 300 are illustrated using GDCS storage instances. Each storage node contains a complete copy of the data.
[0110] As an example and not a limitation, the aforementioned storage nodes can be separately deployed physical devices, or they can be I / O virtualized devices implemented at the granularity of virtual machines or containers, all deployed on a single computing device. In other words, this application does not limit the deployment method of storage nodes.
[0111] By constructing tenant data processing requests into corresponding transactions and / or operators, and sending the transactions and / or operators to the storage system for execution, data association and storage-computation separation in computing cluster 200 and storage cluster 300 are realized. This ensures the statelessness of multiple computing nodes in the computing cluster, enables horizontal scaling of computing cluster 200 and storage cluster 300, and improves the utilization rate of the key-value storage system.
[0112] Because existing key-value storage systems implemented with ZK or ETCD do not support tenant isolation in their underlying storage, multi-tenant sharing of multiple key-value storage systems cannot be achieved.
[0113] To enable multi-tenant sharing of the key-value storage system, this application distinguishes the keys of key-value pairs of each tenant in the storage cluster 300 by adding a tenant identifier as a prefix to the key.
[0114] For example, Figure 3This is a diagram illustrating the data transformation of key-value pairs corresponding to a specific tenant. For example... Figure 3 As shown, taking tenant ID "tenant1" as an example, the data types that this tenant possesses can include user data, large values, tenant information, monitoring, leases, authentication, tasks, and compression.
[0115] The user data “ / data / tenant1” corresponds to two key-value pairs, including the key “ / data / tenant1 / Key1 / versionStamp” and the key “ / data / tenant1 / vindex / versionStamp / Key1”.
[0116] The big value " / splitData / tenant1" corresponds to two key-value pairs, including the key " / splitData / tenant1 / Key1 / uuid / 0" and the key " / splitData / tenant1 / Key1 / uuid / 1".
[0117] In particular, the key-value pairs for user data and large values are all prefixed with "tenant1" to distinguish them, which will not be elaborated further below.
[0118] The value corresponding to the above key is "message StorageValue". The following content will introduce the value "message StorageValue".
[0119] The boolean "DeleteTag" is used to represent the deletion flag for this value;
[0120] The 64-bit integer "createReverslon" is used to represent the cluster's revision identifier when the key was created;
[0121] The 64-bit integer "version" is used to represent the version identifier when the key was created;
[0122] A 64-bit integer "lease" is used to identify the tenant's lease.
[0123] The string "uuid" is used to represent the universally unique identifier (UUID) of the key-value database corresponding to the tenant;
[0124] The unsigned integer type "sum" and the character type "value" are used to represent specific values.
[0125] The key-value pair corresponding to the tenant information " / tenant / tenant1" has the key " / tenant / tenant1 / " and the value "Tenant". "Tenant" is used to describe the specific tenant information of tenant "tenant1".
[0126] The key-value pair corresponding to the monitoring information " / Watch / tenant1" includes the key " / Watch / tenant1 / id / Watchld1", and the value corresponding to this key is "Watch". "Watch" is used to describe the node information corresponding to each node of the subscribed tenant "tenant1" key-value pair.
[0127] Next, the value "Watch" will be introduced in the following content.
[0128] The string "endpoint" is used to describe the monitoring port of the Watch request. In this application, the trigger list of Watch information can be determined through the "endpoint" field.
[0129] The byte-type "Key" is used to describe the key that the Watch request monitors;
[0130] A 64-bit integer "watchId" is used to represent the version identifier of the Watch request;
[0131] The string "uuid" is used to represent the universally unique identifier (UUID) of the key-value database corresponding to the tenant;
[0132] The string "tenant" is used to represent the tenant information of this tenant.
[0133] The key-value pair corresponding to the monitoring information " / Watch / tenant1" also includes a Key.
[0134] The key " / Watch / tenant1 / notify / versionStamp / Key1" has a value of "Notify". "Notify" describes notification information related to data processing requests. In this application, the notification information is used to take effect when the Watch information subscribed to by the tenant fails to trigger.
[0135] Next, the value "Notify" will be introduced in the following content.
[0136] The field “mvccpb.Event.EventType type” is used to describe the type of data processing request. For example, when this field is “1”, it can indicate that the data processing request is a Put request.
[0137] The byte-type "Key" is used to describe the key that the Watch request monitors;
[0138] The byte-type "value" is used to describe the value corresponding to the key monitored by the Watch request;
[0139] The string "tenant" is used to represent the tenant information of this tenant.
[0140] The key-value pair corresponding to the lease " / lease / tenant1" includes the key " / Watch / tenant1 / id / leaseld1" and the key "
[0141] The key " / Watch / tenant1 / binding / leaseld1 / Key1" has a value of "Lease". "Lease" describes the lease information of the tenant "tenant1".
[0142] The value “Lease” will be introduced below.
[0143] A 64-bit integer "ID" is used to represent the length of the lease;
[0144] A 64-bit integer "TTL" is used to represent the duration of the lease (Time To Live, TTL);
[0145] The 64-bit integer "expird" is used to indicate the expiration date of the lease;
[0146] The string "tenant" is used to represent the tenant information of this tenant.
[0147] The key-value pair corresponding to the authentication (Auth) " / auth / tenant1" includes Key " / anth / tenant1 / role / role1", Key
[0148] The key “ / anth / tenant1 / permission / role1 / 0” and the key “ / anth / tenant1 / user / user1” have a value “PemmRecord”. “PemmRecord” is used to describe the authentication and authorization information of the tenant “tenant1”.
[0149] The key-value pair corresponding to Task " / task / tenant1" includes Key " / task / tenant1 / compact" and Key " / task / tenant1 / compact".
[0150] The key " / task / tenant1 / lease" has a value of "Tasklnfo", which describes the task information for executing the key-value pair task. This key-value pair task can include data compression and lease index deletion.
[0151] The key corresponding to the compressed file " / compact / tenant1" is " / compact / tenant1 / index / 10001 / Key1".
[0152] It is worth noting that the prefix tenant ID "tenant1" added to the above Key is intended to facilitate understanding of this application. In actual implementation, other prefixes can also be added, such as the tenant's IP address to distinguish the key-value pairs of each tenant. This application does not limit the type of prefix added to each tenant.
[0153] By adding a prefix to the key-value pairs in the storage cluster, the Server layer 200 can distinguish key-value pairs from multiple tenants, enabling multi-tenant sharing of the key-value storage system, supporting tenant-level traffic control, and improving the utilization of the key-value storage system.
[0154] Firstly, based on the content described above, the data processing method provided in the embodiments of this application will be introduced. It is understood that this method is proposed based on the content described above, and some or all of the content of this method can be found in the description above.
[0155] For example, Figure 4 This is a flowchart illustrating a data processing method for a key-value storage system proposed in this application, as shown below. Figure 4 As shown, data processing of the key-value storage system can be achieved through S410 to S440.
[0156] S410: Receives the data processing request from the first tenant.
[0157] As mentioned above, the key-value storage system includes at least one client layer 100, a service layer 200, and a storage cluster 300.
[0158] The client layer 100 includes at least one client compatible with ZK and ETCD interfaces, such as... Figure 2 The client 110 shown allows multiple tenants to access the server 200. The server layer 200 can be deployed on at least one compute node in the key-value storage system, such as... Figure 2The compute nodes 210, 220, and 230 are shown. The Server layer 200 is communicatively connected to the storage cluster 300.
[0159] In some possible implementations, the storage cluster 300 pre-stores key-value pairs belonging to multiple tenants. Each tenant's key-value pair has its tenant identifier added as a prefix to the key for differentiation.
[0160] Deployed at Server tier 200 Figure 2 Taking the compute node 210 shown as an example, the Server layer 200 can receive data processing requests sent by multiple tenants. These data processing requests include, but are not limited to, Put requests and Delete requests.
[0161] In some possible implementations, taking tenant 10 (the first tenant) as an example, if the data processing request sent by the first tenant through client 110 is of type Put request, the Put request includes the key-value pairs input by the first tenant.
[0162] As one possible implementation, if the data processing request sent by the first tenant is a Delete request, the Delete request includes the key entered by the first tenant.
[0163] By way of illustration and not limitation, this application enables the key-value storage system to support data isolation among tenants and achieve multi-tenant sharing of the key-value storage system by adding a tenant identifier as a prefix to the key. Specifically, tenant 10 can add its tenant identifier to the key when sending a data processing request. Alternatively, the server layer 200 can add the tenant identifier to the key; this application does not limit this approach.
[0164] S420: Create a corresponding transaction based on the data processing request and generate pushdown operators related to the data processing request.
[0165] Server layer 200 can create corresponding transactions based on data processing requests and generate pushdown operators related to the data processing requests. Each pushdown operator includes at least one sub-request, and each sub-request requests the storage cluster 300 to perform a corresponding operation. The aforementioned transaction instructs the storage cluster 300 to process at least one sub-request within the pushdown operator.
[0166] In some possible implementations, when the first tenant 10's data processing request is a Put request, the Server layer 200 can create a corresponding write transaction based on the Put request. This write transaction includes the key-value pairs input by the first tenant, as well as notification information related to the Put request.
[0167] Storage cluster 300 pre-stores second key-value pairs related to Put requests, where the second key-value pairs include the lease information of the first tenant.
[0168] First, the Server layer 200 can send the key in the first key-value pair to the storage cluster 300 so that the storage cluster 300 can return the lease information of the first tenant based on the key in the first key-value pair.
[0169] Then, Server layer 200 can generate pushdown operators related to write transactions based on the lease information of the first tenant.
[0170] Finally, the compute node can send the pushdown operator related to the Put request to the storage cluster 300, so that the storage cluster 300 can process at least one sub-request in the operator using the pushdown operator related to the write transaction, and update the second key-value pair stored in the storage cluster 300 to the first key-value pair input by the user 10.
[0171] Next, we will introduce pushdown operators related to write transactions.
[0172] 1) The first sub-request that requests storage cluster 300 to modify the revision identifier field and version identifier field of the first key-value pair.
[0173] 2) Request storage cluster 300 to replace the lease index associated with the first key-value pair in the second sub-request, the lease index being determined based on the key in the first key-value pair and the lease identifier of the first tenant 10.
[0174] 3) A third sub-request requests storage cluster 300 to store the first key-value pair and generate a multi-version record of the first key-value pair. The key of the first key-value pair in storage cluster 300 is determined based on the version identifier of the write transaction. For example, the original key of the first key-value pair is concatenated with the transaction version identifier for storage.
[0175] 4) Request storage cluster 300 to generate the first key-value pair related Watch information, and generate the version index of the Watch information as a fourth sub-request. The key of the Watch information is determined based on the version identifier of the Watch information write transaction. For example, the original key of the Watch information is concatenated with the transaction version identifier for storage.
[0176] 5) The fifth sub-request requests storage cluster 300 to compress the second key-value pair and generate a compressed index for the second key-value pair. The key of the compressed index is determined based on the version identifier of the write transaction. For example, the key of the second key-value pair is concatenated with the transaction version identifier for storage.
[0177] When the first tenant 10's data processing request is a Delete request, the Server layer 200 can create a corresponding delete transaction based on the data processing request. This delete transaction includes a user-input key, which corresponds to a third key-value pair within a specific range stored in the storage cluster 300. For example, Figure 3 The example shown is a set of key-value pairs prefixed with "tenant1".
[0178] As one possible implementation, if the third key-value pair includes the lease information of the first tenant 10.
[0179] First, the Server layer 200 can create a deletion transaction corresponding to the data processing request based on the key entered by the first tenant 10, and send the key entered by the first tenant 10 to the storage cluster 300 so that the storage cluster 300 can return the lease information of the first tenant based on the key entered by the first tenant.
[0180] Then, Server layer 200 can generate pushdown operators related to the deletion transaction based on the first tenant's lease information.
[0181] Finally, Server layer 200 can send the pushdown operator related to the delete transaction to storage cluster 300, so that storage cluster can use the pushdown operator related to the delete transaction to process at least one sub-request in the pushdown operator to delete the third key-value pair.
[0182] The following sections will introduce pushdown operators related to deletion transactions.
[0183] 1) The sixth sub-request requests storage cluster 300 to add or delete records in the third key-value pair. The key of the third key-value pair is determined based on the version identifier of the deletion transaction; for example, the key of the third key-value pair is concatenated with the transaction version identifier for storage.
[0184] 2) The seventh sub-request of requesting storage cluster 300 to replace the lease index associated with the third key-value pair, the lease index being determined based on the key in the third key-value pair and the lease identifier of the first tenant.
[0185] 3) The eighth sub-request requests storage cluster 300 to compress the third key-value pair and generate a compressed index for the third key-value pair. The key of the compressed index is determined based on the version identifier of the write transaction; the key of the third key-value pair is concatenated with the transaction version identifier for storage.
[0186] 4) The ninth sub-request requests storage cluster 300 to generate the third key-value pair related Watch information and to generate the version index of the Watch information. The key of the Watch information is determined based on the version identifier of the write transaction; for example, the original key of the Watch information is concatenated with the transaction version identifier for storage.
[0187] In some possible implementations, when a tenant enters a key or key-value pair that includes the tenant's tenant identifier, the Server layer 200 can determine the tenant's traffic quota based on the tenant identifier of any tenant.
[0188] To illustrate, taking the first tenant 10 as an example, if the traffic quota of the first tenant 10 exceeds the threshold, the Server layer 200 can return a processing result of data processing request failure to the client connected to the first tenant 10, such as returning the prompt message "Data processing request was aborted because traffic quota exceeded the threshold".
[0189] If the traffic quota of the first tenant 10 does not exceed the threshold, the Server layer 200 can create a transaction corresponding to the data processing request based on the tenant identifier of the first tenant 10, so that the storage cluster 300 can execute the above transaction based on the tenant identifier of the first tenant.
[0190] This method enables key-value storage systems to support tenant-level flow control and data isolation among tenants by adding a tenant identifier as a prefix to the key, thereby achieving multi-tenant sharing of the key-value storage system and improving the utilization rate of the key-value storage system.
[0191] S430: Send the pushdown operator to the storage cluster and commit the transaction to the storage cluster.
[0192] Server layer 200 can send pushdown operators to storage cluster 300 and submit the transaction created in step S420 to storage cluster 300. As mentioned above, the types of transactions mentioned above include, but are not limited to, write transactions with leases and delete transactions.
[0193] Storage cluster 300 can process at least one sub-request in the pushdown operator based on the transaction.
[0194] In some possible implementations, when the type of the above transaction is a write transaction with a lease, the storage cluster 300 can execute the above write transaction, call the underlying data interface to execute multiple sub-requests in the pushdown operator related to the write transaction in step S20, store the first key-value pair in storage node 310, storage node 320 and storage node 330, and return the processing result of the write transaction to the compute node 210 after the write transaction is executed.
[0195] As one possible implementation, when the above transaction type is a deletion transaction with a lease, the storage cluster 300 can execute the above deletion transaction, call the underlying data interface to execute multiple sub-requests in the pushdown operator related to the deletion transaction in step S20, and delete the third key-value pair in storage nodes 310, 320 and 330.
[0196] Storage cluster 300 can return the processing result of a write transaction or delete transaction to the server layer 200 after the write transaction or delete transaction is completed.
[0197] S440: Receives the transaction processing result returned by the storage cluster and sends the processing result to the client accessed by the first tenant.
[0198] Compute node 210 can send the processing result of the data processing request to the client accessed by the first tenant based on the processing result of the transaction. The processing result of the data processing request can include the processing result of either a Put request or a Delete request.
[0199] After the Put request is executed, compute node 210 can also send the Watch information of the first key-value pair to at least one compute node (the first compute node) that has subscribed to the Watch information of the first key-value pair, so that the first compute node can update the index file of the first key-value pair according to the Watch information.
[0200] Understandably, after executing the Delete request, compute node 210 can also send the Watch information of the aforementioned third key-value pair to at least one compute node (second compute node) that has subscribed to the Watch information of the third key-value pair, so that the second compute node can update the index file of the third key-value pair according to the aforementioned Watch information.
[0201] As one possible implementation, when Server layer 200 receives a message from another tenant, such as... Figure 2 After tenant 20 (the second tenant) sends a data processing request, the server layer 200 can also process the data processing request of the second tenant 20 in the same way.
[0202] First, the Server layer 200 can receive data processing requests sent by the second tenant 20. The second tenant is any one of the multiple tenants other than the first tenant 10; the data processing request includes the tenant identifier of the second tenant 20.
[0203] Then, the Server layer 200 can create a transaction corresponding to the data processing request based on the tenant identifier of the second tenant 20, and generate pushdown operators related to the data processing request.
[0204] Then, Server layer 200 can send the above pushdown operator to storage cluster 300 and commit the transaction to storage cluster.
[0205] Finally, Server layer 200 can receive the processing result of the above transaction returned by storage cluster 300, and send the processing result of the transaction to second tenant 20 according to the tenant identifier of second tenant 20.
[0206] This method utilizes the Server layer to receive data processing requests from the first tenant, creates a transaction corresponding to the data processing request, and submits the transaction to the storage cluster so that the storage cluster can execute the transaction. The Server layer then sends the processing result back to the first tenant. By submitting the transaction to the storage cluster for execution, this method enables the storage cluster to support data association and storage-compute separation, ensures the statelessness of multiple compute nodes in the compute cluster, achieves horizontal scaling of both the compute and storage clusters, and improves the utilization rate of the key-value storage system.
[0207] Secondly, based on the content described above, the query method of the key-value storage system provided in the embodiments of this application will be introduced. It is understood that this method is proposed based on the content described above, and some or all of the content of this method can be found in the description above.
[0208] For example, Figure 5 This is a flowchart illustrating a data processing method for a key-value storage system proposed in this application, as shown below. Figure 5 As shown, data processing of the key-value storage system can be achieved through S510 to S530.
[0209] S510: Receives query requests from the first tenant.
[0210] As mentioned above, the key-value storage system includes at least one client layer 100, a service layer 200, and a storage cluster 300.
[0211] The client layer 100 includes at least one client compatible with ZK and ETCD interfaces, such as... Figure 2 The client 110 shown allows multiple tenants to access the server 200. The server layer 200 can be deployed on at least one compute node in the key-value storage system, such as... Figure 2 The compute nodes 210, 220, and 230 are shown. The Server layer 200 is communicatively connected to the storage cluster 300.
[0212] Deployed at Server tier 200 Figure 2 Taking the compute node 210 shown as an example, the Server layer 200 can receive query (Range) requests sent by multiple tenants.
[0213] When the first tenant 10's data processing request is a Range request, the Server layer 200 can create a corresponding read transaction based on the Range request. The read transaction includes a user-inputted Key, which corresponds to multiple key-value pairs within a specific range stored in the storage cluster 300. As mentioned earlier, the user-inputted Key may include a first version identifier corresponding to that Key.
[0214] S520: Send the key entered by the first tenant to the storage cluster in order to retrieve the second version identifier from the storage cluster.
[0215] Server layer 200 can send the key input by first tenant 10 to storage cluster 300 in order to obtain the version identifier (second version identifier) of the higher version key-value pair from multiple versions of key-value pairs within a specific range of storage cluster 300.
[0216] In some possible implementations, when the first version identifier is less than the second version identifier, the Server layer 200 can obtain the version identifier (third version identifier) of the uncompressed lower version key-value pair from multiple versions of key-value pairs within a specific range of the storage cluster 300.
[0217] If the first version identifier is greater than or equal to the third version identifier, the Server layer 200 can create a read transaction corresponding to the query request based on the key entered by the first tenant.
[0218] If the first version identifier is less than the third version identifier, the Server layer 200 can return a query request failure result to the client connected to the first tenant 10. For example, it can return a message such as "The query request was aborted because the query data has been compressed".
[0219] As one possible implementation, when the first version identifier equals the second version identifier, the Server layer 200 can create a read transaction corresponding to the query request based on the key entered by the first tenant 10.
[0220] S530: Submit a read transaction to the storage cluster so that the storage cluster can return the result of the read transaction.
[0221] Server layer 200 can submit the read transaction to the storage cluster 300.
[0222] Storage cluster 300 can execute the above read transaction and return a key-value pair (fourth key-value pair) that matches the first version identifier to Server layer 200.
[0223] This method utilizes the Server layer to determine whether the storage cluster contains key-value pairs within a specific range corresponding to the first version identifier. Based on this, the Server layer can create read transactions to execute from the storage cluster. Compared to key-value storage systems based on election mechanisms, which require compute nodes to participate in key-value pair version maintenance and read transaction execution, this method ensures storage-compute separation in the key-value storage system, enables horizontal scaling of the compute cluster, and improves the utilization rate of the key-value storage system.
[0224] The following section will describe the interactive scenarios when the key-value storage system proposed in this application executes the above-mentioned output processing method.
[0225] For example, Figure 6 This is a schematic diagram illustrating the interaction scenario of a key-value storage system executing leaseless Put and Watch requests. As shown in Figure 6, compared to... Figure 2 The key-value storage system also includes at least one garbage collection server (GC-Server) 400, a metadata cluster 500, and multiple caches. Taking the Server layer 200 deployed on compute node 210 as an example, the key-value storage system also includes a cache 240 connected to compute node 210.
[0226] Reclaim server 400 is used to compress older versions of key-value pairs for a tenant in storage cluster 300 when the tenant's traffic quota exceeds a threshold, in order to perform memory reclamation operations on storage cluster 300.
[0227] Metadata cluster 500 is used to store metadata of the key-value storage system. This metadata is used to represent the configuration information and status data of the key-value storage system.
[0228] Cache 240 is used to store the index file of compute node 210. When multiple tenants access compute node 210, they can quickly query the key-value pairs belonging to each tenant in the storage cluster through the above index file.
[0229] Tenant 10 can send a Put request without a lease to compute node 210 through client 110, such as... Figure 6 As shown in step 1. The Put request includes the key-value pair (first key-value pair) that tenant 10 needs to write.
[0230] Tenant 20 can also send a Put request without a lease to computer node 210 through client 120, such as Figure 6 As shown in step 1`.
[0231] In some possible implementations, a tenant can specify any compute node through a client and send a data processing request to that compute node to request the compute node to perform data operations related to the data processing request.
[0232] As one possible implementation, a tenant can also use an election mechanism to elect a computer node from among multiple computing nodes in the computing cluster, and send a data processing request to that computing node to request the computer node to perform the data operations related to the data processing request.
[0233] Taking the example of tenants 10 and 20 jointly specifying compute node 210 to execute the corresponding Put request, compute node 210 can accept the Put request sent by tenant 10 through client 110 and accept the Put request sent by tenant 20 through client 120.
[0234] For illustration, the first key-value pair includes the key " / data / tenant1 / vindex / versionStamp / Key1" and the corresponding value "StorageValue". The "tenant1" field in the key is the tenant identifier for tenant 10, and the "Key1" field is the version identifier for this key. The value "StorageValue" includes the version identifier, lease identifier, and the specific value of the key " / data / tenant1 / vindex / versionStamp / Key1".
[0235] The key-value pair associated with tenant 20's Put request includes the key " / data / tenant2 / vindex / versionStamp / Key1" and its corresponding value "StorageValue". Specifically, the "tenant2" field in the key is the tenant identifier for tenant 20, and the "Key1" field is the version identifier for that key. The content of the value "StorageValue" will not be elaborated upon here.
[0236] First, compute node 210 can determine the traffic quota for tenant 10 based on the "tenant1" field, and the traffic quota for tenant 20 based on the "tenant2" field, such as... Figure 6 As shown in step 2 of the document.
[0237] In some possible implementations, if tenant 10's traffic quota does not exceed the threshold, compute node 210 can execute... Figure 6 Step 3 continues processing tenant 10's data processing request.
[0238] As one possible implementation, if tenant 20's traffic quota exceeds the threshold, compute node 210 can stop processing tenant 10's data processing requests and return a "Data processing request terminated due to traffic quota exceeding the threshold" message to client 120, such as... Figure 6 As shown in step 3`.
[0239] Then, when the tenant's traffic quota does not exceed the threshold, compute node 210 can create a corresponding write transaction based on the tenant 10's Put request and submit the write transaction to the storage system for execution by storage cluster 300. The write transaction includes the first key-value pair.
[0240] It is understood that any computing node can control the data processing requests of a tenant based on the traffic quota of at least one tenant connected to the node (which can be referred to as flow control). This application will not elaborate on the flow control process of each node.
[0241] As an illustration, when storage cluster 300 executes the above write transaction, it can call the underlying data interface to atomically write the first key-value pair, and... Figure 5 The Notify information associated with the first key-value pair is shown, and Watch information associated with the first key-value pair is generated. The key of this key-value pair is “ / data / tenant1 / vindex / versionStamp / Key1”, and the value is “StorageValue”.
[0242] Storage cluster 300 can return the processing result of the write transaction to compute node 210 after the write transaction is completed, such as... Figure 6 As shown in step 3 of the process.
[0243] Then, compute node 210 can return the processing result of the Put request to client 110 based on the processing result of the above transaction, such as... Figure 6 As shown in step 4 of the document.
[0244] In some possible implementations, if the write transaction corresponding to tenant 10's Put request is executed successfully, compute node 210 can return the successful Put request execution result to client 110 and execute... Figure 6 Step 5 shown
[0245] As one possible implementation, if the write transaction corresponding to the Put request of tenant 10 fails to execute, the compute node 210 can return the processing result of the Put request failure to the client 110.
[0246] Then, compute node 210 can store the Watch information corresponding to tenant 10's Put request in cache 240 to prevent any missed Watch information corresponding to the Put request, such as... Figure 6 As shown in step 5 of the process.
[0247] Then, compute node 210 can traverse the Watch information stored in storage system 300 according to the key " / data / tenant1 / vindex / versionStamp / Key1" to obtain the node information corresponding to each node that has subscribed to the key-value pair of tenant 10, and determine the trigger list of Watch information based on the above node information. Figure 6 Step 6 in the process.
[0248] Then, compute node 210 sends Watch information to the corresponding compute nodes, such as compute nodes 220 and 230, according to the aforementioned trigger list, to trigger Watch subscriptions on the corresponding nodes, thereby achieving data synchronization among storage nodes 310, 320, and 330. Figure 6 As shown in step 7 of the document.
[0249] It is understandable that after receiving the Watch information sent by Compute Node 210, Compute Node 220 and Compute Node 230 can update the index file of the first key-value pair in their respective caches based on the Watch information.
[0250] Finally, once a compute node in the trigger list has successfully triggered the watch subscription, compute node 210 can also delete the Notify information from storage cluster 300 to avoid repeated triggering of the watch subscription.
[0251] The following content will introduce the interaction scenarios of Watch requests.
[0252] Please continue reading. Figure 6 For example, tenant 30 specifies compute node 220 to execute a Watch request.
[0253] First, compute node 220 can receive watch requests sent by tenant 30 through client 120, such as... Figure 6 Step a is shown in the diagram. Illustratively, the Watch request includes... Figure 5 The Watch information shown above describes the key-value pairs that the compute node is monitoring.
[0254] Then, the compute node can create a corresponding write transaction based on the Watch information mentioned above and submit the write transaction to storage cluster 300 so that storage cluster 300 can execute it. Figure 6 As shown in step b.
[0255] As an illustration, when the storage cluster 300 executes the above transaction, it can call the underlying data interface to write Watch information related to the Watch request, and return the processing result of the write transaction to the compute node 220 after the transaction is completed.
[0256] Finally, compute node 220 can return the processing result of the Watch request to client 110 based on the processing result of the write transaction described above, such as... Figure 6 Step c is shown in the diagram.
[0257] In some possible implementations, after completing the Put and Watch requests, compute node 210 and / or storage cluster 300 can also send the configuration information and status data of the key-value storage system to the metadata cluster for storage, which will not be elaborated further below.
[0258] Next, we will describe the interaction scenario of a Put request carrying a lease, taking the Server layer 200 deployed on the compute node 210 as an example.
[0259] For example, Figure 7 This is a flowchart illustrating the process of a key-value store system executing a Put request carrying a lease, as shown below. Figure 7 As shown, the key-value storage system can process Put requests carrying leases through steps 1-11.
[0260] Taking a tenant accessing the key-value storage system through client 110 as an example, the tenant specifies compute node 210 to handle Put requests carrying the lease.
[0261] First, the tenant sends a Put request carrying the lease to compute node 210 via client 110, such as... Figure 7 As shown in step 1, illustratively, storage cluster 300 includes a legacy key-value pair (second key-value pair) related to the Put request; the Put request includes the key-value pair input by the tenant (first key-value pair). The first and second key-value pairs contain the same key.
[0262] Then, compute node 210 can create a corresponding write transaction based on the Put request, initiate the write transaction commit process, and send the Key in the first key-value pair to storage cluster 300 to obtain tenant 10's lease information, such as... Figure 7 Step 2 is shown in the diagram. The write transaction includes the first key-value pair.
[0263] It is understandable that before creating a write transaction, compute node 210 can determine the tenant's traffic quota based on the Key in the first key-value pair. The process for determining the tenant's traffic quota will not be elaborated here.
[0264] Then, the storage cluster 300 can determine the lease information of the first key-value pair based on the key in the first key-value pair, and return the lease information to the compute node 210, such as... Figure 7 As shown in step 3 of the process.
[0265] Then, compute node 210 can determine whether the lease of the first key-value pair is still in existence based on the above lease information.
[0266] If the tenant's lease is no longer valid, the compute node 210 can return a failure message to the tenant via the client 110 stating that "the Put request was aborted due to lease expiration".
[0267] If the tenant's lease is still in effect, compute node 210 can generate pushdown operators related to the Put request, such as... Figure 7 Step 4 is shown in the diagram. The pushdown operator mentioned above includes sub-requests related to the Put request. The following sections will describe these Put request-related sub-requests.
[0268] 1) Request a: Request a is used to request storage cluster 300 to perform a Compare And Swap (CAS) atomic operation to modify the revision identifier field and version identifier field recorded in the value of the first key-value pair.
[0269] Among them, the compare-swap atomic operation allows storage cluster 300 to perform multiple operations in a batch during a single modification. That is, this group of operations is bound into an atomic operation and shares the same revision identifier.
[0270] 2) Request b, which requests the storage cluster 300 to perform a CAS atomic operation to replace the old version of the lease index of the first key-value pair with the key in the first key-value pair and the lease identifier of the tenant, and to perform a lease index based on the key in the first key-value pair and the lease identifier of the tenant so that the compute node 210 can obtain it.
[0271] 3) Request c: Request c is used to request storage cluster 300 to generate multiple versions of the first key-value pair record. In the higher version of the first key-value pair, the key is obtained by concatenating the original key with the transaction version identifier; the value is the value in the first key-value pair.
[0272] 4) Request d: Request d is used to request storage cluster 300 to generate the first key-value pair related Watch information and generate the version index of the aforementioned Watch information so that compute node 210 can obtain it. The key of the Watch information is obtained by concatenating the transaction version identifier with the original key. The value of the Watch information includes the monitoring port of the Watch request, the version identifier, the UUID, and the tenant information of the tenant, such as... Figure 5 The value “Watch” in the table is shown below, and will not be elaborated further here.
[0273] In some possible implementations, the sub-requests associated with the Put request also include request e.
[0274] 5) Request e: Request e requests storage cluster 300 to compress the second key-value pair and generate a compressed version index for compute node 210 to retrieve. The key of the compressed index is obtained by concatenating the transaction version identifier with the original key.
[0275] Then, compute node 210 can send the pushdown operator to storage cluster 300, so that storage cluster 300 can use the pushdown operator to execute write transactions, such as... Figure 7 As shown in step 5 of the process.
[0276] Then, compute node 210 can submit the completed write transaction to storage cluster 300, such as Figure 7 As shown in step 6 of the document.
[0277] Then, storage cluster 300 can execute the above write transaction, calling the underlying data interface to execute requests a to e in the pushdown operator respectively, storing the first key-value pair on multiple storage nodes, and returning the processing result of the write transaction to compute node 210 after the write transaction is completed, such as... Figure 7 As shown in step 7 of the document.
[0278] Then, compute node 210 can return the processing result of the Put request carrying the lease to client 110 based on the processing result of the write transaction described above, such as... Figure 7 As shown in step 8 of the document.
[0279] Finally, compute node 210 can trigger the Watch information through steps 9 to 11, which will not be described in detail here.
[0280] Next, we will describe the interaction scenario of a Delete request carrying a lease, taking the Server layer 200 deployed on the compute node 210 as an example.
[0281] For example, Figure 8 This is a flowchart illustrating the process of a key-value store system executing a delete request carrying a lease, as shown below. Figure 8 As shown, the key-value storage system can process the Delete request carrying the lease through steps 1-11.
[0282] Taking a tenant accessing the key-value storage system through client 110 as an example, the tenant specifies compute node 210 to handle the Delete request carrying the lease.
[0283] First, the tenant sends a Delete request carrying the lease to compute node 210 via client 110, such as... Figure 8 As shown in step 1. Illustratively, the above Delete request includes the key to be deleted entered by tenant 10. The storage cluster 300 pre-stores a third key-value pair, which is a specific range of key-value pairs related to the Delete request.
[0284] Then, compute node 210 can create a corresponding write transaction based on the Put request and initiate the write transaction commit process to storage cluster 300, such as... Figure 8 Step 2 is shown in the diagram. The write transaction includes the key entered by the tenant.
[0285] Then, storage cluster 300 can retrieve the tenant's lease information and the third key-value pair corresponding to the key in the delete request from storage cluster 300 based on the key in the delete transaction, and return the lease information and the third key-value pair to compute node 210, such as... Figure 8 As shown in step 3 of the process.
[0286] Then, computing node 210 can determine whether the tenant's lease is still in effect based on the above lease information.
[0287] If the tenant's lease is no longer valid, the compute node 210 can return a failure message to the tenant via the client 110: "Delete request is suspended due to lease expiration".
[0288] Compute node 210 can generate pushdown operators related to Delete requests, such as Figure 8 Step 4 is shown in the diagram. The pushdown operator mentioned above includes sub-requests related to the Put request. The following section will describe the sub-requests related to the Delete request.
[0289] 1) Request f, which requests storage cluster 300 to generate deletion records for key-value pairs within a specific range corresponding to the aforementioned Key. As mentioned earlier, storage cluster 300 can add a deletion tag "DeleteTag" to the values of the aforementioned key-value pairs to perform the key-value pair deletion operation. The Keys of the aforementioned key-value pairs within the specific range with the added deletion tag are all obtained by concatenating the original Key with the transaction version identifier.
[0290] 2) Request g, which requests the storage cluster 300 to perform a CAS atomic operation to replace the old version of the lease index with the key in the first key-value pair and the lease identifier of the tenant, and to perform a lease index based on the key in the first key-value pair and the lease identifier of the tenant so that the compute node 210 can obtain it.
[0291] 3) Request h: Request h is used to request storage cluster 300 to generate Watch information related to the third key-value pair, and to generate a version index of the aforementioned Watch information so that compute node 210 can obtain it. The key of the Watch information is obtained by concatenating the transaction version identifier with the original key. The value of the Watch information includes the monitoring port of the Watch request, the version identifier, the UUID, and the tenant information of the tenant.
[0292] In some possible implementations, the sub-requests associated with the Delete request also include request d.
[0293] 4) Request i: Request i is used to request storage cluster 300 to compress the third key-value pair and generate a compressed index of key-value pairs within this specific range, so that compute node 210 can obtain it. The key of the compressed index is obtained by concatenating the original key with the transaction version identifier.
[0294] Then, compute node 210 can send the pushdown operator to storage cluster 300, so that storage cluster 300 can use the pushdown operator to perform a deletion transaction, such as... Figure 8 As shown in step 5 of the process.
[0295] Then, compute node 210 can submit a completed deletion transaction to storage cluster 300, such as... Figure 7 As shown in step 6 of the document.
[0296] Then, storage cluster 300 can execute the above deletion transaction, calling the underlying data interface to execute sub-requests f to i in the downcomputation operator respectively, deleting the third key-value pairs in multiple storage nodes, and returning the processing result of the deletion transaction to compute node 210 after the deletion transaction is completed, such as... Figure 8 As shown in step 7 of the document.
[0297] Then, compute node 210 can return the processing result of the Delete request carrying the lease to client 110 based on the processing result of the above transaction, such as... Figure 7 As shown in step 8 of the document.
[0298] Finally, compute node 210 can trigger the Watch information through steps 9 to 11, which will not be described in detail here.
[0299] Next, we will explain the interaction scenario of range requests, taking the Server layer 200 deployed on compute node 210 as an example.
[0300] For example, Figure 9 This is a flowchart illustrating the process of a key-value store system executing a range request, such as... Figure 9As shown, taking a tenant accessing the key-value storage system through client 110 as an example, the tenant specifies compute node 210 to handle range requests.
[0301] First, the tenant sends a range request to compute node 210 via client 110, such as... Figure 9 As shown in step 1, the range request includes the query key input by the tenant. The key in the range request includes the version identifier (first version identifier) corresponding to that key.
[0302] Compute node 210 can query the storage cluster 300 for the version identifier (second version identifier) of the key-value pairs within a specific range corresponding to the Key in the Range request. Then, it determines whether the first version identifier is a higher version identifier based on the second version identifier, and then handles the range request in two different ways, such as... Figure 9 As shown in step 2 of the document.
[0303] If the first version identifier is an older version identifier, compute node 210 can further process range requests in two ways, depending on whether the first version identifier has undergone compression by storage cluster 300: Figure 9 As shown in step 3 of the process.
[0304] If the first version identifier has already been compressed by the storage cluster 300, then compute node 210 can return the range request processing result to client 110. The processing result of the above request may include a message stating "Range request failed because the queried data has been compressed".
[0305] If the first version identifier has not been compressed by storage cluster 300, compute node 210 can create a read transaction based on the key in the range request and submit the read transaction to storage cluster 300 so that storage cluster 300 can execute the read transaction and return key-value pairs within a specific range corresponding to the key that is the same as the first version identifier, such as... Figure 9 As shown in step 5 of the process.
[0306] Finally, compute node 210 can send the processing result of the Range request to client 110, such as... Figure 9 As shown in step 6 of the document.
[0307] If the first version identifier is the version identifier of the new version, compute node 210 can create a read transaction based on the Key in the Range request and submit the read transaction to storage cluster 300. Storage cluster 300 will then execute the read transaction and return key-value pairs within a specific range corresponding to the Key with the same first version identifier, such as... Figure 9 As shown in step 3`.
[0308] Then, compute node 210 can send the Value within the aforementioned specific range to client 110, such as... Figure 9 As shown in step 5`.
[0309] Thirdly, based on the content described above, the service layer of the key-value storage system provided in the embodiments of this application will be introduced. It is understood that this service layer is proposed based on the content described above, and some or all of its content can be found in the description above.
[0310] For example, Figure 10 This is a schematic diagram of the service layer structure of a key-value storage system proposed in this application, as shown below. Figure 10 As shown, the service layer 600 includes a data processing module 601, a lease management module 602, and a monitoring module 603. It is understood that the above service layer can be deployed on at least one computing node in a key-value storage system.
[0311] The following sections will introduce the data processing module 601, the lease management module 602, and the monitoring module 603.
[0312] The data processing module 601 is used to receive a data processing request sent by at least one tenant, and send the processing result of the data processing request to at least one tenant. The data processing request includes a key or key-value pair input by the first tenant.
[0313] The data processing module 601 is further configured to create a corresponding transaction based on the data processing request and generate a pushdown operator related to the data processing request. The pushdown operator includes at least one sub-request, and each sub-request requests the storage cluster 300 to perform a corresponding operation. The aforementioned transaction instructs the storage cluster 300 to process at least one sub-request within the pushdown operator.
[0314] After the storage cluster 300 completes the above transactions, the data processing module 601 can also receive the transaction processing results returned by the storage cluster 300.
[0315] In some possible implementations, the key or key-value pair may also include the tenant identifier of the first tenant. The data processing module 601 is also used to create a corresponding transaction according to the data processing request and submit the transaction to the storage cluster 300 so that the storage cluster 300 executes the transaction according to the tenant identifier of the first tenant.
[0316] The data processing module 601 is also used to determine the traffic quota of the first tenant based on the tenant identifier of the first tenant.
[0317] If the traffic quota of the first tenant exceeds the threshold, the data processing module 601 can return a processing result of data processing request failure to the client connected to the first tenant.
[0318] If the first tenant's traffic quota does not exceed the threshold, the data processing module 601 can create a transaction corresponding to the data processing request based on the key or key-value pair input by the first tenant.
[0319] The lease management module 602 is used to obtain the lease information of the first tenant from the storage cluster 300.
[0320] As one possible implementation, when the data processing request is a write request, the write request includes a first key-value pair input by the first tenant, and the data processing module 601 is also used to create a write transaction corresponding to the data processing request based on the first key-value pair.
[0321] In some possible implementations, the storage cluster 300 pre-stores an older version of the second key-value pair associated with the first key-value pair, when the second key-value pair includes the lease information of the first tenant.
[0322] In this case, when the data processing module 601 creates a write transaction, it is also used to send the key in the first key-value pair to the storage cluster 300 through the lease management module 602, so that the storage cluster 300 can return the lease information of the first tenant based on the key in the first key-value pair.
[0323] The data processing module 601 can generate pushdown operators related to write transactions based on the lease information of the first tenant.
[0324] The above pushdown operators related to write transactions are specifically used for:
[0325] Request storage cluster 300 to modify the revision identifier field and version identifier field recorded in the first key-value pair.
[0326] The request is to replace the lease index associated with the first key-value pair in storage cluster 300. The lease index is determined based on the key in the first key-value pair and the lease identifier of the first tenant.
[0327] The request is made to storage cluster 300 to store the first key-value pair and generate a multi-version record for the first key-value pair. The key of the first key-value pair is determined based on the version identifier of the write transaction.
[0328] The system requests storage cluster 300 to generate monitoring information for the first key-value pair and to generate a version index for this monitoring information. The key of the monitoring information is determined based on the version identifier of the write transaction.
[0329] The request is made to storage cluster 300 to compress the second key-value pair and generate a compressed index for the second key-value pair. The key of the compressed index is determined based on the version identifier of the write transaction.
[0330] The data processing module 601 can send write transaction-related pushdown operators to the storage cluster 300, so that the storage cluster 300 can use the write transaction-related pushdown operators to process at least one sub-request in the pushdown operator and update the second key-value pair to the first key-value pair.
[0331] As one possible implementation, when the data processing request is a data deletion request, the data processing request includes a key entered by the first tenant, and the data processing module 601 is also used to create a deletion transaction corresponding to the data processing request based on the key entered by the first tenant.
[0332] In some possible implementations, the storage cluster 300 pre-stores a specific range of third key-value pairs related to the data processing request, when the third key-value pairs include the lease information of the first tenant.
[0333] In this case, the data processing module 601 is also used to create a deletion transaction corresponding to the data processing request based on the key entered by the first tenant, and send the key entered by the first tenant to the storage cluster 300 so that the storage cluster 300 can return the lease information of the first tenant based on the key entered by the first tenant.
[0334] The data processing module 601 can generate pushdown operators related to the deletion transaction based on the lease information of the first tenant.
[0335] The pushdown operators related to the above-mentioned deletion transactions are specifically used for:
[0336] The request is to add / delete a record in the third key-value pair in storage cluster 300. The key of the third key-value pair is determined based on the version identifier of the deletion transaction.
[0337] The request is to replace the lease index associated with the third key-value pair in storage cluster 300. The lease index is determined based on the key in the third key-value pair and the lease identifier of the first tenant 10.
[0338] The request is to compress the third key-value pair in storage cluster 300 and generate a compressed index of the third key-value pair. The key of the compressed index is determined based on the version identifier of the deletion transaction.
[0339] The storage cluster 300 is requested to generate monitoring information related to a third key-value pair, and to generate a version index for the monitoring information. The key of the monitoring information is determined based on the version identifier of the deletion transaction.
[0340] The data processing module 601 can send the pushdown operator related to the deletion transaction to the storage cluster 300, so that the storage cluster 300 can use the pushdown operator related to the deletion transaction to process at least one sub-request in the pushdown operator and delete the third key-value pair.
[0341] The monitoring module 603 is used to provide an API interface for key-value content changes to at least one tenant.
[0342] In some possible implementations, when the content of the first key-value pair changes, the monitoring module 603 obtains the monitoring information of the first key-value pair from the storage cluster 300 and sends the monitoring information of the first key-value pair to at least one computing node (first computing node) that has subscribed to the monitoring information, so that the first computing node updates the index file of the first key-value pair according to the monitoring information of the first key-value pair.
[0343] When the content of the third key-value pair changes, the monitoring module 603 is also used to obtain the monitoring information of the third key-value pair from the storage cluster 300, and send the monitoring information of the third key-value pair to at least one computing node (second computing node) that has subscribed to the monitoring information, so that the second computing node updates the index file of the third key-value pair according to the monitoring information of the third key-value pair.
[0344] As one possible implementation, service layer 600 can also process data processing requests sent by other tenants.
[0345] by Figure 2 Taking tenant 20 (the second tenant) as an example, the data processing module 601 is used to receive a data processing request sent by the second tenant. The second tenant 20 is any one of multiple tenants other than the first tenant 10, and the data processing request includes the tenant identifier of the second tenant 20.
[0346] The data processing module 601 is used to create a transaction corresponding to the data processing request based on the tenant identifier of the second tenant, and generate pushdown operators related to the data processing request.
[0347] The data processing module 601 is also used to send the pushdown operator to the storage cluster 300 and submit the transaction to the storage cluster 300.
[0348] The data processing module 601 is used to receive the processing result of the above transaction returned by the storage cluster 300, and send the processing result to the second tenant 20 according to the tenant identifier of the second tenant 20.
[0349] In some possible implementations, the service layer 600 may also include a locking module 604, an authentication module 605, and an election module 606, such as... Figure 10 The dashed line portion is shown in the image.
[0350] The locking module 604 provides data locking functionality for at least one tenant. The data processing module 212 can utilize the locking module 604 to ensure data consistency and validity when tenants concurrently access the storage cluster 300. The types of data locks can include global locks, table-level locks, and row-level locks.
[0351] The authentication module 605 is used to provide data security functions to at least one tenant. The data processing module 212 can use the authentication and authorization mechanism of the authentication module 605 to send a query command to the storage cluster 300 in order to obtain the authentication information of the tenant from the storage cluster 300, and determine whether the tenant is legitimate and authorized based on the authentication information, so that the tenant can access the data in the storage cluster 300.
[0352] The election module 606 is used for the election mechanism compatible with the Raft algorithm. The Raft algorithm will not be described in detail here.
[0353] When the above-mentioned module is used as an example of a hardware functional unit, the module may include at least one computing device, such as a server. Alternatively, the module may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0354] The multiple computing devices included in a module can be distributed within the same region or in different regions. Similarly, the multiple computing devices included in a module can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in a module can be distributed within the same Virtual Private Cloud (VPC) or multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0355] It should be noted that, in other embodiments, the data processing module 601 can be used to perform... Figure 4The data processing method of the key-value storage system shown can be executed by any step. The lease management module 602 and the monitoring module 603 can be used to execute any step of the data processing method of the key-value storage system. The steps implemented by the data processing module 601, the lease management module 602, and the monitoring module 603 can be specified as needed. By implementing different steps in the data processing method of the key-value storage system through the data processing module 601, the lease management module 602, and the monitoring module 603, all functions of the service layer 600 can be realized.
[0356] Fourthly, based on the content described above, another service layer of the key-value storage system provided in the embodiments of this application will be introduced. It is understood that this service layer is proposed based on the content described above, and some or all of its content can be found in the description above.
[0357] For example, Figure 11 This is a schematic diagram of another service layer structure of the key-value storage system proposed in this application, such as... Figure 11 As shown, the implementation of service layer 700 can refer to the implementation of service layer 600, including data processing module 601.
[0358] The data processing module 601 is used to receive a query request sent by the first tenant. The query request includes a key entered by the first tenant, and the key entered by the first tenant includes a first version identifier corresponding to the key.
[0359] The data processing module 601 is also used to send the key input by the first tenant to the storage cluster 300 in order to obtain the second version identifier from the storage cluster 300. The storage cluster 300 pre-stores multiple versions of key-value pairs, which are key-value pairs within a specific range related to the query request. The second version identifier is the version identifier of the higher version key-value pair among the multiple versions.
[0360] In some possible implementations, if the first version identifier is less than the second version identifier, the data processing module 601 can obtain the third version identifier from the storage cluster 300. The third version identifier is the version identifier of the lower version key-value pair among multiple versions.
[0361] If the first version identifier is greater than or equal to the third version identifier, the data processing module 601 can create a read transaction corresponding to the query request based on the key input by the first tenant 10.
[0362] As one possible implementation, if the first version identifier is equal to the second version identifier, the data processing module 601 can also create a read transaction corresponding to the query request based on the key input by the first tenant 10.
[0363] In some possible implementations, after creating a read transaction, the data processing module 601 is further used to submit the read transaction to the storage cluster 300 so that the storage cluster 300 can return the processing result of the read transaction. The processing result of the read transaction includes a fourth key-value pair, which is a key-value pair in the storage cluster 300 that conforms to the first version identifier.
[0364] In some possible implementations, if the first version identifier is less than the third version identifier, the data processing module 601 may return a query request failure processing result to the first tenant.
[0365] It should be noted that, in other embodiments, the data processing module 601 can be used to perform... Figure 5 This describes any step in the query method of the key-value storage system. The data processing module 601 implements different steps in the query method of the key-value storage system to achieve all the functions of the service layer 700.
[0366] Fifthly, this application also provides a computing device 800. For example... Figure 12 As shown, the computing device 800 includes a processor 801, a memory 802, a communication interface 803, and a bus 804. The processor 801, memory 802, and communication interface 803 communicate with each other via the bus 804. The computing device 800 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 800.
[0367] The 804 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 12 The bus 104 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 104 may include a path for transmitting information between various components of the computing device 800 (e.g., memory 802, processor 801, communication interface 803).
[0368] Processor 801 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0369] The memory 802 may include volatile memory, such as random access memory (RAM). The processor 801 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0370] The memory 802 stores executable program code, and the processor 801 executes this executable program code to implement the functions of the aforementioned data processing module 601, lease management module 602, and monitoring module 603, thereby achieving... Figure 4 The method shown. That is, the memory 802 stores the information for executing... Figure 5 The instructions for the method shown.
[0371] The communication interface 803 uses modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 800 and other devices or communication networks. In this embodiment, the communication interface 803 may include communication interfaces implemented using the Raft protocol, the remote procedure call (gRPC) protocol, and the HTTP protocol.
[0372] Sixthly, embodiments of this application also provide a computing device cluster. Figure 13 This application provides a computing device cluster 900.
[0373] The computing device cluster 900 may include at least one computing device 800 as described above. This computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0374] like Figure 13 As shown, the computing device cluster 900 includes at least one computing device 800. The memory 802 of one or more computing devices 800 in the computing device cluster may store the same memory for executing... Figure 4 and Figure 5 The instructions for the method shown.
[0375] In some possible implementations, the memory 802 of one or more computing devices 800 in the computing device cluster may also store data for execution. Figure 4 and Figure 5The instructions of the method shown. In other words, a combination of one or more computing devices 800 can jointly execute instructions for performing... Figure 4 and Figure 5 The instructions for the method shown.
[0376] It should be noted that the memory 802 in different computing devices 800 within the computing device cluster can store different instructions, which are used to execute some functions of service layer 600 and service layer 700 respectively. That is, the instructions stored in the memory 802 of different computing devices 800 can implement the functions of one or more modules among data processing module 601, lease management module 602, and monitoring module 603.
[0377] In some possible implementations, one or more computing devices in the computing device cluster 900 can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 14 One possible implementation is shown. For example... Figure 14 As shown, the two computing devices 800A and 800B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this possible implementation, the memory 802 in computing device 800A stores instructions for executing the functions of the data processing module 601 and the lease management module 602. Meanwhile, the memory 802 in computing device 800B stores instructions for executing the functions of the monitoring module 603. Figure 13 The connection method between the computing device clusters shown can be based on the fact that the data processing method of the key-value storage system provided in this application requires a large amount of storage and computing resources (e.g., storing a large amount of data). Therefore, it is considered that the functions implemented by the data processing module 601 and the lease management module 602 are executed by computing device 800A, and the functions implemented by the monitoring module 603 are executed by computing device 800B. It should be understood that... Figure 14 The functions of the computing device 800A shown can also be performed by multiple computing devices 800. Similarly, the functions of the computing device 800B can also be performed by multiple computing devices 800.
[0378] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 13 and Figure 14 The diagram illustrates the connection method of a computing device cluster. The difference is that the memory 802 of one or more computing devices 800 in this cluster can store the same information used for execution. Figure 4 The data processing method of the key-value storage system shown Figure 5 The instructions for the query method of the key-value storage system are shown.
[0379] In some possible implementations, the memory 802 of one or more computing devices 800 in the computing device cluster may also store data for execution. Figure 4 The data processing method of the key-value storage system shown Figure 5 The diagram shows a portion of the instructions for a query method in a key-value storage system. In other words, a combination of one or more computing devices 800 can jointly execute instructions for performing... Figure 4 The data processing method of the key-value storage system shown Figure 5 The instructions for the query method of the key-value storage system are shown.
[0380] In addition to the methods, service layers, and electronic devices described above, embodiments of this application may also provide a computer program product, comprising computer program instructions. When executed by a processor, the computer program instructions cause the processor to perform the steps of the methods described in the "Methods" section of this specification. The computer program product can be written in any combination of one or more programming languages to execute the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The computer program code can be in source code form, object code form, executable file, or some intermediate form. The computer program code can be executed entirely on a user's computing device, partially on a user's device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0381] Furthermore, embodiments of this application may also provide a computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the methods described in the "Method" section of this specification according to the various embodiments of this disclosure. The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. It should be noted that the content contained in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0382] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0383] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0384] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of the various embodiments of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the specific details described above.
[0385] The block diagrams of devices, service layers, equipment, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that connections, arrangements, or service layers must be made in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, service layers, equipment, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0386] It should also be noted that in the service layer, apparatus, and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.
[0387] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
[0388] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. The specific embodiments described above have further detailed the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
[0389] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A data processing method for a key-value storage system, characterized in that, The method is applied to a service layer, which is deployed on at least one compute node in a key-value storage system. The key-value storage system further includes a storage cluster, and the at least one compute node is communicatively connected to the storage cluster. The method includes: Receive data processing requests sent by the first tenant; A corresponding transaction is created based on the data processing request, and a pushdown operator related to the data processing request is generated. The pushdown operator includes at least one sub-request, and each of the at least one sub-request is used to request the storage cluster to perform a corresponding operation. The pushdown operator is sent to the storage cluster, and the transaction is submitted to the storage cluster, the transaction being used to instruct the storage cluster to process the at least one sub-request using the pushdown operator; The system receives the transaction processing result returned by the storage cluster and sends the processing result to the first tenant.
2. The method according to claim 1, characterized in that, The data processing request includes the tenant identifier of the first tenant, and the step of creating a corresponding transaction based on the data processing request and generating a pushdown operator related to the data processing request includes: The traffic quota for the first tenant is determined based on the tenant identifier of the first tenant; If the traffic quota of the first tenant exceeds the threshold, a data processing request failure result is returned to the first tenant; If the traffic quota of the first tenant does not exceed the threshold, a corresponding transaction is created according to the data processing request, and a pushdown operator related to the data processing request is generated.
3. The method according to claim 1, characterized in that, The data processing request includes a first key-value pair, which is a key-value pair input by the first tenant. The storage cluster pre-stores a second key-value pair including the first tenant's lease information. The second key-value pair is an older version of the key-value pair related to the first key-value pair. The step of creating a corresponding transaction based on the data processing request and generating a pushdown operator related to the data processing request includes: A write transaction corresponding to the data processing request is created based on the first key-value pair, and the key in the first key-value pair is sent to the storage cluster so that the storage cluster can return the lease information of the first tenant based on the key in the first key-value pair. The write transaction-related pushdown operator is generated based on the lease information of the first tenant; The write transaction-related pushdown operator is sent to the storage cluster so that the storage cluster can use the write transaction-related pushdown operator to process the at least one sub-request and update the second key-value pair to the first key-value pair.
4. The method according to claim 3, characterized in that, The pushdown operators related to write transactions include: The first sub-request is used to request the storage cluster to modify the revision identifier field and version identifier field recorded in the first key-value pair; The second sub-request requests the storage cluster to replace the lease index associated with the first key-value pair, the lease index being determined based on the key in the first key-value pair and the lease identifier of the first tenant; The third sub-request is used to request the storage cluster to store the first key-value pair and generate a multi-version record of the first key-value pair, wherein the key of the first key-value pair is determined based on the version identifier of the write transaction. The fourth sub-request is used to request the storage cluster to generate monitoring information for the first key-value pair and to generate a version index for the monitoring information, wherein the key of the monitoring information is determined based on the version identifier of the write transaction; The fifth sub-request is used to request the storage cluster to compress the second key-value pair and generate a compressed index of the second key-value pair, wherein the key of the compressed index is determined based on the version identifier of the write transaction.
5. The method according to claim 4, characterized in that, After sending the pushdown operator to the storage cluster and submitting the transaction to the storage cluster, the method further includes: Obtain monitoring information for the first key-value pair from the storage cluster; The monitoring information of the first key-value pair is sent to the first computing node so that the first computing node updates the index file of the first key-value pair according to the monitoring information of the first key-value pair. The first computing node is at least one computing node that has subscribed to the monitoring information of the first key-value pair.
6. The method according to claim 1, characterized in that, The data processing request includes a key input by the first tenant. The storage cluster pre-stores a third key-value pair including the first tenant's lease information. The third key-value pair is a specific range of key-value pairs related to the data processing request. The step of creating a corresponding transaction based on the data processing request and generating a pushdown operator related to the data processing request further includes: A deletion transaction corresponding to the data processing request is created based on the key entered by the first tenant, and the key entered by the first tenant is sent to the storage cluster so that the storage cluster can return the lease information of the first tenant based on the key entered by the first tenant. The pushdown operator related to the deletion transaction is generated based on the lease information of the first tenant; The pushdown operator related to the deletion transaction is sent to the storage cluster so that the storage cluster can use the pushdown operator related to the deletion transaction to process the at least one sub-request and delete the third key-value pair.
7. The method according to claim 6, characterized in that, The pushdown operator related to the deletion transaction includes: The sixth sub-request is used to request the storage cluster to add a deletion record to the third key-value pair, wherein the key of the third key-value pair is determined based on the version identifier of the deletion transaction; The seventh sub-request requests the storage cluster to replace the lease index associated with the third key-value pair, the lease index being determined based on the key in the third key-value pair and the lease identifier of the first tenant; The eighth sub-request is used to request the storage cluster to compress the third key-value pair and generate a compressed index of the third key-value pair, wherein the key of the compressed index is determined based on the version identifier of the deletion transaction; The ninth sub-request is used to request the storage cluster to generate monitoring information related to the third key-value pair and to generate a version index of the monitoring information, wherein the key of the monitoring information is determined based on the version identifier of the deletion transaction.
8. The method according to claim 7, characterized in that, The method further includes: The monitoring information of the third key-value pair is obtained from the storage cluster; The monitoring information of the third key-value pair is sent to the second computing node so that the second computing node updates the index file of the third key-value pair according to the monitoring information of the third key-value pair. The second computing node is at least one computing node that has subscribed to the monitoring information of the third key-value pair.
9. The method according to claim 1, characterized in that, The method further includes: Receive a data processing request sent by a second tenant; the second tenant is any one of a plurality of tenants other than the first tenant, and the data processing request includes the tenant identifier of the second tenant; Create a transaction corresponding to the data processing request based on the tenant identifier, and generate a pushdown operator related to the data processing request; The pushdown operator is sent to the storage cluster, and the transaction is submitted to the storage cluster. The system receives the transaction processing result returned by the storage cluster and sends the processing result to the second tenant according to the tenant identifier.
10. A query method for a key-value storage system, characterized in that, The method is applied to a service layer, which is deployed on at least one compute node in a key-value storage system. The key-value storage system further includes a storage cluster, and the at least one compute node is communicatively connected to the storage cluster. The method includes: Receive a query request sent by a first tenant, the query request including a key entered by the first tenant, the key entered by the first tenant including a first version identifier corresponding to the key; Send the key input by the first tenant to the storage cluster to obtain a second version identifier from the storage cluster. The storage cluster pre-stores multiple versions of key-value pairs, which are key-value pairs of a specific range related to the query request. The second version identifier is the version identifier of the higher version key-value pair among the multiple versions of key-value pairs. If the first version identifier is less than the second version identifier, then the third version identifier is obtained from the storage cluster. The third version identifier is the version identifier of the lower version key-value pair among the multiple versions. If the first version identifier is greater than or equal to the third version identifier, a read transaction corresponding to the query request is created based on the key entered by the first tenant.
11. The method according to claim 10, characterized in that, The method further includes: If the first version identifier is equal to the second version identifier, then a read transaction corresponding to the query request is created based on the key entered by the first tenant.
12. The method according to any one of claims 10 and 11, characterized in that, After creating the read transaction corresponding to the query request based on the key input by the first tenant, the method further includes: The read transaction is submitted to the storage cluster so that the storage cluster returns the processing result of the read transaction, the processing result of the read transaction including a fourth key-value pair, the fourth key-value pair being a key-value pair in the storage cluster that conforms to the first version identifier.
13. The method according to claim 10, characterized in that, The method further includes: If the first version identifier is less than the third version identifier, then a query request failure result is returned to the first tenant.
14. A service layer of a key-value storage system, characterized in that, The service layer is deployed on at least one computing node in the key-value storage system, and the service layer includes: A data processing module is configured to receive a data processing request sent by a first tenant among multiple tenants; create a corresponding transaction based on the data processing request, and generate a pushdown operator related to the data processing request, wherein the pushdown operator includes at least one sub-request, each of the at least one sub-request being used to request the storage cluster to perform a corresponding operation; send the pushdown operator to the storage cluster, and submit the transaction to the storage cluster; the transaction is used to instruct the storage cluster to process the at least one sub-request using the pushdown operator; receive the processing result of the transaction returned by the storage cluster, and send the processing result of the transaction to the first tenant.
15. The service layer according to claim 14, characterized in that, The data processing request includes the tenant identifier of the first tenant, and the data processing module is used for: The traffic quota for the first tenant is determined based on the tenant identifier of the first tenant; If the traffic quota of the first tenant exceeds the threshold, a data processing request failure result is returned to the first tenant; If the traffic quota of the first tenant does not exceed the threshold, then a transaction corresponding to the data processing request is created based on the key or key-value pair entered by the first tenant.
16. The service layer according to claim 14, characterized in that, The service layer also includes: The lease management module is used to obtain the lease information of the first tenant from the storage cluster; The data processing request includes a first key-value pair, which is a key-value pair input by the first tenant. The storage cluster pre-stores a second key-value pair including the first tenant's lease information. The second key-value pair is an older version of the key-value pair related to the first key-value pair. The data processing module is used to: A write transaction corresponding to the data processing request is created based on the first key-value pair. The key in the first key-value pair is sent to the storage cluster through the lease management module so that the storage cluster can return the lease information of the first tenant based on the key in the first key-value pair. The write transaction-related pushdown operator is generated based on the lease information of the first tenant; The write transaction-related pushdown operator is sent to the storage cluster so that the storage cluster can use the write transaction-related pushdown operator to execute the at least one sub-request and update the second key-value pair to the first key-value pair.
17. A service layer of a key-value storage system, characterized in that, The service layer is deployed on at least one computing node in the key-value storage system, and the service layer includes: A data processing module is configured to receive a query request sent by a first tenant; the query request includes a key input by the first tenant, which includes a first version identifier corresponding to the key; send the key input by the first tenant to the storage cluster to obtain a second version identifier from the storage cluster, wherein the storage cluster pre-stores multiple versions of key-value pairs, which are key-value pairs within a specific range related to the query request, and the second version identifier is the version identifier of the higher version key-value pair among the multiple versions; if the first version identifier is less than the second version identifier, obtain a third version identifier from the storage cluster, which is the version identifier of the lower version key-value pair among the multiple versions; if the first version identifier is greater than or equal to the third version identifier, create a read transaction corresponding to the query request based on the key input by the first tenant.
18. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method according to any one of claims 1-9 or 10-13.
19. A computer-readable storage medium, characterized in that, Includes computer program instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method as described in any one of claims 1-9 or 10-13.
20. A computer program product, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1-9 or 10-13.