Rocksdb-based distributed relational database time-to-live table implementation method

By using RocksDB's CompactionFilter and cleanTs mechanisms in distributed relational databases, efficient cleaning of expired data is achieved, the problem of lack of effective TTL cleaning solutions in the existing technology is solved, and the performance and availability of the system are improved.

WO2025124288A1PCT designated stage expired Publication Date: 2025-06-19CHINA TELECOM CLOUD TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/137278
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-12-06
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

In existing distributed relational databases, efficient expiration timetables (TTL) cleaning solutions are lacked, especially in the environment using the RocksDB engine, and the TTL cleaning cannot be effectively used for RocksDB Compaction function.

Method used

A distributed relational database expiration timetable implementation method based on RocksDB is proposed. TTL table is created by the client, the server layer calculates the data range and registers cleaning tasks, and uses CompactionFilter to register cleanTs in RocksDB and pulls cleanTs regularly, and performs Compaction operations to clean up expired data.

Benefits of technology

It realizes efficient cleaning of expired data in distributed relational databases, reduces the complexity of the system and disk IO usage, and improves the usage experience of TTL tables and data cleaning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137278_19062025_PF_FP_ABST
    Figure CN2024137278_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application is a rocksdb-based distributed relational database Time-To-Live (TTL) table implementation method, specifically comprising: creating a TTL table; a server layer calculating the data range of the TTL table, and registering a TTL cleanup task in an engine layer; the server layer starting a timed task, and calculating and advancing cleanTs according to the TTL; the engine layer receiving the TTL task, registering CompactionFilter in rocksdb, and regularly pulling cleanTs of the current task data from a calculating node; by means of a snapshot consistent read method of cleanTs, causing RocksDB to execute Compaction, and calling CompactionFilter to clean data in the range; when a regular check of the server layer based on statistical data requires active Compaction, initiating CompactionRange and CompactionFiles instructions; and, after the engine layer receives an active Compaction instruction, RocksDB carrying out Compaction and deleting the data according to the CompactionFilter, thereby completing data cleaning.
Need to check novelty before this filing date? Find Prior Art

Description

A method for implementing an expiration schedule for a distributed relational database based on rocksdb

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 202311712700.9, filed on December 13, 2023, entitled “A method for implementing an expiration schedule for a distributed relational database based on rocksdb,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] This application belongs to the field of distributed relational database, database kernel, and storage engine design. Specifically, it is a method for implementing an expiration schedule for a distributed relational database based on RocksDB. Background Art

[0004] RocksDB is a KV storage engine with an LSM-Tree architecture. Its advantage of sequential writes greatly improves write performance. In recent years, as a persistent, high-performance storage engine, it has been widely used in the field of distributed relational databases. TiDB and CockroachDB both use RocksDB as their storage engine.

[0005] A Time to Live (TTL) table is a table with a data lifecycle. It allows you to set an expiration date for the data; after the expiration date, the data is automatically deleted. TiDB's TTL is a row-granular table. A background thread periodically scans TTL table data and deletes rows that have exceeded their expiration date. PolarDB-X uses a partitioned table solution that treats TTL tables as partitioned tables, partitioning them by range based on time periods. Once data expires, the partitions are deleted. This is a partitioned table solution with time-granularity.

[0006] For example, a Chinese patent with authorization announcement number CN111400331B discloses a processing method and device based on the TiDB database, which is characterized by receiving a data operation request, the data operation request being used to instruct a first operation to be performed on first data; performing the first operation on the first data according to the data operation request; determining whether the first operation fails to be performed; and if so, retrying the first operation at the application level. The processing method and device based on the TiDB database can retry the first operation at the application level after the data operation fails, thereby increasing the probability of successful execution of the first operation and reducing the possibility of failure to update hot data using the TiDB database under high concurrency conditions; and because the retry of the first operation is performed at the application level, the entire process is imperceptible to the user, thereby achieving an improvement in the success rate of data operations without the user's perception, and optimizing the user experience.

[0007] For example, Chinese patent application publication number CN116107806A discloses a database backup management method, system, device, and storage medium. This application includes: sorting out the database clusters that need to be backed up; classifying the database clusters and developing backup strategies for different types of database clusters; deploying backup software and database backup command scripts on all database node servers; completing the configuration of the backup software and backup command scripts; and utilizing the backup software's unified scheduling function to achieve the goal of integrated database backup management. Compared with related technologies, this application solves the problem of the PolarDB database's original backup command being unable to achieve integrated management. It can achieve integrated management of multiple database instances, greatly reducing the daily operation and maintenance workload of backup administrators.

[0008] The above-mentioned related technologies all have the following problems: 1) Whether it is the TTL table at the partition granularity of PolarDB-X or the TTL table at the row granularity of TiDB, there is a lack of a convenient usage method or an efficient solution for cleaning up expired data; 2) For distributed databases using the RocksDB engine, there is no solution for using RocksDB's compaction to clean up TTL. Summary of the Invention

[0009] According to various embodiments of the present application, the present application proposes a method for implementing an expiration schedule of a distributed relational database based on rocksdb.

[0010] To achieve the above objectives, this application provides the following technical solutions:

[0011] A method for implementing an expiration schedule for a distributed relational database based on RocksDB, comprising:

[0012] Step S1: After the client connects to the database, it creates a TTL table;

[0013] Step S2: The server layer calculates the TTL table data range and registers the TTL cleanup task with the engine layer;

[0014] Step S3: The server layer starts a scheduled task and increases cleanTs based on the TTL calculation. At the same time, the engine layer receives the TTL task, registers the CompactionFilter in RocksDB, and periodically pulls the cleanTs of the current task data from the computing node;

[0015] Step S4: Use the snapshot consistency read method of cleanTs to enable RocksDB to perform compaction and clean up the data within the range by calling CompactionFilter;

[0016] Step S5: The server layer periodically checks whether active compaction is needed based on statistical data, and initiates CompactionRange and CompactionFiles instructions if necessary; and

[0017] Step S6: After the engine layer receives the active compaction instruction, RocksDB performs compaction and deletes data according to the CompactionFilter to complete data cleanup.

[0018] Specifically, the information of the TTL table in step S1 includes: TTL type, data expiration time, and expiration time column.

[0019] Specifically, the specific steps of step S2 include:

[0020] Step S201: The server layer completes table creation, calculates the key range after TTL table data encoding, and writes the start time of the data into the key of the key-value pair;

[0021] Step S202: registering a cleanup task at the engine layer according to the TTL table data range; and

[0022] Step S203: Send an instruction to the engine table indicating that the data within the key range belongs to the TTL table and needs to be cleaned up by compaction.

[0023] Specifically, the specific steps of step S3 include:

[0024] Step S301: define data model and TTL;

[0025] Step S302: Start a scheduled task at the server layer to periodically increase cleanTs based on TTL calculation;

[0026] Step S303: Through the background thread, on each computing node, the cleanTs of the current task data is periodically pulled;

[0027] Step S304: On each computing node, upon receiving a new TTL task, register a CompactionFilter in RocksDB; and

[0028] Step S305: When the data expires, the registered CompactionFilter will be triggered.

[0029] Specifically, the method further includes:

[0030] Trigger CompactionFilter to filter and delete expired data.

[0031] Specifically, triggering CompactionFilter to filter and delete expired data also includes:

[0032] Parse TTL information from keys;

[0033] Compare the TTL information of the data with the size of cleanTs; and

[0034] When TTL is less than cleanTs, the key-value filter is filtered and deleted.

[0035] Specifically, the snapshot consistency reading method of cleanTs in step S4 includes:

[0036] Step S401: When new data is inserted or existing data is updated, the TTL is used as part of the key and stored together with the data in RocksDB;

[0037] Step S402: When the cleanTs value reaches or exceeds the set expiration time, a data cleanup operation is performed;

[0038] Step S403: Merge the data stored in multiple layers into one layer, and delete expired or no longer needed data;

[0039] Step S404: Create a class that inherits from CompactionFilter and rewrite the Filter method to determine whether each key-value pair is expired;

[0040] Step S405: Start the database, register and use the custom CompactionFilter, and automatically call the custom Filter method; and

[0041] Step S406: Use the snapshot function of RocksDB to obtain a timestamp before performing the read operation, and use this timestamp to obtain a data snapshot. If expired data is detected, call the delete API of RocksDB to delete the expired data.

[0042] Specifically, the earliest readable range of the current TTL table is [cleanTs, readTs], where readTs represents the node read by the current snapshot.

[0043] Specifically, the factors for determining whether active compaction is required in step S5 include: data volume statistics exceeding a threshold, large data expiration time statistics, and frequent read and write operations with large data volumes.

[0044] Specifically, the steps of initiating the CompactionRange and CompactionFiles instructions in step S5 include:

[0045] Step S501: Call the CompactionRange and CompactionFiles functions of RocksDB and specify the range to be operated on;

[0046] Step S502: Obtain the relevant range files and the list of files requiring compaction operations by reading the data files and index files in RocksDB; and

[0047] Step S503: Initiate the CompactionRange and CompactionFiles instructions to pass the file to the RocksDB Compaction function for processing.

[0048] Specifically, the method further includes:

[0049] By periodically scanning TTL data keys, the conditions for triggering CompactRange are generated.

[0050] Specifically, the specific steps of RocksDB performing compaction in step S6 include:

[0051] Step S601: Determine the order of compaction operations based on the scores of each layer, start the compaction operation, and wait for thread scheduling.

[0052] Step S602: RocksDB splits the compaction into multiple subcompactions, processes them through child threads, and places them into an input array to form an SST file.

[0053] Step S603: traverse the multi-way SST file through mergeIterator, sort the keys in the multi-way file using the minimum heap method, and then take out the top element of the heap each time;

[0054] Step S604: Create an output file, process subcompactions in a child thread, and merge the results into the final output file; and

[0055] Step S605: Add a task to the thread pool and wait for scheduling. After all subcompactions are processed, the entire compaction process is completed.

[0056] Specifically, the SST file is a file used to persist database data.

[0057] Specifically, a method for implementing an expiration schedule for a distributed relational database based on RocksDB includes: a data storage module, an expiration management module, a compaction module, a consistent read module, and a monitoring and logging module;

[0058] The data storage module uses RocksDB as the backend storage engine and is responsible for storing and reading data;

[0059] The expiration time management module is used to set the expiration policy for each data and manage the expiration time of the data;

[0060] The Compaction module is used to perform the compaction operation of RocksDB;

[0061] The consistent read module uses the snapshot function of RocksDB to handle data pre-fetching and caching; and

[0062] The monitoring and logging module is used to monitor the system's CPU usage and disk I / O indicators in real time, and to detect and handle problems in a timely manner.

[0063] Specifically, the compaction module includes: a compaction trigger unit, a compaction scheduler unit, a compaction worker unit, a compaction filter unit, and a compaction progress monitor unit;

[0064] The compaction trigger unit is used to detect data expiration time and trigger a compaction operation;

[0065] The compaction scheduler unit is used to reasonably arrange the time and priority of compaction tasks according to the system load and data expiration policy;

[0066] The Compaction Worker unit is responsible for performing the actual compaction operation, obtaining the data files to be merged, and performing the merge and delete operations;

[0067] The Compaction Filter unit is used to check the expiration timestamp of the data, identify and delete the expired data, and retain the non-expired data; and

[0068] The Compaction Progress Monitor unit is responsible for monitoring and reporting the progress of the compaction task and regularly updating the task status and progress information. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without creative work.

[0070] FIG1 is a flow chart of a method for implementing an expiration schedule of a distributed relational database based on rocksdb in this application;

[0071] FIG2 is a flowchart of a compaction filter for filtering and cleaning expired data in a distributed relational database expiration schedule implementation method based on RocksDB in this application;

[0072] FIG3 is a flowchart of a method for implementing an expiration schedule of a distributed relational database based on rocksdb in this application to increase cleanTs consistency;

[0073] FIG4 is a diagram of the TTL architecture for clearing expired data in a distributed relational database expiration schedule implementation method based on RocksDB in this application;

[0074] Figure 5 is a system architecture diagram of a method for implementing an expiration schedule of a distributed relational database based on rocksdb in this application. DETAILED DESCRIPTION

[0075] In order to make the technical means, creative features, objectives and effects achieved by this application easy to understand, it should be noted that in the description of this application, the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside" and the like indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, which are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to has a specific direction, is constructed and operated in a specific direction, and therefore cannot be understood as a limitation on this application. In addition, the terms "No. 1", "No. 2" and "No. 3" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance. The following further describes this application in conjunction with specific implementation methods.

[0076] In one embodiment:

[0077] Referring to Figures 1, 2, 3, and 4, this application provides an embodiment: a method for implementing an expiration schedule for a distributed relational database based on RocksDB, comprising the following steps:

[0078] Step S1: After the client connects to the database, it creates a TTL table;

[0079] TTL is a mechanism for setting the lifetime of a key in a key-value pair store. It refers to the survival time of data in the database. When a key-value pair is created, an expiration time can be set for it. After this time, the data will be considered expired and the database will automatically delete the key-value pair.

[0080] Common methods for implementing database expiration schedules are as follows:

[0081] (1) Use date type: Use date type to store expiration time in the database, and determine the expiration time by comparing date values;

[0082] (2) Set the key lifetime: Set the key lifetime in the key-value pair storage. When the key expires, the database will automatically delete the key-value pair.

[0083] (3) Use scheduled tasks: Use scheduled tasks to regularly check whether the data in the database is expired and delete expired data;

[0084] (4) Use triggers: Use triggers in the database to monitor data changes. When data changes, the trigger will determine whether to delete expired data based on preset rules;

[0085] (5) Use embedded cache: Combine the cache with the database and set the expiration time of the cache to delete the data when it expires;

[0086] (6) Use distributed locks: Use distributed locks to control access and deletion of data. When data expires, distributed locks are used to ensure timely deletion of data.

[0087] This application uses the TTL setting key to implement the database expiration schedule. It is simple and easy to use. There is no need to write additional logic in the application to detect whether the data is expired, nor is there a need to maintain a complex expiration schedule in the database. This reduces the burden on the application and database and provides strong predictability.

[0088] Step S2: The server layer calculates the TTL table data range and registers the TTL cleanup task with the engine layer;

[0089] Step S3: The server layer starts a scheduled task and increases cleanTs based on the TTL calculation. At the same time, the engine layer receives the TTL task, registers the CompactionFilter in RocksDB, and periodically pulls the cleanTs of the current task data from the computing node;

[0090] RocksDB is an efficient, high-performance, single-point database engine developed by Facebook based on Google's open-source key-value storage, LevelDB. It utilizes a log-structured database engine. RocksDB is suitable for tuning in various production environments. It can be used directly in memory, Flash, hard disks, or HDFS. It supports various compression algorithms and has a comprehensive suite of tools for production and debugging. It is also an embedded, persistent storage engine, featuring high performance, fast storage, and adaptability.

[0091] CompactionFilter is a pluggable filter that can be attached to the Compaction operation, allowing additional processing of data during the Compaction process, such as deleting expired data.

[0092] Step S4: Use the snapshot consistency read method of cleanTs to enable RocksDB to perform compaction and clean up the data within the range by calling CompactionFilter;

[0093] A snapshot consistency read is a read that retrieves a consistent view of data at a specific point in time to ensure that the data is logically correct. This typically involves prefetching and caching data. This application uses the cleanTs snapshot consistency read method to ensure read consistency, improve read performance, reduce network transmission overhead, and increase database availability.

[0094] In RocksDB, compaction is an important persistence operation responsible for organizing and compressing data stored on disk to free up space and improve query performance. The compaction operation usually scans all data on the disk and organizes it into a continuous data layer. Each layer contains a timestamp, and expired data will be located below newer data, making it easy to identify and clean it up.

[0095] Step S5: The server layer periodically checks whether active compaction is needed based on statistical data. If necessary, it initiates the CompactionRange and CompactionFiles instructions; and

[0096] Step S6: After the engine layer receives the active compaction instruction, RocksDB performs compaction and deletes data according to the CompactionFilter to complete data cleanup.

[0097] RocksDB compaction uses two methods: 1) a general trigger method with general configuration; 2) an active trigger method for TTL table data. The general trigger method is based on LSM-Tree in storage and will trigger the compaction process for data when the compaction conditions are met; the active trigger method is a method that is added to non-TTL tables and triggers compaction only for TTL tables. In the general compaction method, the compaction conditions may not be met, resulting in data not being deleted. This application uses the manual compaction provided by RocksDB to actively trigger the compaction process of TTL tables. Based on statistical information, it achieves refined compaction effects, which can balance functional requirements and performance. Manual compaction can use two methods: CompactRange and compactFiles.

[0098] Compared with the related art, the beneficial effects of this application are:

[0099] 1. This application proposes a method for implementing an expiration schedule for a distributed relational database based on RocksDB, and optimizes and improves the architecture, operating steps, and processes. The system has the advantages of simple processes, low investment and operating costs, and low production work costs. The purpose of clearing expired data is achieved by using the compaction filter function of RocksDB.

[0100] 2. This application proposes a method for implementing an expiration schedule for a distributed relational database based on RocksDB. It adopts the RocksDB compaction solution to clean up expired data. Data is cleaned when RocksDB performs compaction. There is no need to query expired data and then delete it, no need to perform data deletion operations, and no need to maintain expiration time column indexes or partitions. When the distributed database uses a consensus algorithm to achieve consistency among multiple copies, it reduces the consensus log, thereby significantly reducing disk IO usage and reducing the complexity of using TTL tables.

[0101] 3. This application proposes a method for implementing an expiration schedule for a distributed relational database based on RocksDB. It proposes using the RocksDB feature to actively trigger compaction to clean up data based on statistical data for different scenarios of range and sst files. At the same time, to address the problem of asynchronous compaction execution in various RocksDBs, a method using cleanTs to view consistent data from the relational database is used, which optimizes the user experience of the TTL table. It is simple to use, efficient in cleaning expired data, and uses low system resources.

[0102] The information of the TTL table in step S1 includes: TTL type, data expiration time, and expiration time column.

[0103] The specific steps of step S2 include:

[0104] Step S201: The server layer completes table creation, calculates the key range after TTL table data encoding, and writes the start time of the data into the key of the key-value pair;

[0105] Step S202: registering a cleanup task at the engine layer according to the TTL table data range; and

[0106] Step S203: Send an instruction to the engine table indicating that the data within the key range belongs to the TTL table and needs to be cleaned up by compaction.

[0107] The specific steps of step S3 include:

[0108] Step S301: define data model and TTL;

[0109] Step S302: Start a scheduled task at the Server layer to periodically push up cleanTs according to the TTL. Pushing up cleanTs means calculating and pushing up a specific cleanTs based on the TTL of data items, which is used to mark and delete expired data items. In a distributed relational database, each computing node needs to periodically pull the cleanTs of the current task data and register a CompactionFilter in RocksDB. When the data expires, the registered CompactionFilter will be triggered to perform specific operations, such as deleting or updating the expired data. The method of pushing up cleanTs adopted in this application can ensure that the data items in the database can be cleaned and updated in time after expiration, thus ensuring the efficiency and performance of the database.

[0110] Step S303: Through a background thread, on each computing node, periodically pull the cleanTs of the current task data;

[0111] Step S304: On each computing node, when a new TTL task is received, register a CompactionFilter in RocksDB; and

[0112] Step S305: When the data expires, the registered CompactionFilter will be triggered.

[0113] The steps of triggering the CompactionFilter to filter and delete the expired data include: 1) Parse the TTL information from the keys; 2) Compare the TTL of the data with the cleanTs; and 3) If TTL < cleanTs, filter and delete the key-value pair.

[0114] The snapshot consistency read method introducing cleanTs in step S4 includes:

[0115] Step S401: When new data is inserted or existing data is updated, store the TTL as part of the key together with the data in RocksDB;

[0116] Step S402: When the cleanTs value reaches or exceeds the set expiration time, perform data cleaning operations;

[0117] Step S403: Merge the data stored in multiple layers into one layer and delete the expired or unnecessary data;

[0118] Step S404: Create a class inherited from CompactionFilter and override the Filter method to determine whether each key-value pair has expired;

[0119] Step S405: Start the database, register and use the custom CompactionFilter, and automatically call the custom Filter method; and

[0120] Step S406: Use the snapshot function of RocksDB to obtain a timestamp before performing the read operation, and use this timestamp to obtain a data snapshot. If expired data is detected, call the delete API of RocksDB to delete the expired data.

[0121] The significance of cleanTs is that data older than this time is in a cleanable state. The time of cleanup is determined by the compaction of the current data triggered by the storage engine RocksDB. Therefore, data older than cleanTs may have been cleaned by RocksDB, or RocksDB may not have compacted the data yet and thus has not been cleaned yet.

[0122] In the snapshot consistency read of cleanTs, the conditions for determining whether the current TTL can be read are:

[0123] 1) Based on readTs, the current data TTL can be read. readTs represents the node read by the current snapshot.

[0124] 2) TTL is equal to or later than cleanTs;

[0125] Therefore, the earliest readable range of the current TTL table is [cleanTs, readTs]. For the aforementioned inconsistency issues of row records and index data, since the TTL is earlier than cleanTs, they will not be read, thus achieving consistent reading.

[0126] Factors for determining whether active compaction is required in step S5 include: data volume statistics exceeding a threshold, large data expiration time statistics, and frequent read and write operations with large data volumes.

[0127] The steps of initiating the CompactionRange and CompactionFiles instructions in step S5 include:

[0128] Step S501: Call the CompactionRange and CompactionFiles functions of RocksDB and specify the range to be operated on;

[0129] CompactFiles is a granular update method for triggering compactions, which complements CompactRange. Compared with CompactRange, CompactFiles is lightweight and can achieve finer control over compactions.

[0130] The CompactionRange function accepts two parameters, the start key and the end key. Within this range, the files that need to perform the compaction operation will be selected.

[0131] Step S502: Obtain the relevant range files and the list of files requiring compaction operations by reading the data files and index files in RocksDB; and

[0132] Step S503: Initiate the CompactionRange and CompactionFiles instructions to pass the file to the RocksDB Compaction function for processing.

[0133] CompactRange can be triggered based on TTL table statistics or database statistics. For example, by periodically scanning TTL data keys, you can customize the conditions for triggering CompactRange:

[0134] 1) The number of keys that need to be deleted in the scan range reaches a certain threshold;

[0135] 2) Within the scan range, the number of keys to be deleted exceeds a certain percentage of the total number of keys;

[0136] 3) Server-level statistics: data exceeding a certain threshold over the past period of time is deleted;

[0137] 4) Set periodic long time intervals and timed triggering.

[0138] The specific steps of RocksDB performing compaction in step S6 include:

[0139] Step S601: Determine the order of compaction operations based on the scores of each layer, start the compaction operation, and wait for thread scheduling.

[0140] Step S602: RocksDB splits the compaction into multiple subcompactions, processes them through child threads, and places them into an input array to form an SST file.

[0141] In the RocksDB-based distributed relational database expiration schedule implementation method, SST files are used to persist database data. SST files have multiple formats, with the default being BlockBasedTable, which, as the name suggests, is stored based on data blocks.

[0142] The advantages of using the SST file format in this application are: providing persistent storage, sequential writing, binary search, optimized storage space, and support for multiple data structures and auxiliary data blocks, which helps to improve the performance and availability of the database.

[0143] Step S603: traverse the multi-way SST file through mergeIterator, sort the keys in the multi-way file using the minimum heap method, and then take out the top element of the heap each time;

[0144] Step S604: Create an output file, process subcompactions in a child thread, and merge the results into the final output file; and

[0145] Step S605: Add a task to the thread pool and wait for scheduling. After all subcompactions are processed, the entire compaction process is completed.

[0146] In another embodiment:

[0147] Referring to FIG. 5 , another embodiment provided by the present application is a method for implementing an expiration schedule for a distributed relational database based on RocksDB, including:

[0148] Data storage module, expiration time management module, compaction module, consistent reading module, monitoring and logging module;

[0149] The data storage module uses RocksDB as the backend storage engine and is responsible for storing and reading data;

[0150] The expiration time management module is used to set the expiration policy for each data and manage the expiration time of the data;

[0151] The Compaction module is used to perform the compaction operation of RocksDB;

[0152] The consistent read module uses the snapshot function of RocksDB to handle data prefetching and caching;

[0153] The monitoring and logging module is used to monitor the system's CPU usage and disk I / O indicators in real time, and to detect and handle problems in a timely manner.

[0154] The compaction module includes: a compaction trigger unit, a compaction scheduler unit, a compaction worker unit, a compaction filter unit, and a compaction progress monitor unit.

[0155] The compaction trigger unit is used to detect data expiration time and trigger a compaction operation;

[0156] The compaction scheduler unit is used to reasonably arrange the time and priority of compaction tasks according to the system load and data expiration policy;

[0157] The Compaction Worker unit is responsible for performing the actual compaction operation, obtaining the data files to be merged, and performing the merge and delete operations;

[0158] The Compaction Filter unit is used to check the expiration timestamp of the data, identify and delete the expired data, and retain the non-expired data; and

[0159] The Compaction Progress Monitor unit is responsible for monitoring and reporting the progress of the compaction task and regularly updating the task status and progress information.

[0160] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are illustrative rather than restrictive. Under the guidance of this application, ordinary technicians in this field can also change, modify, replace and modify the above-mentioned embodiments without departing from the scope of protection of the purpose of this application and the claims, and all of these are within the protection of this application.

[0161] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0162] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A method for implementing an expiration schedule for a distributed relational database based on rocksdb, characterized in that: include: Step S1: After the client connects to the database, a TTL table is created; Step S2: The server layer calculates the TTL table data range and registers the TTL cleanup task to the engine layer; Step S3: The server layer starts the scheduled task and pushes up the cleanTs according to the TTL calculation. At the same time, the engine layer receives the TTL task, registers the CompactionFilter in rocksdb, and periodically pulls the cleanTs of the current task data from the computing node; Step S4: Use the snapshot consistency read method of cleanTs to enable RocksDB to perform compaction and clean up the data within the range by calling CompactionFilter; Step S5: The server layer periodically checks whether active compaction is needed based on statistical data, and initiates CompactionRange and CompactionFiles instructions if necessary; and Step S6: After the engine layer receives the active compaction instruction, RocksDB performs compaction and deletes data according to the CompactionFilter to complete data cleanup.

2. According to a method for implementing a distributed relational database expiration schedule based on rocksdb according to claim 1, it is characterized in that: The information of the TTL table in step S1 includes: TTL type, data expiration time and expiration time column.

3. According to claim 2, a method for implementing a distributed relational database expiration schedule based on rocksdb is characterized in that: The specific steps of step S2 include: Step S201: The server layer completes table creation, calculates the key range after TTL table data encoding, and writes the start time of the data into the key of the key-value pair; Step S202: registering a cleanup task at the engine layer according to the TTL table data range; and Step S203: Send an instruction to the engine table that the data in the keys range belongs to the TTL table and needs to be cleaned up by compaction.

4. According to claim 3, a method for implementing an expiration schedule of a distributed relational database based on rocksdb is characterized in that: The specific steps of step S3 include: Step S301: define data model and TTL; Step S302: Start a scheduled task at the server layer and periodically increase cleanTs according to TTL calculation; Step S303: Through the background thread, on each computing node, the cleanTs of the current task data is periodically pulled; Step S304: On each computing node, upon receiving a new TTL task, a CompactionFilter is registered in RocksDB; and Step S305: When the data expires, the registered CompactionFilter will be triggered.

5. According to claim 4, a method for implementing an expiration schedule of a distributed relational database based on rocksdb is characterized in that: The method further comprises: Trigger CompactionFilter to filter and delete expired data.

6. A method for implementing an expiration schedule of a distributed relational database based on rocksdb according to claim 5, characterized in that: The triggering of CompactionFilter to filter and delete expired data also includes: Parse TTL information from keys; Compare the TTL information of the data with the size of cleanTs; and When TTL is less than cleanTs, the key-value filter is filtered and deleted.

7. A method for implementing an expiration schedule of a distributed relational database based on rocksdb according to claim 4, characterized in that: The snapshot consistency reading method of cleanTs in step S4 includes: Step S401: When new data is inserted or existing data is updated, TTL is used as part of the key and stored together with the data in RocksDB; Step S402: When the cleanTs value reaches or exceeds the set expiration time, a data cleaning operation is performed; Step S403: merging data stored in multiple layers into one layer, and deleting expired or no longer needed data; Step S404: Create a class that inherits from CompactionFilter and rewrite the Filter method to determine whether each key-value pair is expired; Step S405: Start the database, register and use the custom CompactionFilter, and automatically call the custom Filter method; and Step S406: Use the snapshot function of RocksDB to obtain a timestamp before performing a read operation, and use this timestamp to obtain a data snapshot. When expired data is detected, call the delete API of RocksDB to delete the expired data.

8. A method for implementing an expiration schedule of a distributed relational database based on rocksdb according to claim 7, characterized in that: The earliest readable range of the current TTL table is [cleanTs, readTs], where readTs represents the node read by the current snapshot.

9. A method for implementing a distributed relational database expiration schedule based on rocksdb according to claim 7, characterized in that: The factors for determining whether active compaction is needed in step S5 include: the data volume statistics exceed the threshold, the data expiration time statistics are large, and the read and write operations are frequent and the data volume is large.

10. A method for implementing an expiration schedule of a distributed relational database based on rocksdb according to claim 9, characterized in that: The steps of initiating the CompactionRange and CompactionFiles instructions in step S5 include: Step S501: Call the CompactionRange and CompactionFiles functions of RocksDB and specify the range to be operated; Step S502: Obtain relevant range files and a list of files that need to perform compaction operations by reading data files and index files in RocksDB; and Step S503: Initiate the CompactionRange and CompactionFiles instructions to pass the file to the Compaction function of RocksDB for processing.

11. A method for implementing an expiration schedule of a distributed relational database based on rocksdb according to claim 10, characterized in that: The method further comprises: By periodically scanning TTL data keys, the conditions for triggering CompactRange are generated.

12. A method for implementing an expiration schedule of a distributed relational database based on rocksdb according to claim 10, characterized in that: The specific steps of RocksDB performing compaction in step S6 include: Step S601: Determine the order of compaction operations based on the scores of each layer, start compaction and wait for thread scheduling. Step S602: RocksDB divides the compaction into multiple subcompactions, processes them through child threads and puts them into an input array to form an SST file. Step S603: traverse the multi-path SST files through mergeIterator, sort the keys in the multi-path files using the minimum heap method, and then take out the top element of the heap each time; Step S604: create an output file, process subcompaction in a child thread, and merge the results into the final output file; and Step S605: Add a task to the thread pool and wait for scheduling. After all subcompactions are processed, the entire compaction process is completed.

13. A method for implementing an expiration schedule of a distributed relational database based on rocksdb according to claim 12, characterized in that: The SST file is a file used to persist database data.

14. A method for implementing a distributed relational database expiration schedule based on rocksdb according to claim 12, characterized in that: include: Data storage module, expiration time management module, compaction module, consistent reading module, monitoring and log module; The data storage module uses RocksDB as the backend storage engine and is responsible for storing and reading data; The expiration time management module is used to set the expiration policy for each data and manage the expiration time of the data; The compaction module is used to perform the compaction operation of RocksDB; The consistent read module uses the snapshot function of RocksDB to process data pre-fetching and caching; and The monitoring and logging module is used to monitor the system's CPU usage and disk I / O indicators in real time, and to detect and handle problems in a timely manner.

15. A method for implementing an expiration schedule of a distributed relational database based on rocksdb according to claim 14, characterized in that: The compaction module includes: a compaction trigger unit, a compaction scheduler unit, a compaction worker unit, a compaction filter unit, and a compaction progress monitor unit; The compaction trigger unit is used to detect the data expiration time and trigger the compaction operation; The compaction scheduler unit is used to reasonably arrange the time and priority of compaction tasks according to the system load and data expiration policy; The Compaction Worker unit is responsible for performing the actual compaction operation, obtaining the data files to be merged, and performing the merge and delete operations; The Compaction Filter unit is used to check the expiration timestamp of the data, identify and delete the expired data, and retain the unexpired data; and The Compaction Progress Monitor unit is responsible for monitoring and reporting the progress of the compaction task, and regularly updating the status and progress information of the task.

Citation Information

Patent Citations

  • Stale data removing method and device

    CN108196792A

  • Data processing method and device and computer readable storage medium

    CN116775636A

  • Data processing method and device in storage system based on RocksDB, equipment and medium

    CN116841472A

  • Rocksdb-based distributed relational database expiration schedule implementation method

    CN117874012A

  • Techniques for accelerating compaction

    US20200379775A1