Data deletion method and device, electronic equipment and storage medium

By obtaining the metadata of the target key value, the version data is directly determined, and the deletion marks are propagated in the data compaction process, the problem of low efficiency of composite data deletion is solved, efficient and accurate data deletion is achieved, and resource consumption is reduced.

CN120216486APending Publication Date: 2025-06-27ZUOYEBANG EDUCATION TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510270564.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When facing large-scale composite data, the prior art has low deletion efficiency, large CPU consumption and poor disk I/O performance, making it difficult to meet the high-performance needs of modern data storage and processing.

Method used

By obtaining the metadata of the target key value, the target version data is directly determined without traversing each field to determine the deleted object. Then, based on the target version data, add the version association deletion mark, and propagate and process it layer by layer in the data compaction process to complete the deletion operation.

Benefits of technology

It greatly reduces the steps and times of the deletion operation, improves the deletion efficiency, reduces the consumption of processor and disk resources, ensures the accuracy and completeness of the deletion operation, and maintains data consistency and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216486A_ABST
    Figure CN120216486A_ABST
Patent Text Reader

Abstract

The invention provides a data deletion method and device, electronic equipment and a storage medium, and relates to the technical field of data storage, target version data is directly determined by obtaining metadata of a target key value, and a deletion object is determined without traversing each field. According to the method, after the target version data is determined, deletion can be completed only by adding the version association deletion mark and combining a subsequent data compaction process, so that the operation steps are greatly reduced, and the deletion efficiency is improved. And consumption of system resources such as a processor and a disk can be reduced while the deletion operation frequency is reduced. In a traditional deletion mode, a processor is in a high-load state for a long time due to a large number of scanning and deletion operations, and the burden of a disk can be increased due to frequent disk read-write operations. According to the data processing method and device, after the target version data is determined and the version association deletion mark is added, the data is processed according to the mark in the data compaction process, and the calculation amount of a processor and the disk read-write frequency can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data storage, and in particular, to a method and apparatus for data deletion, an electronic device, and a storage medium. Background Art

[0002] With the explosive growth of data volume and the increasing complexity of data structures, data structures have become increasingly complex and diverse, and efficient data management has become a key challenge. Among them, the management of composite data has become a key factor affecting the performance of storage systems. Composite data refers to a data structure in which a key contains multiple fields. Currently, the common deletion scheme is that when deleting a key, all fields are scanned through the key, and then each field is deleted one by one. This method can maintain a certain efficiency when the number of fields is small, but when the number of fields is huge, problems will emerge. For example, if a key contains millions of fields, the deletion operation needs to be executed millions of times, which will cause a sharp increase in the number of database query and deletion operations, resulting in extremely high CPU consumption. Moreover, since the deletion operation involves writing to the disk, a large number of deletion operations mean frequent disk I / O, which is extremely unfavorable to disk performance, seriously affecting the overall operation efficiency and response speed of the storage system, increasing the system load, and reducing the stability and scalability of the system.

[0003] The composite data deletion technology of related technologies has problems such as low deletion efficiency, high CPU consumption, and poor disk I / O performance when facing large-scale composite data, and it is difficult to meet the high-performance requirements of modern data storage and processing. Therefore, there is an urgent need for a new technical solution to solve these problems and improve the performance and reliability of the storage system.

[0004] Disclosed content

[0005] The present disclosure provides a method and apparatus for data deletion, an electronic device, and a storage medium. Its main purpose is to solve the problems of low deletion efficiency and large resource consumption when deleting large-scale composite data in related technologies.

[0006] According to the first aspect of the present disclosure, a method for data deletion is provided, including:

[0007] In response to a data deletion instruction, determining a target key value corresponding to the data deletion instruction, and based on the target key value, reading target metadata corresponding to the target key value, and obtaining target version data and a target data type in the target metadata;

[0008] Deleting target expired data and target metadata corresponding to the target key value;

[0009] If the target data type is a preset data classification, based on the target version data, a deletion flag is added to the target composite data corresponding to the target key value to generate a version-associated deletion flag.

[0010] When performing the data compaction process, the version-associated deletion flag is propagated layer by layer, and the composite data in the database with the same version as the version-associated deletion flag is deleted.

[0011] In some embodiments, before performing the data compaction process, the data deletion method further includes:

[0012] Obtain the expired version data and expired data type corresponding to the expired key value. When the expired data type is a preset data classification, based on the expired version data, a deletion flag is added to the expired composite data.

[0013] In some embodiments, obtaining the expired version data and expired data type corresponding to the expired key value includes:

[0014] Scan the expired data in the database to obtain the expired key values before the current time;

[0015] Based on the expired key values, query the expired metadata corresponding to the expired key values to obtain the expired version data and expired data type.

[0016] In some embodiments, based on the target version data, adding a deletion flag to the target composite data corresponding to the target key value to generate a version-associated deletion flag includes:

[0017] Based on the target version data, find the target composite data in the database that is the same as the target version data;

[0018] Combine the deletion flag with the version data of the target composite data to generate a version-associated deletion flag.

[0019] In some embodiments, the data compaction process includes: in-memory table flushing and database compaction;

[0020] Propagating the version-associated deletion flag layer by layer includes:

[0021] When performing in-memory table flushing, flush the version-associated deletion flag to the first-level sorted table on disk;

[0022] When performing database compaction, propagate the deletion flag from the first-level sorted table layer by layer until the last layer.

[0023] In some embodiments, deleting the composite data in the database with the same version as the version-associated deletion flag includes:

[0024] Scan the scan key values in the database that are consistent with the version associated with the deletion flag, and obtain the scan composite data corresponding to the scan key values;

[0025] Ignore the scan composite data, continue to perform the writing process on the composite data corresponding to other key values in the database, and mark the scan composite data as logically deleted;

[0026] When merging the multi-level sorting table, mark the composite data corresponding to the version associated with the deletion flag as logically deleted.

[0027] According to a second aspect of the present disclosure, there is provided an apparatus for data deletion, including:

[0028] A first acquisition unit, configured to, in response to a data deletion instruction, determine a target key value corresponding to the data deletion instruction, based on the target key value, read target metadata corresponding to the target key value, and obtain target version data and a target data type in the target metadata;

[0029] A first deletion unit, configured to delete target expired data and target metadata corresponding to the target key value;

[0030] A first addition unit, configured to, if the target data type is a preset data classification, based on the target version data, add a deletion flag to the target composite data corresponding to the target key value to generate a version-associated deletion flag;

[0031] A second deletion unit, configured to, when executing a data compaction process, propagate the version-associated deletion flag layer by layer, and delete the composite data in the database that is consistent with the version of the version-associated deletion flag.

[0032] In some embodiments, the apparatus for data deletion further includes:

[0033] A second acquisition unit, configured to, before executing a data compaction process, acquire expired version data and an expired data type corresponding to an expired key value;

[0034] A second addition unit, configured to, when the expired data type is a preset data classification, based on the expired version data, add a deletion flag to the expired composite data.

[0035] In some embodiments, the second acquisition unit includes:

[0036] A scanning module, configured to scan the expired data in the database to obtain expired key values before the current time;

[0037] A first acquisition module, configured to, based on the expired key value, query the expired metadata corresponding to the expired key value, and obtain the expired version data and the expired data type.

[0038] In some embodiments, the first addition unit includes:

[0039] A search module, configured to search for target composite data identical to the target version data in a database based on the target version data;

[0040] A generation module, configured to combine a deletion flag with the version data of the target composite data to generate a version-associated deletion flag.

[0041] In some embodiments, the data compaction process includes: memory table flushing and database compaction;

[0042] The second deletion unit includes:

[0043] A flushing module, configured to flush the version-associated deletion flag to the first-level sorted table on the disk when performing memory table flushing;

[0044] A propagation module, configured to propagate the deletion flag layer by layer from the first-level sorted table until the last level when performing database compaction.

[0045] In some embodiments, the second deletion unit further includes:

[0046] A second acquisition module, configured to scan the scan key values in the database that are consistent with the version of the version-associated deletion flag, and acquire the scan composite data corresponding to the scan key values;

[0047] A first deletion module, configured to ignore the scan composite data, continue to perform flushing processing on the composite data corresponding to other key values in the database, and mark the scan composite data as logically deleted;

[0048] A second deletion module, configured to mark the composite data corresponding to the version-associated deletion flag as logically deleted when merging multi-level sorted tables.

[0049] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0050] At least one processor; and

[0051] A memory communicatively connected to the at least one processor; wherein,

[0052] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the foregoing first aspect.

[0053] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect.

[0054] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the first aspect above.

[0055] The present disclosure provides a method and device for data deletion, an electronic device and a storage medium, and relates to the technical field of data storage. The present disclosure directly determines the target version data by obtaining the metadata of the target key value, without traversing each field to determine the deletion object. After determining the target version data, the present disclosure only adds a version-related deletion mark and completes the deletion in combination with the subsequent data compaction process, which greatly reduces the operation steps and improves the deletion efficiency. While reducing the number of deletion operations, it can also reduce the consumption of system resources such as processors and disks. In the traditional deletion method, a large number of scanning and deletion operations will cause the processor to be in a high-load state for a long time, and frequent disk read and write operations will increase the disk burden. In the present disclosure, after determining the target version data and adding the version-related deletion mark, the data is processed according to the mark in the data compaction process, which can reduce the amount of processor calculation and the number of disk read and write times. Based on the target key value, the target metadata is read to obtain information such as version and data type, which can ensure the accuracy of the deletion operation. First, the target expired data and target metadata are deleted, and then the deletion mark is added for the preset data classification to avoid accidental deletion or missed deletion. In the data compaction process, version-associated deletion markers are propagated layer by layer and the corresponding composite data is deleted to ensure the integrity of the entire deletion operation, guaranteeing the accuracy and integrity of data management at all levels and maintaining data consistency and reliability.

[0056] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0058] Figure 1 A flowchart of a method for deleting data provided by an embodiment of the present disclosure;

[0059] Figure 2 A flowchart of another method for deleting data provided by an embodiment of the present disclosure;

[0060] Figure 3 A schematic diagram of the structure of a data deletion device provided by an embodiment of the present disclosure;

[0061] Figure 4 A schematic diagram of the structure of another data deletion device provided by an embodiment of the present disclosure;

[0062] Figure 5 Schematic block diagram of an exemplary electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0063] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0064] The following describes a method and apparatus for data deletion, an electronic device, and a storage medium according to embodiments of the present disclosure with reference to the accompanying drawings.

[0065] Figure 1 Flow schematic diagram of a method for data deletion provided by an embodiment of the present disclosure.

[0066] As Figure 1 shown, the method includes the following steps:

[0067] Step 101, in response to a data deletion instruction, determine a target key value corresponding to the data deletion instruction. Based on the target key value, read target metadata corresponding to the target key value, and obtain target version data and a target data type in the target metadata.

[0068] In an embodiment of the present disclosure, when a storage database system receives a data deletion instruction, the primary task is to parse the instruction. Through a specific parsing algorithm and data indexing mechanism, the system can accurately determine the target key value corresponding to this data deletion instruction. This target key value, as the core identifier for subsequent operations, plays a crucial role in the entire data deletion process.

[0069] After determining the target key value, based on the target key value, the system will access a specific storage area storing the target metadata. The storage system usually adopts an efficient data storage structure and indexing method to manage the metadata for quick location and reading. After finding the target metadata corresponding to the target key value, the system will extract key information therefrom, namely the target version data and the target data type. The target version data is a uniquely and sequentially numbered identifier assigned by the system to each data object, used to identify different states and operation sequences of the data, which is crucial for handling data updates, deletions, and consistency maintenance. The target data type, on the other hand, clarifies the structural category of the data corresponding to the key value, such as whether it is a composite data type like hash, list, set, sorted set (zset), etc., or other simple data types, which provides a necessary basis for subsequent execution of different deletion operation logics according to the data type.

[0070] Step 102, delete the target expired data and target metadata corresponding to the target key value.

[0071] In an embodiment of the present disclosure, when executing a data deletion process in a storage system, after determining the target key value corresponding to a data deletion instruction and obtaining the relevant target metadata, the next operation is to delete the target expired data and target metadata corresponding to the target key value. The storage system generally stores expired data and metadata in specific data structures or storage areas respectively for easy management and query. For the target expired data, the system searches and locates in the area storing expired data (such as the expireData structure sorted by expiration time) according to the target key value. Once a target expired data record matching the target key value is found, a deletion operation is performed to remove the record from the storage area and release the storage space it occupies. For the target metadata, the system also retrieves it at the corresponding position storing metadata (such as the area specifically storing metaData) according to the target key value. After locating the target metadata, it is deleted from the storage medium to ensure that the metadata information is consistent with the actual data status. Deleting the target metadata can not only avoid the occupation of system resources by invalid metadata but also prevent abnormal situations caused by incorrect metadata information in subsequent data operations. This series of deletion operations can make the storage system manage data more efficiently and accurately, providing a clear and effective data environment for subsequent data processing processes.

[0072] Step 103, if the target data type is a preset data classification, add a deletion mark to the target composite data corresponding to the target key value based on the target version data to generate a version-associated deletion mark.

[0073] In an embodiment of the present disclosure, during the data deletion operation process of the storage system, after the deletion of the target expired data and the target metadata corresponding to the target key-value is completed, the system will judge the target data type. The system has preset a series of specific data classifications, such as composite data types like hash, list, set, sorted set (zset), etc. as the preset data classifications. If it is determined through judgment that the target data type belongs to one of the above preset data classifications, the system will further perform the operation of adding a deletion mark based on the target version data obtained previously. In the storage structure (such as mixData) for storing the target composite data, based on the target version data, the deletion mark is added in a specific format. Usually, this deletion mark will be closely combined with the target version data to form a specific identification combination. For example, it is added in the format of <target version data + prefix deletion mark, 0>, thus generating a version-associated deletion mark. The generated version-associated deletion mark plays an important role. It provides a clear deletion basis for the subsequent storage system when performing operations such as memtable flush and database compaction. During these subsequent operation processes, the system can accurately identify the target composite data to be deleted according to the version-associated deletion mark, thereby realizing an efficient and accurate data deletion processing flow.

[0074] Step 104, during the data compaction process, propagate the version-associated deletion mark layer by layer, and delete the composite data in the database that has the same version as the version-associated deletion mark.

[0075] In an embodiment of the present disclosure, during the operation of the storage system, data compaction is an important data management operation. When the system starts the data compaction process, the version-associated deletion mark will play a key role in the hierarchical structure of the entire database.

[0076] Databases usually use a multi-level storage structure to manage data to improve storage efficiency and data processing performance. At the beginning of data compaction, the system starts from the top-level data file stored and passes the version-associated deletion marker to the next level along the hierarchical structure. During the data processing of each layer, the system checks the composite data in the database of that layer one by one. For each composite data, the system extracts its version information and compares it with the version in the version-associated deletion marker. If it is found that the composite data version in the database is consistent with the version in the version-associated deletion marker, the system will perform a deletion operation, that is, the composite data will no longer be written to the new data file generated after compaction, which is actually equivalent to deleting it from the database. As the data compaction process gradually advances, the version-associated deletion marker will continue to propagate to the next level, repeating the above-mentioned inspection and deletion operations. Until all levels of the database are traversed, it is ensured that all composite data in the database that are consistent with the version of the version-associated deletion marker have been properly processed. In this way, the version-associated deletion marker realizes the precise deletion of specific composite data in the data compaction process, effectively cleans up invalid data in the database, and improves the storage efficiency and data quality of the database.

[0077] The present disclosure provides a method for data deletion. The present disclosure directly determines the target version data by obtaining the metadata of the target key value, without traversing each field to determine the deletion object. After determining the target version data, the present disclosure only adds a version-related deletion mark and combines it with the subsequent data compaction process to complete the deletion, which greatly reduces the operation steps and improves the deletion efficiency. While reducing the number of deletion operations, it can also reduce the consumption of system resources such as processors and disks. In the traditional deletion method, a large number of scanning and deletion operations will cause the processor to be in a high-load state for a long time, and frequent disk read and write operations will increase the disk burden. In the present disclosure, after determining the target version data and adding the version-related deletion mark, the data is processed according to the mark in the data compaction process, which can reduce the processor's calculation amount and the number of disk reads and writes. Based on the target key value, the target metadata is read to obtain information such as version and data type, which can ensure the accuracy of the deletion operation. First, the target expired data and target metadata are deleted, and then the deletion mark is added for the preset data classification to avoid accidental deletion or missed deletion. In the data compaction process, version-associated deletion markers are propagated layer by layer and the corresponding composite data is deleted to ensure the integrity of the entire deletion operation, guaranteeing the accuracy and integrity of data management at all levels and maintaining data consistency and reliability.

[0078] In the embodiment of the present disclosure, taking hash structure data storage as an example, its storage system includes three types of key information storage mechanisms, which are as follows:

[0079] 1. MetaData: Used to store the metadata of keys, and its storage structure is <key, data type (type) + version data (version) + timestamp>. Among them, type is used to identify the data structure type, such as hash, string, list, set, sorted set (zset), etc. Different data structures correspond to different type values; the version data (version) is an unsigned 64-bit integer (uint64) built into the system. Whenever a new key is written, the version increments. Also, when a key is deleted or expired and then rewritten, a new version will be generated to ensure the uniqueness of the version corresponding to each key and that it does not duplicate the version of the historically deleted key; the timestamp records the relevant time information.

[0080] 2. ExpireData: Mainly stores all keys with an expiration time set. These keys are sorted and stored according to the expiration time, and the storage structure is <timestamp + key, 0>. If a key does not have an expiration time set, the write operation for this step is not performed.

[0081] 3. MixData: Used to store all composite data, and its storage structure is <version + field, value>.

[0082] The data writing process of the hash structure is described in detail. Assume the execution of the instruction hset student-30 name xiaoming:

[0083] First, write the data to MetaData in the format <student-30, hash + version + timestamp>, where the version increments according to the system rules and the timestamp records the current time. Second, determine whether the key has an expiration time set. If an expiration time is set, write <timestamp + student-30, 0> to ExpireData; if no expiration time is set, skip this step. Finally, write <version + name, xiaoming> to MixData, where the version is the version number generated when writing to MetaData.

[0084] Through the above orderly and rigorous storage mechanism and writing process, efficient and accurate storage of hash structure data can be achieved, providing a solid data foundation for subsequent data query, update, and deletion operations.

[0085] To clearly illustrate the embodiments of the present disclosure, this embodiment provides a schematic flowchart of another data deletion method.

[0086] As Figure 2 shown, the method includes the following steps:

[0087] Step 201, in response to a data deletion instruction, determine the target key value corresponding to the data deletion instruction. Based on the target key value, read the target metadata corresponding to the target key value, and obtain the target version data and target data type in the target metadata.

[0088] Specifically, in step 201, when receiving a data deletion instruction, by parsing the instruction, determine the target key value (key) corresponding to the instruction. Based on the target key value, access the corresponding metadata to obtain the (target) version data (version) and (target) data type (type) of the metadata; the target version data is a unique and sequential number assigned by the system to each data object, used to identify different states and operation sequences of the data, which is crucial in handling data updates, deletions, and consistency maintenance. The target data type, on the other hand, clarifies the structural category of the data corresponding to the key value, such as whether it is a composite data type like hash, list, set, sorted set (zset), etc., or other simple data types, which provides a necessary basis for subsequent execution of different deletion operation logics according to the data type.

[0089] In some other embodiments of the present disclosure, while obtaining the target version data and target data type, the (target) timestamp will also be read.

[0090] Exemplarily, when receiving a deletion instruction such as "del student-30", a series of deletion actions will be performed on the target key value (here student-30). Based on the target key value "student-30", retrieve in the area storing metaData. Since metaData is stored in the structure of <key, type + version + timestamp>, the system can accurately locate the metadata record corresponding to the target key value and obtain the corresponding timestamp (timestamp, recording data-related time information), version (an unsigned 64-bit integer built into the system, used to identify the data version to ensure the uniqueness of the version corresponding to each key), and type (data structure type, such as hash, list, set, zset, etc.) from it.

[0091] Step 202, delete the target expired data and target metadata corresponding to the target key value.

[0092] Specifically, in step 202, after obtaining the target key value from the data deletion instruction, it is necessary to first extract relevant information from the target metadata, and then delete the target metadata and the target expired data.

[0093] After obtaining the relevant information, directly delete from the metaData storage area <student-30>This record completes the deletion operation of the target metadata. This step ensures that there is no redundant information related to the target key-value in the metaData, avoiding interference with subsequent data operations.

[0094] Using the previously obtained timestamp and the target key-value "student-30", search for the corresponding <timestamp + student-30, 0> record in the expireData storage area. expireData stores all keys with expiration times sorted by expiration time. In this way, the target expired data record can be quickly located. After finding it, delete it from the expireData storage area to clean up the expired data and release the corresponding storage resources.

[0095] Step 203, if the target data type is a preset data classification, based on the target version data, add a deletion mark to the target composite data corresponding to the target key-value to generate a version-associated deletion mark.

[0096] Furthermore, based on the target version data, adding a deletion mark to the target composite data corresponding to the target key-value to generate a version-associated deletion mark includes: based on the target version data, searching for the target composite data in the database that is the same as the target version data; combining the deletion mark with the version data of the target composite data to generate a version-associated deletion mark.

[0097] Specifically in step 203, after the system receives the deletion instruction, it has determined the target key-value (such as "student-30") according to the instruction and successfully read the corresponding timestamp, version, and type from the metaData. At the same time, the operation on the metaData has been completed. <key>(i.e. <student-30>) and the deletion operation of <timestamp+key> in expireData (i.e., <timestamp+student-30>).

[0098] After completing the above preparatory operations, the system will judge the obtained type. In the embodiments of the present disclosure, the preset data classifications include composite data structure types such as hash, list, set, and zset. If the type belongs to one of these preset data classifications, the subsequent operation of adding deletion marks will be triggered.

[0099] After confirming that the target data type conforms to the preset data classification, based on the target version data (i.e., version) obtained from metaData, the system will search in mixData that stores composite data. mixData stores composite data in the structure of <version+field,value>. The system traverses mixData to find all records whose version data is the same as the target version data. These records are the target composite data that is the same as the target version data. For example, in mixData, search for all records where version is equal to the version corresponding to "student-30" obtained from metaData.

[0100] The system combines the deletion mark (i.e., the prefix deletion mark) with the version data of the found target composite data. Generate a version-associated deletion mark in the format of <version+prefix deletion mark,0>. Here, version is the target version data obtained from metaData, and the prefix deletion mark is a special mark used to indicate that the composite data under this version needs to be deleted. In this way, the deletion mark is closely associated with the version of the target composite data, generating a version-associated deletion mark, providing a basis for accurately deleting the target composite data during the subsequent memtable flush and database compaction processes. During the whole process, only 1 write operation is required when adding the prefix deletion to mixData, effectively reducing the number of write operations and improving the efficiency of the deletion operation.

[0101] Step 204, obtain the expired version data and expired data type corresponding to the expired key value. When the expired data type is a preset data classification, based on the expired version data, add a deletion mark to the expired composite data.

[0102] Further, obtaining the expired version data and expired data type corresponding to the expired key value includes: scanning the expired data in the database to obtain the expired key value before the current time; based on the expired key value, query the expired metadata corresponding to the expired key value to obtain the expired version data and expired data type.

[0103] Specifically in step 204, the system periodically scans the expireData that stores expired data through a background program. The expireData stores all keys with expiration times and is sorted by expiration time. During the scanning process, the system will find all keys whose expiration times are before the current time. These keys are the expired key values. Due to the ordered storage feature of expireData, the system can efficiently locate the expired key values, reducing the scanning time and resource consumption. For example, when the current time is T, the system can quickly filter out all keys in expireData whose expiration times are earlier than T.

[0104] After obtaining the expired key values, the system queries in the metaData area that stores metadata based on these expired key values. The metaData stores the metadata of keys in the structure of <key, type + version + timestamp>. The system locates the corresponding expired metadata records by matching the expired key values with the keys in the metaData. From these expired metadata records, the system obtains the expired version data (i.e., version) and the expired data type (i.e., type). This step ensures that the system obtains the key information required for processing expired data and provides a basis for subsequent operations.

[0105] When the system obtains the expired version data and the expired data type, it will judge the expired data type. The preset data classifications in this patent technical solution include composite data structure types such as hash, list, set, and zset. If the expired data type belongs to one of these preset data classifications, the system will add a deletion mark to the expired composite data based on the expired version data. The specific operation is to add a record in the mixData that stores the composite data in the format of <version + prefix deletion mark, 0>. Here, version is the expired version data obtained from the expired metadata, and the prefix deletion mark is a special mark used to indicate that the composite data under this version needs to be deleted. In this way, it is prepared for thoroughly deleting the expired composite data during the subsequent memtable flush and database compaction processes. Through such a process, the system can automatically and efficiently clean up the expired composite data, ensuring the accuracy of the data and the efficient operation of the storage system.

[0106] In step 205, when performing the memtable flush, the version-associated deletion mark is flushed to the first-level sorted table on the disk.

[0107] Step 206: Scan the scan key values in the database that are consistent with the version associated with the deletion marker, and obtain the corresponding scan composite data for the scan key values.

[0108] Step 207: Ignore the scan composite data, continue to perform the write operation on the composite data corresponding to other key values in the database, and mark the scan composite data as logically deleted.

[0109] Specifically, in steps 205 to 207, the memtable serves as a temporary storage area for data in memory. When specific conditions are met (such as reaching a preset size or time interval), the data in it needs to be persisted to disk, and this process is called memtable flush. When performing the memtable flush operation, the system traverses all the data in the memtable. For the data with version-associated deletion markers (in the format of <version + prefix deletion marker, 0>), the system will flush these markers together with other data to the first-level sorted table on disk. The first-level sorted table on disk is the initial hierarchical structure for data storage on disk. By flushing the version-associated deletion markers to this level, subsequent database operations can identify which data needs to be deleted, providing a basis for subsequent data cleaning.

[0110] After completing the memtable flush and writing the version-associated deletion markers to disk, in order to further process the data marked as deleted, the system needs to perform a scan operation in the database. The basis for the scan is the version information in the version-associated deletion marker. The system starts scanning the database. During the scanning process, it will search for all scan key values (i.e., keys) that are consistent with the version of the version-associated deletion marker. When these scan key values are found, the system will obtain the corresponding scan composite data from the area where the composite data is stored (such as mixData) according to these key values. These composite data are stored in the form of <version + field, value>. By matching the version information, the system can accurately locate the composite data that needs to be deleted, preparing for subsequent deletion operations.

[0111] After obtaining the scanned composite data that is consistent with the version - associated deletion marker version, the system will ignore this data and not write it to the new storage location (such as when performing disk file write operations). At the same time, the system will continue to perform normal flushing processing on the composite data corresponding to other key - values in the database. Depending on the actual situation of the data, these data may be written to the corresponding files, or if these data also carry a prefix deletion marker, the deletion logic will continue to be processed. For the ignored scanned composite data, the system will mark them as logically deleted. This logical deletion is achieved through the previously existing version - associated deletion marker (<version + prefix deletion marker, 0>). In subsequent database compaction and other operations, the system will further process these data based on these markers until all files in the last layer no longer contain the field information corresponding to the version - associated deletion marker, and then the relevant markers and data will be truly deleted from the storage system, completing the entire deletion process.

[0112] Step 208, when performing database compaction, propagate the deletion marker layer by layer from the first - layer sorted table to the last layer.

[0113] Step 209, when merging multi - level sorted tables, mark the composite data corresponding to the version - associated deletion marker as logically deleted.

[0114] Specifically, in Steps 208 to 209, database compaction is an operation to optimize and organize the data stored on the disk. Its purpose is to reduce the number of data files, improve the efficiency of data query and management, and clean up invalid data. In this process, the correct propagation of the deletion marker is crucial. When performing database compaction, the system starts processing from the first - layer sorted table on the disk. In the first - layer sorted table, the system checks whether there is data with a deletion marker (i.e., the version - associated deletion marker <version + prefix deletion marker, 0>). If such data is found, the system will pass this deletion marker to the next - layer sorted table. During the processing of each layer of the sorted table, the system will repeat this operation, that is, check whether there is data related to this deletion marker in the current layer. If there is, the deletion marker will continue to be propagated downwards. In this way, the deletion marker starts from the first - layer sorted table and is propagated layer by layer until it reaches the last - layer sorted table. This process ensures that in the entire database storage structure, the data related to this deletion marker can be correctly identified and processed.

[0115] As data is continuously written, updated, and deleted, multiple levels of sorted tables are formed on the disk. To optimize the storage structure and improve performance, it is necessary to merge these multi-level sorted tables. During the merging process, data that has been marked for deletion needs to be processed. When performing the merging of multi-level sorted tables, the system checks each piece of data one by one. When encountering composite data corresponding to the version-associated deletion marker, the system does not actually delete this data from physical storage but marks it as logically deleted. Specifically, the system determines whether the data needs to be logically deleted based on the version-associated deletion marker propagated previously. If the version of a certain composite data is the same as the version in the deletion marker, then this composite data is marked as logically deleted. In subsequent operations, this data marked as logically deleted will not be written into the newly merged file, which is equivalent to being logically deleted. When all levels of sorted tables have been merged and it is ensured that the last layer of all files no longer contains field information corresponding to the version-associated deletion marker of this version, the system will completely delete the prefix deletion marker of this version, completing the entire data deletion process.

[0116] It should be noted that multiple steps may be included in the embodiments of the present disclosure. For ease of description, these steps are numbered, but these numbers are not intended to limit the execution time slots and execution orders between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not make any limitations in this regard.

[0117] Corresponding to the above data deletion method, the present disclosure also proposes a data deletion device. Since the device embodiments of the present disclosure correspond to the above method embodiments, details not disclosed in the device embodiments can be referred to the above method embodiments, and will not be elaborated in the present disclosure.

[0118] Figure 3 The structural schematic diagram of a data deletion device provided for the embodiments of the present disclosure is as Figure 3 shown, including:

[0119] A first acquisition unit 31, configured to, in response to a data deletion instruction, determine the target key value corresponding to the data deletion instruction, based on the target key value, read the target metadata corresponding to the target key value, and obtain the target version data and target data type in the target metadata;

[0120] A first deletion unit 32, configured to delete the target expired data and target metadata corresponding to the target key value;

[0121] A first addition unit 33, configured to, if the target data type is a preset data classification, based on the target version data, add a deletion marker to the target composite data corresponding to the target key value to generate a version-associated deletion marker;

[0122] The second deleting unit 34 is used to propagate the version-associated deletion mark layer by layer when executing the data compaction process, and delete the composite data in the database that is consistent with the version associated with the deletion mark.

[0123] The present disclosure provides a device for data deletion. The present disclosure directly determines the target version data by obtaining the metadata of the target key value, without traversing each field to determine the deletion object. After determining the target version data, the present disclosure can complete the deletion by only adding a version-related deletion mark and combining it with the subsequent data compaction process, which greatly reduces the operation steps and improves the deletion efficiency. While reducing the number of deletion operations, it can also reduce the consumption of system resources such as processors and disks. In the traditional deletion method, a large number of scanning and deletion operations will cause the processor to be in a high-load state for a long time, and frequent disk read and write operations will increase the disk burden. In the present disclosure, after determining the target version data and adding the version-related deletion mark, the data is processed according to the mark in the data compaction process, which can reduce the processor's calculation amount and the number of disk reads and writes. Based on the target key value, the target metadata is read to obtain information such as version and data type, which can ensure the accuracy of the deletion operation. First, the target expired data and target metadata are deleted, and then the deletion mark is added for the preset data classification to avoid accidental deletion or missed deletion. In the data compaction process, version-associated deletion markers are propagated layer by layer and the corresponding composite data is deleted to ensure the integrity of the entire deletion operation, guaranteeing the accuracy and integrity of data management at all levels and maintaining data consistency and reliability.

[0124] Furthermore, in a possible implementation of this embodiment, as Figure 4 As shown, the data deletion device also includes:

[0125] The second acquisition unit 35 is used to acquire the expired version data and expired data type corresponding to the expired key value before executing the data compaction process;

[0126] The second adding unit 36 ​​is configured to add a deletion mark to the expired composite data based on the expired version data when the expired data type is a preset data classification.

[0127] Furthermore, in a possible implementation of this embodiment, as Figure 4 As shown, the second acquisition unit 35 includes:

[0128] Scanning module 351, used to scan expired data in the database to obtain expired key values ​​before the current time;

[0129] The first acquisition module 352 is used to query the expired metadata corresponding to the expired key value based on the expired key value, and obtain the expired version data and the expired data type.

[0130] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the first adding unit 33 includes:

[0131] A searching module 331, configured to search for target composite data identical to the target version data in the database based on the target version data;

[0132] A generating module 332, configured to combine the deletion flag with the version data of the target composite data to generate a version-associated deletion flag.

[0133] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the data compaction process includes: in-memory table flushing and database compaction;

[0134] The second deletion unit 34 includes:

[0135] A flushing module 341, configured to flush the version-associated deletion flag to the first-level sorting table on the disk when performing in-memory table flushing;

[0136] A propagation module 342, configured to propagate the deletion flag layer by layer from the first-level sorting table until the last layer when performing database compaction.

[0137] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the second deletion unit 34 further includes:

[0138] A second obtaining module 343, configured to scan the scan key values in the database that are consistent with the version of the version-associated deletion flag, and obtain the scan composite data corresponding to the scan key values;

[0139] A first deletion module 344, configured to ignore the scan composite data, continue to perform flushing processing on the composite data corresponding to other key values in the database, and mark the scan composite data as logically deleted;

[0140] A second deletion module 345, configured to mark the composite data corresponding to the version-associated deletion flag as logically deleted when merging multi-level sorting tables.

[0141] It should be noted that the foregoing explanation of the method embodiment also applies to the device in this embodiment. The principles are the same and will not be limited in this embodiment.

[0142] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0143] Figure 5 FIG. 0 shows a schematic block diagram of an exemplary electronic device 400 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0144] As Figure 5 shown, the electronic device 400 includes a computing unit 401 that may perform various appropriate actions and processes in accordance with a computer program stored in a ROM (Read-Only Memory) 402 or a computer program loaded from a storage unit 408 into a RAM (Random Access Memory) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 may also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An I / O (Input / Output) interface 405 is also connected to the bus 404.

[0145] A plurality of components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0146] The computing unit 401 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the method of data deletion. For example, in some embodiments, the method of data deletion may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the methods described above may be executed. Alternatively, in other embodiments, the computing unit 401 may be configured to execute the aforementioned method of data deletion in any other suitable manner (e.g., by means of firmware).

[0147] Various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0148] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0149] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory), or flash memory, optical fibers, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0150] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0151] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0152] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.

[0153] It should be noted that artificial intelligence is a discipline that studies enabling a computer to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0154] The various numerical numbers such as the first, second, etc. involved in this disclosure are only for the convenience of description and are not used to limit the scope of the embodiments of this disclosure, nor do they represent the order of precedence.

[0155] At least one in the present disclosure may also be described as one or more. The plurality may be two, three, four or more, and the present disclosure does not limit this. In the embodiments of the present disclosure, for a technical feature, the technical features in this technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D", etc. There is no order of precedence or order of magnitude between the technical features described by the "first", "second", "third", "A", "B", "C" and "D".

[0156] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and this is not limited herein.

[0157] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure. < / key>

Claims

1. A method for deleting data, characterized in that: The method comprises: In response to a data deletion instruction, determining a target key value corresponding to the data deletion instruction, reading target metadata corresponding to the target key value based on the target key value, and acquiring target version data and target data type in the target metadata; Deleting the target expired data and the target metadata corresponding to the target key value; If the target data type is a preset data classification, then based on the target version data, a deletion mark is added to the target composite data corresponding to the target key value to generate a version-associated deletion mark; When executing the data compaction process, the version-associated deletion mark is propagated layer by layer, and the composite data in the database that is consistent with the version of the version-associated deletion mark is deleted.

2. The method for deleting data according to claim 1, characterized in that: Before executing the data compaction process, the method further includes: Obtain expired version data and expired data type corresponding to the expired key value, and when the expired data type is the preset data classification, add the deletion mark to the expired composite data based on the expired version data.

3. The method for deleting data according to claim 2, characterized in that: The obtaining of expired version data and expired data type corresponding to the expired key value includes: Scan the expired data in the database to obtain the expired key value before the current time; Based on the expired key value, the expired metadata corresponding to the expired key value is queried to obtain the expired version data and the expired data type.

4. The method for deleting data according to claim 1, characterized in that: The step of adding a deletion mark to the target composite data corresponding to the target key value based on the target version data to generate a version-associated deletion mark includes: Based on the target version data, searching a database for the target composite data that is identical to the target version data; The deletion mark is combined with the version data of the target composite data to generate the version-associated deletion mark.

5. The method for deleting data according to claim 1, characterized in that: The data compaction process includes: memory table flushing and database compaction; The step of propagating the version-associated deletion mark layer by layer includes: When executing the memory table flushing, flushing the version associated deletion mark to the first-level sorting table on the disk; When performing the database compaction, the deletion mark is propagated layer by layer from the first-layer sorting table to the last layer.

6. The method for deleting data according to claim 5, characterized in that: Deleting the composite data in the database that is consistent with the version of the deletion mark associated with the version includes: Scan the database for a scan key value that is consistent with the version of the deletion mark associated with the version, and obtain scan composite data corresponding to the scan key value; Ignore the scanned composite data, continue to flush the composite data corresponding to other key values ​​in the database to the disk, and mark the scanned composite data as logically deleted; When merging the multi-level sorting tables, the composite data corresponding to the version-associated deletion mark is marked as logically deleted.

7. A data deletion device, characterized in that: The device comprises: A first acquisition unit is used to respond to a data deletion instruction, determine a target key value corresponding to the data deletion instruction, read target metadata corresponding to the target key value based on the target key value, and acquire target version data and target data type in the target metadata; A first deleting unit, used to delete the target expired data corresponding to the target key value and the target metadata; A first adding unit is used to add a deletion mark to the target composite data corresponding to the target key value based on the target version data to generate a version-associated deletion mark if the target data type is a preset data classification; The second deleting unit is used to propagate the version-associated deletion mark layer by layer when executing the data compaction process, and delete the composite data in the database that is consistent with the version of the version-associated deletion mark.

8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

10. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 6.