Data processing method, computing device, storage medium and computer program product
By obtaining the historical version of the data to be processed in the database and its corresponding data files, using the deletion mark to determine the associated data files, and deleting them when the conditions are met, the problem of large amounts of deletion marks and historical version data occupying storage space is solved, and data query performance is improved.
Patent Information
- Application Number
- CN202510535716.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When users update and delete data frequently or store a lot of data, a large amount of deletion marks and historical versions will accumulate, resulting in an increase in storage space and affecting data query performance.
A data processing method is provided, by obtaining at least two historical version data of the to be processed data and its corresponding data files, determining the associated data file and data association information of the to be processed data based on the deletion mark carried by the historical version data, and deleting the deletion mark and historical version data in the associated data file when the preset deletion conditions are met.
It realizes timely deletion of deleted marks and historical version data, avoids the accumulation of large amounts of undeleted data in the storage space, thereby improving data query performance.
Smart Images

Figure CN120045524A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of database technology, and particularly to data processing methods, computing devices, storage media, and computer program products. Background Art
[0002] In the field of data storage, the data written by users can usually be cached in memory. When the data cached in memory reaches a preset data volume, the data cached in memory is then merged and written to disk. When updating the written data, data update is usually achieved by appending new data. Then, when the written data is updated frequently, a single data will generate multiple historical versions stored in the same or different data files, thus affecting data query efficiency. Based on this, historical versions of the written data can usually be deleted. When deleting data, a deletion mark can be written to the data to be deleted, and when merging data subsequently, the data corresponding to the deletion mark is filtered, thereby achieving the deletion of historical version data.
[0003] However, when users frequently update and delete data or there is a large amount of stored data, a large number of deletion marks and historical version data will accumulate, thus causing a large amount of deletion marks and historical version data to occupy storage space, and further affecting data query performance. Therefore, an effective technical solution is urgently needed to solve the above problems. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a data processing method. One or more embodiments of this specification simultaneously relate to a data processing device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a data processing method is provided, which is applied to a database based on the LSM storage engine and includes: Obtain at least two historical version data of the data to be processed, and the data files corresponding to the at least two historical version data; According to the deletion marks carried by any historical version data among the at least two historical version data, determine the associated data file of the data to be processed and the data association information corresponding to the associated data file in the data files corresponding to the at least two historical version data; When it is determined according to the data association information that the associated data file meets the preset deletion condition, delete the deletion marks and historical version data included in the associated data file.
[0006] According to a second aspect of the embodiments of the present specification, a data processing device is provided, which is applied to a database based on an LSM storage engine and includes: An acquisition module, configured to acquire at least two historical version data of the data to be processed, and data files corresponding to the at least two historical version data; A determination module, configured to determine, according to the deletion flag carried by any historical version data in the at least two historical version data, an associated data file of the data to be processed and data association information corresponding to the associated data file in the data files corresponding to the at least two historical version data; A deletion module, configured to delete the deletion flag and historical version data included in the associated data file when it is determined according to the data association information that the associated data file meets a preset deletion condition.
[0007] According to a third aspect of the embodiments of the present specification, a computing device is provided, including: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above method are implemented.
[0008] According to a fourth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above method are implemented.
[0009] According to a fifth aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above method are implemented.
[0010] An embodiment of the present specification provides a data processing method, including: acquiring at least two historical version data of the data to be processed, and data files corresponding to the at least two historical version data; determining, according to the deletion flag carried by any historical version data in the at least two historical version data, an associated data file of the data to be processed and data association information corresponding to the associated data file in the data files corresponding to the at least two historical version data; and deleting the deletion flag and historical version data included in the associated data file when it is determined according to the data association information that the associated data file meets a preset deletion condition.
[0011] The above method obtains at least two historical version data of the data to be processed, as well as the data files corresponding to each historical version data. According to the deletion marks carried by any version data in the at least two historical version data, the associated data file of the data to be processed and the data association information corresponding to the associated data file are determined from the data files corresponding to the at least two historical version data. When it is determined that the associated data file meets the preset deletion condition according to the data association information, the deletion marks and historical version data included in the associated data file are deleted, realizing the timely deletion of the deletion marks and historical version data, avoiding the accumulation of a large number of undeleted deletion marks and historical version data in the storage space, and further avoiding affecting the data query performance. Description of the Drawings Figure 1 is a schematic diagram of an application scenario of a data processing method provided by an embodiment of this specification; Figure 2 is a flowchart of a data processing method provided by an embodiment of this specification; Figure 3 is a schematic diagram of the process of determining an associated data file in a data processing method provided by an embodiment of this specification; Figure 4 is a flowchart of the processing process of a data processing method provided by an embodiment of this specification; Figure 5 is a schematic diagram of the structure of a data processing device provided by an embodiment of this specification; Figure 6 is a block diagram of the structure of a computing device provided by an embodiment of this specification. Detailed Description of the Embodiments
[0012] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.
[0013] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0014] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0015] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0016] First, the noun terms involved in one or more embodiments of this specification are explained.
[0017] LSM: Log-Structured Merge Tree, a data structure used in storage systems, especially suitable for scenarios that require efficient processing of write operations. The LSM tree improves write efficiency by converting random write operations into sequential write operations and manages data through a hierarchical storage structure.
[0018] KV: Key-Value, a key-value pair is a simple data storage model, where each data item consists of a unique identifier (key) and associated data (value). This model is very suitable for fast lookup and update operations because it allows direct access to the corresponding value through the key.
[0019] map: A data structure or mechanism used to manage and optimize data storage and access.
[0020] Bloom filter: Placed in the metadata area of each data file, used to quickly check whether a KV exists in the current data file. False positives may occur, indicating that the KV exists in the current data file, but false negatives cannot occur, indicating that the KV does not exist in the current data file.
[0021] In practical applications, in a database designed based on LSM, the data written by users is first cached in memory and then sorted and written to disk. The data that has been written to disk cannot be modified anymore. Therefore, for the update and deletion operations of data (i.e., KV data), new data needs to be written to implement data update and deletion. During query, the Bloom filter in each data file can be used to determine whether to scan the data file. For all data files to be scanned, the Bloom filter creates an iterator, and a priority queue is used to manage the iterator so that the data with a newer version will be scanned first. Based on this, all historical versions of the scanned data can be queried.
[0022] Then, when the data on disk is updated frequently, new data will be continuously written. As a result, multiple versions of a KV data will be scattered in different data files, which will seriously affect the data query efficiency. Based on this, the database can periodically trigger a data sorting process (minor compaction) to merge and sort the data files. Based on this data sorting process, several data files of similar sizes can be selected, and the data in these data files can be cleaned, sorted, and merged into a new data file. In addition, for larger data files, data sorting is usually not frequently performed based on this data sorting process, but a data merging process (major compaction) is triggered. The triggering time interval of this data merging process is longer than that of the above data sorting process. This data merging process can select all data files and merge them into a large data file. Usually, the triggering time interval of this data merging process is relatively long, such as 20 days, 30 days, etc.
[0023] In this database, for the deletion operation of data, a deletion marker (deletemarker) will be written to the data. During the data query process, if the data with this deletion marker is scanned, then all previous historical versions of this data will be ignored, thus achieving the deletion effect. However, these deletion markers can usually only be deleted through the above data merging process because the data sorting process cannot guarantee that all data files containing all historical versions of a piece of data can be selected. If the deletion marker is cleared based on the data sorting process, then some historical version data that has not been deleted will be queried in subsequent data queries, which is contrary to the previously called deletion operation. If users update and delete data in the database frequently, then a large number of deletion markers and historical version data will accumulate in the database, which will greatly affect the query performance. Therefore, an effective method is urgently needed to solve the above technical problems.
[0024] In this specification, a data processing method is provided. This specification also relates to a data processing device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0025] See Figure 1 , Figure 1 FIG. shows a schematic diagram of an application scenario of a data processing method provided according to an embodiment of this specification. The data processing method includes: Obtain at least two historical version data of the data to be processed, and the data files corresponding to the at least two historical version data; According to the deletion marks carried by any historical version data in the at least two historical version data, determine the associated data file of the data to be processed and the data association information corresponding to the associated data file in the data files corresponding to the at least two historical version data; When it is determined that the associated data file meets the preset deletion condition according to the data association information, delete the deletion marks and historical version data included in the associated data file.
[0026] Specifically, Figure 1 includes an edge device 102 and a distributed database 104.
[0027] In specific implementation, the user can send a data query request to the distributed database 104 through the edge device 102. The distributed database 104 responds to the data query request, determines the data to be processed corresponding to the data query request, obtains at least two historical version data of the data to be processed, and the data files corresponding to the at least two historical version data. According to the deletion marks carried by any historical version data in the at least two historical version data, determine the associated data file of the data to be processed and the data association information corresponding to the associated data file in the data files corresponding to the at least two historical version data. When it is determined that the associated data file meets the preset deletion condition according to the data association information, delete the deletion marks and historical version data included in the associated data file, so as to realize the timely deletion of the deletion marks and historical version data, avoid accumulating a large number of undeleted deletion marks and historical version data in the storage space, and further avoid affecting the data query performance. The edge device 102 may include a browser, an APP (Application), or a web application such as an H5 (Hyper Text Markup Language 5) application, or a light application (also known as a mini-program, a lightweight application program), or a cloud application, etc. The edge device may be developed based on the software development kit (SDK) of the corresponding service provided by the server side, such as developed based on the real-time communication (RTC) SDK. The edge device may be deployed in an electronic device and needs to rely on the device or certain APPs in the device to run, etc. The electronic device may have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications may usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0028] See Figure 2 , Figure 2 FIG. shows a flowchart of a data processing method according to an embodiment of the present specification, which is applied to a database based on the LSM storage engine and specifically includes the following steps.
[0029] Step 202: Obtain at least two historical version data of the data to be processed, and the data files corresponding to the at least two historical version data.
[0030] Specifically, the data processing method provided in the embodiments of the present specification may be applied to a database, specifically to a distributed database, which may be constructed based on the LSM storage engine.
[0031] Then, the data to be processed can be understood as the data stored in a distributed database. Further, the data to be processed can be understood as the data stored on a disk that cannot be modified. At least two historical version data of the data to be processed can be understood as all versions of the data to be processed after being updated. For example, for the data to be processed A, the user writes data A1 on January 2, writes data A2 on January 3, and writes data A3 on January 5. The data A1, A2, and A3 are all historical version data of the data A. It can be understood that there is a writing order relationship between at least two historical version data of the data to be processed. The data version written earlier is less than the data version written later. When querying the data, the new version of the data to be processed will be scanned first. The data files corresponding to at least two historical version data can be understood as the data files storing each historical version data among at least two historical version data. For example, data A1 is stored in data file F1, data A2 is stored in data file F2, and data A3 is stored in data file F3. Then, the data files F1, F2, and F3 are the data files corresponding to at least two historical version data. It can be understood that among at least two historical version data, multiple historical version data can also be stored in the same data file. For example, data A1 and data A2 are stored in data file F1, and data A3 is stored in data file F2.
[0032] Based on this, the distributed database can be scanned to obtain at least two historical version data of the data to be processed, and the data files storing each historical version data among at least two historical version data can be obtained.
[0033] In practical applications, before obtaining at least two historical version data of the data to be processed and the data files corresponding to at least two historical version data, it further includes: In response to a data query request, determine the data to be processed corresponding to the data query request.
[0034] Among them, the data query request can be understood as a request for querying the data to be processed received by the distributed database.
[0035] Specifically, in order to ensure the deletion marks that affect the data reading performance are cleared, in response to the data query request, the data to be processed that the data query request wants to query can be determined, and then at least two historical version data of the data to be processed and the corresponding data files are obtained.
[0036] In summary, in response to a data query request, by performing subsequent statistical operations on the deletion marks of the data to be processed that the user wants to query and determining the associated data files, the statistical operations of the deletion marks are placed at the moment when the user reads, so that the deletion marks that may affect the user's reading performance can be cleared, without deleting all deletion marks and wasting computing resources.
[0037] Step 204: According to the deletion marks carried by any historical version data in the at least two historical version data, in the data files corresponding to the at least two historical version data, determine the associated data file of the data to be processed and the data association information corresponding to the associated data file.
[0038] Among them, the deletion mark can be understood as a mark written to the data stored in the distributed database. If the data carries this deletion mark, it means that the data is marked for deletion, and all historical version data of this data in the distributed database will be deleted. When querying data, if it is scanned that the data carries this deletion mark, then all historical version data of this data will not be queried. For example, data A1, A2, and A3 are all historical version data of data A, and the data version of A3 is greater than the data version of A2 which is greater than the data version of A1, that is, A3 is the new version data of data A. In response to a data query request for data A3, if it is determined that data A3 carries a deletion mark, then data A1 and data A2 will not be returned as query results. The associated data file of the data to be processed can be understood as a data file that stores historical version data that can be deleted by the deletion mark. For example, data A1 is stored in data file F1, data A2 is stored in data file F2, and data A3 is stored in data file F3. When data A3 carries a deletion mark, it means that data A2 and data A1 will both be deleted by this deletion mark. Then the associated data files of data file F3 are data files F1 and F2, the associated data file of data file F1 is data file F3, and the associated data file of data file F1 is data file F3. Then, data files F1, F2, and F3 are all associated data files of the data to be processed A. The data association information corresponding to the associated data file can be understood as the file information and the number of files of the data files that have an association relationship with the associated data file. For example, for data file F3, its associated data files are F1 and F2, then the data association information of this data file F3 can be "F1+1, F2+1", the file information is F1 and F2, and the number of files is 2.
[0039] Specifically, based on the deletion marks carried by any one or more of the at least two historical version data, the data files with an association relationship can be determined among the data files corresponding to the at least two historical version data, so as to determine the associated data file of the data to be processed and the data association information corresponding to the associated data file.
[0040] Specifically, when implementing, determining the associated data file of the data to be processed and the data association information corresponding to the associated data file based on the deletion marks carried by any historical version data among the at least two historical version data in the data files corresponding to the at least two historical version data includes: Based on the deletion marks carried by any historical version data among the at least two historical version data, determine the associated data file of the data to be processed in the data files corresponding to the at least two historical version data; Based on the deletion marks, determine the association relationship between the any historical version data and other historical version data; Based on the association relationship, record the data association information corresponding to the associated data file.
[0041] Among them, other historical version data can be understood as other historical version data among the at least two historical version data except the historical version data carrying the deletion mark. The association relationship between any historical version data and other historical version data can be understood as the relationship of whether the other historical version data can be deleted by the deletion mark carried by the any historical version data.
[0042] Specifically, based on the deletion marks carried by any historical version data among the at least two historical version data, the associated data file of the data to be processed can be determined in the data files corresponding to the at least two historical version data, and based on the deletion marks, the association relationship between the any historical version data and other historical version data can be determined, and the data association information corresponding to the associated data file can be recorded based on the association relationship.
[0043] Further, when determining the associated data file, based on the sequence relationship between the data versions of the at least two historical version data, it can be determined in turn whether the data file corresponding to each historical version data is an associated data file. If so, the associated data file with an association relationship to the associated data file is further determined. For example, it can be determined starting from the data file corresponding to the historical version data of the latest version.
[0044] For example, data A1, A2, and A3 are all historical version data of data A, and the data version of A3 is greater than that of A2, which is greater than that of A1. That is, A3 is the new version data of data A. Data A1 is stored in data file F1, data A2 is stored in data file F2, and data A3 is stored in data file F3. When A3 carries a deletion mark, the deletion mark carried by data A3 can delete data A1 and A2. Then, there is an association relationship between data A3 and data A1, A2. Then, there is an association relationship between data file F3 where data A3 is stored and data files F1 and F2 where data A1 and A2 are stored. Then, this data file F3 is used as an associated data file, and the data association information corresponding to this associated data file F3 can include the file information and the number of files of data file F1 and data file F2, that is, "F1 + 1, F2 + 1".
[0045] In summary, by counting the deletion marks of the data to be processed, the associated data files associated with the data to be processed are obtained, providing information for the subsequent deletion of the deletion marks.
[0046] Further, the determining the associated data file of the data to be processed in the data files corresponding to the at least two historical version data according to the deletion mark carried by any historical version data in the at least two historical version data includes: When it is determined that any historical version data in the at least two historical version data carries a deletion mark, the data file corresponding to the any historical version data is determined as the first associated data file; According to the deletion mark, the second associated data file associated with the first associated data file is determined from the data files corresponding to the at least two historical version data; The first associated data file and the second associated data file are determined as the associated data files of the data to be processed.
[0047] Specifically, after determining the data file corresponding to any historical version data as the first associated data file, the first associated data file can be placed in a queue for temporary storage, and then the determined second associated data file is also placed in this queue for temporary storage. All the associated data files temporarily stored in the queue are the associated data files of the data to be processed.
[0048] For example, data A1, A2, and A3 are all historical version data of data A, and the data version of A3 is greater than that of A2, and the data version of A2 is greater than that of A1. That is, A3 is the new version data of data A. Data A1 is stored in data file F1, data A2 is stored in data file F2, and data A3 is stored in data file F3. When A3 carries a deletion mark, the data file F3 in which data A3 is stored is determined as the first associated data file. The deletion mark carried by data A3 can delete data A1 and A2. Then, there is an association relationship between data A3 and data A1, A2. Then, there is an association relationship between the data file F3 in which data A3 is stored and the data files F1 and F2 in which data A1 and A2 are stored. Then, data files F1 and F2 can be determined as the second associated data files, and the first associated data file F3 and the second associated data files F1 and F2 are determined as the associated data files of the data to be processed.
[0049] In summary, by determining the associated data files of the data to be processed based on the deletion mark, it is convenient to record the data association information according to the associated data files subsequently, and further convenient to determine whether to delete the deletion mark and historical version data included in the associated data files subsequently.
[0050] Step 206: When it is determined that the associated data file meets the preset deletion condition according to the data association information, delete the deletion mark and historical version data included in the associated data file.
[0051] Specifically, it can be determined whether the associated data file meets the preset deletion condition according to the data association information. When it is determined that the associated data file meets the preset deletion condition, delete the deletion mark and historical version data included in the associated data file.
[0052] In specific implementation, it can be determined that the associated data file meets the preset deletion condition when it is determined that the data association information meets the preset association threshold and / or the number of files of the associated data file meets the preset number threshold.
[0053] In practical applications, there are multiple associated data files; After determining the associated data files of the data to be processed and the data association information corresponding to the associated data files, it further includes: Create an associated data file set according to the multiple associated data files of the data to be processed and the data association information corresponding to each associated data file; Calculate the set association information of the associated data file set according to the data association information corresponding to each associated data file; When it is determined that the associated data file meets the preset deletion condition according to the data association information, deleting the deletion mark and historical version data included in the associated data file includes: When it is determined that the set of associated data files meets the preset deletion condition according to the set association information, deleting the deletion marks and historical version data included in the multiple associated data files.
[0054] Among them, the set of associated data files may include associated data files and the data association information corresponding to the associated data files. The set association information can be understood as the sum of the data association information corresponding to each associated data file included in the set of associated data files.
[0055] Specifically, a set of associated data files can be created according to multiple associated data files of the data to be processed and the data association information corresponding to each associated data file in the multiple associated data files, and the sum of the data association information of the multiple associated data files can be calculated as the set association information of the set of associated data files according to the data association information of the multiple associated data files in the set of associated data files. When it is determined that the set of associated data files meets the preset deletion condition according to the set association information, all the deletion marks and historical version data included in the multiple associated data files in the set of associated data files are deleted.
[0056] In practical applications, the set of associated data files can be a file queue, and a map can be used to store the data association information. See Figure 3 , Figure 3 shows a schematic flow chart of determining associated data files in a data processing method provided according to an embodiment of the present specification. As Figure 3 shown, taking the example of a user querying data KV during a certain data query, when querying data, four historical version data V4, V3, V2, and V1 of the data KV to be processed are scanned, where the version of V4 is greater than the version of V3, which is greater than the version of V2, which is greater than the version of V1.
[0057] First, a map is used to store the data association information, and a file queue is used to temporarily store the data files of the historical version data. At this time, the file queue is empty. First, the historical version data V4 is scanned, and it is determined that the historical version data V4 carries a deletion mark. Then, the data file D corresponding to the historical version data V4 is put into the file queue, indicating that the data file D is an associated data file of the data KV to be processed.
[0058] Next, the historical version data V3 is scanned. The historical version data V3 does not carry a deletion mark. At this time, the file queue is checked and it is found that there is a data file D in the file queue. Since the deletion mark carried by the historical version data V4 will delete the historical version data V3, there is an association relationship between the historical version data V4 and V3. And the historical version data V3 is stored in the data file C. That is to say, there is also an association relationship between the data file D and the data file C. Then, the data file C is also put into the file queue. The association of the historical version data V3 is recorded in the map of the data file D, that is, it is recorded that the data file D is associated with the data file C, and the number of associated files is 1. Then, the data association information of the data file D is "C+1" at this time. At the same time, the association of the historical version data V4 is recorded in the map of the data file C. At this time, the data association information of the data file C is "D+1".
[0059] Next, the historical version data V2 stored in the data file B is scanned. The historical version data V2 carries a deletion mark. The file queue is checked and it is found that there are a data file D and a data file C in the file queue. And the deletion mark carried by the historical version data V4 stored in the data file D will delete the historical version data V2. That is to say, there is an association relationship between the historical version data V4 and the historical version data V2. That is to say, there is also an association relationship between the data file D and the data file B. Then, the data file B is also put into the file queue. The association of the historical version data V2 is recorded in the map of the data file D. Then, the data association information of the data file D is "C+1; B+1" at this time. The association of the historical version data V4 is recorded in the map of the data file B. Then, the data association information of the data file B is "D+1" at this time.
[0060] Next, the historical version data V1 stored in the data file A is scanned. The historical version data V1 does not carry a deletion mark. The file queue is checked, and it is found that there are data files D, C, and B in the file queue. Among them, the deletion mark carried by the historical version data V4 stored in the data file D will delete the historical version data V1. That is to say, there is an association relationship between the historical version data V4 and the historical version data V1. That is to say, there is also an association relationship between the data file D and the data file A. Then, the data file A is put into the file queue, and the association of the historical version data V1 is recorded in the map of the data file D. At this time, the data association information of the data file D is "C+1; B+1; A+1". The association of the historical version data V4 is recorded in the map of the data file A. At this time, the data association information of the data file A is "D+1"; the deletion mark carried by the historical version data V2 stored in the data file B will also delete the historical version data V1. That is to say, there is an association relationship between the historical version data V2 and the historical version data V1. That is to say, there is also an association relationship between the data file B and the data file A. Then, the association of the historical version data V1 is recorded in the map of the data file B. At this time, the data association information of the data file B is "D+1; A+1". The association of the historical version data V2 is recorded in the map of the data file A. At this time, the data association information of the data file A is "D+1; B+1".
[0061] Then, the data files A, B, C, and D contained in the file queue are the associated data files of the data KV to be processed, and the set of associated data files is the file queue.
[0062] In summary, by scanning the historical version data in sequence, the statistics of the association relationships between the historical version data are realized, which is convenient for subsequent determination of whether a purge operation is required according to the data association information, and further ensures that the data query performance is not affected.
[0063] In specific implementation, after determining the above-mentioned associated data files, a data file selection strategy can be executed according to the associated data files and the data association information. This data file selection strategy is the process of determining the associated data files as described above. It can be understood that the larger the sum of the association values of the data association information recorded in the map of an associated data file, the more deletion marks and historical version data are associated with the associated data file, and the higher the cleaning priority of the associated data file. Since each associated data file maintains a map, all the data files in the distributed database can be grouped through the association of these maps. Combining the above Figure 3In the example shown, the associated data files A, B, C, and D can belong to the same group, that is, the set of associated data files. The associated data files included in the group with the largest sum of associated values can be marked for deletion and the historical version data can be cleared. The specific implementation method is as follows.
[0064] When it is determined that the set of associated data files meets the preset deletion condition according to the set association information, deleting the deletion marks and historical version data included in the multiple associated data files includes: When it is determined that the set association information meets the preset association threshold, it is determined that the set of associated data files meets the preset deletion condition, and the deletion marks and historical version data included in the multiple associated data files are deleted; or When it is determined that the number of files of the associated data files meets the preset number threshold, it is determined that the set of associated data files meets the preset deletion condition, and the deletion marks and historical version data included in the multiple associated data files are deleted.
[0065] In an embodiment of this specification, the sum of the data association information of all the associated data files in the set of associated data files can be calculated as the set association information. When it is determined that the set association information meets the preset association threshold, it is determined that the set of associated data files meets the preset deletion condition, and the deletion marks and historical version data included in the multiple associated data files in the set of associated data files are deleted.
[0066] In another embodiment of this specification, the number of files of the associated data files included in the set of associated data files can be determined. When it is determined that the number of files meets the preset number threshold, it can be determined that the set of associated data files meets the preset deletion condition, and the deletion marks and historical version data included in the multiple associated data files in the set of associated data files are deleted.
[0067] In practical applications, the preset association threshold can be obtained by multiplying the total number of KV data in all the associated data files included in the set of associated data files by the ratio threshold R, and the ratio threshold R can be set manually and adjusted flexibly. When it is determined that the user frequently updates the data in the distributed database, the ratio threshold R can be set larger, and vice versa, the ratio threshold R can be set smaller.
[0068] Then, further, for a distributed database, the total associated value score of the maps of all data files in the distributed database can be calculated periodically. When it is determined that the total associated value score is greater than the preset associated threshold, the data file selection strategy is executed according to the above process. The data files are grouped by the maps of all data files, and multiple associated data file sets are obtained according to the above process. The associated data file set with the highest associated value score (i.e., the set association information) is selected, and the deletion marks and historical version data of all associated data files in the associated data file set are deleted.
[0069] In summary, by selecting the associated data file set with the most associated deletion marks, the deletion of the deletion marks in the associated data file set can be realized in a timely manner, without waiting until the data merging process is triggered to clear the deletion marks, ensuring the data reading performance of the distributed database.
[0070] Further, the deletion of the deletion marks and historical version data included in the multiple associated data files includes: Obtain all data files in the data partition where the data to be processed is stored; According to the data to be processed, filter the all data files to obtain a data filtering result, and obtain a reference data file containing the data to be processed according to the data filtering result; When it is determined that the associated data file set includes the reference data file, delete the deletion marks and historical version data included in the multiple associated data files.
[0071] Among them, the data partition can be understood as the basic unit of data management in the distributed database.
[0072] Specifically, after determining the associated data file set, before deleting the deletion marks and historical version data included in the associated data files in the associated data file set, the associated data file set can be checked once to ensure that the data files corresponding to all historical version data of the data to be processed are included in the associated data file set, avoiding the omission of historical version data. Then, all data files in the data partition where the data to be processed is stored can be obtained, and according to the data to be processed, the all data files are filtered to obtain a data filtering result, and a reference data file containing the data to be processed is obtained according to the data filtering result. When it is determined that the associated data file set includes all reference data files, it means that there is no omission of historical version data, and at this time, the deletion operation of the deletion marks and historical version data can be performed.
[0073] For example, the associated data file set includes associated data file A, associated data file B, associated data file C, and associated data file D. Data filtering is performed on all data files in the data partition, and the data filtering result determines that the reference data files containing the data to be processed are reference data file A, reference data file B, and reference data file C. At this time, the associated data file set contains all reference data files, indicating that there is no omission. At this time, a deletion operation is performed; if the data filtering result determines that the reference data files containing the data to be processed are reference data file A, reference data file B, and reference data file E, since the reference data file E is not included in the associated data file set, it indicates that there is an omission at this time, and the deletion mark and the deletion operation of the historical version data are not performed.
[0074] In practical applications, the Bloom filters of all data files in the data partition can be checked, and the data files that return positive (i.e., reference data files) are recorded. It is checked whether these positive data files are included in the associated data file set. If so, the deletion marks and historical version data in the associated data file set can be cleared. If not, the deletion marks and historical version data in the associated data file set cannot be cleared because it indicates that there is historical version data of the data to be processed that is not included in the associated data file set.
[0075] In summary, by performing data checks before the deletion operation, it is ensured that the associated data file set to be deleted can contain all historical version data of the data to be processed, avoiding the omission of historical version data, and thus avoiding the error situation of querying historical version data caused by the deletion of the deletion mark while the historical version data is not deleted.
[0076] In addition, after deleting the deletion marks and historical version data included in the multiple associated data files, it further includes: Obtain the remaining data included in the multiple associated data files; Merge the remaining data to obtain a target data file.
[0077] Among them, the remaining data can be understood as other KV data in the associated data file except for the historical version data.
[0078] Specifically, after deleting the deletion marks and historical version data included in the multiple associated data files, the other KV data in the multiple associated data files except for the historical version data can be merged to obtain a target data file, realizing file merging.
[0079] In practical applications, the method further includes: Determine the data partition where the data to be processed is stored; In the case where it is determined that the data storage amount of the data partition is greater than the preset storage threshold, perform a splitting process on the data partition to obtain the target data partitions obtained after splitting.
[0080] Among them, the target data partitions obtained after splitting can be at least two data partitions. For example, one data partition can be split into two data partitions.
[0081] Specifically, it is possible to determine the data partition in which the data to be processed is stored, monitor the data storage amount of this data partition, and in the case where it is determined that the data storage amount of the data partition is greater than the preset storage threshold, split the data partition to obtain the target data partitions obtained after splitting, so that the distributed database can respond to massive data with high elasticity and ensure data storage performance.
[0082] However, in practical applications, the preset storage threshold can be a threshold of a fixed size, that is, whenever the data storage amount of a data partition reaches a threshold of a fixed size (for example, it can be 8GB), the data partition is split into two new data partitions. This method is suitable for storing larger data tables, that is, using multiple data partitions to perform data sharding storage on a larger data table. However, for relatively small data tables, such a method will result in a small number of data partitions storing the relatively small data table, which is not conducive to the load balancing of the database. Moreover, if this relatively small data table is updated frequently, even if the above data processing method is used to clear the deletion marks and historical version data, it will also result in a larger clearing granularity, affecting resources and performance. Based on this, the following data partition splitting strategy can be used to dynamically calculate the preset storage threshold according to the number of partitions of the data partition corresponding to the data table. The specific implementation method is as follows.
[0083] After determining the data partition in which the data to be processed is stored, it further includes: Determine the data table corresponding to the data to be processed, and determine the number of partitions of the data partition in which the data table is stored; Calculate the preset storage threshold according to the number of partitions.
[0084] Among them, the data table corresponding to the data to be processed can be understood as the data table to which the data to be processed belongs, and this data table can be understood as the data table that has enabled the above data file selection strategy.
[0085] Specifically, for the data table corresponding to the data to be processed, it is possible to determine the number of partitions of the data partition in the data storage node of the data table in the distributed database, and calculate the preset storage threshold according to this number of partitions, which is convenient for subsequent splitting operations of the data partition according to this preset storage threshold.
[0086] In practical applications, the goal of this data partition splitting strategy is to quickly split the data partitions of a data table into a certain number, and the preset storage threshold can be distinguished according to the above data processing method. If the data file selection strategy of the above data processing method is adopted, the preset storage threshold can be set relatively large so that the data partitions have a smaller granularity, and further make the deletion mark and the clearing operation of historical version data more rapid. Specifically, the number of partitions of the data partitions of the data table in the current data storage node can be determined, and the preset storage threshold is obtained by multiplying the number of partitions by a custom constant and then multiplying by 128MB. When it is determined that the data storage amount of the data partition is greater than the preset storage threshold, it is determined that the data partition needs to be split.
[0087] It can be understood that the monitoring of the above storage amount can be performed on the data partitions stored in all data tables in the distributed database.
[0088] In summary, by calculating the preset storage threshold according to the number of partitions, it is convenient to perform the splitting operation of the data partition according to the preset storage threshold later, and ensure that the computing resources and data storage performance of the distributed database are not affected.
[0089] Further, the determination of the number of partitions of the data partitions stored in the data table includes: Determine the number of partitions of the data partitions stored in the data table at a preset time interval; The calculation of the preset storage threshold according to the number of partitions includes: Calculate a first storage threshold according to the number of partitions at the preset time interval; Obtain the second storage threshold calculated in the previous time interval, and determine the preset storage threshold according to the comparison result between the first storage threshold and the second storage threshold.
[0090] Among them, the preset time interval can be understood as the time interval for calculating the preset storage threshold. For example, the distributed database can calculate the preset storage threshold every 10 minutes, or it can also calculate the preset storage threshold every 1 hour or 1 day. The embodiments of this specification do not limit this. Then, the first storage threshold can be understood as the storage threshold calculated in the current time interval, and the second storage threshold can be understood as the storage threshold calculated in the previous time interval. Specifically, the number of partitions of the data partition stored in the current time interval data table can be determined at a preset time interval. In the current time interval, the first storage threshold is calculated according to the number of partitions, and the second storage threshold calculated in the previous time interval of the current time interval is obtained. The first storage threshold and the second storage threshold are compared to obtain a comparison result. When it is determined according to the comparison result that the first storage threshold is greater than the second storage threshold, the first storage threshold is determined as the preset storage threshold; when it is determined according to the comparison result that the first storage threshold is less than the second storage threshold, the second storage threshold is determined as the preset storage threshold. That is to say, the larger storage threshold among the first storage threshold and the second storage threshold is determined as the preset storage threshold.
[0091] It can be understood that in practical applications, the smaller storage threshold among the first storage threshold and the second storage threshold can also be determined as the preset storage threshold. That is, when it is determined according to the comparison result that the first storage threshold is greater than the second storage threshold, the second storage threshold is determined as the preset storage threshold; when it is determined according to the comparison result that the first storage threshold is less than the second storage threshold, the first storage threshold is determined as the preset storage threshold.
[0092] In addition, since the preset storage threshold is positively correlated with the number of partitions of the data partition in the data storage node, when the nodes of the distributed database are expanded, some data partitions will be moved to the expanded data storage nodes, resulting in a decrease in the preset storage threshold. Based on this, when calculating the preset storage threshold each time, the calculated preset storage threshold can be recorded in a variable. Then, the second storage threshold calculated in the previous time interval can be obtained from this variable. Moreover, when calculating the preset storage threshold for the first time, it indicates that the data partition is restarted, that is, the data partition has undergone a load balancing migration or a migration caused by expanding the data storage node. At this time, the preset storage threshold can be calculated according to the following formula: (data storage amount of the data partition / 128MB × constant + 1.0) × 128MB × constant.
[0093] In summary, by dynamically calculating the preset storage threshold, it is possible to prevent the distributed database from splitting and avalanching after node expansion, and ensure the load balance of the data partition.
[0094] In practical applications, a scenario simulation of high-frequency updates was conducted on a distributed database that executes the above data processing method. Data queries were continuously performed on the secondary index table in the distributed database, and the data therein was updated. At this time, a large number of deletion marks would be generated in the secondary index table. Since the initial data volume of the secondary index table was 10 million rows, when the above data processing method was not executed, these initial data volumes would be concentrated in a 600MB data file. For a long time thereafter, there would be multiple deletion marks in this data file. However, due to the large size of this data file, the data sorting process could not select this data file, and it could only wait until the data merging process triggered after a period of time to clear and merge the deletion marks in this data file. However, when the above data processing method was executed and the ratio threshold R was set to 2, the deletion marks and historical version data would be found and deleted every half hour, and the data reading response time (i.e., read RT) was controlled. When the above data processing method was executed and the ratio threshold R was set to 1, the deletion marks and historical version data would be found and deleted every 6 minutes, the time to trigger the deletion operation was faster, and the data reading response time was controlled lower, thereby ensuring the data reading performance.
[0095] In summary, the above method determines at least two historical version data of the data to be processed, as well as the data files corresponding to each historical version data. According to the deletion marks carried by any version data in the at least two historical version data, the associated data file of the data to be processed and the data association information corresponding to the associated data file are determined from the data files corresponding to the at least two historical version data. When it is determined that the associated data file meets the preset deletion condition according to the data association information, the deletion marks and historical version data included in the associated data file are deleted, realizing the timely deletion of the deletion marks and historical version data, avoiding the accumulation of a large number of undeleted deletion marks and historical version data in the storage space, and further avoiding affecting the data query performance. The following combines the attached Figure 4 , taking the application of the data processing method provided in this specification in a distributed database as an example, to further illustrate the data processing method. Among them, Figure 4 shows the processing process flowchart of a data processing method provided by an embodiment of this specification, which specifically includes the following steps.
[0096] Step 402: In response to a data query request, determine the data to be processed corresponding to the data query request.
[0097] Step 404: Obtain at least two historical version data of the data to be processed, as well as the data files corresponding to the at least two historical version data.
[0098] Step 406: When it is determined that any of the at least two historical version data carries a deletion mark, determine the data file corresponding to the any historical version data as the first associated data file.
[0099] Step 408: According to the deletion mark, determine the second associated data file associated with the first associated data file from the data files corresponding to the at least two historical version data.
[0100] Step 410: Determine the first associated data file and the second associated data file as the associated data files of the data to be processed.
[0101] Step 412: According to the deletion mark, determine the association relationship between any historical version data and other historical version data, and record the data association information corresponding to the associated data file according to the association relationship.
[0102] Step 414: When it is determined that the data association information meets the preset association threshold, determine that the associated data file meets the preset deletion condition, and delete the deletion mark and historical version data included in the associated data file.
[0103] Step 416: Obtain the remaining data included in the multiple associated data files, and merge the remaining data to obtain the target data file.
[0104] In summary, the above method obtains at least two historical version data of the data to be processed, and the data files corresponding to each historical version data. According to the deletion mark carried by any version data in the at least two historical version data, the associated data file of the data to be processed and the data association information corresponding to the associated data file are determined from the data files corresponding to the at least two historical version data. When it is determined that the associated data file meets the preset deletion condition according to the data association information, the deletion mark and historical version data included in the associated data file are deleted, realizing the timely deletion of the deletion mark and historical version data, avoiding the accumulation of a large number of undeleted deletion marks and historical version data in the storage space, and further avoiding affecting the data query performance. Corresponding to the above method embodiment, this specification also provides an embodiment of a data processing device, which is applied to a database based on the LSM storage engine. Figure 5 The structural schematic diagram of a data processing device provided by an embodiment of this specification is shown. As Figure 5 shown, the device includes: An acquisition module 502, configured to acquire at least two historical version data of the data to be processed, and the data files corresponding to the at least two historical version data; A determination module 504, configured to determine, in data files corresponding to the at least two historical version data, an associated data file of the data to be processed and data association information corresponding to the associated data file according to a deletion flag carried in any of the at least two historical version data; A deletion module 506, configured to delete the deletion flag and historical version data included in the associated data file when it is determined according to the data association information that the associated data file meets a preset deletion condition.
[0105] In an optional embodiment, the determination module 504 is further configured to: Determine, in data files corresponding to the at least two historical version data, an associated data file of the data to be processed according to a deletion flag carried in any of the at least two historical version data; Determine an association relationship between any historical version data and other historical version data according to the deletion flag; Record data association information corresponding to the associated data file according to the association relationship.
[0106] In an optional embodiment, the determination module 504 is further configured to: When it is determined that any of the at least two historical version data carries a deletion flag, determine a data file corresponding to the any historical version data as a first associated data file; Determine, from data files corresponding to the at least two historical version data, a second associated data file associated with the first associated data file according to the deletion flag; Determine the first associated data file and the second associated data file as associated data files of the data to be processed.
[0107] In an optional embodiment, there are multiple associated data files; The determination module 504 is further configured to: Create an associated data file set according to multiple associated data files of the data to be processed and data association information corresponding to each associated data file; Calculate set association information of the associated data file set according to the data association information corresponding to each associated data file; The deletion module 506 is further configured to: Delete the deletion flag and historical version data included in the multiple associated data files when it is determined according to the set association information that the associated data file set meets a preset deletion condition.
[0108] In an alternative embodiment, the deletion module 506 is further configured to: When it is determined that the set association information meets a preset association threshold, determine that the associated data file set meets a preset deletion condition, and delete the deletion marks and historical version data included in the multiple associated data files; or When it is determined that the number of files in the associated data file set meets a preset number threshold, determine that the associated data file set meets a preset deletion condition, and delete the deletion marks and historical version data included in the multiple associated data files.
[0109] In an alternative embodiment, the deletion module 506 is further configured to: Obtain all data files in the data partition where the data to be processed is stored; Perform data filtering on all the data files according to the data to be processed to obtain a data filtering result, and obtain a reference data file containing the data to be processed according to the data filtering result; When it is determined that the associated data file set includes the reference data file, delete the deletion marks and historical version data included in the multiple associated data files.
[0110] In an alternative embodiment, the apparatus further includes a merging module, configured to: Obtain the remaining data included in the multiple associated data files; Merge the remaining data to obtain a target data file.
[0111] In an alternative embodiment, the obtaining module 502 is further configured to: In response to a data query request, determine the data to be processed corresponding to the data query request.
[0112] In an alternative embodiment, the apparatus further includes a splitting module, configured to: Determine the data partition where the data to be processed is stored; When it is determined that the data storage amount of the data partition is greater than a preset storage threshold, perform a splitting process on the data partition to obtain a target data partition obtained after splitting.
[0113] In an alternative embodiment, the splitting module is further configured to: Determine the data table corresponding to the data to be processed, and determine the number of partitions of the data partition where the data table is stored; Calculate a preset storage threshold according to the number of partitions.
[0114] In an alternative embodiment, the splitting module is further configured to: Determine the number of partitions of the data partition stored in the data table at a preset time interval; Calculate a first storage threshold according to the number of partitions at the preset time interval; Obtain a second storage threshold calculated in the previous time interval, and determine a preset storage threshold according to the comparison result between the first storage threshold and the second storage threshold.
[0115] In summary, the above device obtains at least two historical version data of the data to be processed, and the data files corresponding to each historical version data. According to the deletion marks carried in any version data of the at least two historical version data, the associated data file of the data to be processed and the data association information corresponding to the associated data file are determined from the data files corresponding to the at least two historical version data. When it is determined that the associated data file meets the preset deletion condition according to the data association information, the deletion marks and historical version data included in the associated data file are deleted, so as to realize the timely deletion of the deletion marks and historical version data, avoid the accumulation of a large number of undeleted deletion marks and historical version data in the storage space, and further avoid affecting the data query performance. The above is a schematic solution of a data processing device according to this embodiment. It should be noted that the technical solution of the data processing device and the technical solution of the above data processing method belong to the same concept. For the details not described in the technical solution of the data processing device, reference can be made to the description of the technical solution of the above data processing method.
[0116] Figure 6 FIG. shows a structural block diagram of a computing device 600 according to an embodiment of the present specification. The components of the computing device 600 include but are not limited to a memory 610 and a processor 620. The processor 620 is connected to the memory 610 through a bus 630, and a database 650 is used to store data.
[0117] The computing device 600 further includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interfaces (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0118] In one embodiment of the present application, the above components of the computing device 600 and Figure 6 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 6 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.
[0119] The computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 600 can also be a mobile or stationary server.
[0120] Wherein, the processor 620 is configured to execute the following computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the above data processing method are implemented.
[0121] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the computing device, since it is basically similar to the embodiment of the data processing method, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the embodiment of the data processing method.
[0122] An embodiment of this specification also provides a computer-readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the above data processing method.
[0123] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the computer-readable storage medium, since it is basically similar to the embodiment of the data processing method, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the embodiment of the data processing method.
[0124] An embodiment of this specification also provides a computer program product including computer programs / instructions, which, when executed by a processor, implement the steps of the above data processing method.
[0125] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the description of the technical solution of the above data processing method.
[0126] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0127] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0128] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0129] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0130] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A data processing method, applied to a database based on an LSM storage engine, comprising: Acquire at least two historical versions of data to be processed, and data files corresponding to the at least two historical versions of data; According to the deletion mark carried by any historical version data of the at least two historical version data, in the data files corresponding to the at least two historical version data, determining the associated data file of the data to be processed and the data association information corresponding to the associated data file; When it is determined according to the data association information that the associated data file meets a preset deletion condition, the deletion mark and the historical version data contained in the associated data file are deleted.
2. The method according to claim 1, wherein determining, in the data files corresponding to the at least two historical versions of data, the associated data files of the data to be processed and the data association information corresponding to the associated data files according to the deletion mark carried by any of the at least two historical versions of data, comprises: Determine, according to a deletion mark carried by any historical version data of the at least two historical version data, in the data files corresponding to the at least two historical version data, a data file associated with the data to be processed; Determine, based on the deletion mark, the association relationship between the arbitrary historical version data and other historical version data; According to the association relationship, data association information corresponding to the associated data file is recorded.
3. The method according to claim 2, wherein determining, in the data files corresponding to the at least two historical versions of data, the associated data files of the data to be processed according to the deletion mark carried by any historical version of data in the at least two historical versions of data, comprises: In the case where it is determined that any historical version data among the at least two historical version data carries a deletion mark, determining a data file corresponding to the any historical version data as a first associated data file; Determining, according to the deletion mark, a second associated data file associated with the first associated data file from data files corresponding to the at least two historical versions of data; The first associated data file and the second associated data file are determined as associated data files of the data to be processed.
4. The method according to claim 1, wherein the associated data files are multiple; After determining the associated data file of the data to be processed and the data association information corresponding to the associated data file, the method further includes: Creating a set of associated data files according to the multiple associated data files of the data to be processed and the data association information corresponding to each associated data file; Calculating the set association information of the set of associated data files according to the data association information corresponding to each associated data file; The step of deleting the deletion mark and the historical version data contained in the associated data file when it is determined according to the data association information that the associated data file meets the preset deletion condition comprises: When it is determined according to the set association information that the set of associated data files meets a preset deletion condition, the deletion marks and the historical version data included in the plurality of associated data files are deleted.
5. The method according to claim 4, wherein when it is determined according to the set association information that the set of associated data files meets a preset deletion condition, deleting the deletion marks and historical version data contained in the plurality of associated data files comprises: In the case where it is determined that the set association information meets a preset association threshold, determining that the set of associated data files meets a preset deletion condition, and deleting the deletion marks and historical version data included in the plurality of associated data files; or When it is determined that the number of the associated data files meets a preset number threshold, it is determined that the associated data file set meets a preset deletion condition, and the deletion marks and historical version data included in the plurality of associated data files are deleted.
6. The method according to claim 4, wherein deleting the deletion marks and historical version data contained in the plurality of associated data files comprises: Acquire all data files in the data partition of the data storage to be processed; According to the data to be processed, data filtering is performed on all the data files to obtain data filtering results, and according to the data filtering results, a reference data file containing the data to be processed is obtained; In the case where it is determined that the associated data file set includes the reference data file, deletion marks and historical version data included in the plurality of associated data files are deleted.
7. The method according to any one of claims 4 to 6, after deleting the deletion marks and historical version data contained in the plurality of associated data files, further comprising: Obtaining remaining data contained in the plurality of associated data files; The remaining data are merged to obtain a target data file.
8. The method according to any one of claims 1 to 6, before obtaining at least two historical versions of the data to be processed and the data files corresponding to the at least two historical versions of the data, further comprising: In response to a data query request, to-be-processed data corresponding to the data query request is determined.
9. The method according to any one of claims 1 to 6, further comprising: Determine the data partition of the data storage to be processed; When it is determined that the data storage capacity of the data partition is greater than a preset storage threshold, the data partition is split to obtain a target data partition obtained after the split.
10. The method according to claim 9, after determining the data partition of the data storage to be processed, further comprising: Determine a data table corresponding to the data to be processed, and determine the number of partitions of the data partitions stored in the data table; A preset storage threshold is calculated according to the number of partitions.
11. The method according to claim 10, wherein determining the number of partitions of the data partition stored in the data table comprises: Determine the number of partitions of the data partition stored in the data table according to a preset time interval; Calculating a preset storage threshold according to the number of partitions includes: Calculating a first storage threshold according to the preset time interval and the number of partitions; A second storage threshold calculated at a previous time interval is obtained, and a preset storage threshold is determined according to a comparison result between the first storage threshold and the second storage threshold.
12. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 11 are implemented.
13. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.
14. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Data eliminating method, device and system
CN103678337A
Data storage method and device and data query method and device
CN114595224A
Metadata storage method and device and metadata query method and device
CN115221165A
Data storage management system and method based on K8s cluster and medium
CN118819390A
Data processing method and device and computer readable storage medium
CN119067093A