Data Block Splitting Method, Apparatus, Device, Medium and Product
By dividing the data files into intermediate files and end files during the data block splitting process, and only migrating the end files and modifying meta-information, the problem of read and write amplification in the existing technology is solved, and more efficient data block splitting is achieved.
Patent Information
- Application Number
- CN202410917850.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-09
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-07-09
AI Technical Summary
The prior art will lead to large read and write amplification when performing data block splitting, requiring reading all data files and recursive merging and redundant deletion.
By determining the data file combination and split key interval of the data block to be split, the data file is divided into intermediate files and end files, and only the end files are migrated and meta-information modified to avoid reading intermediate files.
It reduces read and write amplification during data block splitting, and improves the speed and efficiency of data block splitting.
Smart Images

Figure CN118885113B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of data processing, and in particular, to a data block splitting method, apparatus, device, medium and product. Background Art
[0002] In existing key-value storage systems, such as those based on RocksDB, the entire data range is divided into a series of continuous data blocks according to the prefix range of keys. Each data block is described by a left-closed right-open interval, such as [starting key, splitting key), and each data block includes several data files, and each data file corresponds to a key interval. When the amount of data in a data block increases to a certain extent, a splitting operation needs to be performed on the data block, that is, a splitting key is found in the middle of the data block, and the data between the starting key and the splitting key is split out from the original data block.
[0003] However, in the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0004] When performing data block splitting, the prior art needs to read all data files within all [starting key, splitting key), and then perform operations such as recursive merging and redundant deletion on the data within this range to obtain a new data block. Therefore, a large read-write amplification will occur in this process. Summary of the Invention
[0005] The embodiments of the present invention provide a data block splitting method, apparatus, device, medium and product to reduce the read-write amplification generated during data block splitting.
[0006] In a first aspect, the embodiments of the present invention provide a data block splitting method, the method comprising:
[0007] Determine a first data block to be split in the current data layer, and a first data file combination and a first split key interval corresponding to the first data block to be split;
[0008] For the first data file combination, use the data files with all keys distributed within the first split key interval as first intermediate files, and use the data files with some keys distributed within the first split key interval as first end files;
[0009] Modify the data block identifier in the meta-information of the first intermediate file to a first target data block identifier;
[0010] Migrate the part of the keys within the first split key interval and the data corresponding to the part of the keys in the first end file to the corresponding file in the next data layer, and set the data block identifier in the meta-information of this file to the first target data block identifier.
[0011] Second aspect, an embodiment of the present invention further provides a data block splitting device, including:
[0012] A first response module, configured to determine a first data block to be split in the current data layer, as well as a first data file combination and a first split key range corresponding to the first data block to be split;
[0013] A first file determination module, configured to, for the first data file combination, use a data file in which all keys are distributed within the first split key range as a first intermediate file, and use a data file in which some keys are distributed within the first split key range as a first end file;
[0014] A first identification modification module, configured to modify the data block identification in the meta information of the first intermediate file to a first target data block identification;
[0015] A first compaction module, configured to migrate some keys within the first split key range and the data corresponding to the some keys in the first end file to the corresponding file in the next data layer, and set the data block identification in the meta information of the file to the first target data block identification.
[0016] Third aspect, an embodiment of the present invention provides an electronic device, the electronic device includes:
[0017] One or more processors;
[0018] A memory, configured to store one or more programs;
[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the data block splitting method provided in any embodiment of the present invention.
[0020] Fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the data block splitting method provided in any embodiment of the present invention.
[0021] Fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the data block splitting method provided in any embodiment of the present invention.
[0022] The embodiments in the above-mentioned invention have the following advantages or beneficial effects:
[0023] After the first data file combination corresponding to the first data block to be split and the first split key range are determined, for each data file in the first data file combination, if all its keys are distributed within the first split key range, it is regarded as an intermediate file and retained in the current data layer, and only the data block identifier in its metadata is modified to the target data block identifier; since the intermediate file does not need to be read during the metadata modification process of the intermediate file, there will be no read-write amplification; if only some keys in the data file are distributed within the first split key range, the part of the keys distributed within the partial range of the first split key and the data corresponding to the part of the keys are migrated to the corresponding file in the next data layer, and the data block identifier in the metadata of the file is modified to the first target data block identifier. Since each data file in the key-value storage system is relatively small, the first end file is also small, and the data volume corresponding to the data within the split key range in the first end file is even smaller. Compared with the existing data block splitting method that needs to read all split data, in this embodiment, only the data within the split key range in the first end file needs to be compacted, so the read-write amplification during the data block splitting process can be greatly reduced, and at the same time, the splitting speed of the data block can be increased. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic flowchart of the data block splitting method provided by an embodiment of the present invention;
[0025] Figure 2 is another schematic flowchart of the data block splitting method provided by an embodiment of the present invention;
[0026] Figure 3 is a schematic structural diagram of the data block splitting device provided by an embodiment of the present invention;
[0027] Figure 4 is another schematic structural diagram of the data block splitting device provided by an embodiment of the present invention;
[0028] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings, rather than all the structures.
[0030] Figure 1The flowchart of the data block splitting method provided by the embodiments of the present invention. This embodiment is applicable to the situation of automatically completing the data block splitting operation of a key-value storage system. This method can be executed by a data block splitting device integrated in an electronic device, and this device can be implemented in a software and / or hardware manner. As Figure 1 shown, the method specifically includes the following steps:
[0031] S110. Determine the first data block to be split in the current data layer, as well as the first data file combination and the first split key range corresponding to the first data block to be split.
[0032] The key-value storage system stores data as a set of key-value pairs, where the key serves as the unique identifier of the value. Both the key and the value can be objects of any type. The key-value storage system usually includes multiple data layers, such as layer L0, layer L1, layer L2, etc.
[0033] In this embodiment, one data layer is used to store one or more complete data blocks and / or partial data in one or more complete data blocks. The data of one data block can be stored in one or more data layers. One data block usually includes multiple data files, and the key ranges corresponding to different data files are different. Among them, the data file is the file used by the data block to store data. The data storage amount of any data file cannot exceed the corresponding file data amount threshold, and the data storage amount of each data layer cannot exceed the corresponding layer data amount threshold. The layer data amount thresholds of different data layers are different. For example, as the data layer identifier is larger, the corresponding layer data amount threshold of the data layer is larger.
[0034] The split instruction used to trigger the corresponding data block splitting operation can be manually triggered by the user or automatically triggered by the key-value storage system. For the latter, the existing data block splitting trigger conditions can be used to trigger the split instruction, and this embodiment does not make specific limitations here.
[0035] The split key range can be understood as the key range corresponding to the data that needs to be split out from the first data block to be split. The split key range can be expressed as [start key, end key).
[0036] After the first data block to be split is determined, determine the first data file combination and the first split key range corresponding to the first data block to be split. The first data file combination includes at least one data file. The larger the split key range, the more the corresponding number of data files, and vice versa, the corresponding number of data files is usually less.
[0037] S120. For the first data file combination, use the data files with all keys distributed within the first split key range as the first intermediate files, and use the data files with some keys distributed within the first split key range as the first end files.
[0038] In a key-value storage system, the keys in each data file are stored in order, such as in ascending order. Therefore, some keys are distributed in the first split key range, which means that there are still some keys not distributed in the first split key range, and there is a clear split key between this part of the keys and the remaining keys, and this split key is the start key or the end key of the first split key range.
[0039] In one embodiment, the first end file includes a first start file and / or a first end file; determine the start key and the end key of the first split key range; use the data file whose corresponding key range spans the start key as the first start file, and use the data file whose corresponding key range spans the end key as the first end file.
[0040] Herein, "span" can be understood as crossing a set boundary. In this embodiment, it means crossing the start key or the end key. Taking the start key as an example, for a data file including the start key or the end key, if the start key is located at the start endpoint of the corresponding key range of a certain data file, then this data file should not be considered to span the start key, so it is the first intermediate file rather than the first start file. If the start key is located in the middle of the corresponding key range of a certain data file, then it is determined that this data file spans the start key and is the first start file.
[0041] In one embodiment, determine the start key and the end key of the first split key range, as well as the minimum key and the maximum key of each data file; use the data file whose minimum key is less than the start key and whose maximum key is greater than the end key as the first start file; use the data file whose minimum key is less than the end key and whose maximum key is greater than the end key as the first end file.
[0042] Specifically, use the minimum key of the first split key range as the start key, and use the maximum key of the first split key range as the end key. Determine the file identifier of each data file and the corresponding key range of each data file. Use the left-end key of the key range as the minimum key, and use the right-end key of the key range as the maximum key. Take the file identifier of each data file, the minimum key and the maximum key of the corresponding key range of each data file as an element, summarize the elements corresponding to each data file in the first data block to be split into an array, and use this array as the target array. Exemplarily, the target array can be expressed as [(1, key11, key12), (2, key21, key21)…(n, keyn1, keyn2)]. Wherein, (n, keyn1, keyn2) is an element, in this element, n is the data file identifier, keyn1 represents the minimum key of the key range corresponding to the data file with the identifier n, and keyn2 represents the minimum key of the key range corresponding to the data file with the identifier n.
[0043] Traverse the target array. If the minimum key of the data file corresponding to any file identifier is less than the start key and the maximum key is greater than the start key, it indicates that the key range corresponding to the data file spans the start key. This also makes all keys between the start key and the maximum key belong to the first split key range. Therefore, this data file is used as the first start file. If the minimum key of the data file corresponding to any file identifier is less than the end key and the maximum key is greater than the end key, it indicates that the key range corresponding to the data file spans the end key. This also makes all keys between the minimum key and the end key belong to the first split key range. Therefore, this data file is used as the first end file.
[0044] In one embodiment, a data file with both the minimum key and the maximum key distributed between the start key and the end key is used as the first intermediate file.
[0045] Exemplarily, the current file identifier is n, the corresponding minimum key is keyn1, and the maximum key corresponding to the current file identifier is keyn2. If start key ≤ keyn1 < keyn2 ≤ end key is satisfied, it is determined that the data file with identifier n is an intermediate file.
[0046] By summarizing the file identifiers of each data file in the first data block to be split, the minimum key and the maximum key of each data file corresponding to the key range into a target array, the query speed of the minimum key and the maximum key of each data file is improved; by comparing the size relationships between the maximum key and the minimum key corresponding to each data file in the first data block to be split with the start key and the end key of the first split key range respectively, the first start file and / or the first end file in the first data block to be split are determined, and when necessary, the first intermediate file is determined, which improves the determination speed and accuracy of the first intermediate file and the first end file.
[0047] S130. Modify the data block identifier in the meta-information of the first intermediate file to the first target data block identifier.
[0048] Since the key range corresponding to the first intermediate file is within the first split key range, before and after the data block split, except for the change in the data block identifier, the first intermediate file does not change in other aspects. For this reason, in this embodiment, the intermediate file is not read, and only the data block identifier in the meta-information of the first intermediate file is modified to the first target data block identifier to reduce unnecessary data reading and writing during the data block split, thereby reducing the read-write amplification during the data block split.
[0049] It can be understood that the larger the first split key interval is, the more first intermediate files it usually corresponds to. And the more first intermediate files there are, compared with the prior art, the greater the reduction in read-write amplification during the data block splitting process, because the prior art needs to read all the data that needs to be split out from the first data block to be split, that is, it needs to read all the first intermediate files and the first end file.
[0050] Among them, read-write amplification includes read amplification and write amplification. Read amplification can be understood as: when performing a read operation in a flash memory, additional operations may be involved to meet specific read requirements. For example, when a certain data needs to be read, if there is other irrelevant data in the storage unit where the data is located, in order to read the data, the entire storage unit must be read simultaneously, and subsequent operations also need to process the redundant irrelevant data. In this case, read amplification occurs. Among them, the storage unit is the smallest data storage unit corresponding to the data processing scenario.
[0051] Write amplification can be understood as: when performing a write operation in a flash memory, additional operations may need to be performed before actually writing a data. For example, before writing a new data in a flash memory, the old data to be replaced must be erased first. In this way, multiple erase and write operations may be involved before actually writing a new data, resulting in the problem of write amplification.
[0052] In one embodiment, the data block identifier in the first intermediate file metadata is set to the target data block identifier through the following steps:
[0053] Step a1: Delete the metadata of the first intermediate file from the manifest file.
[0054] The manifest file records the metadata of each data file. The metadata of each data file includes the latest status information of the corresponding data file, such as the data layer where the data file is located, the corresponding data block identifier, the latest modification time, etc. In this embodiment, the metadata of the first intermediate file is first deleted from the manifest file to update the manifest file. It can be understood that the updated manifest file does not include the metadata of the first intermediate file.
[0055] Step a2: Modify the data block identifier in the metadata of the first intermediate file to the first target data block identifier to update the metadata of the first intermediate file.
[0056] Since the first split data block identifier in the metadata of the first intermediate file is modified to the first target data block identifier without reading the first intermediate file, there is no read-write amplification in this process.
[0057] Step a3: Add the updated metadata of the first intermediate file to the manifest file to update the manifest file.
[0058] Since the meta information of the updated first intermediate file includes the target data block identifier, the meta information of the updated first intermediate file is added to the manifest file and persisted in the manifest file. In this way, the manifest file records the latest status of the first intermediate file, facilitating the status traceability of the first intermediate file and subsequent data usage.
[0059] S140. Migrate the partial keys within the first split key range in the first end file and the data corresponding to the partial keys to the corresponding file in the next data layer, and set the data block identifier in the meta information of the file to the first target data block identifier.
[0060] Among them, the keys within the first split key range in the first end file can be understood as the keys in the first end file that belong to the range between the starting key and the maximum key. Among them, the starting key is the minimum key of the first key range, and the maximum key is the maximum key in the first end file.
[0061] In one embodiment, the keys within the first split key range in the first end file are used as the key combination to be migrated; the data file in the next data layer corresponding to the current data layer that has key overlap with the key combination to be migrated is used as the target file; the merge sort algorithm is used to perform a merge operation on the keys to be migrated and all the keys in the target file to obtain a merge result; the merge result is stored in the target file in the next data layer. In this embodiment, the merge sort algorithm is used to perform a merge operation on the keys to be migrated and all the keys in the target file, so that the data to be migrated is stored in the target file in an orderly manner, which helps to improve the key query speed and thus improve the key value reading speed.
[0062] It should be noted that if the first data file combination only includes the first end file and does not include the first intermediate file, the data within the first split key range in the first end file is directly migrated to the corresponding next data layer.
[0063] In this embodiment, the data block identifiers of both the first intermediate file and the first end file are the target data block identifiers. The first intermediate file is located in the current data layer, and the first end file is located in the next data layer, completing the data splitting operation of splitting the data block corresponding to the target data identifier from the above-mentioned first data block to be split.
[0064] In one embodiment, the target data block identifier is determined based on the first data block to be split identifier. The target data block identifier includes a first field and a second field. The first field inherits the target field in the first data block to be split identifier, and the target field includes the second field in the first data block to be split identifier. Therefore, the data block identifier can reflect the relationship between data blocks or the sequence relationship between the generation times of different data blocks. For example, if the target data block is split from the first data block to be split, it can be regarded as a sub-data block of the first data block to be split, and the generation time of the target data block is later than that of the first data block to be split.
[0065] In the technical solution provided by the embodiment of the present invention, after determining the first data file combination and the first split key interval corresponding to the first data block to be split, for each data file in the first data file combination, if all its keys are distributed within the first split key interval, it is used as an intermediate file and retained in the current data layer, and only the data block identifier in its metadata is modified to the target data block identifier; since the intermediate file does not need to be read during the modification of the metadata of the intermediate file, there is no read / write amplification; if only some keys in the data file are distributed within the first split key interval, the part of the keys distributed within the first partial split key interval and the data corresponding to the part of the keys are migrated to the corresponding file in the next data layer, and the data block identifier in the metadata of the file is modified to the first target data block identifier; since each data file in the key-value storage system is relatively small, the first end file is also small, and the data volume corresponding to the data within the first split key interval in the first end file is even smaller. Compared with the existing data block splitting method that needs to read all split data, in this embodiment, only the data within the first split key interval in the first end file needs to be compacted, so the read / write amplification during the data block splitting process can be greatly reduced, and at the same time, the splitting speed of the data block can be improved.
[0066] Figure 2 It is another flowchart of the data block splitting method provided by the embodiment of the present invention. In this embodiment, on the basis of the foregoing embodiment, the data block splitting operation of the next data layer corresponding to the current data layer is automatically completed, as Figure 2 shown. The method includes:
[0067] S210. Determine the first data block to be split in the current data layer, and the first data file combination and the first split key interval corresponding to the first data block to be split.
[0068] S220. For the first data file combination, use the data file with all keys distributed within the first split key interval as the first intermediate file, and use the data file with some keys distributed within the first split key interval as the first end file.
[0069] S230. Modify the data block identifier in the meta information of the first intermediate file to the first target data block identifier.
[0070] S240. Migrate the partial keys and the data corresponding to the partial keys within the first split key range in the first end file to the corresponding file in the next data layer, and set the data block identifier in the meta information of this file to the first target data block identifier.
[0071] S250. If the data migrated to the next data layer causes the next data layer to meet the predetermined data block splitting condition, determine the second data block to be split in the next data layer, as well as the second data file combination and the second split key range corresponding to the second data block to be split. The second split key range is greater than or equal to the partial key range, and the partial key range is the key range of the partial keys distributed within the first split key range in the first end file.
[0072] Among them, the set data block splitting condition is the existing data block splitting trigger condition, such as the arrival of the scheduled splitting time, the data layer score reaching the set threshold, etc. Among them, the score of each data block is based on the number of files in each data layer or the amount of stored data.
[0073] During the data block splitting process of the current data layer, some data is merged into the next data layer, which will inevitably cause changes in the number of files or the amount of data stored in the next data layer. This change may cause the data score of the next data layer to reach the set threshold. In one embodiment, when it is detected that the data score of the next data layer reaches the set threshold, it is determined that the next data layer meets the set data splitting condition, and thus a second splitting instruction for the second data block in the next data layer is generated. The second splitting key range corresponding to the second splitting instruction is greater than or equal to the partial key range, and the partial key range is the key range of the partial keys distributed within the first split key range in the first end file. Among them, the split key range corresponding to any splitting instruction needs to make the two data blocks obtained after splitting the corresponding data block meet the set data block conditions, such as the set data volume, the set number of data files, etc. For this reason, when necessary, the split key range corresponding to the received data needs to be enlarged so that each data block after splitting meets the set data block conditions.
[0074] S260. For the second data file combination, use the data file with all keys distributed within the second split key range as the second intermediate file, and use the data file with partial keys distributed within the second split key range as the second end file.
[0075] S270. Modify the second data block identifier in the meta information of the second intermediate file to the second target data block identifier.
[0076] S280. Migrate the partial keys within the second split key range in the second end file and the corresponding data of the partial keys to the corresponding file in the next data layer, and set the data block identifiers in the meta-information of this file to the second target data block identifiers.
[0077] In this embodiment, the specific implementation manners of steps S260 - S280 are the same as those of S120 - S140 in the foregoing embodiment.
[0078] It can be understood that if the data split of the first data block to be split does not cause an obvious change in the data volume of the next data layer, for example, it does not make the data layer score of the next data layer reach the set threshold, etc., then there is no need to perform a data block split operation on the corresponding data block in this next data layer.
[0079] In the embodiment of the present invention, after a data block split operation is performed on a data block in any data layer, if the data block split result causes the next data layer to meet the set data block split condition, the split operation of the corresponding data block in the next data layer is automatically triggered, and the data split operation of this next data layer is completed based on the same data split method, ensuring the accuracy and thoroughness of the data block split.
[0080] The following is an embodiment of the data block splitting device provided by the embodiment of the present invention. This device and the data block splitting method in the above embodiment belong to the same inventive concept. For the details not described in detail in the embodiment of the data block splitting device, reference can be made to the content of the above embodiments.
[0081] Figure 3 It is a schematic structural diagram of the device provided by the embodiment of the present invention. This device includes:
[0082] A first response module 310, configured to determine a first data block to be split in the current data layer, and a first data file combination and a first split key range corresponding to the first data block to be split;
[0083] A first file determination module 320, configured to, for the first data file combination, use a data file in which all keys are distributed within the first split key range as a first intermediate file, and use a data file in which partial keys are distributed within the first split key range as a first end file;
[0084] A first identifier modification module 330, configured to modify the data block identifier in the meta-information of the first intermediate file to a first target data block identifier;
[0085] The first compaction module 340 is configured to migrate the partial keys within the first split key range in the first end file and the data corresponding to the partial keys to the corresponding file in the next data layer, and set the data block identifier in the meta-information of the file to the first target data block identifier.
[0086] In one embodiment, as Figure 4 shown, the apparatus further includes:
[0087] The second response module 350 is configured to, if the data migrated to the next data layer causes the next data layer to meet the predetermined data block splitting condition, determine the second data block to be split in the next data layer, and the second data file combination and the second split key range corresponding to the second data block to be split, where the second split key range is greater than or equal to the partial key range, and the partial key range is the key range of the partial keys distributed within the first split key range in the first end file;
[0088] The second file determination module 360 is configured to, for the second data file combination, use the data file in which all keys are distributed within the second split key range as the second intermediate file, and use the data file in which some keys are distributed within the second split key range as the second end file;
[0089] The second identifier modification module 370 is configured to modify the second data block identifier in the meta-information of the second intermediate file to the second target data block identifier;
[0090] The second compaction module 380 is configured to migrate the partial keys within the second split key range in the second end file and the data corresponding to the partial keys to the corresponding file in the next data layer, and set the data block identifiers in the meta-information of the file to the second target data block identifier.
[0091] In one embodiment, the first compaction module 340 is specifically configured to:
[0092] Use the keys within the first split key range in the first end file as the key combination to be migrated;
[0093] Use the data file in the next data layer corresponding to the current data layer, which has key overlap with the key combination to be migrated, as the target file;
[0094] Use the merge sort algorithm to perform a merge operation on the keys to be migrated and all the keys in the target file to obtain a merge result;
[0095] Store the merge result in the target file in the next data layer.
[0096] In one embodiment, the first file determination module 320 is specifically configured to:
[0097] Determine the start key and end key of the first split key range, as well as the minimum key and maximum key of each data file;
[0098] Use the data file whose minimum key is less than the start key and whose maximum key is greater than the end key as the first start file;
[0099] Use the data file whose minimum key is less than the end key and whose maximum key is greater than the end key as the first end file.
[0100] In one embodiment, the first file determination module 320 is specifically configured to:
[0101] Use the data file whose minimum key and maximum key are both distributed between the start key and the end key as the first intermediate file.
[0102] In one set of embodiments, the identifier modification module 330 is specifically configured to:
[0103] Delete the meta information of the first intermediate file from the manifest file;
[0104] Modify the data block identifier in the meta information of the first intermediate file to the first target data block identifier to update the meta information of the first intermediate file;
[0105] Add the updated meta information of the first intermediate file to the manifest file to update the manifest file.
[0106] In the technical solution provided by the embodiment of the present invention, after determining the first data file combination corresponding to the first data block to be split and the first split key interval, for each data file in the first data file combination, if all its keys are distributed within the first split key interval, it is regarded as an intermediate file and retained in the current data layer, and only the data block identifier in its metadata is modified to the target data block identifier; since the intermediate file does not need to be read during the modification of the metadata of the intermediate file, there will be no read / write amplification; if only some keys in the data file are distributed within the first split key interval, the part of the keys distributed within the partial interval of the first split key and the data corresponding to the part of the keys are migrated to the corresponding file in the next data layer, and the data block identifier in the metadata of the file is modified to the first target data block identifier. Since each data file in the key-value storage system is relatively small, the first end file is also relatively small, and the data volume corresponding to the data within the split key interval in the first end file is even smaller. Compared with the existing data block splitting method that needs to read all the split data, in this embodiment, only the data within the split key interval in the first end file needs to be compacted, so the read / write amplification during the data block splitting process can be greatly reduced, and at the same time, the splitting speed of the data block can be improved.
[0107] The data block splitting device provided by the embodiment of the present invention can execute the data block splitting method provided by the embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the data block splitting method.
[0108] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Figure 5 It shows a block diagram of an exemplary server 12 suitable for implementing the embodiments of the present invention. Figure 5 The shown server 12 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0109] As Figure 5 shown, the server 12 is presented in the form of a general-purpose computing device. The components of the server 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0110] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0111] Server 12 typically includes a variety of computer system readable media. These media can be any available media accessible to server 12, including volatile and non-volatile media, removable and non-removable media.
[0112] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Server 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 5 not shown, commonly referred to as a "hard disk drive"). Although Figure 5 not shown in the figure, a disk drive for reading and writing on removable non-volatile disks (such as a "floppy disk") and an optical disk drive for reading and writing on removable non-volatile optical disks (such as a CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to bus 18 through one or more data media interfaces. System memory 28 can include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0113] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods in the embodiments described in the present invention.
[0114] Server 12 can also communicate with one or more external devices 14 (such as keyboards, pointing devices, monitors 24, etc.), and can also communicate with one or more devices that enable users to interact with the server 12, and / or communicate with any device that enables the server 12 to communicate with one or more other computing devices (such as network cards, modems, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Moreover, the server 12 can also communicate with one or more networks (such as local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the server 12 through the bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the server 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0115] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28. For example, it implements the steps of a data block splitting method provided by an embodiment of the present invention. The method includes:
[0116] Determine the first data block to be split in the current data layer, and the first data file combination and the first split key interval corresponding to the first data block to be split;
[0117] For the first data file combination, use the data files with all keys distributed within the first split key interval as the first intermediate files, and use the data files with some keys distributed within the first split key interval as the first end files;
[0118] Modify the data block identifier in the meta-information of the first intermediate file to the first target data block identifier;
[0119] Migrate the part of the keys within the first split key interval and the corresponding data in the first end file to the corresponding file in the next data layer, and set the data block identifier in the meta-information of this file to the first target data block identifier.
[0120] Of course, those skilled in the art can understand that the processor can also implement the technical solutions of the data block splitting method provided by any embodiment of the present invention.
[0121] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the data block splitting method provided by the foregoing embodiments of the present invention. The method includes:
[0122] Determine the first data block to be split in the current data layer, as well as the first data file combination and the first split key range corresponding to the first data block to be split;
[0123] For the first data file combination, use the data files with all keys distributed within the first split key range as the first intermediate files, and use the data files with some keys distributed within the first split key range as the first end files;
[0124] Modify the data block identifier in the meta-information of the first intermediate file to the first target data block identifier;
[0125] Migrate the partial keys within the first split key range and the data corresponding to the partial keys in the first end file to the corresponding files in the next data layer, and set the data block identifier in the meta-information of this file to the first target data block identifier.
[0126] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0127] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0128] The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0129] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0130] Embodiments of the present invention also provide a computer program product, including a computer program which, when executed by a processor, implements the... method provided in any embodiment of the present application.
[0131] In the process of implementing the computer program product, computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0132] Those of ordinary skill in the art should understand that the above-mentioned modules or steps of the present invention may be implemented using a general-purpose computing device. They may be concentrated on a single computing device or distributed over a network composed of multiple computing devices. Optionally, they may be implemented using program code executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they may be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them may be fabricated into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0133] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A data block splitting method, characterized in that: The method comprises: Determine a first data block to be split in the current data layer, and a first data file combination and a first split key interval corresponding to the first data block to be split; For the first data file combination, a data file in which all keys are distributed in the first split key interval is used as a first intermediate file, and a data file in which some keys are distributed in the first split key interval is used as a first end file, and the keys in the data files are stored in sequence; Modify the data block identifier in the meta information of the first intermediate file to the first target data block identifier; Migrate some keys in the first end file that belong to the first split key interval and the data corresponding to the some keys to the corresponding file of the next data layer, and set the data block identifier in the metadata of the file to the first target data block identifier.
2. The method according to claim 1, characterized in that: After migrating the partial keys and the data corresponding to the partial keys in the first end file that belong to the first split key interval to the corresponding files of the next data layer, the method further includes: If the data migrated to the next data layer makes the next data layer meet the predetermined data block splitting condition, then determine a second data block to be split in the next data layer, and a second data file combination and a second split key interval corresponding to the second data block to be split, wherein the second split key interval is greater than or equal to a partial key interval, and the partial key interval is a key interval of a partial key in the first end file that is distributed in the first split key interval; For the second data file combination, a data file in which all keys are distributed in the second split key interval is used as a second intermediate file, and a data file in which some keys are distributed in the second split key interval is used as a second end file; Modify the second data block identifier in the meta information of the second intermediate file to the second target data block identifier; Migrate some keys in the second split key interval and the data corresponding to the some keys in the second end file to the corresponding file of the next data layer, and set the data block identifiers in the metadata of the file to the second target data block identifiers.
3. The method according to claim 1, characterized in that The step of migrating the data in the first split key interval in the first end file to a corresponding file in the next data layer includes: Using the keys in the first split key interval in the first end file as key combinations to be migrated; Taking a data file in a next data layer corresponding to the current data layer, which has a key overlap with the key combination to be migrated, as a target file; Using a merge sort algorithm to merge the key to be migrated with all the keys in the target file to obtain a merge result; The merge result is stored in the target file of the next data layer.
4. The method according to claim 1, characterized in that: The first end file includes a first start file and / or a first end file, and the data file in which some keys are distributed in the first split key interval as the first end file includes: Determine the start key and the end key of the first split key interval, and the minimum key and the maximum key of each data file; The data file whose minimum key is smaller than the starting key and whose maximum key is larger than the starting key is used as the first starting file; The data file whose minimum key is smaller than the end key and whose maximum key is larger than the end key is used as the first end file.
5. The method according to claim 4, characterized in that The data file in which all keys are distributed in the first split key interval is used as a first intermediate file, including: A data file in which both the minimum key and the maximum key are distributed between the start key and the end key is used as the first intermediate file.
6. The method according to claim 1, characterized in that The step of modifying the data block identifier in the meta information of the first intermediate file to the first target data block identifier includes: Deleting the meta information of the first intermediate file from the manifest file; Modify the data block identifier in the meta information of the first intermediate file to the first target data block identifier, so as to update the meta information of the first intermediate file; The updated meta information of the first intermediate file is added to the manifest file to update the manifest file.
7. A data block splitting device, characterized in that: include: A first response module, configured to determine a first data block to be split in a current data layer, and a first data file combination and a first split key interval corresponding to the first data block to be split; A first file determination module is used for, for the first data file combination, taking a data file in which all keys are distributed in the first split key interval as a first intermediate file, taking a data file in which some keys are distributed in the first split key interval as a first end file, wherein the keys in the data file are stored in sequence; A first identifier modification module, used to modify the data block identifier in the meta information of the first intermediate file to a first target data block identifier; The first compaction module is used to migrate some keys in the first split key interval and the data corresponding to the some keys in the first end file to the corresponding file of the next data layer, and set the data block identifier in the metadata of the file to the first target data block identifier.
8. An electronic device, characterized in that: The electronic device comprises: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data block splitting method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the data block splitting method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that When executed by a processor, the computer program implements the data block splitting method as described in any one of claims 1-6.
Citation Information
Patent Citations
MongoDB data migration monitoring method and device based on log analysis
CN110147353A
Split-key estimation method for table partition in disbtributed data storage systems
WO2020199192A1