Data compression method and device, storage medium and electronic equipment
By monitoring and rationally compressing data in MemTable, SSTable micro-blocks, SSTable small blocks, and SSTable macro-blocks, the problems of read/write amplification and storage space waste in the database system are solved, thereby improving the performance and storage efficiency of the database system.
Patent Information
- Application Number
- CN202510873715.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-24
AI Technical Summary
In existing technologies, database systems suffer from read/write amplification when writing data from memory to disk, leading to wasted storage space and performance degradation. This is especially true in LSM-Tree structures, where the storage efficiency of SSTable microblocks is relatively low.
By monitoring the MemTable data in memory, when the first compression condition is met, the data is written as merged data into SSTable microblocks and stored in the disk; by monitoring the SSTable microblock data, when the second compression condition is met, the data is incrementally compressed in SSTable small blocks; by monitoring the SSTable small block data, when the third compression condition is met, the data is updated to the target value in SSTable macroblocks.
It effectively reduces read/write amplification, saves storage space, and improves the performance of the database system. In particular, it optimizes storage efficiency by reasonably layering and compressing SSTable micro-blocks, SSTable small blocks, and SSTable macro-blocks.
Smart Images

Figure CN120832093A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the technical field of computer, and particularly relates to a data compression method and device, a storage medium and an electronic device. BACKGROUND
[0002] At present, databases have been widely applied in various industries. In order to improve the performance of writing data of a database system and optimize the data structure, many database systems use LSM-Tree algorithm mechanism to write data.
[0003] Generally, the database system will first store the data written into the database in the memory, and then merge and write the data in the memory into the disk. The data in the disk is compressed in the form of SSTable, and the database system will continue to compress and press down the data in the form of SSTable in a hierarchical manner from the upper layer to the lower layer. In this way, from the memory to the disk, the data constitutes the tree structure organization of the LSM-Tree.
[0004] In the above process, how to effectively compress each SSTable in the disk to improve the performance of the database system is a problem to be solved. SUMMARY
[0005] Embodiments of the present specification provide a data compression method, device, storage medium and electronic device to partially solve the problems existing in the prior art.
[0006] Embodiments of the present specification adopt the following technical solutions: The present specification provides a data compression method, which comprises: Monitoring data written into a MemTable in the memory; When the data written into the MemTable meets a first compression condition, reading each data written into the MemTable; Writing each read data as a merged data into a SSTable microblock in the disk; the size of the SSTable microblock is not less than the MemTable; Storing the SSTable microblock.
[0007] The present specification provides a data compression method, which comprises: Monitoring data written into each SSTable microblock in the disk; When the data written into each SSTable microblock meets a second compression condition, for each data in each SSTable microblock, determining the key of the data and the value of the data in the SSTable microblock; According to the key of the data, the data corresponding to the key is queried in the SSTable microblock as target data; The value of the data in the SSTable microblock is added to the target data stored in the SSTable microblock.
[0008] The data compression method provided in the specification is for any data written in the SSTable microblock of the disk of the database system, the key of the data corresponding to at least one version of value in the SSTable microblock; the method comprises: Monitoring the data written in each SSTable microblock in the disk; When the data written in each SSTable microblock meets the third compression condition, for each data in each SSTable microblock, the key of the data is determined; Among each version of value corresponding to the key of the data, the target value is determined, and the data corresponding to the key of the data is queried in the SSTable macroblock as target data; wherein each version of value corresponding to the key of the data includes each version of value corresponding to the key of the data in the SSTable microblock and the value corresponding to the key of the data in the SSTable macroblock; The value in the target data stored in the SSTable macroblock is updated to the target value.
[0009] The data compression device provided in the specification comprises: The monitoring module is used for monitoring the data written in the MemTable in the memory; The reading module is used for reading each data written in the MemTable when the data written in the MemTable meets the first compression condition; The compression module is used for writing each read data as a merged data in the SSTable microblock in the disk; the size of the SSTable microblock is not less than the MemTable; The storage module is used for storing the SSTable microblock.
[0010] The data compression device provided in the specification comprises: The monitoring module is used for monitoring the data written in each SSTable microblock in the disk; The determination module is used for determining the key of each data and the value of the data in the SSTable microblock when the data written in each SSTable microblock meets the second compression condition; According to the key of the data, the data corresponding to the key is queried in the SSTable microblock as target data; a compression module, configured to add the value of the data in the SSTable microblock to the target data stored in the SSTable microblock.
[0011] The data compression apparatus provided in the specification is configured to write any data in a SSTable microblock in a disk of a database system, the key of the data corresponding to values of at least one version in the SSTable microblock; the apparatus comprises: a monitoring module, configured to monitor data written in each SSTable microblock in the disk; a first determination module, configured to determine the key of each data in each SSTable microblock when the data written in each SSTable microblock meets a third compression condition; a second determination module, configured to determine a target value from the values of each version corresponding to the key of the data; wherein the values of each version corresponding to the key of the data include values of each version corresponding to the key of the data in the SSTable microblock and values corresponding to the key of the data in the SSTable macroblock; a query module, configured to query data corresponding to the key of the data in the SSTable macroblock as target data; a compression module, configured to update the value in the target data stored in the SSTable macroblock to the target value.
[0012] The computer readable storage medium provided in the specification stores a computer program, and the computer program is executed by a processor to implement the data compression method.
[0013] The electronic device provided in the specification comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the data compression method when executing the program.
[0014] The above at least one technical solution adopted by the embodiments of the specification can achieve the following beneficial effects: The data compression method disclosed in the embodiments of the specification writes data in a MemTable as a merged data in a SSTable microblock when compressing the MemTable into the SSTable microblock in a disk, and then stores the SSTable microblock, which can effectively reduce read-write amplification caused by compressing the MemTable, reduce storage space consumed for storing the SSTable microblock in a database system, and thus improve performance of the database system. BRIEF DESCRIPTION OF DRAWINGS
[0015] The accompanying drawings, which are included to provide a further understanding of the present description and constitute a part of the present description, illustrate the illustrative embodiments of the present description and serve to explain the present description together with the description. In the drawings: Figure 1 A first data compression method flow chart provided for the embodiments of the present description; Figure 2 A second data compression method flow chart provided for the embodiments of the present description; Figure 3 A third data compression method flow chart provided for the embodiments of the present description; Figure 4 A structure diagram of an SSTable file provided for the embodiments of the present description; Figure 5 A first data compression device schematic diagram provided for the embodiments of the present description; Figure 6 A second data compression device schematic diagram provided for the embodiments of the present description; Figure 7 A third data compression device schematic diagram provided for the embodiments of the present description; Figure 8 A structure diagram of an electronic device provided for the embodiments of the present description. DETAILED DESCRIPTION
[0016] In order to make the purposes, technical solutions and advantages of the present description clearer, the technical solutions of the present description will be described below in conjunction with the specific embodiments of the present description and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present description, not all the embodiments. Based on the embodiments in the present description, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present description.
[0017] LSM-Tree is a high-performance data structure optimization mechanism, which mainly writes data to the disk of the database system through MemTable and SSTable files. MemTable is a file in the memory of the database system, and SSTable is a file in the disk of the database system. LSM-Tree writes the newly added data in the database system to MemTable when working, sorts the data in MemTable by key when the accumulated data in MemTable reaches a certain amount, and compresses MemTable into SSTable and writes it to the disk, and then empties MemTable. The LSM-Tree mechanism further layers the data in the disk, such as L0, L1, and L2 layers. After the data in MemTable is compressed into SSTable, it is first written to the L0 layer, and when the accumulated SSTable in the L0 layer reaches a certain amount, the SSTable in the L0 layer is compressed into a larger SSTable and written to the L1 layer. In this way, the data in the disk can finally form a tree structure.
[0018] It can be seen that the higher the layer of the data in the disk, the smaller the data block of SSTable, and the lower the layer, the larger the data block of SSTable. In this specification, the smallest SSTable data block (located in the L0 layer) is referred to as a SSTable microblock, the second smallest SSTable data block (located in the L1 layer) is referred to as a SSTable small block, and the largest SSTable data block (located in the L2 layer) is referred to as a SSTable macroblock. The specific size of the SSTable microblock, the SSTable small block, and the SSTable macroblock can be set as needed, as long as the SSTable microblock is smaller than the SSTable small block, and the SSTable small block is smaller than the SSTable macroblock.
[0019] In order to effectively compress data into the above three kinds of data blocks in the LSM-Tree mechanism and save the disk space of the database system as much as possible, different methods can be used for compression according to the size relationship of the above three kinds of data blocks in the embodiments of the present specification.
[0020] The technical solutions provided by the embodiments of the present specification will be described in detail below with reference to the accompanying drawings.
[0021] I. Compressing MemTable in memory into SSTable microblock in disk, as shown in Figure 1
[0022] Figure 1 The first data compression method flowchart provided by the embodiments of the present specification specifically includes the following steps: S100: Monitor the data written to MemTable in memory.
[0023] In the embodiments of the present specification, the method shown inFigure 1 The subject of the method shown for compressing data can be a database system, and specifically a device carrying the database system, such as a server or a server cluster.
[0024] The database system can monitor the data written in the MemTable in its own memory, where the data written in the MemTable can specifically include any changed data in the database system, such as newly added data in the database, changed data in the database, and deleted data in the database.
[0025] S102: When the data written in the MemTable meets a first compression condition, read each data written in the MemTable.
[0026] The first compression condition described in the embodiments of the present specification can specifically be that the data written in the MemTable reaches a first preset threshold. When the data written in the MemTable reaches the first preset threshold, the database system can read the data written in the MemTable in the memory.
[0027] S104: Write each read data as a merged data in an SSTable microblock in the disk.
[0028] Wherein, the size of the SSTable microblock is not less than the size of the MemTable.
[0029] In the database system, the smallest unit of storing data is a system data block, and the size of each system data block is fixed. The size of the system data block is irrelevant to the database system, and only related to the operating system. For example, when the operating system of the database system is a Linux system, the system data block is a Linux system data block.
[0030] The database system needs to store the data read from the MemTable in system data blocks first, and then compress the system data blocks into SSTable microblocks and write them to the disk. Since the size of the system data block is fixed, storing the data read from the MemTable in the system data block will cause read-write amplification. For example, assuming that there are two data in the MemTable, one data is 9 KB in size, and the other data is 10 KB in size, and the size of one system data block is fixed at 4 KB, storing the first data requires 3 system data blocks, which wastes 3 KB of storage space, and storing the other data block requires 3 system data blocks, which wastes 2 KB of storage space. The data amount written to the SSTable microblock is 6 system data blocks, which is 24 KB of data amount, but the actual data amount is only 19 KB. Therefore, 5 KB of data amount is wasted in the SSTable microblock, which is larger than the size of one system data block.
[0031] In the LST-Tree mechanism, when reading data, the database system usually first tries to read in the MemTable in the memory. If it fails to read, it continues to read in the SSTable in each level from top to bottom in order, that is, it first tries to read from the SSTable microblock in the L0 layer, and if it still fails to read, it tries to read from the SSTable microblock in the L1 layer. Therefore, the SSTable microblock in the L0 layer is read most frequently among all the data stored in the disk, and the storage space of the SSTable microblock wasted due to read-write amplification is unacceptable.
[0032] Therefore, in the embodiments of the present specification, the database system can write each data read from the MemTable to the SSTable microblock as a merged data, that is, write each data read from the MemTable as a continuous data to the SSTable microblock.
[0033] Specifically, the database system can store the merged data read from the MemTable as a continuous merged data using a plurality of system data blocks; the total size of the system data blocks storing the merged data is not less than the size of the merged data, and then compress the system data blocks storing the merged data into an SSTable microblock.
[0034] Continuing with the above example, after the data of 9 KB and 10 KB in size are taken as a continuous merged data, the continuous merged data becomes a larger data of 19 KB in size and containing the above two data. Since the size of one system data block is 4 KB, only 5 system data blocks are needed to store the merged data, and the 5 system data blocks are compressed into a 20 KB SSTable microblock, which only wastes 1 KB of storage space.
[0035] Therefore, writing the data in the MemTable as a continuous merged data into the SSTable microblock can control the storage space waste caused by read-write amplification within a range of not more than one system data block.
[0036] S106: store the SSTable microblock.
[0037] The layers in the disk described in the specification (such as L0, L1, L2) are not only logical layers, and the SSTable files in different layers are completely separated on the disk. The smaller the data block of the SSTable, the more suitable it is for random reading, and the larger the data block of the SSTable, the more suitable it is for sequential reading. After the database compression obtains the SSTable microblock, the SSTable microblock can be stored in the storage space corresponding to the L0 layer of the disk.
[0038] It should be noted that the first data compression method described above can be applied to any one or several servers in a server or server cluster carrying a database system, and can be specifically implemented by executing a computer program in the storage medium (including memory, disk, or any storage medium) of the server or in the processing chip (such as FPGA, ASIC, etc.) connected to the server. The embodiments of the present specification do not limit the hardware structure of the database system. The memory of the database system described above can include any form of power loss storage medium, such as DRAM, SRAM, etc. The disk of the database system described above can include any form of power lossless storage medium, such as ROM, SSD, etc.
[0039] II. Compressing the SSTable microblock in the disk into the SSTable small block in the disk, as shown in Figure 2
[0040] Figure 2 The second data compression method provided by the embodiments of the present specification includes the following steps: S200: Monitor the data written in each SSTable microblock in the disk.
[0041] In the embodiments of the present specification, as the data written in the MemTable in the memory is continuously compressed into the SSTable microblock in the L0 layer of the disk, the SSTable microblock in the L0 layer will also gradually accumulate and increase, and the database system can monitor the data written in each SSTable microblock in its own disk.
[0042] S202: When the data written in each SSTable microblock meets the second compression condition, for each data in each SSTable microblock, determine the key of the data and the value of the data in the SSTable microblock.
[0043] The second compression condition described in the embodiments of the present specification can be that the number of each SSTable microblock in the L0 layer reaches a second preset threshold. When the number of each SSTable microblock reaches the second preset threshold, the database system can first determine the key (i.e., key) of each data in each SSTable microblock and the value (i.e., value) of the data in the SSTable microblock.
[0044] Since the data in the SSTable microblock comes from the data written by the MemTable, the data written by the MemTable includes not only the newly added data in the database, but also the data that has been stored in the database but has changed and the data that has been deleted. Therefore, for any data in the SSTable microblock, the version of the data before being changed or deleted can have been compressed into the SSTable small block in the L1 layer or even the SSTable macroblock in the L2 layer. In order to improve the utilization rate of the SSTable small block in the L1 layer, the database system needs to compress according to the value of the data in the SSTable microblock and the value of other versions in the SSTable small block in the L1 layer.
[0045] S204: According to the key of the data, query the data corresponding to the key in the SSTable small block as target data.
[0046] For each data in each SSTable microblock, after the database system determines the key of the data in step S202, it can query the data corresponding to the key in each SSTable small block in the disk L1 layer as target data according to the key of the data.
[0047] For example, in step S202, one data in the SSTable microblock is “Zhang San, age 22”, then the database system can determine that the key of the data is “Zhang San” and the value in the SSTable microblock is “age 22”. In step S204, according to the key “Zhang San” of the data, the data corresponding to “Zhang San” is queried in each SSTable small block in the disk L1 layer as target data, and the target data queried is “Zhang San, age 21”.
[0048] S206: Add the value of the data in the SSTable microblock to the target data stored in the SSTable small block.
[0049] Since the read frequency of the SSTable microblock is lower than that of the SSTable small block but higher than that of the SSTable macroblock, for a piece of data, no matter whether the data is changed or deleted, the values of each version corresponding to the key of the data or the deletion mark set when the data is deleted need to be stored in the L1 layer, so that the data can be traced back through the SSTable small block of the L1 layer in the future.
[0050] However, if the data in the SSTable microblock is directly stored in a new SSTable small block in the L1 layer when the data is changed or deleted, the storage space of the L1 layer will be greatly wasted. Therefore, in order to save the storage space of the L1 layer, the method of incremental compression is adopted in the embodiment of the present specification, that is, the value of the data in the SSTable microblock is added to the target data stored in the SSTable small block as another version of the value of the data.
[0051] Specifically, the database system can keep the key of the target data stored in the SSTable small block unchanged, add the value of the target data stored in the SSTable small block as the first version value corresponding to the key, and add the value of the data in the SSTable microblock as the second version value corresponding to the key to the target data.
[0052] Continuing with the above example, if the data "Zhang San, age 22" in the SSTable microblock is directly written into a new SSTable small block in the L1 layer, the two pieces of data with the key "Zhang San" are contained in each SSTable small block in the L1 layer, one is the data with the value "age 21" that has been compressed into the L1 layer, and the other is the data with the value "age 22" that is currently compressed.
[0053] However, in the embodiment of the present specification, the database determines the target data "Zhang San, age 21" in the SSTable small block, keeps the key of the data unchanged, that is, "Zhang San", and keeps the "age 21" in the SSTable small block as the first version value of the key "Zhang San", and adds the value "age 22" of the key "Zhang San" in the SSTable microblock as the second version value corresponding to the key "Zhang San" to the target data in the SSTable small block, so that the target data in the SSTable small block becomes "Zhang San, age 21, age 22", and there are no two pieces of data with the key "Zhang San", which can effectively save the storage space of the L1 layer disk.
[0054] It should be noted that the second data compression method described above can be applied to any one or several servers in a server or server cluster carrying a database system, and can be specifically implemented by executing a computer program in a storage medium (including any storage medium such as memory, disk, etc.) of the server or a processing chip (such as FPGA, ASIC, etc.) connected to the server. The embodiments of the present specification do not limit the hardware structure of the database system. The memory of the database system described above can include any form of power loss storage medium, such as DRAM, SRAM, etc. The disk of the database system described above can include any form of power lossless storage medium, such as ROM, SSD, etc.
[0055] III. Compressing the SSTable small blocks in the disk into SSTable macro blocks in the disk, as shown in Figure 3
[0056] Figure 3 The third data compression method provided by the embodiments of the present specification includes the following steps: S300: Monitor the data written in each SSTable small block in the disk.
[0057] Similar to the above two cases, the SSTable small blocks in the L1 layer of the disk of the database system will also accumulate over time, and therefore, the database system also needs to monitor the data written in each SSTable small block in the L1 layer of the disk.
[0058] S302: When the data written in each SSTable small block meets the third compression condition, determine the key of each data in each SSTable small block.
[0059] The third compression condition described in the embodiments of the present specification can be that the number of SSTable small blocks in the disk reaches a third preset threshold. When the number of SSTable small blocks in the L1 layer of the disk reaches the third preset threshold, each SSTable small block needs to be compressed into an SSTable macro block and enter the L2 layer of the disk. At this time, the database system first determines the key (i.e., key) of each data in each SSTable small block in the L1 layer of the disk.
[0060] S304: Determine the target value among the values of each version corresponding to the key of the data; and query the data corresponding to the key of the data in the SSTable macro block as the target data.
[0061] The values of each version corresponding to the key of the data include the values of each version corresponding to the key of the data in the SSTable small block and the values corresponding to the key of the data in the SSTable macro block.
[0062] Similar to the second case, on one hand, for any data in the SSTable chunk, there can be more than one version of the value of the data in the SSTable chunk, and there can also be a value corresponding to the key of the data in the SSTable macro chunk, and the key of the data in the SSTable macro chunk is an earlier version of the data that has been compressed into the L2 layer of the disk. Therefore, the above-mentioned each version of the value corresponding to the key of the data in the specification includes not only each version of the value corresponding to the key of the data in the SSTable chunk, but also the value corresponding to the key of the data in the SSTable macro chunk. The database system needs to determine the target value from all versions of the value corresponding to the key of the data.
[0063] On the other hand, the database system also needs to query the data corresponding to the key of the data in the SSTable macro chunk of the L2 layer of the disk as the target data.
[0064] It should be noted that the execution order of the above-mentioned database system to determine the target value and query the target data is not divided, and can be executed synchronously.
[0065] S306: updating the value in the target data stored in the SSTable macro chunk to the target value.
[0066] Since the SSTable macro chunk of the L2 layer of the disk has the lowest reading frequency, the SSTable macro chunk of the L2 layer does not need to save all versions of the value of the data to save the storage space of the L2 layer of the disk as much as possible, so that the database system can directly update the value in the target data stored in the SSTable macro chunk determined in step S304 to the target value determined in step S304.
[0067] In step S304, the method for the database system to determine the target value from all versions of the value of the data can be: determining the time length when each version of the value corresponding to the key of the data is written into the database system, and determining the target value from each version of the value corresponding to the key of the data according to the time length. The time length when a version of the value is written into the database system can be calculated from when the version of the value is written into the MemTable, that is, the time length from when the version of the value is written into the MemTable to the current time is the time length when the version of the value is written into the database system. If the time length exceeds the preset time length, it means that the version of the value has expired and does not need to be stored, and if the time length does not exceed the preset time length, it means that the version of the value has not expired and needs to be stored, and the version of the value is the target value. It should be noted that the target value can be more than two versions of the value of the data.
[0068] In addition, for the case that the data stored in the database system is deleted, from the MemTable, to the SSTable microblock of the disk L1 layer, the deleted data is not directly deleted from the MemTable, the SSTable microblock, and the SSTable microblock, but the data is written in the MemTable, the SSTable microblock, and the SSTable microblock, and a deletion mark is set for the written data. When the data written in the SSTable microblock meets the third compression condition, and the database system compresses the SSTable microblock into the SSTable macroblock, the key of the data with the deletion mark in each SSTable microblock can be determined as the deletion key, and the data corresponding to the deletion key is deleted from the data stored in the SSTable macroblock. That is, the deleted data is not actually deleted in the MemTable, the SSTable microblock, and the SSTable microblock, but only a deletion mark is set, and the data is actually deleted in the SSTable macroblock.
[0069] It should be noted that the third data compression method described above can be applied to any one or several servers in the server or server cluster carrying the database system, and can be specifically implemented by executing a computer program in the storage medium (including any storage medium such as memory, disk, etc.) of the server or in the processing chip (such as FPGA, ASIC, etc.) connected to the server. The embodiments of the present specification do not limit the hardware structure of the database system described above. The memory of the database system described above can include any form of power loss storage medium, such as DRAM, SRAM, etc. The disk of the database system described above can include any form of power lossless storage medium, such as ROM, SSD, etc.
[0070] The above is the data compression method provided by the embodiments of the present specification in three cases. Those skilled in the art should understand that, since the three data compression methods described above are respectively for the compression of three different specifications of data blocks, the three data compression methods can be used individually, or two or more methods can be combined.
[0071] In addition, the above three data compression methods are all explained by taking the data in the disk of the database system as divided into three layers: L0, L1, and L2. Those skilled in the art should be able to understand that the SSTable microblocks, SSTable small blocks, and SSTable macroblocks described in the above embodiments are only logical concepts located in different layers and of different sizes. When the data in the disk is divided into m layers (m is greater than 3), the SSTable data blocks in the upper n1 layer can be called SSTable microblocks in turn, and the SSTable microblocks of each layer in the n1 layer gradually increase as they go down. The SSTable data blocks in the middle n2 layer are called SSTable small blocks, and the SSTable small blocks of each layer in the n2 layer gradually increase as they go down. The SSTable data blocks in the lower n3 layer are called SSTable macroblocks, and the SSTable macroblocks of each layer in the n3 layer gradually increase as they go down. n1+n2+n3=m. Then when compressing the MemTable in the memory into SSTable microblocks, the following can be used: Figure 1 The first data compression method shown is used, and when the SSTable microblock is further compressed downward into a larger SSTable microblock in the n1 layer (i.e., not exceeding the n1 layer), compression can continue to be performed according to the first data compression method mentioned above or by using the compression method in the prior art. When the SSTable microblock is compressed into an SSTable small block (i.e., compressed from the n1 layer to the n1+1 layer), the second data compression method mentioned above can be used, and when the SSTable small block is further compressed downward into a larger SSTable small block in the n2 layer (i.e., not exceeding the n2 layer), compression can continue to be performed according to the second data compression method mentioned above or by using the compression method in the prior art. When the SSTable small block is compressed into an SSTable macroblock (i.e., compressed from the n2 layer to the n2+1 layer), the third data compression method mentioned above can be used, and when the SSTable macroblock is further compressed downward into a larger SSTable macroblock in the n3 layer, compression can continue to be performed according to the third data compression method mentioned above or by using the compression method in the prior art.
[0072] Furthermore, in the embodiments of this specification, in order to facilitate the processing of data in the SSTable files (SSTable microblocks, SSTable small blocks, SSTable macroblocks) on the disk, the SSTable microblocks in this specification can store data in a mixed row and column manner, such as Figure 4 shown.
[0073] Figure 4The structure of the SSTable file provided by the embodiments of the present specification is shown in the following figure. Since the SSTable microblock is usually the minimum unit of data read by the database system from the disk, and other larger SSTable files are essentially SSTable files containing more SSTable microblocks, therefore, Figure 4 In the present specification, only two concepts of SSTable microblock and SSTable macroblock are shown, and the so-called SSTable small block is no longer shown. The so-called SSTable small block is only a SSTable macroblock containing less SSTable microblocks.
[0074] In the present specification, Figure 4 In the present specification, in addition to containing several SSTable microblocks, one SSTable macroblock also contains the macroblock header information, offset information, microblock index information and meta information of the SSTable macroblock. Among them: The macroblock header information is used to describe the block identification and size of the SSTable macroblock; The offset information is used to describe the offset of the SSTable macroblock relative to a specified position in the disk; The microblock index information is used to index the SSTable microblocks in the SSTable macroblock; The meta information is used to describe the statistical information of the SSTable macroblock, such as the number of rows.
[0075] And for any SSTable microblock in the SSTable macroblock, the SSTable microblock contains microblock header information, column header information of each data column of the data contained in the SSTable microblock, meta information of each data column of the data contained in the SSTable microblock, data rows contained in the SSTable microblock, and data row index information. Among them: The microblock header information is used to describe the block identification and size of the SSTable microblock; The column header information can be the column name of each data column of the data contained in the SSTable microblock; The meta information of the data column is the statistical information of the values in each data column. For example, assuming that a data column is gender, the meta information of the column can be the statistical information of the values Male and Female; The data row is a complete data organized by row, such as “Zhang San, male, 22 years old”; The data row index information is used to index the data rows contained in the SSTable microblock.
[0076] Through Figure 4The row-column mixed storage format can facilitate processing based on data in the SSTable microblock in the disk, for example, when data is analyzed, the column header information of the data column in the SSTable microblock and the meta information of the data column can be directly used for analysis, and when a transaction is processed based on data, the data row in the SSTable microblock can be used for processing.
[0077] The above is a data compression method provided by an embodiment of the present specification, based on the same idea, the present specification also provides a corresponding device, a storage medium and an electronic device.
[0078] Figure 5 The first data compression device provided by an embodiment of the present specification is a schematic diagram of the device, which is applied to a database system, and the device comprises: The monitoring module 501 is configured to monitor data written in the MemTable in the memory; The reading module 502 is configured to read each data written in the MemTable when the data written in the MemTable meets a first compression condition; The compression module 503 is configured to write each data read as a merged data in an SSTable microblock in the disk; the size of the SSTable microblock is not less than the MemTable; The storage module 504 is configured to store the SSTable microblock.
[0079] Optionally, the compression module 503 is specifically configured to write each data read as a merged data, and store the merged data by using a plurality of system data blocks; the total size of each system data block storing the merged data is not less than the size of the merged data; and each system data block storing the merged data is compressed into an SSTable microblock.
[0080] Optionally, the compression module 503 is further configured to monitor data written in each SSTable microblock in the disk; when the data written in each SSTable microblock meets a second compression condition, for each data in each SSTable microblock, determine the key of the data and the value of the data in the SSTable microblock; according to the key of the data, query the data corresponding to the key of the data in the SSTable microblock as a target data; and add the value of the data in the SSTable microblock to the target data stored in the SSTable microblock.
[0081] Optionally, the compression module 503 is specifically configured to keep the key of the target data stored in the SSTable microblock unchanged, add the value of the target data stored in the SSTable microblock as the first version value corresponding to the key, and add the value of the data in the SSTable microblock as the second version value corresponding to the key to the target data.
[0082] Optionally, for any data written in the SSTable microblock in the disk, the key of the data corresponds to at least one version of value in the SSTable microblock. The compression module 503 is further configured to monitor the data written in each SSTable microblock in the disk, determine the key of each data in each SSTable microblock when the data written in each SSTable microblock meets a third compression condition, determine a target value from each version of value corresponding to the key of the data, and query data corresponding to the key of the data in the SSTable macroblock as target data, wherein each version of value corresponding to the key of the data includes each version of value corresponding to the key of the data in the SSTable microblock and the value corresponding to the key of the data in the SSTable macroblock, and update the value in the target data stored in the SSTable macroblock to the target value.
[0083] Optionally, the compression module 503 is specifically configured to determine the time length of each version of value corresponding to the key of the data written in the database system, and determine a target value from each version of value corresponding to the key of the data according to the time length.
[0084] Optionally, the compression module 503 is further configured to determine the key of the data with a deletion mark in the SSTable microblock as a deletion key when the data written in each SSTable microblock meets the third compression condition, and delete the data corresponding to the deletion key from the data stored in the SSTable macroblock.
[0085] Optionally, for any SSTable microblock, the SSTable microblock includes column header information of data columns of each data in the SSTable microblock, meta information of the data columns of each data in the SSTable microblock, data rows of each data in the SSTable microblock, and data row index information.
[0086] Figure 6 A second data compression device schematic diagram provided by an embodiment of the present specification is provided, the device is applied to a database system, and the device includes: A monitoring module 601 is configured to monitor the data written in each SSTable microblock in the disk. The determining module 602 is configured to determine, for each data in each SSTable microblock, a key of the data and a value of the data in the SSTable microblock when the data written in the SSTable microblock meets a second compression condition. The querying module 603 is configured to query, according to the key of the data, data corresponding to the key in the SSTable microblock as target data. The compressing module 604 is configured to add the value of the data in the SSTable microblock to the target data stored in the SSTable microblock.
[0087] Optionally, the compressing module 604 is specifically configured to keep the key of the target data stored in the SSTable microblock unchanged, add the value of the target data stored in the SSTable microblock as a first version value corresponding to the key, and add the value of the data in the SSTable microblock as a second version value corresponding to the key to the target data.
[0088] Figure 7 A third data compression device provided by an embodiment of the present specification is shown in a schematic diagram. The device is applied to a database system, and is used for any data written in an SSTable microblock in a disk of the database system. The key of the data corresponds to at least one version of value in the SSTable microblock. The device comprises: The monitoring module 701 is configured to monitor data written in each SSTable microblock in the disk. The first determining module 702 is configured to determine, for each data in each SSTable microblock, a key of the data when the data written in the SSTable microblock meets a third compression condition. The second determining module 703 is configured to determine a target value from each version of value corresponding to the key of the data. Each version of value corresponding to the key of the data includes each version of value corresponding to the key of the data in the SSTable microblock and a value corresponding to the key of the data in the SSTable macroblock. The querying module 704 is configured to query, in an SSTable macroblock, data corresponding to the key of the data as target data. The compressing module 705 is configured to update the value in the target data stored in the SSTable macroblock to the target value.
[0089] Optionally, the second determining module 703 is specifically configured to determine a time length of each version of value corresponding to the key of the data written in the database system, and determine the target value from each version of value corresponding to the key of the data according to the time length.
[0090] Optionally, the compression module 705 is also used to, when the data written in each SSTable small block meets the third compression condition, determine the key of the data marked with deletion in the SSTable small block as the deletion key; and delete the data corresponding to the deletion key from the data stored in the SSTable macro block.
[0091] This specification also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can be used to execute the data compression method provided above.
[0092] based on Figures 1-3 The data compression method shown in the embodiment of this specification also provides Figure 8 The structural diagram of the electronic device shown in FIG. Figure 8 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its operations. The processor reads the corresponding computer program from the non-volatile storage into the memory and then runs it to implement the above-mentioned data compression method.
[0093] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A data compression method, the method comprising: monitoring data written in a MemTable in memory; when the data written in the MemTable meets a first compression condition, reading each data written in the MemTable; writing the read data as a merged data in an SSTable microblock in disk; the size of the SSTable microblock is not less than the MemTable; storing the SSTable microblock. 2.The method of claim 1, wherein writing the read data as a merged data in an SSTable microblock in disk specifically comprises: writing the read data as a merged data, and storing the merged data using a plurality of system data blocks; the total size of each system data block storing the merged data is not less than the size of the merged data; compressing each system data block storing the merged data into an SSTable microblock. 3.The method of claim 1, further comprising: monitoring data written in each SSTable microblock in disk; when the data written in each SSTable microblock meets a second compression condition, for each data in each SSTable microblock, determining the key of the data and the value of the data in the SSTable microblock; querying, according to the key of the data, the data corresponding to the key in the SSTable microblock as a target data; adding the value of the data in the SSTable microblock to the target data stored in the SSTable microblock. 4.The method of claim 3, wherein adding the value of the data in the SSTable microblock to the target data stored in the SSTable microblock specifically comprises: keeping the key of the target data stored in the SSTable microblock unchanged, adding the value of the target data stored in the SSTable microblock as a first version value corresponding to the key to the target data, and adding the value of the data in the SSTable microblock as a second version value corresponding to the key to the target data. 5.The method of claim 1, wherein for any data written in an SSTable microblock in disk, the key of the data corresponds to at least one version of value in the SSTable microblock; the method further comprising: monitoring data written in each SSTable microblock in disk; when the data written in each SSTable microblock meets a third compression condition, for each data in each SSTable microblock, determining the key of the data; determining a target value from each version of value corresponding to the key of the data; and querying, in an SSTable macroblock, the data corresponding to the key of the data as a target data; wherein each version of value corresponding to the key of the data includes each version of value corresponding to the key of the data in the SSTable microblock and the value corresponding to the key of the data in the SSTable macroblock; updating the value in the target data stored in the SSTable macroblock to the target value. 6.The method of claim 5, wherein the target value is determined from the values of the versions corresponding to the key of the data, and specifically comprising: determining time lengths of the values of the versions corresponding to the key of the data written into the database system; and determining the target value from the values of the versions corresponding to the key of the data according to the time lengths. 7.The method of claim 5, wherein when the data written into each of the SSTable microblocks satisfies a third compression condition, the method further comprises: determining a key of the data with a deletion mark in the SSTable microblock as a deletion key; and deleting the data corresponding to the deletion key from the data stored in the SSTable macroblock. 8.The method of any one of claims 1-7, wherein for any SSTable microblock, the SSTable microblock comprises column header information of data columns of each data in the SSTable microblock, meta information of the data columns of each data in the SSTable microblock, data rows of each data in the SSTable microblock, and data row index information. 9.A data compression method, the method comprising: monitoring data written into each of SSTable microblocks in a disk; when the data written into each of the SSTable microblocks satisfies a second compression condition, determining, for each data in each of the SSTable microblocks, a key of the data and a value of the data in the SSTable microblock; querying, according to the key of the data, data corresponding to the key of the data in a SSTable microblock as target data; and adding the value of the data in the SSTable microblock to the target data stored in the SSTable microblock. 10.The method of claim 9, wherein the value of the data in the SSTable microblock is added to the target data stored in the SSTable microblock, and specifically comprising: keeping the key of the target data stored in the SSTable microblock unchanged, adding a first version value of the target data stored in the SSTable microblock as the key and a second version value of the data in the SSTable microblock as the key to the target data. 11.A data compression method, wherein for any data written into a SSTable microblock in a disk of a database system, a key of the data corresponds to at least one version of value in the SSTable microblock; the method comprising: monitoring data written into each of SSTable microblocks in the disk; when the data written into each of the SSTable microblocks satisfies a third compression condition, determining, for each data in each of the SSTable microblocks, a key of the data; determining a target value from values of versions corresponding to the key of the data; and querying, in a SSTable macroblock, data corresponding to the key of the data as target data; wherein the values of the versions corresponding to the key of the data include values of versions corresponding to the key of the data in the SSTable microblock and a value corresponding to the key of the data in the SSTable macroblock; and updating a value in the target data stored in the SSTable macroblock to the target value. 12.The method of claim 11, wherein the target value is determined from the values of the versions corresponding to the key of the data, and the determining the target value comprises: determining time lengths at which the values of the versions corresponding to the key of the data are written into the database system; and determining the target value from the values of the versions corresponding to the key of the data according to the time lengths. 13.The method of claim 11, wherein when the data written into each of the SSTable chunks satisfies a third compression condition, the method further comprises: determining a key of data with a deletion mark in the SSTable chunk as a deletion key; and deleting data corresponding to the deletion key from data stored in the SSTable macro chunk. 14.A data compression apparatus, comprising: a monitoring module configured to monitor data written into a MemTable in a memory; a reading module configured to read each of the data written into the MemTable when the data written into the MemTable satisfies a first compression condition; a compression module configured to write each of the read data as a merged data into an SSTable micro chunk in a disk; wherein a size of the SSTable micro chunk is not less than the MemTable; and a storage module configured to store the SSTable micro chunk. 15.A data compression apparatus, comprising: a monitoring module configured to monitor data written into each of SSTable micro chunks in a disk; a determining module configured to, when the data written into each of the SSTable micro chunks satisfies a second compression condition, determine, for each of the data in each of the SSTable micro chunks, a key of the data and a value of the data in the SSTable micro chunk; a querying module configured to query, according to the key of the data, data corresponding to the key of the data in an SSTable chunk as target data; and a compression module configured to add the value of the data in the SSTable micro chunk to the target data stored in the SSTable chunk. 16.A data compression apparatus, wherein for any data written into an SSTable chunk in a disk of a database system, a key of the data corresponds to at least one version of a value in the SSTable chunk, and the apparatus comprises: a monitoring module configured to monitor data written into each of SSTable chunks in the disk; a first determining module configured to, when the data written into each of the SSTable chunks satisfies a third compression condition, determine, for each of the data in each of the SSTable chunks, a key of the data; a second determining module configured to determine a target value from values of versions corresponding to the key of the data, wherein the values of the versions corresponding to the key of the data include values of the versions corresponding to the key of the data in the SSTable chunk and a value corresponding to the key of the data in an SSTable macro chunk; a querying module configured to query, in the SSTable macro chunk, data corresponding to the key of the data as target data; and a compression module configured to update a value in the target data stored in the SSTable macro chunk to the target value. 17. A computer readable storage medium, the storage medium having stored thereon a computer program which, when executed by a processor, implements the method of any one of claims 1-13.
18. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor implementing the method of any one of claims 1-13 when executing the program.
Citation Information
Patent Citations
Writing and block granularity compressing and combining method and system of key value storage system based on OCSSD
CN112346666A
Data management method and device of key value storage system
CN114253908A
Key value storage engine merging strategy conversion method and system, medium and equipment
CN116088762A
Software design method, device and equipment for mobile database
CN118377463A
Data system, data management method and device and data query method and device
CN119621717A