Data Compression Method, Apparatus and Computer Device
By maintaining recommended records in the database, and using the encoding rule information of historical compressed objects to quickly select compression encoding rules, the problem of time-consuming and resource-consuming selection of compression encoding rules in the database is solved, and the compression efficiency is improved.
Patent Information
- Application Number
- CN202010656511.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-07-09
AI Technical Summary
The prior art requires a lot of computing resources and time to select compression coding rules in a database, resulting in inefficient compression.
By maintaining a recommendation record, record the compression coding rules and corresponding compression rate information of the historical compressed object. First, find out whether there are encoding rules that meet the compression rate conditions in the recommendation record. If not, start the regular encoding process and directly use the recommendation rules for compression.
It significantly improves compression efficiency and reduces the time and resource consumption for calculating appropriate encoding rules for compressed objects.
Smart Images

Figure CN111817722B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of data processing, and particularly to data compression methods, devices, and computer equipment. Background Art
[0002] Data compression can be understood as representing more information with fewer encodings. When performing data compression, it is necessary to select an appropriate compression encoding rule to obtain better compression efficiency. Currently, there are more and more compression encoding rules, and a relatively large computational cost is required to select an appropriate compression encoding rule. Summary of the Invention
[0003] To overcome the problems existing in the related art, this specification provides a data compression method, device, and computer equipment.
[0004] According to a first aspect of an embodiment of this specification, a data compression method is provided, including:
[0005] Obtain an object to be compressed;
[0006] Check whether there is a recommended compression encoding rule that meets the compression rate condition in the recommended records. The recommended records are used to record: the compression encoding rules of historical compression objects and the corresponding compression rate information, and the historical compression objects are of the same type as the object to be compressed;
[0007] If there is, compress the object to be compressed using the recommended compression encoding rule;
[0008] If not, start a conventional compression encoding process to obtain the estimated compression rates of multiple compression encoding rules for the object to be compressed, select a target compression encoding rule at least based on the estimated compression rates, and compress the object to be compressed using the target compression encoding rule.
[0009] Optionally, the object to be compressed includes: a data unit to be compressed obtained by splitting the data to be compressed; the historical compression objects include: other historical compressed data units obtained by splitting the data to be compressed.
[0010] Optionally, the object to be compressed includes: a data unit to be compressed obtained by splitting the data to be compressed in a data table; the historical compression objects include: the historical compressed data units corresponding to the data table.
[0011] Optionally, the object to be compressed includes: a column of data to be compressed in a data table; the historical compression objects include: the historical compressed data in the same column as the column of data to be compressed in the data table.
[0012] Optionally, the object to be compressed includes: a copy of a data table; the historical compression objects include: the main data table corresponding to the copy of the data table.
[0013] Optionally, the compression rate condition at least includes: the compression rate information of the recommended compression coding rule is higher than a set threshold.
[0014] Optionally, after compressing the object to be compressed, it further includes: updating the recommendation record based on the compression coding rule used by the object to be compressed.
[0015] Optionally, the updating the recommendation record based on the compression coding rule used by the object to be compressed includes:
[0016] Obtaining the actual compression rate of the object to be compressed, and updating the recommendation record at least based on the actual compression rate and the compression coding rule used by the object to be compressed.
[0017] Optionally, the method further includes: obtaining access requirement information of the object to be compressed, where the access requirement information is related to the decompression efficiency of the object to be compressed;
[0018] The compression rate condition includes: the compression rate information of the recommended compression coding rule matches the access requirement information.
[0019] Optionally, the selecting the target compression coding rule at least based on the estimated compression rate includes:
[0020] Selecting the target compression coding rule based on the estimated compression rate and the access requirement information.
[0021] Optionally, the access requirement of the object to be compressed is determined by obtaining historical access data of the object to be compressed and / or the historical compressed object.
[0022] Optionally, the historical access data includes: historical access frequency.
[0023] Optionally, the updating the recommendation record at least based on the actual compression rate and the compression coding rule used by the object to be compressed includes:
[0024] Updating the recommendation record based on the actual compression rate, the access requirement information, and the compression coding rule used by the object to be compressed.
[0025] Optionally, the compression rate information of the compression coding rule includes: a confidence level representing the actual compression rate level of the compression coding rule;
[0026] The updating the recommendation record based on the actual compression rate, the access requirement information, and the compression coding rule used by the object to be compressed includes:
[0027] When the compression encoding rule used by the object to be compressed is the recommended compression encoding rule recorded in the recommendation record, if the actual compression rate of the object to be compressed matches the access requirement information, increase the confidence level of the recommended compression encoding rule; otherwise, decrease the confidence level of the recommended compression encoding rule.
[0028] When the compression encoding rule used by the object to be compressed is different from the recommended compression encoding rule recorded in the recommendation record, if the actual compression rate of the object to be compressed matches the access requirement information, replace the recommended compression encoding rule in the recommendation record with the compression encoding rule used by the object to be compressed.
[0029] According to the second aspect of the embodiments of this specification, a data compression device is provided, and the device includes:
[0030] An acquisition module, configured to: acquire an object to be compressed;
[0031] A search module, configured to: search whether there is a recommended compression encoding rule that meets the compression rate condition in the recommendation record, where the recommendation record is used to record: the compression encoding rule of the historical compression object and the corresponding compression rate information, and the historical compression object has the same type as the object to be compressed;
[0032] A first compression module, configured to: if there is a recommended compression encoding rule that meets the compression rate condition, use the recommended compression encoding rule to compress the object to be compressed;
[0033] A second compression module, configured to: if there is no recommended compression encoding rule that meets the compression rate condition, start a conventional compression encoding process to obtain the estimated compression rates of multiple compression encoding rules for the object to be compressed, select a target compression encoding rule at least based on the estimated compression rates, and use the target compression encoding rule to compress the object to be compressed.
[0034] According to the third aspect of the embodiments of this specification, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the embodiments of the foregoing data compression method when executing the program.
[0035] The technical solutions provided by the embodiments of this specification may include the following beneficial effects:
[0036] In the embodiments of the present specification, since the compression coding rules and the corresponding compression ratio information of the historical compression object are recorded in the recommendation record, it is possible to search in the recommendation record for a suitable compression coding rule based on the compression ratio condition. If not, start the conventional coding compression process to compress the object to be compressed; if there is one, it can directly use this compression coding rule for compression. Therefore, there is no need to spend time and resources calculating a suitable compression coding rule for the current object to be compressed, thus significantly improving the compression efficiency.
[0037] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. Brief Description of the Drawings
[0038] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.
[0039] Figure 1 is a flowchart of a data compression method shown according to an exemplary embodiment of this specification.
[0040] Figure 2 is a schematic diagram of a data compression method shown according to an exemplary embodiment of this specification.
[0041] Figure 3 is a hardware structure diagram of a computer device where a data compression device is located shown according to an exemplary embodiment of this specification.
[0042] Figure 4 is a block diagram of a data compression device shown according to an exemplary embodiment of this specification. Detailed Embodiments
[0043] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. On the contrary, they are only examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0044] The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit this specification. The singular forms "a", "the", and "said" used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0045] It should be understood that although terms such as first, second, and third may be used in this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0046] For relational databases, data query processing and data storage have always been the two most core elements. Generally speaking, data query processing consumes more CPU and memory resources, while data storage mainly consumes storage hardware resources such as disks. However, to date, the development of storage hardware has always lagged behind the development speed of other system hardware. On the one hand, the growth rate of storage hardware space always fails to keep up with the storage needs of users who hope to store a wider range and longer-term data; on the other hand, the access latency of current storage hardware always lags behind that of CPU and memory by at least one order of magnitude. Generally speaking, the problems of space and access latency have become the two major difficulties hindering the rapid development of database storage.
[0047] To solve the above problems, various database manufacturers in the industry have actively introduced general data compression technologies, which compress data at the page / block level based on their respective databases. This not only greatly reduces the storage space occupancy, but also improves the storage access latency during large-scale data scanning queries because the amount of data IO (input / output) is reduced. The price paid is more CPU and memory consumption during the compression and decompression processes. From the overall perspective of the database, the introduction of compression technology can make the overall resource usage more balanced, and under the same hardware conditions, it can meet more demanding business requirements of users.
[0048] Compression can be understood as representing more information with fewer encodings. A common compression method in the mixed row and column compression of databases is to compress different data using different encoding rules according to the data distribution.
[0049] The theoretical limit of data compression is the corresponding information entropy, that is, the more data information is obtained, the higher the final data compression rate will be. In order to support more usage scenarios, general data compression technology generally only builds a sliding dictionary for byte streams for encoding and compression, and will use context information within a certain range. Relational databases naturally have more background knowledge of internal data. For example, the type of data column can infer the value range distribution of data, and data in the same column will have stronger clustering. Therefore, more and more database vendors will split data according to compression units, store them in column storage format inside each compression unit, and use the context relationship within the column and the data relationship between columns to find the most matching encoding rules for compression, such as dictionary encoding, run-length encoding, incremental encoding, etc. Compared with general compression, these encoding rules can bring higher compression rates because they use built-in data information. At the same time, using built-in storage formats also enables databases to have the ability to directly query based on encoded data. Especially for databases with LSM-Tree (Log-StructuredMerge-Tree) architecture, since the stored data needs to be continuously merged, and the merge operation involves the overall rewriting of the data, it is natural to add row-column mixed encoding compression during the data reorganization process.
[0050] Although the database's internal support for row-column mixed coding compression can bring higher compression rates, it is accompanied by complex issues such as compression coding rule selection and detection. The database needs to traverse the column data in all compression units to extract summary data, and then evaluate the final compression size based on different compression rules and decide on the final coding rule. The entire coding rule detection process requires a lot of pre-calculation and matching comparison. In addition, the current compression coding rule detection generally only targets all data contexts of the same compression unit. Once the scope is expanded, the detection complexity will increase exponentially.
[0051] The encoding rule compression of each database is basically implemented based on the row-column mixed storage mode. A certain number of data record tuples are selected as a whole compression unit, and all data in this compression unit are organized and stored by column. The database will traverse and scan each column of data to obtain the basic distribution and data characteristics of the column data, and then select appropriate rules for encoding and compression based on the specific data characteristics. Different databases implement different encoding compression rules. Generally speaking, they basically include dictionary encoding, run-length encoding, incremental encoding, numerical encoding, and some open source general compression encoding. In this implementation method, the compression encoding rule selection speed is slow. Before the database determines the final compression encoding rule for a column of data, it must go through several stages of data scanning analysis, pre-calculation of compression rates of different encoding rules, encoding rule selection, and data compression. Each stage requires a large computational cost.
[0052] Based on this, the embodiments of this specification provide a data compression solution. The solution in this embodiment relates to recommendation records, which are used to record the compression coding rules and the corresponding compression rate information of historical compression objects. The historical compression objects are of the same type as the objects to be compressed. Therefore, when compressing the objects to be compressed in the embodiments of this specification, it is possible to first check in the recommendation records whether there is a suitable compression coding rule based on the compression rate condition. If not, start the conventional coding compression process to compress the objects to be compressed; if there is one, it can directly use this compression coding rule for compression. Therefore, there is no need to spend time and resources calculating a suitable compression coding rule for the current object to be compressed, thus significantly improving the compression efficiency.
[0053] Next, the embodiments of this specification will be described in detail. As Figure 1 shown, Figure 1 is a flowchart of a method shown in accordance with an exemplary embodiment of this specification, including the following steps:
[0054] In step 102, obtain the objects to be compressed;
[0055] In step 104, check whether there is a recommended compression coding rule in the recommendation records that meets the compression rate condition. The recommendation records are used to record: the compression coding rules of historical compression objects and the corresponding compression rate information, and the historical compression objects are of the same type as the objects to be compressed.
[0056] In step 106, if there is one, use the recommended compression coding rule to compress the objects to be compressed
[0057] In step 108, if there is none, start the conventional compression coding process to obtain the estimated compression rates of multiple compression coding rules for the objects to be compressed, select the target compression coding rule at least based on the estimated compression rates, and use the target compression coding rule to compress the objects to be compressed.
[0058] This embodiment is applicable to various data compression scenarios. The object to be compressed can be various types of objects such as data tables, video files, audio files, or images, and this embodiment does not limit this. The historical compressed object is an object of the same type as the object to be compressed and has been compressed before the object to be compressed; for different objects to be compressed, there can be various different implementation methods for the historical compressed object to be of the same type as the object to be compressed. As an example, the object to be compressed and the historical compressed object can be data of different times and different versions, such as program source files, etc.; the object to be compressed and the historical compressed object can also be data of the same format, such as data of the same video format; the object to be compressed and the historical compressed object can also be each sub-data belonging to the same parent data, such as multiple sub-data obtained by splitting the original data, or multiple data units belonging to the same data table, etc.
[0059] As an example, assume that currently it is necessary to compress the object to be compressed, Data-A1; in the first compression process, since there is no historical compressed object of the same type as the object to be compressed for reference, the recommendation record is empty, and there is no recommended compression coding rule that meets the compression rate condition. Therefore, the conventional compression coding process is started, and the estimated compression rates of multiple compression coding rules for the object to be compressed can be obtained. At least based on the estimated compression rates, a target compression coding rule is selected, and the object to be compressed is compressed using the selected target compression coding rule.
[0060] After the first compression is completed, the recommendation record is updated based on the compression coding rule used by the object to be compressed, Data-A1. For example, the situation of this compression can be recorded in the recommendation record. Specifically, the target compression coding rule and the compression rate information of this target compression coding rule can be written in the recommendation record. In some examples, the estimated compression rate of the compression coding rule for the object to be compressed is not much different from the actual compression rate, and the compression rate information can be information related to the estimated compression rate; in other examples, the estimated compression rate of the compression coding rule for the object to be compressed may have a certain difference from the actual compression rate. It may be that the actual compression rate is lower than the estimated compression rate, or it may be higher than the estimated compression rate. Therefore, in this embodiment, the actual compression rate of the object to be compressed can also be obtained, and the recommendation record is updated at least based on the actual compression rate and the compression coding rule used by the object to be compressed; then the compression rate information can be information related to the actual compression rate, and the actual compression rate can be calculated after Data-A1 is compressed. For example, it is determined according to the ratio of the actual size of the object to be compressed after compression to the original size of the object to be compressed. Among them, the compression rate information can directly adopt the value of the compression rate, or it can be related information obtained by converting the compression rate in a set manner according to needs.
[0061] Subsequently, due to inserting new data or other reasons, Data-A2 is newly generated based on Data-A1; in this compression process, Data-A2, as a new object to be compressed, has the same type as Data-A2, and Data-A1 is used as a historical compression object to reference its compression process.
[0062] During the compression process of Data-A2, the compression encoding rule can be selected by referring to the compression situation of Data-A1; for example, the recommended record records the compression encoding rule for compressing Data-A1 and the corresponding compression rate information, and it can be determined whether to select the compression encoding rule of Data-A1 for encoding according to the compression rate condition; if the compression encoding rule of Data-A1 is selected, the Data-A2 can be directly compressed using this compression encoding rule. Among them, the compression rate condition in this embodiment can be flexibly configured according to needs. For example, some objects to be compressed have certain requirements for the compression rate, and different compression encoding rules have different compression effects on the compression object. This embodiment can set the compression rate condition to select a suitable recommended compression encoding rule. As an example, the compression rate condition can include: the compression rate information of the recommended compression encoding rule is higher than the set threshold, and the specific threshold can be flexibly configured according to actual needs, and this embodiment does not limit this.
[0063] It can be understood that during the compression process of Data-A2, since the compression encoding rule of the historical compression object can be selected for compression, there is no need to spend time and resources calculating a suitable compression encoding rule for the current object to be compressed, thus significantly improving the compression efficiency.
[0064] Of course, if the compression encoding rule of Data-A1 does not meet the compression rate condition of Data-A2, Data-A2 can also start a conventional compression encoding process to obtain the estimated compression rates of multiple compression encoding rules for the object to be compressed, and at least select a target compression encoding rule based on the estimated compression rate, and use the target compression encoding rule to compress the object to be compressed.
[0065] After compressing Data-A2, the recommended record can also be updated based on the compression situation of Data-A2; for example, in actual business, it may occur that due to inserting new data into Data-A1, Data-A2 may not be applicable to the original compression encoding rule of Data-A1. Therefore, in this embodiment, the actual compression rate of Data-A2 can be obtained, and the recommended record can be updated at least based on the actual compression rate and the compression encoding rule used by Data-A2. For example, the recommended record is updated to the compression encoding rule used by Data-A2, and the compression rate information is updated according to the actual compression rate of Data-A2.
[0066] In an actual compression business scenario, there may be a problem of large data to be compressed. Therefore, the data to be compressed can be divided into multiple data units to be compressed as needed. By adopting the data compression method of this embodiment in such a business scenario, a better improvement in compression efficiency can be obtained. As an example, the object to be compressed may include: the data units to be compressed obtained by dividing the data to be compressed; the historical compression object includes: other historical compressed data units obtained by dividing the data to be compressed.
[0067] As an example, as Figure 2 shown, Data-B can be divided into 4 data units to be compressed, namely Data-b1 to Data-b4.
[0068] For Data-b1, in the first compression process, since there is no historical compression object for reference, the recommendation record is empty, and there is no recommended compression coding rule that meets the compression rate condition. Therefore, the conventional compression coding process is started, and multiple compression coding rules can be obtained for the estimated compression rate of the object to be compressed. At least based on the estimated compression rate, the target compression coding rule is selected, and the object to be compressed is compressed using the selected target compression coding rule.
[0069] After the first compression is completed, the situation of this compression can be recorded in the recommendation record; for example, the target compression coding rule and the compression rate information of the target compression coding rule are recorded in the recommendation record. Subsequently, during the compression process of Data-b2, the compression situation of Data-b1 can be referred to for selecting the compression coding rule; for example, the compression coding rule for compressing Data-b1 and the corresponding compression rate information are recorded in the recommendation record, and it can be determined whether to select the compression coding rule of Data-b1 for coding according to the compression rate condition; if the compression coding rule of Data-b1 is selected, then this compression coding rule can be directly used to compress the Data-b2. After the compression of Data-b2, it is determined whether to update the recommendation record according to the compression situation. Subsequently, Data-b3 and Data-b4 can also achieve fast compression and update the recommendation record using the same process.
[0070] In some other examples, taking the compression scenario of a data table as an example, in a distributed relational database such as the LSM-Tree architecture, the data table may face multiple merges, and the new data in each merge needs to be compressed. Using the solution of the embodiment of this specification, there is no need to repeatedly estimate the compression rates of multiple compression coding rules. The recommendation record can be constructed using the compression knowledge of the historical compression object, and in subsequent other compression processes, the compression coding rule can be quickly obtained based on the recommendation record.
[0071] As an example, currently, the data to be compressed in tableC needs to be compressed. Since the data to be compressed in tableC is large, according to the set data unit size, the data to be compressed in tableC is divided into 5 data units to be compressed, namely table-c1 to table-c5.
[0072] For table-c1, in the first compression process, since there is no historical compression object for reference and the recommendation record is empty, the conventional compression encoding process is started. Multiple compression encoding rules can be obtained for the estimated compression rate of the object to be compressed. At least based on the estimated compression rate, a target compression encoding rule is selected, and the object to be compressed is compressed using the selected target compression encoding rule.
[0073] After the first compression is completed, the current compression situation can be recorded in the recommendation record; for example, the target compression encoding rule and the compression rate information of the target compression encoding rule are recorded in the recommendation record. Subsequently, during the compression process of table-c2, the compression situation of table-c1 can be referred to for selecting the compression encoding rule; for example, the compression encoding rule and the corresponding compression rate information for compressing table-c1 are recorded in the recommendation record. It can be determined whether to select the compression encoding rule for table-c1 for encoding according to the compression rate condition; if the compression encoding rule for table-c1 is selected, the table-c2 can be directly compressed using this compression encoding rule. After table-c2 is compressed, it is determined whether to update the recommendation record according to the compression situation. Subsequently, table-c3 and table-c4 can also achieve fast compression and update the recommendation record using the same process.
[0074] In some other examples, a data table usually contains multiple columns of data with different attributes, and the data in each column may vary greatly. Different compression encoding rules can be used for the data in different columns of the data table. For example, in a user data table, a column of data related to user age is integer data, and a column of data related to user names is string data. Appropriate compression encoding rules can be selected for these two columns of data respectively. Based on this, the object to be compressed in this embodiment may include: a column of data to be compressed in the data table; the target compression object may include: the historical compressed data in the same column as the column of data to be compressed in the data table. Based on this, this embodiment can distinguish column data, and the compression encoding rules recorded in the recommendation record correspond to the column data. When a column of data in the data table is compressed and then needs to be compressed again after an update, the compression encoding rule of the historical compressed data in the same column of the data table can be used for compression, thereby improving the compression efficiency.
[0075] In some other examples, the object to be compressed includes: a copy of a data table; the historical compressed object includes: the main data table corresponding to the copy of the data table.
[0076] Similar to the previous examples, the content of the copy of the data table is exactly the same as that of the main data table, and the copy of the data table is consistent with the main data table. Usually, after the main data table is updated, the copy of the data table is updated according to the update operation of the main data table. If the main data table involves compressed storage after update, the copy of the data table also needs to be compressed and stored after update. Based on this, during the compression process of the copy of the data table, the compression encoding rule of the main data table can be directly used to perform compression. Optionally, when compressing the copy of the data table, search for the compression encoding rule of the main data table and the corresponding compression rate information in the recommended records, and directly use the compression encoding rule of the main data table for compression. Of course, in actual business, there may also be a situation where the actual compression rate of the compression encoding rule selected for the main data table is not high. In the search for recommended records, it can be determined according to needs that there is no recommended compression encoding rule that meets the compression rate condition in the recommended records, and then the conventional compression encoding process is started to compress the copy of the data table.
[0077] In actual business, the evaluation criteria for compression encoding rules are single, and the main basis for selecting compression encoding rules is the level of the final actual compression rate, without comprehensively considering the additional overhead brought by different compression encoding rules to different data access modes under the actual database business load. For example, assuming that the data compression rate is high, the corresponding data decompression process will take longer, and the compression rate is negatively correlated with the decompression efficiency. If the data needs to be accessed frequently and the data compression rate is high, and data decompression processing is required for each access, it will inevitably bring a large additional decompression overhead. Based on this, in this embodiment, the access requirement information of the object to be compressed is obtained. The access requirement information may be related to the decompression efficiency. For example, the greater the access requirement, the faster the corresponding decompression efficiency should be. This access requirement information can be configured by the business party, or can be configured by the data access party, or can be determined by collecting the historical access data of the object to be compressed and / or the historical compressed object; for example, assuming that the historical compressed object is frequently accessed, the data access requirement is large, and the corresponding decompression efficiency requirement is large. Based on this, a higher access requirement can be set; if the historical compressed object is rarely accessed, the data access requirement is small, and the corresponding decompression efficiency requirement is small. Based on this, a lower access requirement can be set. Optionally, the historical access information of the object to be compressed can be obtained based on the historical compression data of the same type as the object to be compressed. Optionally, the historical access information may include the operation type of the data, such as insert, delete, update, and query; or the number of operations on the data within a certain time period, or the historical access frequency, etc.
[0078] Correspondingly, the compression ratio condition includes: the compression ratio information of the recommended compression coding rule matches the access requirement, so that an appropriate compression coding rule can be selected based on the compression ratio requirement and the access requirement information, rather than being selected only based on a single factor of the compression ratio, so that the finally compressed data can still meet the access requirements of the service.
[0079] Based on this, in the case of obtaining access requirement information, the recommendation record can also be updated based on the actual compression ratio, the access requirement information, and the compression coding rule used by the object to be compressed. As an example, at the initial stage of compression, the compression coding rule used by the object to be compressed is written into the recommendation record at the beginning. Subsequently, as the content of the data to be compressed changes, or the access situation of the data to be compressed changes, the compression coding rule used by the subsequent object to be compressed may change. Based on this, dynamic update of the recommendation record needs to be implemented. Optionally, in this embodiment, the compression ratio information of the compression coding rule is implemented by a confidence level, and this confidence level represents the actual compression ratio level of the compression coding rule, so as to dynamically adjust the credibility of the compression coding rule during the continuous compression process; optionally, in this embodiment, updating the compression coding rule and the corresponding compression ratio information in the recommendation record based on the actual compression ratio and the access requirement may include: when the compression coding rule used by the object to be compressed is the recommended compression coding rule recorded in the recommendation record, if the actual compression ratio of the object to be compressed matches the access requirement information, increase the confidence level of the recommended compression coding rule, otherwise decrease the confidence level of the recommended compression coding rule; when the compression coding rule used by the object to be compressed is different from the recommended compression coding rule recorded in the recommendation record, if the actual compression ratio of the object to be compressed matches the access requirement information, replace the recommended compression coding rule in the recommendation record with the compression coding rule used by the object to be compressed.
[0080] As can be seen from the above embodiment, when the object to be compressed is compressed using the recommended compression coding rule recorded in the recommendation record, after the compression is completed, the recommended compression coding rule is checked for suitability through the actual compression rate and access requirement information. If the actual compression rate of the object to be compressed matches the access requirement information, the confidence of the recommended compression coding rule is increased so that the compression coding rule can continue to be used in subsequent compression; otherwise, the confidence of the recommended compression coding rule is reduced so that when the confidence of the recommended compression coding rule is lower than the threshold, the conventional compression coding process can be started to obtain a new compression coding rule. If there is no suitable compression coding rule in the recommendation record, the conventional compression coding process is started to obtain a new compression coding rule, which is different from the recommended compression coding rule recorded in the recommendation record. Similarly, after the compression is completed, the recommended compression coding rule is checked for suitability through the actual compression rate and access requirement information. If the actual compression rate of the object to be compressed matches the access requirement information, the recommended compression coding rule in the recommendation record can be replaced with the compression coding rule used by the object to be compressed, so that the latest suitable compression coding can be recorded in the recommendation record for subsequent compression. If the actual compression ratio of the object to be compressed does not match the access requirement information, the compression encoding rule is not applicable and the recommendation record may not be updated.
[0081] Next, the data compression method of the present application will be described again through an embodiment. Distributed databases such as the LSM-Tree (The Log-Structured Merge-Tree) architecture face many data merging needs. Taking the relational database of the LSM-tree architecture as an example, the LSM-tree is a multi-layer structure, similar to a tree structure, small at the top and large at the bottom. The first is the C0 layer of the memory, which stores all the recently written data tuple records. This memory structure is ordered and can be updated in situ at any time, and supports queries at any time. The remaining C1 to Cn layers are all on the disk.
[0082] The data writing process is as follows: when a data writing operation comes, it is first appended to the Write Ahead Log (that is, the log recorded before the actual writing), and then added to the C0 layer. When the data in the C0 layer reaches a certain size, the C0 layer and the C1 layer are merged, similar to merge sort. This process is called Compaction. The merged new-C1 will be written to the disk sequentially, replacing the original old-C1. When the C1 layer reaches a certain size, it will continue to merge with the lower layer. After the merge, all old files can be deleted, leaving new files.
[0083] It can be seen that when a merge occurs, the data needs to be reorganized and written into a new storage file, and the new storage file needs to be compressed.
[0084] When merging and writing, the data will be sliced into independent compression units according to the fixed size specified by the user when creating the table, and the data will be compressed by row-column hybrid coding within each compression unit.
[0085] In addition, in the scenario of a distributed database, each data table will have multiple data replicas. The main data table and the data table replicas need to stagger the merging time to alternately provide services to ensure the availability and stability of user business access.
[0086] As an example, the statistical information module configured inside the database can obtain the access data of the affiliated table. Based on this access data, the access requirements of the data table can be quickly determined. For example, according to the operation type of historical access and the historical access frequency, the more point query operations on the data table and the higher the number of data table accesses, the higher the requirement for the access latency of the data table. For these data, the storage resources saved by compression cannot make up for the additional computational resource overhead, and a compression coding rule with a lower compression ratio needs to be used to ensure the user's access speed; while for data with fewer point query operations and lower data table access times, where users do not have many queries and have a lower requirement for access latency, a compression rule with a higher compression ratio can be used.
[0087] During the merging process of the data table, a large number of data units need to be compressed, and a recommendation record containing the current best compression coding rule is maintained. After each compression of the object to be compressed with a selected compression coding rule, the compression situation of the compression coding rule will be updated to the recommendation library, always keeping the recommendation record maintaining the latest compression coding rule in the current state.
[0088] Before compressing each data unit to be compressed in the data table, two input messages are received: the access requirement information of the data table and the latest recommendation record.
[0089] When the recommendation record is empty or the confidence level of the compression coding rule in the recommendation record does not reach the specified threshold, the conventional coding compression process is started to detect the data unit to be compressed. The data characteristics of each column of data are analyzed using the conventional coding compression process, and the compression ratios are sorted based on the estimated compression ratios of different compression coding rules. Finally, based on the access requirement information of the data table, among multiple compression coding rules with different estimated compression ratios, the compression coding rule that meets the access requirements of the data table is selected as the compression coding rule used for this compression; among them, the data unit to be compressed may involve multiple columns of different types of data, and corresponding compression coding rules can be selected for each column of data, that is, the data unit to be compressed can correspond to two or more compression coding rules.
[0090] When the confidence level of the compression coding rule in the recommendation record exceeds the specified threshold, the compression coding rule in the recommendation record can be directly used as the compression coding rule for the data unit to be compressed this time; among them, the data unit to be compressed may involve multiple columns of different types of data, and the corresponding compression coding rule can be selected for each column of data; in some examples, it is possible that some data columns can select the compression coding rule from the recommendation record, while some data columns cannot be selected, then the conventional compression coding process can be started for detection and selection.
[0091] After selecting one or more compression coding rules required for the data unit to be compressed this time, the data is scanned by column within the data unit to be compressed, and the selected compression coding rules are started to be used for compression. After each data column is compressed, the actual compression rate information of each data column is obtained.
[0092] After the overall compression of the data unit is completed, the recommendation record is updated based on the actual compression rate information of each data column:
[0093] In the case where the compression coding rule used by the object to be compressed is the recommended compression coding rule recorded in the recommendation record, if the actual compression rate of the object to be compressed matches the access requirement information, the confidence level of the recommended compression coding rule is increased, otherwise the confidence level of the recommended compression coding rule is decreased;
[0094] In the case where the compression coding rule used by the object to be compressed is different from the recommended compression coding rule recorded in the recommendation record, if the actual compression rate of the object to be compressed matches the access requirement information, the recommended compression coding rule in the recommendation record is replaced with the compression coding rule used by the object to be compressed.
[0095] For the compression process of the data table copy, the compression coding rule of the main data table with the same data can be directly used to compress the data in the data table copy.
[0096] As can be seen from the above embodiments, the data compression method in the embodiments of this specification adjusts the compression coding rule in the recommendation record in real time by making full use of more historical context information through a semi-supervised learning method.
[0097] For the main data table, during the merging process, the data units to be compressed in the early stage of compression need to start the conventional compression encoding process to detect each compression encoding rule, so that the latest recommended records can be trained. When the confidence level exceeds the specified threshold, the subsequent data units to be compressed can directly use the compression encoding rules of the recommended records to start compression, and no longer need to perform complex rule detection work; for the data table copy, since the data is exactly the same as the main data table, the final recommended encoding rules of the main data table can be directly used for compression, and all the detection overhead of the compression encoding rules can be completely saved.
[0098] In this embodiment, for the selection of the compression encoding rule corresponding to each data unit to be compressed, on the one hand, the tuning of the access requirements of the data table to which it belongs is added, and the selection of the final compression encoding rule is adjusted according to the access latency requirements; on the other hand, the selection of the compression encoding rule corresponding to each data unit to be compressed is not only based on the data of this unit for decision-making, but the best compression encoding rule is comprehensively considered through the actual compression rate results of more compression units.
[0099] The above embodiment proposes a data compression method for the LSM-Tree architecture distributed relational database. Combining the row-column hybrid compression encoding rule of the database, a semi-supervised multi-level feedback mechanism is implemented. The access requirement information of the data table is used to correct and adjust the selection of the specific compression encoding rule for each data unit. At the same time, the appropriate compression rule can be continuously learned during the compression process to update the recommended records, and the encoding rule selection in the entire merging process is accelerated. For the data table, the detection speed of the compression encoding rule can be greatly accelerated according to other data copies in the distributed database cluster and the historical compression data knowledge of the local machine. At the same time, the appropriate compression encoding rule is adaptively selected according to the access requirement characteristics of the data, so as to improve the compression rate and at the same time improve the access speed of the data as much as possible.
[0100] Especially for the data table copy, the recommended records of the main data table can be directly used, so that the rule recommendation and training process that consumes computing resources can be skipped, and overall, more context information can be fully utilized to accelerate and optimize the detection of the compression encoding rule.
[0101] Corresponding to the embodiment of the foregoing data compression method, this specification also provides an embodiment of a data compression device and a computer device to which it is applied.
[0102] Embodiments of the data compression device in this specification can be applied to computer devices, such as servers or terminal devices. Embodiments of the device can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor in the file processing where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From a hardware perspective, as Figure 3 shown, it is a hardware structure diagram of the computer device where the data compression device in this specification is located. In addition to Figure 3 the processor 310, memory 330, network interface 320, and non-volatile memory 340 shown, for the server or electronic device where the device 331 is located in the embodiment, usually according to the actual functions of the computer device, it may also include other hardware, which will not be elaborated here.
[0103] As Figure 4 shown, Figure 4 is a block diagram of a data compression device shown according to an exemplary embodiment of this specification. The device includes:
[0104] An acquisition module 41, configured to: acquire an object to be compressed;
[0105] A search module 42, configured to: search whether there is a recommended compression coding rule that meets the compression rate condition in the recommended records, where the recommended records are used to record: the compression coding rules of historical compression objects and the corresponding compression rate information, and the historical compression objects are of the same type as the object to be compressed;
[0106] A first compression module 43, configured to: if there is a recommended compression coding rule that meets the compression rate condition, use the recommended compression coding rule to compress the object to be compressed;
[0107] A second compression module 44, configured to: if there is no recommended compression coding rule that meets the compression rate condition, start a conventional compression coding process to obtain the estimated compression rates of multiple compression coding rules for the object to be compressed, select a target compression coding rule at least based on the estimated compression rates, and use the target compression coding rule to compress the object to be compressed.
[0108] Optionally, the object to be compressed includes: a data unit to be compressed obtained by splitting the data to be compressed; the historical compression objects include: other historical compressed data units obtained by splitting the data to be compressed.
[0109] Optionally, the object to be compressed includes: a data unit to be compressed obtained by splitting the data to be compressed in a data table; the historical compression objects include: the historical compressed data units corresponding to the data table.
[0110] Optionally, the object to be compressed includes: a column of data to be compressed in a data table; the historical compression object includes: historical compressed data in the same column as the column of data to be compressed in the data table.
[0111] Optionally, the object to be compressed includes: a copy of a data table; the historical compression object includes: the main data table corresponding to the copy of the data table.
[0112] Optionally, the compression rate condition at least includes: the compression rate information of the recommended compression coding rule is higher than a set threshold.
[0113] Optionally, the apparatus further includes an update module for: after compressing the object to be compressed, updating the recommended record based on the compression coding rule used by the object to be compressed.
[0114] Optionally, the update module is further configured to:
[0115] Obtain the actual compression rate of the object to be compressed, and update the recommended record at least based on the actual compression rate and the compression coding rule used by the object to be compressed.
[0116] Optionally, the obtaining module is further configured to: obtain access requirement information of the object to be compressed, where the access requirement information is related to the decompression efficiency of the object to be compressed;
[0117] The compression rate condition includes: the compression rate information of the recommended compression coding rule matches the access requirement information.
[0118] Optionally, the second compression module is further configured to:
[0119] Select a target compression coding rule based on the estimated compression rate and the access requirement information.
[0120] Optionally, the access requirement of the object to be compressed is determined by obtaining historical access data of the object to be compressed and / or the historical compression object.
[0121] Optionally, the historical access data includes: historical access frequency.
[0122] Optionally, the update module is further configured to:
[0123] Update the recommended record based on the actual compression rate, the access requirement information, and the compression coding rule used by the object to be compressed.
[0124] Optionally, the compression rate information of the compression coding rule includes: a confidence level indicating the actual compression rate level of the compression coding rule;
[0125] The update module is further configured to:
[0126] When the compression coding rule used by the object to be compressed is the recommended compression coding rule recorded in the recommendation record, if the actual compression rate of the object to be compressed matches the access requirement information, increase the confidence level of the recommended compression coding rule; otherwise, decrease the confidence level of the recommended compression coding rule.
[0127] When the compression coding rule used by the object to be compressed is different from the recommended compression coding rule recorded in the recommendation record, if the actual compression rate of the object to be compressed matches the access requirement information, replace the recommended compression coding rule in the recommendation record with the compression coding rule used by the object to be compressed.
[0128] Correspondingly, this specification also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the program, it implements the embodiments of the foregoing data compression method.
[0129] The implementation processes of the functions and roles of each module in the above data compression device are specifically described in the implementation processes of the corresponding steps in the above data compression method, and will not be elaborated here.
[0130] For the embodiments of the data compression device, since it basically corresponds to the embodiments of the data compression method, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this specification. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0131] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0132] Those skilled in the art will readily conceive of other embodiments of the present specification after considering the specification and practicing the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations of the specification that follow the general principles of the specification and include known common general knowledge or conventional technical means in the technical field not claimed in this specification. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the specification are pointed out by the following claims.
[0133] It should be understood that this specification is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this specification is only limited by the appended claims.
[0134] The above are only the preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification shall be included within the scope of protection of this specification.
Claims
1. A data compression method, comprising: Taking each column of data in a data table as a data unit to be compressed, and performing the following compression processing: Obtaining access requirement information of the data unit to be compressed, where the access requirement information is obtained through the historical access frequencies of the data unit to be compressed and its historical compressed data units in the same column; Checking whether there is a recommended compression coding rule that meets the compression rate condition in the recommended records, where: The recommended records are used to record the actual compression rates and confidence levels of the compression coding rules used by the historical compressed data units in the same column; The compression rate condition includes: the confidence level of the recommended compression coding rule is higher than a set threshold, and the actual compression rate matches the access requirement information; If there is, using the recommended compression coding rule to compress the data unit to be compressed, and adjusting the confidence level of the recommended compression coding rule based on whether the actual compression rate matches the access requirement information after compression, where if it matches, the confidence level is increased, and if it does not match, the confidence level is decreased; If not, selecting a target compression coding rule from multiple compression coding rules based on the access requirement information and compressing the data unit to be compressed, and replacing the compression coding rule used by the historical compressed data unit with the target compression coding rule when it is determined that the actual compression rate matches the access requirement information after compression.
2. The method according to claim 1, the method further comprising: After the data table is compressed and stored, compressing and storing a data table copy of the data table according to the recommended records.
3. The method according to claim 1, the access requirement information of the data unit to be compressed is determined by obtaining the historical access data of the data unit to be compressed and / or the historical compressed data unit.
4. A data compression device, the device comprising: A determination module, configured to take each column of data in a data table as a data unit to be compressed; A compression module, configured to perform the following compression processing on each data unit to be compressed: Obtaining access requirement information of the data unit to be compressed, where the access requirement information is obtained through the historical access frequencies of the data unit to be compressed and its historical compressed data units in the same column; Checking whether there is a recommended compression coding rule that meets the compression rate condition in the recommended records, where: The recommended records are used to record the actual compression rates and confidence levels of the compression coding rules used by the historical compressed data units in the same column; The compression rate condition includes: the confidence level of the recommended compression coding rule is higher than a set threshold, and the actual compression rate matches the access requirement information; If there is, using the recommended compression coding rule to compress the data unit to be compressed, and adjusting the confidence level of the recommended compression coding rule based on whether the actual compression rate matches the access requirement information after compression, where if it matches, the confidence level is increased, and if it does not match, the confidence level is decreased; If not, a target compression encoding rule is selected from multiple compression encoding rules based on the access requirement information, and the data unit to be compressed is compressed. When it is determined that the actual compression rate matches the access requirement information after compression, the compression encoding rule used by the historical compressed data unit is replaced with the target compression encoding rule.
5. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, the method according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Data compression method, device and system and server
CN102761540A
Encoding and decoding method and apparatus, and encoding and decoding device
CN108322220A